A LLM-enabled collaborative platform for data extraction, generation, and evaluation

The LLM-enabled collaborative platform addresses misinterpretations and inefficiencies in traditional systems by using contextual modifiers and relative truth parameters to enhance knowledge extraction and generation, ensuring accurate and scalable processing of complex text data.

WO2026019904A1PCT designated stage Publication Date: 2026-01-22ALLSCI CORP

Patent Information

Application Number
PCT/US2025/037888
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-14
Filing Date
2025-07-16
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Traditional collaborative platforms face challenges in processing complex text data due to misinterpretations caused by rule-based natural language processing models, reliance on pre-defined ontologies that introduce bias and subjectivity, and inefficiencies in constructing and querying knowledge graphs.

Method used

A large language model (LLM) is used to adapt to various text styles and formats, eliminating the need for pre-defined ontologies and enabling efficient extraction and generation of structured knowledge, with contextual modifiers and relative truth parameters to manage knowledge graph connections.

Benefits of technology

The LLM-based platform enhances accuracy and relevance of search results, provides real-time adaptation, and improves scalability by handling complex queries and dynamic datasets, facilitating efficient knowledge extraction and generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025037888_22012026_PF_FP_ABST
    Figure US2025037888_22012026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein are system, method, and computer program product aspects for textual data extraction, generation, and evaluation. Text is input into a first fine-tuned large language model (LLM) to generate an atom (e.g., a textual phrase in a particular category). The atom is input into a second LLM that has been fine-tuned for structured output corresponding to the particular category of information. A logical structure is generated based on a structured output of the second LLM, wherein the logical structure represents the textual phrase of the atom and contains a contextual attribute associated with the textual phrase. The embodiment then stores the logical structure into a knowledge graph as a modifier node having a time-variant attribute (e.g., a timestamp associated with the textual phrase and / or the contextual attribute).
Need to check novelty before this filing date? Find Prior Art

Description

A LLM-ENABLED COLLABORATIVE PLATFORM FOR DATA EXTRACTION, GENERATION, AND EVALUATIONCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of provisional U.S. Patent Application No. 63 / 672,631, filed July 17, 2024 and provisional U.S. Patent Application No. 63 / 772,227 filed March 14, 2025, both of which are hereby incorporated by reference in its entirety.BACKGROUND

[0002] This disclosure is generally directed to a collaborative, multi-source system, and more particularly to a large language model (LLM) enabled collaborative platform for textual data extraction, generation, and evaluation.BRIEF DESCRIPTION OF THE FIGURES

[0003] The accompanying drawings are incorporated herein and form a part of the specification.

[0004] FIG. 1 A is an example illustrating atom generation, according to some embodiments.

[0005] FIG. IB is another example illustrating atom generation, according to some embodiments.

[0006] FIG. 1C is an example illustrating a structured prompt input and a LLM output, according to some embodiments.

[0007] FIG. 2A is an example illustrating logical structure generation, according to some embodiments.

[0008] FIG. 2B is an example illustrating logical structure, according to some embodiments.

[0009] FIG. 3 is a block diagram of a collaborative platform, according to some embodiments.

[0010] FIG. 4A is a block diagram of a system engine, according to some embodiments.

[0011] FIG. 4B is a block diagram of a quantifier, according to some embodiments.

[0012] FIG. 4C is a block diagram of a visualizer, according to some embodiments.

[0013] FIG. 5 is a flowchart illustrating a method for textual data extraction, generation, and evaluation, according to some embodiments.

[0014] FIG. 6 illustrates an example computer system useful for implementing various embodiments.

[0015] FIGs. 7-15 provide additional features and benefits of a collaborative platform, according to some embodiments.

[0016] FIG. 16 is an example illustrating semantic similarity map generation, according to some embodiments.

[0017] FIGs. 17, and 18A-B are examples illustrating a semantic similarity map, according to some embodiments.

[0018] FIG. 19 is an example illustrating evidence generation, according to some embodiments.

[0019] FIGs. 20A-F are examples illustrating a chat bot interface and its corresponding functionalities, according to some embodiments.

[0020] FIGs. 21 A-B show an example flowchart illustrating generation of a unique user identifier, according to some embodiments.

[0021] FIGs. 22A-B show an example illustrating an expertise model, according to some embodiments.

[0022] In the drawings, like reference numbers generally indicate identical or similar elements. Additionally, generally, the left-most digit(s) of a reference number identifies the drawing in which the reference number first appears.DETAILED DESCRIPTION

[0023] Provided herein are system, apparatus, device, method and / or computer program product embodiments, and / or combinations and sub-combinations thereof, for data extraction, generation, and evaluation.

[0024] Embodiments described herein may allow users to perform a variety of actions including but not limited to publishing, extracting, generating, and / or evaluating different data independently, directly in a collaborative platform. The data may be any information or message conveyed in written or printed form including but not limited to books, articles, surveys, transcripts, and / or scholarly and scientific literatures. The scholarly and scientific literatures may include but are not limited to academic papers presentingempirical research and theoretical contributions that span various disciplines within the natural (including physical) and social sciences. In some embodiments, the collaborative platform may extract and / or generate different categories of data from the data input. The different categories of data may be referred to as ‘atoms’ - namely, textual phrases having features that characterize the textual phrase as being in a particular category of information. The particular categories of information may include, for example and without limitation, hypotheses, research questions, and / or research findings from literature, such as scientific literature.

[0025] Traditional collaborative platforms suffer from various technological problems and challenges associated with processing complex text data including extracting and / generating value information from the text data, and / or representing the text data in a structured way. Traditional collaborative platforms may rely on natural language processing (NLP) models to process complex text data. A NLP model may extract different entities (also referred to herein as “phrases,” or “atoms”) from unstructured text and then classify the extracted entities into different categories. The NLP may then generate relationships between the entities extracted from the unstructured text. However, many NLP models struggle with context and ambiguity, leading to misinterpretations due to the rule-based nature for extracting and / or generating categories of data.

[0026] Some traditional collaborative platforms also rely on a pre-defined ontology to represent text data and / or data categories in a structured way. Such an ontology can be complex and time-consuming to traverse, especially for large and dynamic domains. Additionally, ontologies may require expert knowledge and continuous updates to stay relevant. Ontologies are also typically based on human interpretation and understanding of a domain, which can introduce bias and subjectivity into the representation of knowledge. Scaling ontologies to handle large and diverse datasets can also be challenging.

[0027] Embodiments herein solve these technological challenges associated with processing complex text data through the use of a large language model (LLM) in conjunction with a unique logical structure. The LLM may handle the complex text data since LLMs are trained on vast datasets from diverse text sources. This approach may allow the LLM to adapt to various text styles and formats, and tackle various language tasks without needing specific training for each task. For example, some embodiments of the collaborative platform described herein may use a LLM for sequence-to-sequencetasks including but not limited to extracting a research hypothesis, generating and / or answering general factual knowledge questions, performing simple arithmetic, translating between languages, summarizing text, chat bot applications, and / or extracting research findings. The LLM-based collaborative platform may also be adept at complex scientific information extraction and generation in which the text input may be used to seed a text output (e.g., atoms) from the model. With the LLM, the collaborative platform may not need to use a pre-defined ontology because the predicates are extracted from the text data itself. The use of a LLM may resolve difficulties related to the complex rules for predefining an ontology to represent the complex scientific rules. In some aspects, the use of LLM may also supplement the value of a pre-defined ontology since the pre-defined ontology may be subjective when it considers how reality is perceived through individual perspectives and experiences.

[0028] Traditional collaborative platforms also suffer from technological challenges associated with identifying factual knowledge from scientific literature. Traditional collaborative platforms may only take a statement and state the extracted and / generated knowledge, including a hypothesis, as if it is true. However, the statement extracted by the traditional collaborative platforms may be taken out of context, or may otherwise be inaccurate, making the statement useless because the relationship and / or semantic structure including the nodes and edges of knowledge graph built upon the statement may need to be updated over time. Such a process was largely ineffective because the extracted and / or generated knowledge could be irrelevant or contradict what the original source material was trying to state. In addition, traditional knowledge bases will often display information as true in all contexts, when in fact information may only be true in some contexts. Yet traditional knowledge bases are not capable of maintaining such context with regard to information stored therein.

[0029] Traditional collaborative platforms also suffer from technological challenges associated with constructing a knowledgebase to store and / or traverse the extracted and / or generated knowledge. Previously, it would have been time-consuming to construct a knowledge graph since traditional collaborative platforms had to address the complex rules for extracting different entities (e.g., the nodes of the knowledge graph) and / or extract the relationships between the entities (e.g., the edges of the knowledge graph). Before extracting these entities and their relationships, the traditional collaborative platforms also need to define an ontology that captures knowledge that exists in thedomain and the properties that can be used to describe them. After constructing a knowledgebase, traversing and / or querying the information stored in the knowledge graph is also challenging. The knowledge graph may often be composed of millions or billions of nodes and edges, and stored in a graph database. For example, a knowledge graph may be stored in a resource description framework (RDF) data store, with nodes stored as triples. The RDF triple data store is a type of graph database that may store semantic facts and supports optional schema models that allow for formal description of the data. Alternatively or additionally, a knowledge base may be stored using other graph database solutions such as, for example and without limitation, GraphDB available from TigerGraph, Inc. of Redwood City, CA, Neo4J Graph Database available from Neo4j, Inc. of San Mateo, CA, and / or Amazon Neptune available from Amazon.com, Inc. of Seattle, WA.

[0030] However, querying such a knowledge graph / graph database requires specialized knowledge in query languages as well as a deep understanding of the underlying structure of these graphs, making a wide range of end-users without deep knowledge of these technical concepts excluded from querying these knowledge graphs effectively. Also, even an expert in the field may also find it challenging to effectively query the information they want in the traditional knowledge graph in an efficient manner, due to a large volume of extracted and / generated knowledge and information having been stored into the knowledge graph.

[0031] Embodiments described herein solve these technological challenges associated with stating factual knowledge from scientific literature by linking (also referred to herein as connecting) supporting or refuting evidence to the extracted and / or generated knowledge (e.g., hypotheses) using LLMs and modifiable nodes in a knowledge graph built upon the linking of the extracted and / or generated data. For example, a node may be modifiable in that the node’s connections to other nodes (e.g., the supporting or refuting evidence) may be changed (e.g., updated or added) over time. When such connections incorporate a time-variant attribute such as a timestamp, such changes may be monitored and / or provided by the collaborative platform, and such changes give context to the extracted and / or generated knowledge. For example, such connections allow the determination of how many times a hypothesis is supported and / or refuted from the evidence. For example, the collaborative platform may assign a timestamp to every new result (e.g., a vote) in the database so that the collaborative platform can roll back andforth in time to see the support or refutation of a particular hypothesis having been taken place over time. This linking may also be added over time, creating an effective representation of the extracted and / or generated knowledge in a structured way.

[0032] Embodiments described herein also solve these technological challenges by incorporating one or more contextual modifiers into the underlying logical structure of the node to provide a context for the information contained in the logical structure (also referred to herein as a symbolic logical structure). For example, a document may provide a particular statement (e.g., “Aspirin reduces heart disease.”), only to then refute that statement or limit the applicability of that statement (e.g., “However, such correlation has only been determined in rats.”). Without context, the first statement may be preserved in a logical structure simply as a fact. By adding one or more contextual modifiers to the logical structure, however, knowledge can be extracted from a document in a way that is relevant and supported by the context of that knowledge (e.g., the first statement with the context of the second statement).

[0033] In some embodiments, the collaborative system may improve triples by adding contextual modifiers to the triples. For example, some embodiments may represent the triples in a Neo4j graph database using JSON as an intermediary datatype between the LLM and the knowledge graph, although a skilled artisan would recognize that other databases may alternatively be used. The collaborative platform may refine the triple by adding a “relative truth parameter” for the predicate and extracting information including a modifier — that is a word, phrase, or clause that modifies and / or gives information about that predicate — that protects the relative truth parameter of that predicate.

[0034] The collaborative platform may generate explicit logical structure and connections between the extracted and / or generated knowledge, populate a knowledge graph with them, and / or represent the extracted and / or generated knowledge in a way that is relevant and supported by the scholarly and scientific literatures. The collaborative platform may populate the knowledge graph by storing the logical structure into the knowledge graph as a modifiable node having a time-variant attribute (e.g., the assigned timestamp).

[0035] Once the logical structure is stored in the knowledge graph, connections may be made between the knowledge graph and other logical structures in the knowledge graph. For example, the various logical structures may be nodes in the knowledge graph, and the connections may be edges between the nodes. In some embodiments, a relative truth parameter of the logical structure may be added or updated to identify, for example andwithout limitation, a number of votes of support or refutation given to the logical structure over time. In the example of scientific literature where the logical structure represents a hypothesis, the relative truth parameter may be updated based on a number of votes provided by support or refutation evidence contained in other logical structures in the knowledge graph. In some embodiments, the relative truth parameter gives an indication of the probability of whether a hypothesis is likely to be true by quantifying the number of votes for support or refutation over time, instead of simply stating whether it is a true hypothesis or not. The votes may be identified via a connection between one logical structure (e.g., the extracted and / or generated knowledge) and another logical structure that identifies, for example and without limitation, a variant, equivalent, contradictory, and / or related relationship between the logical structures. The relative truth parameter may be stored as a modifier of the logical structure. In some embodiments, each new connection creates a new (updated) relative truth parameter that is stored as a new modifier of the logical structure. New connections and / or modifiers for the logical structure may each contain their own time-variant attribute. This allows a user to roll back and forth in time to monitor the status of the extracted and / or generated knowledge over time (e.g., the support and / or refutation of a particular hypothesis over time).

[0036] The collaborative platform may use the relative truth parameter and / or the modifier to query different entities in the knowledge graph. By using the relative truth parameter and / or the modifier, the collaborative platform may provide a shortcut connection between different entities in the knowledge graph, such that when traversing the knowledge graph, the user may efficiently query the related entities given a located entity. The collaborative system may also use the connection between different atoms including but not limited to a supporting or refuting evidence to an extracted and / or generated knowledge to traverse a knowledge graph and / or between different knowledge graphs. The connection between atoms may also provide a shortcut connection between different entities within a knowledge graph and / or across different knowledge graphs for any querying purposes. The short cut provided in a knowledge graph based on the relative truth parameter and the modifier from the logical structure, and / or the connection between the atoms, may resolve difficulties related to the inefficiency of traversing the knowledge graph, making the collaborative system more scalable and faster to extract and / or generate knowledge than traversing a knowledge graph formed from traditional semantic triples.

[0037] The collaborative platform may use an agentic atom search and refine engine when querying (and refining) atoms in the knowledge database (or knowledge graph). This agentic search engine may offer at least an increased accuracy and relevance of search results, a real-time adaptation to user queries, and a proactive insight of actions.

[0038] This agentic engine may better interpret the intent of the end user behind a search query (usually a complex search query) and deliver query responses that are more contextually relevant. The agentic engine may handle complex queries that involve multiple steps or require reasoning. For example, the agentic system may decompose the complex user queries into manageable sub-questions, and then gather pertinent information from each sub-query. This agentic engine may continuously learn and adjust its responses based on real-time user feedbacks and the current situation. The agentic engine may not only find relevant data but also interpret and contextualize the data based on its understanding of the current situation, providing a more dynamic and relevant response. The agentic engine, beyond merely providing search or refine responses, may also proactively suggest related information, tasks, or actions based on the behaviors or contexts of the end user.

[0039] The agentic engine may also orchestrate multiple actions in a specific order or simultaneously based on at least the user query, any current situations, and / or computational resources. This agentic engine may then assign the multiple actions to other specialized agents or tools to carry out individual actions. This action assignment can maximize the efficiency based on a coordination from a network of agents, ensuring that the all actions are properly timed and aligned with each other to achieve a more focused functionality and the maximum efficiency. In addition, unlike a traditional collaborative platform that relies on a static data set, in some embodiments this agentic engine processes the user query in real-time by actively querying various data sources such as databases, APIs, and the web. This fresh information may be consistently pulled in and analyzed on the fly. This continuous data digest allows the agentic engine to continuously monitor and react to changes as they happen, making decisions based on the latest available data.

[0040] Embodiments herein may expand the amount of science published and the speed at which it is communicated. The collaborative platform may create incentives to replicate scientific work and publish negative results. The collaborative platform may allow for streamlined and more extensive peer-review and credit for participating in thepeer-review process earlier in the researcher workflow. The collaborative platform may allow the users to add value to each other’s contributions through at least peer review and / or linking hypotheses to questions, findings to hypotheses. The collaborative platform may measure, compare, and / or analyze researcher contributions to science. For example, the collaborative platform may help researchers and organizations monitor their own and other’s contributions, impact, and performance through a variety of measures. The collaborative platform may also provide artificial intelligence (Al) tools for researchers to improve their hypotheses, questions, and / or research design. The collaborative platform may provide a marketplace between peers, funders, and / or suppliers in which researchers, funders, products, and services can connect in long-term. The collaborative platform may also help researchers determine where they should specialize to add the most value to the scientific community.Atom Generation

[0041] Embodiments of the collaborative platform herein may extract and generate different atoms. For example, when the input text is scientific literature, the atoms may represent hypotheses, research questions, and / or research findings from the scientific literature. While scientific literature will be referred to herein as one example use case, a skilled artisan will recognize that similar categorization techniques and logical storage as a logical structure are applicable in other use cases for other types of texts as well.

[0042] In the scientific literature example, a hypothesis may be extracted for identification and / or display by the collaborative platform. That is, from an input scientific article, data corresponding to a hypothesis statement of the article may be identified and / or generated by the collaborative platform as an atom - a unit of scientific knowledge. Such hypotheses may help design studies, allow for practical testing, and add to scientific knowledge. One role of the hypotheses to the collaborative platform may be to organize research projects, making them purposeful, focused, and valuable to the scientific community. A hypothesis can also help discover similar hypotheses and identify gaps in the existing hypotheses. A hypothesis can also help users to formulate a specific and testable research problem, and to design an appropriate method to collect and analyze data. In addition, the hypotheses can be tested to answer a research question. Using a collection of hypotheses, the collaborative platform can group similar hypotheses together, and generate related research questions associated with the group of hypotheses.A well-defined and specific research question is more likely to guide users in making informed decisions on research design and population and subsequently indicates what data will be collected and analyzed in the research projects.

[0043] In some embodiments, the collaborative platform may extract and / or generate atoms based on an NLP system for structuring an existing body of scientific literature. The NLP system may rely on named entity recognition (NER) to extract the atoms. NER is a subtask of information extraction that seeks to locate and classify named entities mentioned in unstructured text into pre-defined categories including but not limited to fields, backgrounds, and / or properties. The NLP system may also use a robust relation extraction and / or generation technique to accurately extract and / or generate the relationships between the named entities from the unstructured text.

[0044] In some embodiments, the collaborative platform may extract and / or generate the atoms based on a LLM. The collaborative platform may be particularly adept at sequence- to-sequence tasks where the text input may be used to seed a text output (e.g., atoms) from the model. For example, the use cases of a LLM for sequence-to-sequence tasks are broad and may include but are not limited to extracting research hypotheses, generating and / or answering general factual knowledge questions, performing simple arithmetic, translating between languages, summarizing text, chat bot applications, and / or extracting research findings. It may stand to reason that the collaborative platform based on a LLM may also be adept at complex scientific information extraction and generation. In some embodiments, the LLM has been fine-tuned to output atoms from a new text based on atoms generated from other texts of a similar type. For example, in the example of scientific literature, the LLM may have been fine-tuned to output atoms from a newly input scientific document based on atoms generated from other scientific documents.

[0045] In some embodiments, the collaborative platform may record or publish the extracted or generated atoms. The platform may make the process of creating and extracting those atoms transparent in which the sources of the atoms (e.g., source paragraph, source documents, etc.) are visible to the users of the collaborative platform. The platform, in some embodiments, may also share the involved steps and data used for the atom extraction and generation (e.g., via a chat bot or a user interface), including the data used, algorithms applied, and decision-making behind the atom creation, making the atom extraction or generation understandable and accessible to the end user.

[0046] In some embodiments, the collaborative platform may automatically refresh or update its database with new data on a regular basis (e.g., daily, weekly, monthly, etc.) for the atom extraction and generation. This update may happen automatically at a designated time each period. For example, if the update occurs daily, the update may happen automatically at a designated time each day (e.g., overnight) to minimize the service disruption to the end user. In some embodiments, the new data may come from various sources and / or APIs, either internal and / or external, which accumulates throughout the day or a period of time. The new data may be cleaned and validated before being ingested into the collaborative platform. In some embodiments, the platform may use a database management system, including but not limited to relational databases (e.g., MySQL, PostgreSQL, any other SQL queries such as Query DSL in addition to SQL, and the like) and / or non-relational databases (e.g., NoSQL databases, OpenSearch, and the like) to store and manage the new data, and sometimes, transform the data into a format compatible with the existing database.Hypothesis generation from scientific article example

[0047] FIG. 1 A shows an example illustrating atom generation, according to some embodiments. Atom generation method 100 may be performed by an inference pipeline of knowledge generation module 424 of system engine 420 in FIG. 4A. However, atom generation method 100 is not limited to that example implementation.

[0048] In some embodiments, text including scientific article 102 is received. At 104, the text may be tokenized to obtain data chunks 106. A data chunk is a set of text from the source that is smaller than the source. In some embodiments, the data chunks may include a paragraph. In some embodiments, the data chunks may include a plurality of sentences. For example, in some embodiments, the data chunks may include three sentences. In some embodiments, the data chunks may include a single sentence. In some embodiments, portions of data chunks may overlap. For example, in some embodiments where the data chunks include three sentences, a first sentence and a third sentence of the data chunk may be overlapped with an adjacent data chunk obtained in tokenization.

[0049] At 108, keyword searching within data chunks 106 may be performed to obtain a set of ranked lists 110 of the data chunks, wherein each ranked list within the set of ranked lists 110 is associated with each of a plurality of keywords known to be indicativeof the particular category of information. An element of the ranked list may include a relevance score between a keyword and one of data chunks 106.

[0050] At 112, a set of ranked lists 110 may be fused to obtain a fused ranked list. In some embodiments, the fusion 112 may be performed by at least a reciprocal rank fusion (RRF) and / or other fusion algorithms that would be appreciated by a person of ordinary skill in the art. In some embodiments, the fusing performed at 112 may include combining a set of lists of relevance scores associated with set of ranked lists 110. For example, in some embodiments the relevance scores may include but are not limited to BM25 scores and / or other classic similarity scores. Top-k data chunks 114 may then be extracted from data chunks 106 based on the fused ranked list.

[0051] At 116, a structured prompt containing top-k data chunks 114 may be input into a first LLM 118. In some embodiments, first LLM 118 may include a fine-tuned Mistral Instruct model (e.g., a decoder model having 7 billion parameters under the APACHE 2.0 license). First LLM 118 may include multiple LLMs of a same or different type. In some embodiments, the multiple LLMs may be chained and / or parallelized. The LLM chaining and / or parallelizing is a process of connecting multiple LLMs or connecting one or more LLMs to other applications, tools, and services to produce the best possible output from a given input. First LLM 118 may be fine-tuned in a supervised manner for a specific kind of structured output (e.g., hypotheses). The weights of first LLM 118 may be quantized during fine-tuning. Instead of training whole first LLM 118, a quantization and low-rank adapters (QLoRA) may be used to train an adapter, enabling the fine-tuning of a massive LLM with billions of parameters on one or more relatively small, highly available GPUs and in a more efficient way. During the fine-tuning, more weight may be placed on monitoring the recall.

[0052] In addition, when generating synthetic training data, other LLMs may be leveraged and the schema defined at the end before prompting. The outputs may be inspected and the augmentation and prompts reiterated to get the best output. Small batches of outputs of LLM 118 may be vetted and the prompt and ranking schemes finetuned accordingly. The prompt may then be inputted into the model after augmentation. Outputs 120, including but not limited to relevance metrics for the top-k data chunks, main data category (e.g., hypothesis), and / or specific data categories (e.g., specific hypotheses), may be generated.Research question generation from hypothesis example

[0053] FIG. IB shows another example illustrating atom generation, according to some embodiments. Atom generation method 150 may be performed by knowledge generation module 424 of system engine 420 in FIG. 4A. However, atom generation method 150 is not limited to that example implementation.

[0054] In some embodiments, an atom 152 generated from the text and / or from another atom is embedded into a vector space 152. For instance, referring back to the scientific article example in atom generation 100, atom generation 100 may generate atom 120 for this embedding. The atom 120 generated from the text in atom generation 100 may refer to atom 152 in this embodiment, but they might not be tied in other embodiments. In some embodiments, atom 152 may be generated from one or more other atoms. In this example, atom 152 may include a hypothesis and vector space 152 may include a hypothesis space. The embedding may be performed by using, for example, a pre-trained BAAI / bge model and / or other text embedding techniques that would be appreciated by a person of ordinary skill in the art.

[0055] At 154, clustering may be performed within vector space 152 to generate one or more preliminary clusters 156 of the atoms. The clustering may include, for example, performing a k-means clustering and / or other clustering techniques that would be appreciated by a person of ordinary skill in the art. The number of clusters in the k-means clustering may be determined by, for example, an elbow method, a silhouette score, and / or a LLM capability.

[0056] At 158, one or more atoms from each of the generated preliminary clusters of atoms 156 may be sampled. An example of a cluster of sampled atoms is shown in 160.

[0057] At 162, a structured prompt containing a cluster of sampled atoms 160 may be input into a first LLM. In some embodiments, the first LLM may include a fine-tuned Llama-3 model. For example, the model may be trained as an adapter using Quantized Low-Rank Adaptation (QLoRA), enabling the efficient fine-tuning of a massive LLM with billions of parameters on one or more relatively small GPUs (e.g., an NVIDIA RTX 4090 GPU available from NVIDIA Corp, of Santa Clara, CA). In some embodiments, the first LLM may include multiple LLMs of a same or different type.

[0058] Outputs 164, including but not limited to one or more refined clusters of the atoms and / or research questions associated with the one or more refined clusters, may begenerated based on an output of the one or more first LLMs. An example of generated research questions is shown in 164.

[0059] FIG. 1C shows an example illustrating a structured prompt input and a LLM output, according to some embodiments. Structured prompt input and LLM output 170 may correspond to the input and output, respectively, of atom generation method 150 in FIG. IB. However, structured prompt input and LLM output 170 are not limited to that example embodiment.

[0060] In some embodiments, a structured prompt input 172 may include an instruction 172a, a raw atom cluster 172b, and / or a structure and object definition 172c. Raw atom cluster 172b may be, for example, a raw hypothesis cluster. LLM output 174 may include a cluster of atoms 176. Cluster of atoms 176 may include an intermediate research question 176a and / or a main research question 176b.Logical Structure Generation

[0061] Embodiments of the collaborative platform herein may generate symbolic logical structures from the atoms. Continuing the scientific literature example, the atoms may include but not be limited to hypotheses, research questions, and / or research findings from the scientific literature. In some embodiments, a logical structure may include a semantic triple and as many as three modifiers — predicate modifier, subject modifier, and object modifier. A semantic triple may be a sequence of three entities that codify a statement about semantic data in the form of subject-predicate-object expressions. The modifier may be a word, phrase, or clause that modifies — that is, gives information about — another word in the same sentence. A modifier can be an adjective but it can also be an adverb. A modifier can also be a phrase or clause. In some embodiments, the modifier is a relative truth parameter or relative truth statement.

[0062] In some embodiments, the collaborative platform may use the logical structure to query different entities in the knowledge graph. By providing the relative truth parameter and / or the modifier, the collaborative platform may provide a shortcut connection between different entities (e.g., nodes) in the knowledge graph, such that when traversing the knowledge graph, the user may efficiently query the related entities given a located entity (e.g., node).

[0063] In some embodiments, the logical structure may be generated and / or extracted based on NLP techniques. The NLP techniques may separate the generation process intotwo tasks including but not limited to NER to extract entities and / or identify relationships between the named entities. Sequence labeling may first be used to identify all entities in a sentence, and then relation detection performed through various classification networks.

[0064] In some embodiments, the logical structure may be generated and / or extracted based on a neural-network-based approach. Such an approach may include but is not limited to recurrent neural network (RNN), long short-term memory (LSTM), and / or other encoder-decoder models. Such a model may use a parameter sharing mechanism to extract entities and relations from a network and / or encoder-decoder models. The collaborative platform may treat a logical structure as a token sequence and may employ the encoder-decoder framework to generate triple elements like machine translation.

[0065] In some embodiments, the logical structure may be generated and / or extracted based on a LLM. Such an LLM may be a same or different from the LLM used to extract and / or generate the atom. The logical structure may be extracted and / or generated via zero-shot or few-shot learning. By doing this, the collaborative platform may perform well in a variety of downstream tasks on relational triple extraction and generation, even when provided with only a few examples as instructions.

[0066] FIG. 2A shows an example illustrating logical structure generation, according to some embodiments. Logical structure generation method 200 may be performed by knowledge generation module 424 of system engine 320 in FIG. 4A. However, logical structure generation 200 is not limited to that example implementation.

[0067] In some embodiments, at 204, an input atom 202 may be parsed into a plurality of phrases. Continuing the scientific literature example, atom 202 may include a sentence including but not limited to a hypothesis from a received text. In some embodiments, parsing 204 may assist in reporting any errors with syntax of the atom 202. Parsing 204 may assist in recovering from errors that are often experienced, allowing the processing of the remaining portions performed in the collaborative platform to continue uninterrupted. Parsing 204 may also generate a symbol table and / or create intermediate representations of atom 202. Parsing 204 may include but is not limited to recursive descent parsing, shift-reduce parsing, chart parsing, regular expression parsing, and / or dependency parsing.

[0068] At 206, one or more text entities may be extracted from the plurality of phrases parsed from 204. Entity extraction 206 is the process of breaking down textual phrases or sentences (e.g., atoms 202) into their proper parts of speech (e.g., nouns, verbs,adjectives, adverbs, etc.) based on word definitions and context. With this information, noun phrases can be identified which, in turn, help to identify the specific pieces of information or objects within a text that carry particular significance.

[0069] At 208, a structured prompt for one or more second LLMs may be generated, where the structured prompt may contain the one or more text entities extracted at 206. In some embodiments, the structured prompt may include one or more prompt examples to one or more second LLMs, where the prompt examples may include the expected outputs and / or responses from the one or more second LLMs.

[0070] At 210, generated structure prompt from 208 may be input to the one or more second LLMs. In some embodiments, the one or more second LLMs may include a finetuned Mistral Instruct model. In some embodiments, the model may be trained as an adapter using QLoRA, enabling the efficient fine-tuning of a massive LLM with billions of parameters on one or more relatively small, highly available GPUs. In some embodiments, a rotary position embedding (RoPE) may be used to embed the sequence input of the model when the sequence length of the model used may be longer than the graph output. The RoPE is a type of position encoding and / or embedding technique for NLP that may encode absolute positional information with a rotation matrix and may naturally incorporate explicit relative position dependency in self-attention formulation of the model.

[0071] In some embodiments, the one or more second LLMs may demonstrate zero-shot capabilities. In some embodiments, few-shot prompting may be used as a technique to enable in-context learning where demonstrations are provided in the prompt to steer the one or more second LLMs to better performance on more complex tasks than using the zero-shot setting. In few-shot prompting, the demonstrations may serve as conditioning for subsequent examples where the collaborative platform would like the LLMs to generate a response. In addition, logical structure 212 may be generated based on the structured output of the one or more second LLMs.FIG. 2B illustrates an example logical structure, according to some embodiments. Logical structure 222 may be generated by logical structure generation method 200 in FIG. 2A. However, logical structure 222 are not limited to that example embodiment. In some embodiments, logical structure 222 may be generated from an atom 242, such as a hypothesis. Logical structure 222 may include a semantic triple and a modifier. The semantic triple may include a subject 244, an object 248, and a predicate 246 that relatessubject 244 to object 248. The modifier may identify a contextual attribute for the semantic triple including, for the semantic triple, each of a subject modifier 254, a predicate modifier 256, and an object modifier 258. Atom 242 may be generated from a text data 232 including but not limited to an article. The collaborative platform may generate and / or extract other metadata from text data 232 including but not limited to an author 234a, a topic 234b, and / or a journal 234c. Some metadata may be generated and / or extracted from other metadata. For example, the collaborative platform may also generate and / or extract an institution 236 from the author 234a.Semantic Similarity Map Generation

[0072] Embodiments of the collaborative platform herein may generate a semantic similarity map for the generated or extracted atoms (e.g., hypotheses). This semantic similarity map may be a visualization of how closely related the atoms are. This semantic similarity map can group atoms that are closely related and space apart atoms that are less related. A pixel or a pixel grid may be used in the semantic similarity map to represent an active semantic feature, for example, a hypothesis in each hypotheses cluster. The distance between two pixels or pixel grids on the map may often be used as a quantitative measure of semantic similarity, with shorter distances signifying higher similarity. The semantic similarity of two items may also be compared in the semantic similarity map based on their pixel grids. Unlike simple word matching, the semantic comparison may focus on the actual meaning conveyed by a term, not merely the literal wording.

[0073] In some embodiments, the semantic similarity can be measured using a variety of methods, including but not limited to, statistical, topological, and / or machine learning approaches. For example, vector space models, including not but limited to Word2Vec and GloVe, may be used to embed words or phrases as vectors in a high-dimensional space. A distance metric may then be used to calculate the similarity between the two vectors. The distance metric may include, but is not limited to, a cosine similarity, a Euclidean distance, a Manhattan distance, and / or a Jaccard similarity. The semantic similarity may also use a tool to measure the distance between two pixels or pixel grids on the map to calculate a numerical representation of their semantic similarity.

[0074] FIG. 16 shows an example flowchart illustrating semantic similarity map generation, according to some embodiments. Semantic similarity map generation method 1600 may be performed by, for example, knowledge generation module 424 of systemengine 320 in FIG. 4A. However, semantic similarity map generation method 1600 is not limited to that example implementation.

[0075] At 1610, for each hypothesis of hypotheses 1605, a nearest neighbor clustering may be applied in which a predefined number of nearest neighbors may be obtained and returned as a cluster. In this step, different distance metrics may be used to calculate the distance between the hypothesis and each of its nearest neighbors.

[0076] At 1615, the hypotheses clusters obtained from the nearest neighbor clustering 1610 may be embedded. This embedding may transform the hypotheses clusters into a vector representation. The embedding may be performed by using Word2Vec, GloVe, or other word embedding approaches. The embedding model may be a custom trained model optimized for a particular objective using, for example, representation learning. For example, where the hypotheses relate to particular drugs or compounds, this custom trained model may bring vectors which mention the same drug name closer in the embedding space. This embedding may also be performed by a pre-trained BAAI / bge model and / or other text embedding techniques that would be appreciated by a person of ordinary skill in the art.

[0077] At 1620, a dimension reduction may then be applied to the embedding of the hypotheses clusters. In this step, uniform manifold approximation and projection (UMAP), t-SNE, or any other linear or non-linear dimension reduction algorithms may be applied. In some embodiments, the dimension may be reduced to two-dimensional plane or three-dimensional space for visualization, but the dimension can also be reduced to any dimensions for different purposes.

[0078] In some embodiments, at 1625, the reduced embedding may be ranked within each hypotheses cluster. This ranking of reduced embedding within each hypotheses cluster may order the hypotheses data points (each with a linking to the original hypothesis) within a specific cluster, where each data point may be assigned a relative position based on its similarity to other points within that cluster. This ranking may be used to identify the most representative data points within a cluster and then allow for deeper analysis of the cluster’s characteristics and potentially revealing sub-patterns within the cluster. The ranking may be based on a cosine similarity or other semantic similarity measures that would be appreciated by a person of ordinary skill in the art.

[0079] At 1630, after obtaining the ranked reduced embedding, a mapping of the original hypotheses to the ranked reduced embedding (of each hypothesis) may be applied. Insome embodiments, a mapping between each original hypothesis and its reduced embedding may be stored in a structured format, such as a dictionary or database that supports efficient similarity search or retrieval. Using such mapping, the original hypothesis may be easily retrieved given the reduced embedding.

[0080] At 1635, after obtaining the original hypotheses mapped to the ranked reduced embedding, an LLM may be queried to generate the topic that summarizes the hypotheses within each cluster. Depending on the length and complexity of the list and the capabilities of the chosen LLM, the whole hypotheses may be provided in a context window, to direct the LLM to condense the list of hypotheses into a concise summary or topic. In some embodiments, this may instead be done by summarizing each hypothesis individually and then combining those summaries to create a final summary.

[0081] At 1640, in response to the LLM query in 1635, the LLM can generate a response, for example, the topic or summary for each hypothesis cluster. In some embodiments, the LLM can also be used to refine, either automatically or interactively with the end user, the topics being generated. Once the topic modeling is satisfactory to the end user or has achieved a pre-defined quality threshold, the topics may be provided as a label (e.g., with different colors) to each of the hypothesis clusters on the semantic similarity map.

[0082] FIGs. 17 and 18A-B show examples illustrating a semantic similarity map, according to some embodiments. However, the semantic similarity map is not limited to that example implementation.

[0083] As illustrated in FIG. 17, in some embodiments, the semantic similarity map may provide a user interface that supports the end user’s search for certain hypotheses or other atoms within that map. For example, the end user may provide a product name or any keywords they want to search. This search may be performed, for example, through a visible search bar, enabling the end user to quickly navigate and access relevant content by searching for keywords. The user interface may also reflect such keyword searching within the semantic similarity map and visualize the connection between the search keywords and any related information associated with such keyword.

[0084] In some embodiments, the end user may navigate the semantic similarity map using the user interface, for example, by moving a curser into a particular location (e.g., a pixel or a pixel grid), as illustrated in FIG. 18 A. The actual hypothesis at that location may be displayed (e.g., pop up) to the end user, as illustrated in FIG. 18B. In some embodiments, the hypothesis may be displayed with a relevance level to other hypotheseswithin a cluster. In some embodiments, a window may appear on top of the main screen, displaying information or requiring a user action, such as a notification, a confirmation message, or an error alert. The display of information may be triggered by a specific interaction between the end user and the user interface.

[0085] In some embodiments, as illustrated in FIG. 18 A, the semantic similarity map may be visualized in terms of time. For example, a visual representation of different distributions or clustering of the hypotheses may be created over a period of time. Timeinvariant visualization can be shown using different at least color-coding, animations, or layers to illustrate how the overall semantic similarity (or the semantic similarity at a pixel or pixel grid) changes or evolves throughout a specific timeframe. In some aspects, the user interface may include a slider to show the visual representation at different times over a period of time.Atom Linking

[0086] Embodiments of the collaborative platform herein may use Al models to link (i.e., connect) different atoms extracted and / or generated from the input text. A goal of atom linking is to identify and retrieve relevant atoms from a collection of different atoms. The atom linking can be performed using algorithms that search through the atom collection. Atom linking may be useful in the collaborative platform because it may allow for finding patterns in the extracted and / or generated atoms. A relationship between atoms may be identified as between at least one of a hypothesis and a variant, equivalent, contradictory, or related hypothesis; the hypothesis and a supporting, refuting, inconclusive, or unrelated research finding; and / or the hypothesis and a related research question. The linking between atoms may also provide a shortcut connection between different atoms and / or entities (entities may be extracted from the atoms) within a knowledge graph and / or across different knowledge graphs for any querying purposes. The shortcut provided in a knowledge graph based on the linking (relationship) between the atoms may resolve difficulties related to the inefficiency of traversing the knowledge graph with a large volume of entities (e.g., nodes) and relationships (e.g., edges), making the collaborative system more scalable and faster to extract and / or generate knowledge than traversing a knowledge graph formed from atoms without this linking.

[0087] A user may provide the collaborative platform with a query, where the query may include but is not limited to a few words or a particular sentence. The collaborativeplatform may then index the query with the related data and / or match the terms in the query with the related data. The collaborative platform may rely on machine learning algorithms that discover data patterns through supervised and / or unsupervised learning approaches to find relevant data patterns between different atoms.

[0088] In some embodiments, the collaborative platform may use supervised learning approaches to link different atoms together. For example, the collaborative platform may use some characteristics of the atoms, including but not limited to entities, attributes, keywords and / or other textual information, to find relevant atoms from a database storing the atoms and / or a knowledge graph. A supervised learning model may be trained based on these characteristics to identify whether the atoms are matched to each other. Supervised learning models to link the atoms may include but are not limited to handcrafted retrieval models, semantic-based models, term dependency-based models, and / or learning to rank models. Deep learning approaches, on the other hand, may include but are not limited to methods of representation learning, methods of matching function learning, and / or methods of relevance learning.

[0089] In some embodiments, the collaborative platform may use unsupervised learning approaches to link different atoms. For example, unsupervised learning problems exist where there are no specific keywords or attributes by which the collaborative platform can find relevant atoms from the database storing the atoms; instead, the collaborative platform may look at patterns of the atoms in various data sources to find relevant information. Some features of such unsupervised learning models include embedding of the different atoms into a representation, identifying related atoms in the database within a distance metric, and ranking techniques for ranking different atoms based on their relevancy score and / or distance metric regarding the embedding of the atoms.Evidence Engine

[0090] Embodiments of the collaborative platform herein may use an evidence engine (e.g., based on the LLM or any other Al models) to link a hypothesis to the supporting or refuting evidence. This evidence finding for hypotheses may allow the end user to determine whether a hypothesis is likely to be supported or refuted based on the collected evidence from multiple sources, rather than the platform providing a definitive assertion that the hypothesis is true or false. In some embodiments, an end user may gather relevant data or evidence from those multiple sources through various methods, including but notlimited to, surveys, interviews, observations, or controlled experiments, depending on the hypothesis. The collected evidence may be from multiple articles and / or article snippets belong to different articles. The evidence may also be organized in terms of time, indicating in which year the supported or refuted evidence was released or published. In some embodiments, the evidence engine may operate in a real-time manner, in which the evidence may be collected simultaneously when the end user enters the hypothesis.

[0091] FIG. 19 shows an example flowchart illustrating evidence generation, according to some embodiments. Evidence generation method 1900 may be performed by, for example, knowledge generation module 424 of system engine 320 in FIG. 4A. However, evidence generation method 1900 is not limited to that example implementation.

[0092] At search 1910, a hypothesis 1905 may be used to search for any supporting or refuting evidence from one or more internal or external sources or databases, including but not limited to, scientific literatures (e.g., articles), websites, practice guidelines, data sets, social media, clinical trials, and / or any other sources having relevant information to the hypothesis. In some embodiments, to augment the search performance, a null hypothesis or the opposite hypotheses associated with hypothesis 1905 may be generated and used, making sure the negatives hypotheses will also be considered to elicit as many relevant articles as possible.

[0093] At search 1910, the keywords of hypothesis 1905 may be extracted. The keyword extraction of hypothesis 1905 (including the positive, negative, opposite, and null hypotheses) may be conducted by at least one of an unsupervised method, a graph-based method, and / or a supervised method. Unsupervised methods may include, for example, and without limitation, rapid automatic keyword extraction (RAKE) and yet another keyword extraction (YAKE) in which RAKE may analyze word co-occurrence patterns in the hypotheses to identify keywords and YAKE may use local text features and statistical information of the hypotheses to extract keywords. Graph-based methods may include, for example and without limitation, selectivity -based keyword extraction that extracts nodes from a graph representation network built upon the hypotheses as keyword candidates. Supervised methods, for example, KeyBERT, may utilize bidirectional encoder representations from transformers (BERT) to convert words or phrases in the hypothesis into high-dimensional vectors or embedding and then apply cosine similarity or another clustering method to identify a list of the most important words or phrases (ranked by their relevance or importance) in the hypothesis.

[0094] In some aspects, the keyword extractions of hypothesis 1905 may be conducted using LLMs, for example KeyLLM. KeyLLM may be combined with KeyBERT for keyword extraction, but KeyLLM may also work without KeyBERT. For example, KeyLLM may create keywords for each hypothesis 1905. This creation of keywords may be implemented by asking LLMs to come up with a plurality of keywords (e.g., reference keywords) for each related hypothesis. In some aspects, if the reference keywords are already pre-extracted or available at the database, this keyword creation may use those keywords. KeyLLM may then ask the LLMs to check whether these reference keywords actually appear in hypothesis 1905 and limit the keywords search to those that are found in the hypothesis. After obtaining a list of keywords from hypothesis 1905, this list of keywords may be fine-tuned by asking the LLMs.

[0095] After extracting the keywords for hypothesis 1905 at search 1910, those extracted keywords may be used to find any articles (or article snippets) 1915a-n, from the internal or external sources, that are relevant to the hypothesis. The keyword search of relevant articles may be performed using, for example, semantic search API endpoint. But for granular search (e.g., using semantic similarity on top of the semantic search API endpoint) may also be leveraged in some of the embodiments. Matching conditions between the hypothesis and the articles (or article snippets) may include an exact match, a phrase, and / or a broad match to dictate how closely the search query needs to align with the keyword to trigger a match. An exact match may refer to a search query that must precisely match the keyword, including word order and variations like plurals. The phrase match may indicate that the search query must include the exact keyword phrase, but can have additional words before or after the phrase. The broad match allows the search query to include related terms, synonyms, and variations of the keyword. In some embodiments, based on a suitable matching condition (which, in some embodiments, may be changed based on any systematic or environmental factors), a pre-defined number of matches may be used to shortlist the candidate pool of the articles.

[0096] After locating the shortlisted articles associated with hypothesis 1905, at search 1910, those shortlisted articles or article snippets may then be re-ranked. The re-ranking may be based on a semantic similarity measure (e.g., a cosine similarity or other semantic similarity measures) of the shortlisted articles with regard to the hypothesis. Other reranking models would be appreciated by a person of ordinary skill in the art. A subset of the re-ranked articles (e.g., ranked in the top N) may be identified as the relevant articlesto the hypothesis (including the positive, negative, opposite, and null hypotheses), for example, relevant articles 1915a-n.

[0097] After re-ranking and identifying the top N relevant articles 1915a-n, a separate evidence mining process 1920a-n may be applied to each of the relevant articles, respectively. A certain content or a certain article snippet may be identified within each of the top N articles. In some embodiments only part of the articles, for example, a title, a background section, or other sections of interest, may be used for this evidence mining for higher mining efficiency.

[0098] After evidence mining 1920a-n, the mined evidence associated with hypothesis 1905 (including the positive, negative, opposite, and null hypotheses) may be provided, as part of a prompt, to query one or more LLMs 1925. In some embodiments, multiple instances of a multi -threaded LLM operation may be initiated to incorporate the collected article snippets and / or metadata associated with the article snippets into the LLM. As part of the prompt, the actual number of examples of the mined evidence from the top N relevant articles may be determined based on the complexity of the task, the specific LLM model, and / or any computation resources that can be used. The LLM may then be queried to quote the exact or a related string from the text evidence it has been provided that supports or refutes the hypotheses, forming an evidence set 1930 (e.g., a list of supporting and / or refuting evidence) associated with input hypothesis 1905.

[0099] In some embodiments, during evidence mining 1920a-n, once the supporting or refuting evidence has been collected, the collaborative platform may apply a statistical test to determine whether the items of evidence in the evidence set are statistically significant enough to support or refute the hypothesis. This hypothesis test may involve calculating a test statistic and comparing it to a critical value or evaluating a p-value to make a decision (e.g., whether this evidence should be provided to the LLM) based on the chosen significance level.

[0100] After forming an item of evidence in evidence set 1930 that incorporates the supporting or refuting evidence, this evidence may then be provided, as part of a prompt, to query another one or more LLMs 1935. In some embodiments, the one or more LLMs 1935 may generate a final output 1940 in which the final output may include, but is not limited to, a summary of the supporting or refuting evidence with a shorter length and / or in any format suitable to the end user. The final output may also include another summary of the supporting or refuting evidence over a period of time. The one or more LLMs 1935can support this functionality by at least incorporating, for example, a starting or ending time or a closed time range, to generate this summary. The final output may then be controlled and generated to reflect specific information or context, from the supporting or refuting evidence, as indicated by the end user.Agentic Atom Searching and Refining

[0101] Embodiments of the collaborative platform herein may use an agentic atom searching and refining system to search and then refine the atoms. This agentic atom searching and refining system may use Al and / or the evidence engine discussed above to actively search any sources for information based on specific criteria set by a user. The agentic system may interpret natural language, understand complex queries, and filter results based on context and relevance. The agentic system may also leverage Al algorithms to learn from user behavior and improve search results over time by analyzing user preferences and past searches. The agentic system may tailor results to individual needs for different end users regarding different user inputs.

[0102] In some embodiments, the agentic search engine may include, but is not limited to, a perception and learning component, a decision-making component, an execution component, a learning system, and / or tools. The perception and learning component may gather and process information from multiple sources or environments. The decisionmaking component may analyze the data and make informed decisions in which any reasoning capabilities of agents can be leveraged to optimize the decision-making process. The execution component may execute the tasks based on the decisions made in which the agents may use the equipped tools to enhance the execution. The learning system may enhance performance through trial and error, feedback, and reinforcement learning. The agentic search engine may also leverage other Al systems or models to actively seek out information and make autonomous decisions while searching for any evidence. In some embodiments, the agentic search engine may not only find relevant data but also interpret and contextualize it based on its understanding of the situation, rather than merely passively retrieving results based on keywords or vector embedding of the search query alone.

[0103] In some embodiments, the agentic search engine of the collaborative platform may be equipped with a set of tools, including but not limited to, NLP tools, information retrieval tools, knowledge graph integration tool, reasoning and inference tools, taskmanagement tools, external integration APIs, and / or feedback tools. The NLP tools may support capabilities for understanding complex queries. The information retrieval tools may retrieve relevant data from various sources. For example, knowledge graph integration tools may allow the agentic search engine to combine data from various sources, creating a unified knowledge graph where entities and their relationships are interconnected. The task management tools may orchestrate multiple actions, including but not limited to, creating action lists, assigning actions or tasks to different agents, setting deadlines, tracking progress, collaborating on projects (or tasks), providing visual boards, and / or offering customizable fields to organize the tasks. The external integration APIs may provide interface and functionalities to interact with external services for this agentic search engine. The feedback tools may continuously improve the performance of the agentic search engine based on any interactions between the collaborative platform and the end users.

[0104] To enhance the capability of the LLMs in understanding and interpreting the user specific information or the up-to-date database, the agentic search engine may use agentic retrieval augmented generation (RAG) to incorporate a user specific database (with the most current atoms and other user specific information) into a chat bot interface. This agentic RAG may extend traditional RAG by incorporating any autonomous agents that can perform a list of actions, including but not limited to, reasoning and planning, utilizing tools, filtering datasets, reflecting, and adapting. The agentic RAG may decompose the complex user queries into manageable sub-questions, and employ various tools, such as agentic search engines or user specific databases, to gather pertinent information. The agentic RAG may also run, for example, a database management system, including but not limited to relational databases (e.g., MySQL, PostgreSQL, any other SQL queries such as Query DSL in addition to SQL, and the like) and / or nonrelational databases (e.g., NoSQL databases, OpenSearch, and the like) to filter down any datasets before running a vector search. This filtering may ensure only relevant data is passed through to the LLMs. In addition, the agentic RAG may assess the relevance of the retrieved data from the user specific dataset and adjust the strategies accordingly.

[0105] After locating the atoms (e.g., hypotheses) in the user specific database, the agentic search engine may gather the supporting or refuting evidence associated with the located hypotheses. The agentic search engine may analyze any internal states or action history, predefined or saved in the memory or any databases, to identify such supportingor refuting evidence. Searching of such supporting or refuting evidence associated with hypotheses can be performed in a real-time manner. For example, for some of the hypotheses, the supporting or refuting evidence may be stored in a database with a linking to the certain hypotheses in which such linking can be retrieved by the agentic search engine from the database. The supporting or refuting evidence may also be retrieved from a memory storing all these linking and association. The agentic search engine can actively analyze and interpret complex queries in real-time, not only retrieving relevant information but also making autonomous decisions and taking actions based on the context of the search. When handling a new search, this agentic search engine can also adapt its approach based on any historical actions or approaches, and incorporate this new search information to boost any future search processes. In some embodiments, the agentic search engine may dynamically react and refine its search strategy as it operates.

[0106] In some embodiments, the search atoms, for example the hypotheses, can be further refined in which an agent system of the collaborative system may actively improve and optimize the quality of the located (e.g., searched) or generated hypotheses or other atoms. The Al agent may iteratively evaluate, adjust, and refine the hypotheses based on its understanding of the contexts, goals, and feedback mechanisms within the collaborative platform. For example, an agent system may analyze the content and identify areas for improvement on its own, making adjustments based on its internal logic and defined parameters. The adjustments from the agent system may also be driven by the desired outcome, such as making the content more informative, depending on a specific goal (e.g., a quantified metric) which is predefined in the agent system.

[0107] In some embodiments, the Al agent may only refine the hypotheses or other atoms after receiving a confirmation of the refinement action provided by an end user via a chat bot interface. The agent may continuously receive other feedback from the environment, such as user interactions or system metrics, which may further inform its refinement process and allow the agent system to adapt itself over time.Chat Bot Interface

[0108] Embodiments of the collaborative platform herein may use a chat bot interface to support an end user to query and then refine atoms on the fly. The chat bot interface allows an end user to have textual or spoken conversions (e.g., in natural language) with the collaborative platform. The response from the chat bot interface may be based on oneor more messages or queries from the end user and any analysis for phrases, keywords, and contexts related to issues and solutions common to the industry. In some embodiments, the chat hot interface, may provide a recommendation of the atoms to the end user in addition to providing a response to a user query. For example, the chat hot interface may analyze the user data and preferences to suggest any atoms that the end user may be looking for, such as a research article providing supporting or refuting evidence regarding an ongoing research experiment or a hypothesis under study. The chat bot interface may also provide any context in support of the inquiry from the end user. For example, the chat bot interface may, based on the input from the end user, provide some key concepts explaining how the response should be generated (including a plurality of steps) and what sources of information this response may rely on. In addition, the chat bot interface may streamline routine tasks relating to the atoms, including but not limited to any pre-processing, post-processing, and visualization of the atoms.

[0109] In some embodiments, the chat bot interface may allow the end user to update the settings, for example the configuration and parameters that determine how the chat bot should interact with the end user, including but not limited to personality, response style or tone, available information, and the context within which the chat bot may operate, defining a specific user-centric behavior and capabilities within a specific environment or application. The chat bot may also support a chat resume functionality to re-engage with a conversation on the interface that the end user was previously using. In some embodiments, the chat resume functionality may prompt the end user to start a new conversation or pick up where they left off, allowing the end user to continue the previous dialogue with the chat bot interface. This resume functionality may be implemented by simply typing again in the chat bot interface to re-initiate the conversation. The chat bot interface may then identify the context of the previous interaction and respond (or resume) accordingly.

[0110] FIGs. 20A-F show examples illustrating the chat bot interface and its corresponding functionalities, according to some embodiments. The chat bot interface may be performed by, for example, a user interface 350 of collaborative platform 300 in FIG. 3. However, the chat bot interface is not limited to that example implementation.[OHl] In some embodiments, the chat bot interface may support functionalities in conjunction with any other LLMs with a support of both natural-language input and output. In some embodiments, this chat bot interface may be built on top of the agenticsearching and refining system discussed above to understand human language, analyze user input, and generate relevant responses. The chat hot interface may include a natural language understanding component that analyzes the user’s input and identifies key words, phrases, and intent to understand the meaning behind the message. The chat bot interface may also keep track of the conversation flow, remembering previous interactions to provide contextually relevant responses. In addition, the chat bot interface may generate a response, which may be a simple text answer, a menu of options, or other actions. As an extension or a plug-in application of the agentic searching and refining system, the chat bot interface may, in some embodiments, use any other Al models or tools to search the internet and provide answers to user queries. The chat bot interface may also remember context, allowing a user to ask follow-up questions without repeating the original query. These follow-up questions may include a user question or intent to obtaining a more precise result, for example the refined search results of the atoms, or providing additional context to their initial search item.

[0112] In some embodiments, the chat bot interface may also support functionalities in conjunction with any other LLMs to provide a gap analysis for the user. This gap analysis may identify the difference between the user’s or an organization’s current state and their desired future state. For example, the chat bot interface may allow the user to articulate what the user wants to do, what level of performance or metrics the user expects, and what specific capabilities are essential for the user. The chat bot interface (e.g., with a separate analysis component) may then analyze the current LLM capabilities, identify source of the gap, and / or investigate any reasons why the gap exists (e.g., data limitations, model architecture, training process, etc.). This gap analysis can help the users to pinpoint research areas, novelty, or any processing of atoms (e.g., hypotheses) that may need improvement, and eventually, develop strategies to bridge those gaps.Unique User Identifier

[0113] Embodiments of the collaborative platform herein may use a unique user identifier that maps profiles of the authors (e.g., researchers or any individuals contributing to the atoms) to their articles, clinical trials, grants, and / or any other publications or works with which they may be associated. The user identifier may be assigned to a group of authors / individuals or a certain institution or organization.

[0114] FIGs. 21 A-B show an example flowchart illustrating generation of a unique user identifier, according to some embodiments.

[0115] Generating the unique user identifier may use a user-driven system in which the collaborative platform may give an end user a best guess or suggestion as to which works (e.g., articles) are theirs. The end user can then confirm or reject the suggestion. The user- driven system may start by normalizing and storing reference names in a structured format to create a clean and standardized baseline for comparison and to avoid issues with variations in formatting. In some embodiments, to store the reference names in a structured format, the user-driven system may use a relational database management system in which the reference names data may be organized into one or more tables with rows and columns, each column representing a specific attribute with a defined data type. This may allow for efficient querying and manipulation using, for example, a database management system, including but not limited to relational databases (e.g., MySQL, PostgreSQL, any other SQL queries such as Query DSL in addition to SQL, and the like) and / or non-relational databases (e.g., NoSQL databases, OpenSearch, and the like) to define the schema and access data within the database. This database management system may set up a predefined structure for the reference names data with clear relationships between different names attributes or data points.

[0116] In some embodiments, the input of the user-driven system may include a user entry or query of a name, alias, affiliation (e.g., associated institution), and / or topic. The user-driven system may retrieve the search criteria based on the user entry. In some embodiments, the search criteria may include, but are not limited to, two conditions. A first example condition may be that the search results to be returned must not contain two repeating last names, since those repeating names may disrupt the ranking algorithms to be performed. A second example condition may be that at least one of the search criteria must be satisfied to return the search result. The search criteria within the second condition may include, but are not limited to, a fuzzy match full name with scalar boost, an exact last name match with scalar boost, a name alternative match with scalar boost, an institution match with scalar boost, and / or a research topics match with scalar boost. This second condition may curate a list of candidates by scaling the ranking score with features that may indicate a good match. A skilled artisan will understand that other search conditions may be used instead of or in addition to those listed here.

[0117] After obtaining a list of candidates, the user-driven system may perform any processing of the names of the candidates useful to ensure a consistent comparison, such as by removing any variations that may potentially affect the matching. For example, the system may convert the author’s name to lowercase, remove whitespace, handle special characters, remove known prefixes and suffixes, and / or convert the names to a standard format (e.g., first, middle, last). The system may then use the last name to find initial potential matches since the last name is typically the most stable name component that provides a good initial filtering.

[0118] The system may then apply, for example, a three-tier matching strategy to further identify matches between the authors and works with which they may be associated. For example, this three-tier matching strategy may include, but is not limited to, an exact match comparison, a fuzzy match comparison if the exact match fails, and / or an enhanced match comparison if fuzzy match fails as well. The exact match comparison may compare standard strings exactly to identify a perfect match without complex comparisons. The fuzzy match comparison may be performed by setting up a weighted score for each name part, for example, 50% weight on last name, 30% weight on first name, and 20% weight on middle name. Other possible weighting scores for each name part would be appreciated by a person of ordinary skill in the art. The fuzzy match comparison may then calculate an overall score incorporating all the weighted scores of the name parts. In some embodiments, the fuzzy match comparison may only return the match results when the overall score is over a first threshold. The fuzzy match comparison may handle minor variations or typos while maintaining accuracy with weighted importance. In addition, when the overall score does not reach the first threshold (fuzzy match comparison fails) but is above a second threshold (which is lower or less strict than the first threshold), an enhanced match comparison may apply a more aggressive weighting to add more weights to the last name and apply a bonus point for any exact last name matches. This weighting change in the enhanced match comparison can catch harder cases while maintaining a reasonable confidence level.

[0119] In some embodiments, after obtaining the output from the three-tier matching, the user-driven system may apply a final filtering process to identify the matches between the authors and any works with which they may be associated. This filtering may use the output of the three-tier matching to determine whether the quality of the match is satisfied. This filtering may also calculate how many matches each user identifier has andthen calculate an average score for each user identifier. In addition, this filtering may incorporate both the quality and quantity of the matches to return the final matching outputs. For example, the final outputs may be obtained when both quality and quantity criteria have been satisfied (e.g., the quality being over a threshold and the quantity being over another threshold) or either one of the quality or quantity achieves a higher individual threshold (e.g., higher than both the quality and quantity thresholds). This final filtering may strike a balance between the quantity and quality of matches.Feedback Loop

[0120] Embodiments of the collaborative platform herein may have a feedback loop (also known as closed-loop learning) that describes the process of leveraging the output of an Al system and corresponding end-user actions in order to retrain and improve models over time. The Al-generated outputs, including but not limited to the extracted and / or generated atoms and the atom linking, are compared against the final decision and provide feedback to the Al models, allowing it to learn from its mistakes.

[0121] In some embodiments, the collaborative platform may leverage feedback from subject matter experts through collaboration during technical workshops, conducting field tests on live Al systems, and monitoring performance over time. The inputs from these subject matter experts may be ultimately fed back to the Al model to improve future performance. For example, the collaborative platform may have training samples from a survey that collects feedback from the subject matter experts. The collaborative platform may plan to perform training of the LLMs and / or Al models using these training samples based on a reinforcement learning from human feedback (RLHF). The RLHF may seek to train a “reward model” for the collaborative platform directly from feedback of the subject matter experts. The RLHF applied in the collaborative platform may include creating a preference dataset, training a reward function with supervised learning using this preference dataset, and / or using the reward learning in the reinforcement learning loop to fine-tune the LLMs and / or the Al models.

[0122] In some embodiments, the collaborative platform may integrate a work order management system to the closed-loop feedback mechanism. The system may enable feedback in which Al outputs are presented to end-users and their corresponding actions are recorded and sent back to the collaborative platform. For example, these actions may include but are not limited to providing human feedback on updating the extracted and / orgenerated atoms and / or their linking. In particular, updating the linking may include but is not limited to inserting and / or deleting a link between a hypothesis and a variant, equivalent, contradictory, or related hypothesis, inserting and / or deleting a link between a hypothesis and a supporting, refuting, inconclusive, or unrelated research finding, and / or inserting and / or deleting a link between a hypothesis and a related research question.

[0123] In some embodiments, the collaborative platform may provide machine learning operational tools to retrain and deploy Al and / or LLM models in a feedback loop. Models can be retrained automatically or on demand, based on model performance drift or availability of additional training data. New models can then be deployed along with existing models to track performance on live data before promoting models into production.User Rating

[0124] Embodiments of the collaborative platform herein may provide a variety of measurement metrics to help users, including individual researchers, research groups, and organizations, to quantify and monitor their own and other’s contributions, collaboration, impact, and / or performance. The collaborative platform may also provide a recommendation of, for example and without limitation, content, atoms, collaborators, and / or grants to such users based on associated measurement metrics and / or measured ratings. For example, the measurement metrics may address behavior that the user exhibits when they are interacting with the collaborative platform. The measurement metrics may account for the activities of the users which can be seen, heard, or inferred through what they have done on the collaborative platform. The individual measurement metrics may include but are not limited to a user contribution metric, a user performance metric, a user collaborative metric, and / or a user score. The measurement metrics of a user may also be combined with different users within a research group and an organization to form a group measurement metric.

[0125] In some embodiments, the user contribution metric may include a number of contributions made by the user to at least one of the text input, the atoms, the updated atoms, the linking between atoms, and / or the updated linking.

[0126] In some embodiments, the user performance metric may include at least one of a number of citations (e.g., citation by field in terms of time), an h-index, a number of publications, a number of patents issued and / or applied, a number of clinical trialsperformed, a number and / or an amount of grants received, a number of comments, a number of times a particular term is used by the user as compared to published literature, and / or a number of peer reviews made by the user. For example, the number of citations may be based on the number of times a publication has been cited, including self-citations and citations from other authors, either within or outside the same field. The h-index is an index that may attempt to measure both the scientific productivity and the apparent scientific impact of a researcher. The index may be based, for example, on a set of the researcher’s most cited papers and a number of citations that they have received in other researchers’ publications. So, the more citations a work has and / or the higher an h-index is, potentially the greater the impact of the publication and / or research.

[0127] In some embodiments, the collaborative platform may calculate the user performance metric for a user at least over a period of time, within a specific field, within a geographical location, and / or at a certain organization / institution. Tracking user performance metrics over time may gauge how engaged a user is with the collaborative platform, and may help to identify, based on the temporal trends and patterns of the user performance metrics, any further cross-fields and / or cross-locations cooperation between different individuals or institutions. The user performance metric may also be normalized and / or weighted between different fields across different publications and / or across different organizations. In some embodiments, the performance metric may be based on an aggregation of a plurality of contributions by the user. In some embodiments, the user performance metric may also quantify the performance of a non-user participant in the field, including but not limited to an individual researcher, a research group, and / or an organization. Similarly, the performance metrics used in quantifying the performance of a non-user participant may include a number of citations, an h-index, a number of publications, a number of patents issued and / or applied, a number of clinical trials performed, a number and / or an amount of grants received, a number of comments, and / or a number of peer reviews made by the non-user participant.

[0128] In some embodiments, the user collaborative metric may include at least one of a number of collaborations, a number of joint-authored publications, a number of joint- invented patents issued and / or applied, a number of joint-performed clinical trials, a number of co-contributions, a number and / or an amount of joint grants received, a number of co-citations, a number of citations received by joint publications, patents, and / or clinical trials, a number of mutual comments, or a number of mutual peer reviewsbetween the user and peers of the user (e.g., non-user participants). The user collaborative metric may be used to quantify the user collaborations and / or contributions within a collaborative network of the collaborative platform.

[0129] In some embodiments, the collaborative network of the system may be a contribution network between scholars, including a variety of entities including but not limited to organizations and / or individual researchers. Such scholars may be autonomous, geographically distributed, and / or heterogeneous in terms of their professional environment, expertise areas, cultures, social capital and / or goals. The discipline of the contribution network of the collaborative platform may focus on the behavior of the entities that collaborate to better achieve goals including updating the atoms and / or their linking.

[0130] In some embodiments, the collaborative network of the collaborative platform may be a citation network that is related to a user performance among the different scholars with heterogeneous backgrounds and research expertise. The citation network may be a directed graph that describes the citations within a collection of documents. Each vertex (or node) in the graph may represent a document in the collection, and each edge may be directed from one document toward another that citing the document or being cited by the document. Citation graphs may be utilized in various ways, including but not limited to forms of citation analysis and / or academic search tools. The discipline of the citation network of the collaborative platform may focus on the behavior of the entities that communicate better with each other based on the citing and / or the cited documents.

[0131] In some embodiments, other types of network graphs, such as a co-citation network, may be closely related to the citation network. The co-citation network is a graph between documents as nodes, where two documents are connected if they share a common citation. Other related networks may be formed using other information present in the document, including but not limited to the metadata and the atoms extracted and / or generated from the document. For example, in a collaboration graph such as a coauthorship network, the nodes may the authors of documents, linked if they have coauthored the same document. The link weights between two authors in co-authorship networks can increase over time if they have further collaboration.

[0132] In some embodiments, the collaborative platform may calculate a user score based on the user contribution metric, the user performance metric, and / or the collaborativemetric. The user score may be calculated as a weighted sum model that multiplies each of the number of user metrics by a specified weight. The user score may be indicative for individual users and researchers to monitor their contributions, impact, and performance.

[0133] In some embodiments, the collaborative platform may combine the individual user scores to form a group measurement metric for different research organizations and / or institutions. The combining may weight different users within an organization based on multiple factors including but not limited to seniority, research fields, and / or research interest. The weighted users within the organization may then be summarized to form a group metric that is indicative of quantifying and monitoring the contributions, impact, and / or performance of the group of users.Visualization Features

[0134] Embodiments of the collaborative platform herein may provide visualization features that allow the user to add value to each other researchers’ contributions through peer review of the texts including but not limited to scholarly literatures, atoms, and / or linking between the atoms. The visualization feature may be interactive and the collaborative platform may provide rich visualization capabilities, including customizable dashboards, user-friendly interfaces, and / or comparison tools. The visualization features may also provide tools based on LLM or Al models to facilitate users with actions including but not limited to formulating hypothesis and / or other atoms, designing experiments, and / or asking related research questions. The visualization features may navigate the users through the scientific universe including the atoms, the linking of the atoms, and / or the knowledge graph or database that store the connections and / or other knowledge.

[0135] In some embodiments, the customizable dashboards of the collaborative platform may include but is not limited to tabs for contributors, triples, contributions, and the linkages between different atoms of an uploaded scientific literature. On the contributors tab, the user may see different users who have contributed to the literature and / or the extracted / generated atoms from the literature. The user may get the detailed information of these contributors on the contributors tab. On the triples tab, the users may see a semantic triple and / or a modifier related to the atoms of the scientific literature. On the contributions tab, the user may answer questions that get at peer review, where thequestions may be about the quality of the extracted and / or generated atoms (e.g., hypothesis) and whether the hypothesis can be refuted or supported by testing.

[0136] In some embodiments, the peer review articles or literatures may be suggested by the atom linking that has been established, providing a structured post-publication peer review process that maximizes the transparency, credit, and support for the peer reviewers. The peer review process may record the names and reviews provided by the peer reviewers, making the peer reviewers and their reviews visible to the community. Other researchers or scholars within the community may be able to reply to and comment on those peer reviews. The peer review statistics may then be updated in the public researcher profiles associated with the peer reviewers. The peer reviews may also be provided in an anonymous manner if the peer reviewers prefer not to release their names (or any other identity information) to the public.

[0137] For example, as part of the peer review process, the customizable dashboard of the collaborative platform may provide a first question to an end user (e.g., peer reviewer) as: “is the hypothesis clearly stated (e.g., are variables and the objective unambiguously defined?” In an example embodiment, the peer reviewer may select, as one of the answers: “very unclear”, “unclear”, “neutral (uncertain)”, “clear”, or “very clear”. The customizable dashboard of the collaborative platform may then provide a second question to the end user, such as, for example: “do you expect the hypotheses to be supported or refuted?” In an example embodiment, the peer reviewer may select, as one of the answers: “no hypotheses are stated or the hypotheses are untestable”, “strongly refuted”, “somewhat refuted”, “neutral / uncertain”, “somewhat supported”, or “strongly supported”. The customizable dashboard of the collaborative platform may further provide a third question to request additional clarifications (based on the answer from the second question), such as, for example: “what level of impact would this hypothesis have on the field if it were strongly supported by results?” In an example embodiment, the peer reviewer may select, as one of the answers: “no impact”, “low impact”, “moderate impact”, “high impact”, or “very high impact”. Other possible questions and answers, as would be appreciated by a person of ordinary skill in the art, may alternatively or additionally be provided. Furthermore, an end user may provide any comments associated with the questions or the answers. Such comments may include, but are not limited to, any opinions, analysis, or interpretation of the question or answer themselves, rather than simply answering it directly. These comments may also represent any user perspectiveand potentially highlight any key aspects or complexities within the question or the answer provided by the customizable dashboards of the collaborative platform.

[0138] In some embodiments, the answers from different users may accumulate over time, resulting certain hypotheses may be more influential than others. The ranking of different hypotheses may include but is not limited to: limited, moderate, reasonable, high, and / or exceptional. A score may then be assigned to the hypothesis to reflect this ranking. When ranking different hypotheses, the hypotheses which are clearly defined, well-supported or refuted by existing evidence, testable with available methods, and / or have a strong logical connection to the research articles (or any other atoms) may be prioritized for visualization. The customizable dashboard(s) of the collaborative platform may place those hypotheses with strong directions (supports or refutes) at the top of the page, followed by non-directional hypotheses, and then less well-supported or difficult- to-test hypotheses at the bottom. On the linkages tab, the user may see the different interactions within the extracted and / or generated atoms. Such interactions may include but is not limited to the relationship between a hypothesis and a variant, equivalent, contradictory, or related hypothesis, a hypothesis and a supporting, refuting, inconclusive, or unrelated research finding, and / or a hypothesis and a related research question. The different users may also vote for the relationship and the interaction graph may record and visualize different number of votes in different colors.

[0139] In some embodiments, the comparison tool of the collaborative platform system may allow the users to compare their analytics with other users’ analytics based on the variety of measurement metrics including the user contribution metric, the user performance metric, the user collaborative metric, and / or the user score calculated based on the user contribution metric, the user performance metric, and / or the user collaborative metric. The comparison of the analytics between different users may be visualized by at least a chart, a graph, a table and / or other plots that would be appreciated by a person of ordinary skill in the art.

[0140] In some embodiments, the comparison tool may allow adding filter on a particular field, a technology, a keyword, a geographical location, a specific expertise area, and / or over a specific period of time. FIGs. 22A-B show an example illustrating an expertise model, according to some embodiments. The expertise model illustrates example subjects having been published, and also illustrates how those subjects change over a period of time. For example, the subjects may include, but are not limited to, health sciences, lifesciences, physical sciences, and / or social sciences. Within each subject, there are several subdisciplines shown. This expertise model also shows an impact analysis (e.g., citations) over time in each of the subjects listed above. Other impact analytics which are changed over time can be shown, as would be appreciated by a person of ordinary skill in the art.

[0141] The comparison tool support one-to-one comparison between different individual researchers, one-to-multiple comparison between an individual researcher and a group of researchers, and / or multiple-to-multiple comparison between different groups of researchers. For example, when performing comparison between different groups, the comparison tool may take each group, summarize their contributions and skills, and / or quantify the group metrics / criteria into a score. The score may then be used by the collaborative platform as a recommendation of at least contents, atoms, collaborators, and / or grants to peers of the user or third-party agencies to make more informed decisions including but not limited to strengthening the peer collaboration and / or providing funding supports.

[0142] In some embodiments, the interface of the collaborative platform may populate the atoms page to the uses when they publish a new contribution. For example, when the users choose to add a research question as their contributions, the LLM and / or Al models of the collaborative platform may provide the users with related research questions across the database of the collaborative platform. These related research questions may be generated from the text and / or inputted by the users. The LLM and / or Al models may also provide the users with suggestions for prompts when the research questions from the user input are not specific enough. For example, the collaborative platform may apply a rephrase and respond (RaR) strategy to refine the question asking from the users. The RaR strategy may enable the LLM to rephrase the question and incorporate additional details that may be retrieved from the extracted and / or generated atoms in the database and / or the input text. The collaborative platform may then ask the users to link the refined questions to other extracted and / or generated atoms including the research results and / or the hypothesis. The interface may be updated as the atoms and their linking populate to the collaborative platform.Example System Diagram

[0143] FIG. 3 is a block diagram of a collaborative platform 300, according to some embodiments. In some embodiments, the collaborative platform 300 may include asystem engine 320, a quantifier 330, a visualizer 340, and / or a user interface 350. System engine 320 may include one or more processors, buffers, servers, routers, modems, antennae, and / or circuitry configured to interface with quantifier 330, visualizer 340, and / or user interface 350.

[0144] In some embodiments, a data source 310 may be part of collaborative platform 300. In some embodiments, data source 310 may be a separate computing platform, including but not limited to smartphones, tablet computers, laptop computers, desktop computers, web browsers, and / or other computing devices, apparatuses, systems, or platforms. In some embodiments, data source 310 may transmit information to collaborative platform 300 either in a wired or wireless manner and may be connected via, for example, the Internet, a Local Area Network, or a Wide Area Network. The transmission may utilize a network protocol, such as, for example, a hypertext transfer protocol (HTTP), a TCP / IP protocol, Ethernet, or an asynchronous transfer mode.

[0145] In some embodiments, collaborative platform 300 receives data from data source 310. Data from data source 310 may include text data. The text data may refer to any information or message conveyed in written or printed form including but not limited to books, articles, surveys, transcripts, and / or scholarly and scientific literatures. The scholarly and scientific literatures may include but are not limited to academic papers presenting empirical research and theoretical contributions that span various disciplines within the natural and social sciences.

[0146] After collaborative platform 300 receives the data, system engine 320 may be triggered by the data characteristics that match predefined criteria in system engine 320. These criteria may be determined based on a list of factors including but not limited to types of data input, the system capabilities, the computational resource, and / or any transmission effects. System engine 320 may process the data from data source 310 based on the criteria to extract and / or generate valuable information from the data. The valuable information may include different atoms, wherein the atoms may include but are not limited to hypotheses, research questions, and / or research findings extracted and / or generated from the scientific literatures.

[0147] After the data is processed at system engine 320, system engine 320 may transmit the processed data to quantifier 330. The processed data may include but is not limited to the atoms and / or the processed scientific literatures themselves. Quantifier 330 may generate measurement metrics to quantify the processed data from system engine 320.The measurement metrics may include, for example and without limitation, a user contribution metric, a user performance metric, a user collaborative metric, and / or a user score. The measurement metrics of a user may also be combined with different users to form a group measurement metric. Quantifier 330 may transmit this variety of measurement metrics back to system engine 320. System engine 320 may then transmit the variety of measurement metrics generated from quantifier 330 to visualizer 340 for display. Additionally, system engine 320 may also transmit the processed data to visualizer 340 for display.

[0148] In addition, user interface 350 may interface with system engine 320, quantifier 330, and / or visualizer 340. User interface 350 may provide instructions to the collaborative platform 300 to update the transmitted results from system engine 320, quantifier 330, and / or visualizer 340 within a feedback loop. The transmitted results may include, for example and without limitation, the transmitted processed data from system engine 320 to quantifier 330, the transmitted measurement metrics from quantifier 330 to visualizer 340, and / or the transmitted processed data from system engine 320 to visualizer 340. In some embodiments, the instructions provided by user interface 350 may be generated by the Al models and / or LLM, based on at least the historical actions / behaviors of system engine 320, quantifier 330, and / or visualizer 340. In some embodiments, the instructions may be received as part of data source 310. For example, the instructions may include feedback and / or evaluations from other users. These instructions may include values assigned by the other users to update the transmitted results from system engine 320, quantifier 330, and / or visualizer 340.

[0149] In some embodiments, after the users provide instructions at user interface 350, user interface 350 may transmit the instructions back to quantifier 330. Quantifier 330 may also generate measurement metrics, including but not limited to a user contribution metric, a user performance metric, a user collaborative metric, and / or a user score to quantify the user behaviors (e.g., instructions) provided from user interface 350. Quantifier 330 may also transmit these measurement metrics back to system engine 320. System engine 320 may then transmit the measurement metrics quantifying the user behaviors generated from quantifier 330 to visualizer 340 for display. After receiving the processed data from system engine 320 and the measurement metrics quantifying the processed data and / or the user behaviors generated from quantifier 330, visualizer 330 may prioritize the visualization results based on the predefined criteria in system engine320. The prioritization conducted in visualizer 330 may provide a comprehensive and dynamic user interface for collaborative platform 300.

[0150] FIG. 4A is a block diagram of system engine 420, according to some embodiments. In some embodiments, system engine 420 may include but is not limited to a data processing module 422, a knowledge generation module 424, a knowledge retrieval module 426, and / or a knowledge connection module 428.

[0151] In some embodiments, when collaborative platform 300 receives data from data source 310, collaborative platform 300 may transmit the data to system engine 320. The data from data source 310 may include text data, wherein the text data may refer to any information or message conveyed in written or printed form, including but not limited to books, articles, surveys, transcripts, and / or scholarly and scientific literatures. Data processing module 422 may be configured to process the data based on predefined rules. Data processing module 422 may perform a variety of data processing techniques, including but not limited to data preprocessing and / or data embedding. The data preprocessing may include but is not limited to text cleaning, tokenization, chunking, stemming, and / or lemmatization. The data embedding may include performing at least one-hot encoding, Bag of Words (BOW), Term Frequency and Inverse document Frequency (TF-IDF), Word2Vec, Skip-Gram, pre-trained word-embedding using embedding layers, and / or any general embedding models (e.g., BAAI / bge model).

[0152] In some embodiments, knowledge generation module 424 may be configured to capture meaningful knowledge, including but not limited to information and patterns from the processed data from data processing module 422. Continuing the scientific literature example, the knowledge may include but is not limited to hypotheses, research questions, research findings, logical structures, and / or other valuable information from the text data received from data source 310. For example, the knowledge generation of the processed data may be performed by finding expressions that refer to a specific entity, linking named entity, extracting and / or generating relationships between named entities, and / or storing the knowledge extraction and / or generation results into a knowledge graph. Within the knowledge graph, the nodes may represent the specific entities and the edges may represent the relationships between the entities. Knowledge generation module 424 may be implemented by using at least a NLP model and / or a LLM.

[0153] In some embodiments, knowledge retrieval module 426 may be configured to retrieve knowledge from a database that stores the extracted and / or generated knowledgefrom knowledge generation module 424. Knowledge retrieval module 426 may seek to return relevant information to the extracted and / or generated knowledge from knowledge generation module 424 in a structured form. For example, the retrieved knowledge may include at least a variant, equivalent, contradictory, or related hypothesis to the hypothesis extracted and / or generated from knowledge generation module 424. The retrieved knowledge may include at least a supporting, refuting, inconclusive, or unrelated research finding to the hypothesis extracted and / or generated from knowledge generation module 424. The retrieved knowledge may also include a related research question to the hypothesis extracted and / or generated from knowledge generation module 424. In some embodiments, knowledge retrieval module 426 may use supervised learning and / or unsupervised learning approach to retrieve relevant knowledge from the database.

[0154] In some embodiments, knowledge connection module 428 may be configured to connect the extracted and / or generated knowledge from knowledge generation module 424 and the retrieved knowledge from knowledge retrieval module 426. Knowledge connection module 428 may add a new node and / or a new edge to the knowledge graph constructed from knowledge generation module 424. Knowledge connection module 428 may also provide a shortcut connection between different entities of a knowledge graph based on the relative truth value and / or the modifier of the logical structure. Using such a shortcut, when traversing the knowledge graph, the user may efficiently query the related entities through a shortest path when given an entity. In addition, the atoms may be represented as nodes in the knowledge graph, and the connections between the atoms may be represented as edges. When a connection is made it may be represented by creating a new edge in the knowledge graph between the two nodes representing the atoms. The knowledge graph may be updated with new nodes and / or connections to represent newly added atoms when they are made. Knowledge connection module 428 may use the formed linking between different atoms to provide a short cut connection between different entities within a knowledge graph and / or across different knowledge graphs for any querying purposes. For example, linking between different atoms, including but not limited to supporting or refuting evidence for an extracted and / or generated knowledge, may speed up the traversing of knowledge graph. A user may skip redundant and / or unrelated entities in traversing of the knowledge graph by using the relative truth value and / or the modifier of the logical structure, and / or the formed linking between atoms.

[0155] FIG. 4B is a block diagram of quantifier 430, according to some embodiments. In some embodiments, quantifier 430 may include but is not limited to a user contribution module 432, a user performance module 434, a user collaborative module 436, and / or a user rating module 438. In some embodiments, system engine 320 retrieves and / or connects the extracted and / or generated knowledge. Continuing the scientific literature use case example, the knowledge may include but is not limited to hypotheses, research questions, research findings, logical structures, and / or other valuable information extracted and / or generated from the text data received from data source 310. System engine 320 may transmit processed knowledge related to the user to quantifier 330. User interface 350 may also transmit user behaviors (e.g., user contributions and / or user instructions) to quantifier 330.

[0156] In some embodiments, user contribution module 432 may be configured to quantify the user behaviors received from user interface 350. The user behaviors may include but are not limited to feedback from the user. User contribution module 432 may provide measurement metrics, including but not limited to a number of contributions made by the user to at least one of the text, the atom, the updated atom, the connection, and / or the updated connection to quantify the user behaviors from user interface 350.

[0157] In some embodiments, user performance module 434 may be configured to quantify the processed knowledge related to the user from system engine 320. The processed knowledge related to the user may include but is not limited to any evaluations from other users regarding the extracted and / or generated knowledge from system engine 320. User performance module 434 may provide measurement metrics including but not limited to a number of citations, an h-index, a number of publications, a number of patents issued and / or applied, a number of clinical trials performed, a number and / or an amount of grants received, a number of comments, and / or a number of peer reviews made by the user to quantify the knowledge related to the user from system engine 320. In some embodiments, the number of citations may also include but are not limited to a citation average by field (percentile), citations received by interdisciplinary fields, and / or a highly cited number (top 1% or 10%) in field. In some embodiments, user performance module 434 may also provide measurement metrics made by a non-user participant to quantify their performance in the field. The non-user participant may include but is not limited to individual researchers, research groups, and / or organizations.

[0158] In some embodiments, user collaborative module 436 may be configured to quantify the collaboration from the processed knowledge related to the user and peers of the user from system engine 320. User collaborative module 436 may provide measurement metrics, including but not limited to a number of collaborations, a number of joint-authored publications, a number of joint-invented patents issued and / or applied, a number of joint-performed clinical trials, a number of co-contributions, a number and / or an amount of joint grants received, a number of co-citations, a number of citations received by joint publications, patents, and / or clinical trials, a number of mutual comments, or a number of mutual peer reviews between the user and peers of the user.

[0159] In some embodiments, user rating module 438 may be configured to provide a comprehensive rating (e.g., user score) that weights the quantified metric from one or more of user contribution module 432, user performance module 434, or use collaborative module 436. The comprehensive rating may be calculated as a weighted sum that multiplies each of the number of quantification metrics determined from the user and / or the connection between the user and their peers by a specified weight. The weight estimation may be adaptive to different factors of the user and other participants in the field, including but not limited to professional environment, expertise areas, cultures, social capital and / or goals.

[0160] FIG. 4C is a block diagram of visualizer 440, according to some embodiments. In some embodiments, visualizer 440 may include but is not limited to a knowledge visualization module 442, a connection visualization module 444, and / or a metric visualization module 446. In some embodiments, visualizer 440 may visualize the processed knowledge from system engine 320 and the measurement metrics from quantifier 330.

[0161] In some embodiments, knowledge visualization module 442 may be configured to visualize the extracted and / or generated knowledge from system engine 320. The extracted and / or generated knowledge may include but is not limited to hypotheses, research questions, research findings, logical structures, and / or knowledge graphs constructed based on the extracted and / or generated knowledge.

[0162] In some embodiments, connection visualization module 444 may be configured to visualize the connected knowledge from system engine 320. The connected knowledge may include but is not limited to connections between hypotheses, connections betweenhypotheses and research findings, and / or connections between hypotheses and research questions.

[0163] In some embodiments, metric visualization module 446 may be configured to visualize the user metrics and comparison of user metrics between different users and / or organizations. The user metrics may include but are not limited to user contribution metrics, user performance metrics, and / or user collaborative metrics. The user metrics and comparison of user metrics between different users may be visualized, for example, over a period of time, within a specific field, within a geographical location, and / or at a certain organization.

[0164] FIG. 5 is a flowchart illustrating a method 500 for textual data extraction, generation, and evaluation, according to some embodiments. Method 500 can be performed by processing logic that can comprise hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executing on a processing device), or a combination thereof. It is to be appreciated that not all steps may be needed to perform the disclosure provided herein. Further, some of the steps may be performed simultaneously, or in a different order than shown in FIG. 5, as will be understood by a person of ordinary skill in the art.

[0165] Method 500 shall be described with reference to FIG. 3. However, method 500 is not limited to that example embodiment. In 502, system engine 320 receives text from a user from data source 310. In some embodiments, the text may include academic literature. In 504, system engine 320 inputs the text into one or more first LLMs. The one or more first LLMs have been fine-tuned to output atoms based on atoms generated from other texts. In 506, system engine 320 generates an atom from the text using the one or more first LLMs. The atom may include a textual phrase having features that characterize the textual phrase as being in a particular category of information. The particular category of information may be, for example and without limitation, a hypothesis, a research question, or a research finding. In 508, system engine 320 inputs the atom into one or more second LLMs. The one or more second LLMs have been fine-tuned for structured output corresponding to the particular category of information based on textual phrases corresponding to the particular category of information from other texts. In 510, system engine 320 generates a logical structure based on a structured output of the one or more second LLMs. The logical structure may represent the textual phrase of the atom and contains one or more contextual attributes associated with the textual phrase. In 512,system engine 320 stores the logical structure into a knowledge graph as a modifiable node having a time-variant attribute. The time-variant may include a timestamp. In some embodiments, the generating the logical structure may include generating a semantic triple based on the structured output of the one or more second LLMs, generating a modifier based on the structured output of the one or more second LLMs, and generating a symbolic triple as the logical structure by adding the modifier to the semantic triple. The semantic triple may include a subject, an object, and a predicate that relates the subject to the object. The modifier may identify the one or more contextual attributes for the semantic triple.Example Computer System

[0166] Various embodiments may be implemented across, executed by, and / or deployed in a cloud computing environment containing physical and / or virtual resources and applications (including one or more physical machines, physical storage, virtual machines, virtual storage, or hypervisors). The cloud computing environment may provide computation, software, data access, storage, and / or other services that do not require end-user knowledge of a physical location and configuration of a system and / or a device that delivers the services. For example, one or more LLMs used by various embodiments may be hosted in a cloud network. In another example, a knowledge graph generated by various embodiments may be stored in a cloud network. The cloud computing system may include one or more computer resources, such as personal computers, workstations, computers, server devices, or other types of computation and / or communication devices. The cloud resources may include compute instances executing in the cloud computing resources, which may in turn communicate with other cloud computing resources via wired connections, wireless connections, or a combination of wired and wireless connections. Individual users may interact with the cloud network via one or more computer systems in communication with compute resources within the cloud network.

[0167] Some embodiments of the collaborative platform described herein may be hosted by one or more servers coupled to the cloud for distributed execution and / or storage of the platform. Some embodiments may be implemented, for example, using one or more well-known computer systems, such as computer system 600 shown in FIG. 6 or multiple of such computer systems 600 connected via a network and / or cloud. For example,embodiments herein using the collaborative platform may be implemented using combinations or sub-combinations of computer system 600. Also or alternatively, one or more computer systems 600 may be used, for example, to implement any of the embodiments discussed herein, as well as combinations and sub-combinations thereof. A “module,” as the term is used herein, is a computational element that performs one or more functions according to computer readable instructions stored on one or more memories or other non-transitory computer-readable media.

[0168] Computer system 600 may include one or more processors (also called central processing units, or CPUs), such as a processor 604. Processor 604 may be connected to a communication infrastructure or bus 606.

[0169] Computer system 600 may also include user input / output device(s) 603, such as monitors, keyboards, pointing devices, etc., which may communicate with communication infrastructure 606 through user input / output interface(s) 602.

[0170] One or more of processors 604 may be a graphics processing unit (GPU). In an embodiment, a GPU may be a processor that is a specialized electronic circuit designed to process mathematically intensive applications. The GPU may have a parallel structure that is efficient for parallel processing of large blocks of data, such as mathematically intensive data common to computer graphics applications, images, videos, etc.

[0171] Computer system 600 may also include a main or primary memory 608, such as random access memory (RAM). Main memory 608 may include one or more levels of cache. Main memory 608 may have stored therein control logic (i.e., computer software) and / or data.

[0172] Computer system 600 may also include one or more secondary storage devices or memory 610. Secondary memory 610 may include, for example, a hard disk drive 612 and / or a removable storage device or drive 614. Removable storage drive 614 may be a floppy disk drive, a magnetic tape drive, a compact disk drive, an optical storage device, tape backup device, and / or any other storage device / drive.

[0173] Removable storage drive 614 may interact with a removable storage unit 618. Removable storage unit 618 may include a computer usable or readable storage device having stored thereon computer software (control logic) and / or data. Removable storage unit 618 may be a floppy disk, magnetic tape, compact disk, DVD, optical storage disk, and / any other computer data storage device. Removable storage drive 614 may read from and / or write to removable storage unit 618.

[0174] Secondary memory 610 may include other means, devices, components, instrumentalities or other approaches for allowing computer programs and / or other instructions and / or data to be accessed by computer system 600. Such means, devices, components, instrumentalities or other approaches may include, for example, a removable storage unit 622 and an interface 620. Examples of the removable storage unit 622 and the interface 620 may include a program cartridge and cartridge interface (such as that found in video game devices), a removable memory chip (such as an EPROM or PROM) and associated socket, a memory stick and USB or other port, a memory card and associated memory card slot, and / or any other removable storage unit and associated interface.

[0175] Computer system 600 may further include a communication or network interface 624. Communication interface 624 may enable computer system 600 to communicate and interact with any combination of external devices, external networks, external entities, etc. (individually and collectively referenced by reference number 628). For example, communication interface 624 may allow computer system 600 to communicate with external or remote devices 628 over communications path 626, which may be wired and / or wireless (or a combination thereof), and which may include any combination of LANs, WANs, the Internet, etc. Control logic and / or data may be transmitted to and from computer system 600 via communications path 626.

[0176] Computer system 600 may also be any of a personal digital assistant (PDA), desktop workstation, laptop or notebook computer, netbook, tablet, smart phone, smart watch or other wearable, appliance, part of the Internet-of-Things, and / or embedded system, to name a few non-limiting examples, or any combination thereof.

[0177] Computer system 600 may be a client or server, accessing or hosting any applications and / or data through any delivery paradigm, including but not limited to remote or distributed cloud computing solutions; local or on-premises software (“onpremise” cloud-based solutions); “as a service” models (e.g., content as a service (CaaS), digital content as a service (DCaaS), software as a service (SaaS), managed software as a service (MSaaS), platform as a service (PaaS), desktop as a service (DaaS), framework as a service (FaaS), backend as a service (BaaS), mobile backend as a service (MBaaS), infrastructure as a service (laaS), etc.); and / or a hybrid model including any combination of the foregoing examples or other services or delivery paradigms.

[0178] Any applicable data structures, file formats, and schemas in computer system 600 may be derived from standards including but not limited to JavaScript Object Notation (JSON), Comma Separated Values (CSV), Extensible Markup Language (XML), Yet Another Markup Language (YAML), Extensible Hypertext Markup Language (XHTML), Wireless Markup Language (WML), MessagePack, XML User Interface Language (XUL), or any other functionally similar representations alone or in combination. Alternatively, proprietary data structures, formats or schemas may be used, either exclusively or in combination with known or open standards.

[0179] In some embodiments, a tangible, non-transitory apparatus or article of manufacture comprising a tangible, non-transitory computer useable or readable medium having control logic (software) stored thereon may also be referred to herein as a computer program product or program storage device. This includes, but is not limited to, computer system 600, main memory 608, secondary memory 610, and removable storage units 618 and 622, as well as tangible articles of manufacture embodying any combination of the foregoing. Such control logic, when executed by one or more data processing devices (such as computer system 600 or processor(s) 604), may cause such data processing devices to operate as described herein.

[0180] Based on the teachings contained in this disclosure, it will be apparent to persons skilled in the relevant art(s) how to make and use embodiments of this disclosure using data processing devices, computer systems and / or computer architectures other than that shown in FIG. 6. In particular, embodiments can operate with software, hardware, and / or operating system implementations other than those described herein.

[0181] It is to be appreciated that the Detailed Description section, and not any other section, is intended to be used to interpret the claims. Other sections can set forth one or more but not all exemplary embodiments as contemplated by the inventor(s), and thus, are not intended to limit this disclosure or the appended claims in any way.

[0182] While this disclosure describes exemplary embodiments for exemplary fields and applications, it should be understood that the disclosure is not limited thereto. Other embodiments and modifications thereto are possible, and are within the scope and spirit of this disclosure. For example, and without limiting the generality of this paragraph, embodiments are not limited to the software, hardware, firmware, and / or entities illustrated in the figures and / or described herein. Further, embodiments (whether or notexplicitly described herein) have significant utility to fields and applications beyond the examples described herein.

[0183] Embodiments have been described herein with the aid of functional building blocks illustrating the implementation of specified functions and relationships thereof. The boundaries of these functional building blocks have been arbitrarily defined herein for the convenience of the description. Alternate boundaries can be defined as long as the specified functions and relationships (or equivalents thereof) are appropriately performed. Also, alternative embodiments can perform functional blocks, steps, operations, methods, etc. using orderings different than those described herein.

[0184] References herein to “one embodiment,” “an embodiment,” “an example embodiment,” or similar phrases, indicate that the embodiment described may include a particular feature, structure, or characteristic, but every embodiment may not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it would be within the knowledge of persons skilled in the relevant art(s) to incorporate such feature, structure, or characteristic into other embodiments whether or not explicitly mentioned or described herein. Additionally, some embodiments can be described using the expression “coupled” and “connected” along with their derivatives. These terms are not necessarily intended as synonyms for each other. For example, some embodiments can be described using the terms “connected” and / or “coupled” to indicate that two or more elements are in direct physical or electrical contact with each other. The term “coupled,” however, can also mean that two or more elements are not in direct contact with each other, but yet still co-operate or interact with each other.

[0185] The breadth and scope of this disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.

Claims

WHAT IS CLAIMED IS:

1. A computer-implemented method comprising: receiving, by a computer processor, text from a user; inputting, by the computer processor, the text into one or more first large language models (LLMs) that have been fine-tuned to output atoms based on atoms generated from other texts; generating, by the computer processor, an atom from the text using the one or more first LLMs, wherein the atom comprises a textual phrase having features that characterize the textual phrase as being in a particular category of information; inputting, by the computer processor, the atom into one or more second LLMs that has been fine-tuned for structured output corresponding to the particular category of information based on textual phrases corresponding to the particular category of information from other texts; generating, by the computer processor, a logical structure based on a structured output of the one or more second LLMs, wherein the logical structure represents the textual phrase of the atom and contains one or more contextual attributes associated with the textual phrase; and storing, by the computer processor, the logical structure into a knowledge graph as a modifiable node having a time-variant attribute.

2. The computer-implemented method of claim 1, wherein the generating the logical structure comprises: generating a semantic triple based on the structured output of the one or more second LLMs, wherein the semantic triple comprises a subject, an object, and a predicate that relates the subject to the object; generating a modifier based on the structured output of the one or more second LLMs, wherein the modifier identifies the one or more contextual attributes for the semantic triple; and generating a symbolic triple as the logical structure by adding the modifier to the semantic triple.

3. The computer-implemented method of claim 2, wherein the modifier comprises each of a subject modifier, a predicate modifier, and an object modifier for the semantic triple.

4. The computer-implemented method of claim 1, wherein the text comprises academic literature and the particular category of information is at least one of a hypothesis, a research question, or a research finding.

5. The computer-implemented method of claim 1, further comprising: modifying the modifiable node by adding to the logical structure a connection to another logical structure in the knowledge graph, wherein the connection has a new timevariant attribute while the time-variant attribute of the modifiable node is maintained.

6. The computer-implemented method of claim 1, wherein the time-variant attribute comprises a timestamp.

7. The computer-implemented method of claim 1, further comprising: generating, by the computer processor, a connection between the atom and at least one atom generated from a different text based on a similarity relationship between the atom and at least one atom generated from a different text; updating, by the computer processor, the atom based on feedback from the user to generate an updated atom; and generating, by the computer processor, an updated connection between the updated atom and the at least one atom generated from a different text.

8. The computer implemented method of claim 7, further comprising: receiving one or more evaluations from other users, wherein the one or more evaluations comprise values assigned by the other users to at least one of the text, the atom, the updated atom, the connection, or the updated connection.

9. The computer implemented method of claim 8, further comprising: measuring a rating of the user based on the received one or more evaluations of the other users.

10. The computer-implemented method of claim 1, wherein the inputting the text comprises: tokenizing the text to obtain data chunks;performing keyword searching within the data chunks to obtain a ranked list of the data chunks associated with each of a plurality of keywords known to be indicative of the particular category of information; fusing a set of ranked lists associated with the each of the plurality of keywords to obtain a fused ranked list; extracting top-k data chunks from the data chunks based on the fused ranked list; and inputting a structured prompt containing the top-k data chunks into the one or more first LLMs.

11. The computer-implemented method of claim 10, wherein one data chunk comprises three sentences.

12. The computer-implemented method of claim 11, wherein a first sentence and a third sentence of the data chunk are overlapped with an adjacent data chunk obtained in tokenization.

13. The computer-implemented method of claim 10, wherein an element of the ranked list comprises a relevance score between a keyword and a data chunk.

14. The computer-implemented method of claim 10, wherein the fusing comprises combining a set of lists of relevance scores associated with the set of ranked lists.

15. The computer-implemented method of claim 10, wherein the one or more first LLMs comprises a fine-tuned Mistral Instruct model.

16. The computer-implemented method of claim 10, further comprising generating relevance metrics for the top-k data chunks.

17. The computer-implemented method of claim 1, wherein the generating an atom comprises: embedding the atom generated from the text into a vector space, wherein the atom comprises a hypothesis;performing clustering within the vector space to generate one or more preliminary clusters of the atom; inputting a structured prompt containing the one or more preliminary clusters into the one or more first LLMs; generating one or more refined clusters of the atom based on an output of the one or more first LLMs; and generating another atom from the atom based on the output of the one or more first LLMs, wherein the other atom comprises a research question associated with the one or more refined clusters.

18. The computer-implemented method of claim 17, wherein the embedding is performed by using at least a pre-trained BAAI / bge model.

19. The computer-implemented method of claim 17, wherein the clustering comprises performing a k-means clustering.

20. The computer-implemented method of claim 19, wherein a number of clusters in the k- means clustering is determined by at least an elbow method, a silhouette score, or a LLM capability.

21. The computer-implemented method of claim 17, wherein the one or more first LLMs comprises a fine-tuned Llama-3 model.

22. The computer-implemented method of claim 1, wherein the inputting the atom comprises: parsing the atom into a plurality of phrases, wherein the atom comprises one sentence from the received text; extracting a plurality of text entities from the plurality of phrases; and inputting a structured prompt containing the plurality of text entities into the one or more second LLMs.

23. The computer-implemented method of claim 22, wherein the inputting a structured prompt comprises inputting the plurality of text entities as prompt examples into a fewshot learning model.

24. The computer-implemented method of claim 22, wherein the one or more second LLMs comprises a fine-tuned Mistral Instruct model.

25. The computer-implemented method of claim 1, wherein the generating an atom comprises generating a plurality of atoms from the text, and the method further comprises: identifying a relationship between at least two atoms from the plurality of atoms; and generating a connection between the at least two atoms within the plurality of atoms.

26. The computer-implemented method of claim 25, wherein the identified relationship from the plurality of atoms comprise a relationship at least between: a hypothesis and a variant, equivalent, contradictory, or related hypothesis; the hypothesis and a supporting, refuting, inconclusive, or unrelated research finding; or the hypothesis and a related research question.

27. The computer-implemented method of claim 25, wherein the generated connection within the plurality of atoms comprises at least one of: a first connection between a hypothesis and a variant, equivalent, contradictory, or related hypothesis; a second connection between the hypothesis and a supporting, refuting, inconclusive, or unrelated research finding; or a third connection between the hypothesis and a related research question.

28. The computer-implemented method of claim 7, wherein the atom comprises a hypothesis, and the updating comprises performing at least one of: inserting a first connection between the hypothesis and a variant, equivalent, contradictory, or related hypothesis; updating the first connection between the hypothesis and the variant, equivalent, contradictory, or related hypothesis; inserting a second connection between the hypothesis and a supporting, refuting, inconclusive, or unrelated research finding;updating the second connection between the hypothesis and the supporting, refuting, inconclusive, or unrelated research finding; inserting a third connection between the hypothesis and a related research question; or updating the third connection between the hypothesis and the related research question.

29. The computer-implemented method of claim 9, wherein the measuring comprises: determining a user contribution metric from the evaluation; determining a user performance metric from the evaluation; determining a user collaborative metric from the evaluation; and determining a user score based on at least the user contribution metric, the user performance metric, or the user collaborative metric.

30. The computer-implemented method of claim 29, wherein the user contribution metric comprises a number of contributions made by the user to at least one of the text, the atom, the updated atom, the connection, or the updated connection.

31. The computer-implemented method of claim 29, wherein the user performance metric comprises at least one of a number of citations, an h-index, a number of publications, a number of patents issued, a number of clinical trials performed, a number of grants received, a number of comments, a number of times a particular term is used by the user as compared to published literature, or a number of peer reviews made by the user.

32. The computer-implemented method of claim 29, wherein the user collaborative metric comprises at least one of a number of collaborations, a number of co-contributions, or a number of co-citations between the user and peers of the user.

33. The computer-implemented method of claim 27, further comprising: generating a first output visualizing the first connection; generating a second output visualizing the second connection; and generating a third output visualizing the third connection.

34. The computer-implemented method of claim 9, wherein the measuring comprises generating a user performance metric based on an aggregation of a plurality of contributions by the user.

35. The computer-implemented method of claim 9, further comprising: generating a first output visualizing the measured rating of the user; and generating a second output visualizing a comparison of the measured rating between the user and peers of the user.

36. The computer-implemented method of claim 35, wherein the first output comprises: a first chart to display a user contribution metric for the user over a time period; a second chart to display a user performance metric for the user over the time period; and a third chart to display a rating score for the user over the time period.

37. The computer-implemented method of claim 35, wherein the second output comprises: a first chart to display a comparison of a user contribution metric between the user and the peers of the user over a time period; a second chart to display a comparison of a user performance metric between the user and the peers of the user over the time period; and a third chart to display a comparison of a rating score between the user and the peers of the user over the time period.

38. The computer-implemented method of claim 33, further comprising providing a recommendation to peers of the user or third-party agencies based on at least the measured rating of the user or the comparison of the measured rating between the user and the peers of the user.

39. A system comprising: one or more memories; at least one processor each coupled to at least one of the memories and configured to perform operations comprising: receiving text from a user;inputting the text into one or more first large language models (LLMs) that have been fine-tuned to output atoms based on atoms generated from other texts; generating an atom from the text using the one or more first LLMs, wherein the atom comprises a textual phrase having features that characterize the textual phrase as being in a particular category of information; inputting the atom into one or more second LLMs that has been fine-tuned for structured output corresponding to the particular category of information based on textual phrases corresponding to the particular category of information from other texts; generating a logical structure based on a structured output of the one or more second LLMs, wherein the logical structure represents the textual phrase of the atom and contains one or more contextual attributes associated with the textual phrase; and storing the logical structure into a knowledge graph as a modifiable node having a time-variant attribute.

40. The system of claim 39, wherein the generating the logical structure comprises: generating a semantic triple based on the structured output of the one or more second LLMs, wherein the semantic triple comprises a subject, an object, and a predicate that relates the subject to the object; generating a modifier based on the structured output of the one or more second LLMs, wherein the modifier identifies the one or more contextual attributes for the semantic triple; and generating a symbolic triple as the logical structure by adding the modifier to the semantic triple.

41. The system of claim 40, wherein the modifier comprises each of a subject modifier, a predicate modifier, and an object modifier for the semantic triple.

42. The system of claim 39, wherein the text comprises academic literature and the particular category of information is at least one of a hypothesis, a research question, or a research finding.

43. The system of claim 39, wherein the operations further comprise: modifying the modifiable node by adding to the logical structure a connection to another logical structure in the knowledge graph, wherein the connection has a new timevariant attribute while the time-variant attribute of the modifiable node is maintained.

44. The system of claim 39, wherein the time-variant attribute comprises a timestamp.

45. The system of claim 39, wherein the operations further comprise: generating a connection between the atom and at least one atom generated from a different text based on a similarity relationship between the atom and at least one atom generated from a different text; updating the atom based on feedback from the user to generate an updated atom; and generating an updated connection between the updated atom and the at least one atom generated from a different text.

46. The system of claim 45, wherein the operations further comprise: receiving one or more evaluations from other users, wherein the one or more evaluations comprise values assigned by the other users to at least one of the text, the atom, the updated atom, the connection, or the updated connection.

47. The system of claim 46, wherein the operations further comprising: measuring a rating of the user based on the received one or more evaluations of the other users.

48. The system of claim 39, wherein the inputting the text comprises: tokenizing the text to obtain data chunks; performing keyword searching within the data chunks to obtain a ranked list of the data chunks associated with each of a plurality of keywords known to be indicative of the particular category of information; fusing a set of ranked lists associated with the each of the plurality of keywords to obtain a fused ranked list;extracting top-k data chunks from the data chunks based on the fused ranked list; and inputting a structured prompt containing the top-k data chunks into the one or more first LLMs.

49. The system of claim 48, wherein one data chunk comprises three sentences.

50. The system of claim 49, wherein a first sentence and a third sentence of the data chunk are overlapped with an adjacent data chunk obtained in tokenization.

51. The system of claim 48, wherein an element of the ranked list comprises a relevance score between a keyword and a data chunk.

52. The system of claim 48, wherein the fusing comprises combining a set of lists of relevance scores associated with the set of ranked lists.

53. The system of claim 48, wherein the one or more first LLMs comprises a fine-tuned Mistral Instruct model.

54. The system of claim 48, wherein the operations further comprise generating relevance metrics for the top-k data chunks.

55. The system of claim 39, wherein the generating an atom comprises: embedding the atom generated from the text into a vector space, wherein the atom comprises a hypothesis; performing clustering within the vector space to generate one or more preliminary clusters of the atom; inputting a structured prompt containing the one or more preliminary clusters into the one or more first LLMs; generating one or more refined clusters of the atom based on an output of the one or more first LLMs; and generating another atom from the atom based on the output of the one or more first LLMs, wherein the other atom comprises a research question associated with the one or more refined clusters.

56. The system of claim 55, wherein the embedding is performed by using at least a pretrained BAAI / bge model.

57. The system of claim 55, wherein the clustering comprises performing a k-means clustering.

58. The system of claim 57, wherein a number of clusters in the k-means clustering is determined by at least an elbow method, a silhouette score, or a LLM capability.

59. The system of claim 55, wherein the one or more first LLMs comprises a fine-tuned Llama-3 model.

60. The system of claim 39, wherein the inputting the atom comprises: parsing the atom into a plurality of phrases, wherein the atom comprises one sentence from the received text; extracting a plurality of text entities from the plurality of phrases; and inputting a structured prompt containing the plurality of text entities into the one or more second LLMs.

61. The system of claim 60, wherein the inputting a structured prompt comprises inputting the plurality of text entities as prompt examples into a few-shot learning model.

62. The system of claim 60, wherein the one or more second LLMs comprises a fine-tuned Mistral Instruct model.

63. The system of claim 39, wherein the generating an atom comprises generating a plurality of atoms from the text, and the method further comprises: identifying a relationship between at least two atoms from the plurality of atoms; and generating a connection between the at least two atoms within the plurality of atoms.

64. The system of claim 63, wherein the identified relationship from the plurality of atoms comprise a relationship at least between:a hypothesis and a variant, equivalent, contradictory, or related hypothesis; the hypothesis and a supporting, refuting, inconclusive, or unrelated research finding; or the hypothesis and a related research question.

65. The system of claim 63, wherein the generated connection within the plurality of atoms comprises at least one of: a first connection between a hypothesis and a variant, equivalent, contradictory, or related hypothesis; a second connection between the hypothesis and a supporting, refuting, inconclusive, or unrelated research finding; or a third connection between the hypothesis and a related research question.

66. The system of claim 45, wherein the atom comprises a hypothesis, and the updating comprises performing at least one of: inserting a first connection between the hypothesis and a variant, equivalent, contradictory, or related hypothesis; updating the first connection between the hypothesis and the variant, equivalent, contradictory, or related hypothesis; inserting a second connection between the hypothesis and a supporting, refuting, inconclusive, or unrelated research finding; updating the second connection between the hypothesis and the supporting, refuting, inconclusive, or unrelated research finding; inserting a third connection between the hypothesis and a related research question; or updating the third connection between the hypothesis and the related research question.

67. The system of claim 47, wherein the measuring comprises: determining a user contribution metric from the evaluation; determining a user performance metric from the evaluation; determining a user collaborative metric from the evaluation; anddetermining a user score based on at least the user contribution metric, the user performance metric, or the user collaborative metric.

68. The system of claim 67, wherein the user contribution metric comprises a number of contributions made by the user to at least one of the text, the atom, the updated atom, the connection, or the updated connection.

69. The system of claim 67, wherein the user performance metric comprises at least one of a number of citations, an h-index, a number of publications, a number of patents issued, a number of clinical trials performed, a number of grants received, a number of comments, a number of times a particular term is used by the user as compared to published literature, or a number of peer reviews made by the user.

70. The system of claim 67, wherein the user collaborative metric comprises at least one of a number of collaborations, a number of co-contributions, or a number of co-citations between the user and peers of the user.

71. The system of claim 65, wherein the operations further comprise: generating a first output visualizing the first connection; generating a second output visualizing the second connection; and generating a third output visualizing the third connection.

72. The system of claim 67, wherein the measuring comprises generating a user performance metric based on an aggregation of a plurality of contributions by the user.

73. The system of claim 47, wherein the operations further comprise: generating a first output visualizing the measured rating of the user; and generating a second output visualizing a comparison of the measured rating between the user and peers of the user.

74. The system of claim 71, wherein the first output comprises: a first chart to display a user contribution metric for the user over a time period; a second chart to display a user performance metric for the user over the time period; anda third chart to display a rating score for the user over the time period.

75. The system of claim 71, wherein the second output comprises: a first chart to display a comparison of a user contribution metric between the user and the peers of the user over a time period; a second chart to display a comparison of a user performance metric between the user and the peers of the user over the time period; and a third chart to display a comparison of a rating score between the user and the peers of the user over the time period.

76. The system of claim 73, wherein the operations further comprise providing a recommendation to peers of the user or third-party agencies based on at least the measured rating of the user or the comparison of the measured rating between the user and the peers of the user.

77. A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising: one or more memories; at least one processor each coupled to at least one of the memories and configured to perform operations comprising: receiving text from a user; inputting the text into one or more first large language models (LLMs) that have been fine-tuned to output atoms based on atoms generated from other texts; generating an atom from the text using the one or more first LLMs, wherein the atom comprises a textual phrase having features that characterize the textual phrase as being in a particular category of information; inputting the atom into one or more second LLMs that has been fine-tuned for structured output corresponding to the particular category of information based on textual phrases corresponding to the particular category of information from other texts; generating a logical structure based on a structured output of the one or more second LLMs, wherein the logical structure represents the textual phrase of theatom and contains one or more contextual attributes associated with the textual phrase; and storing the logical structure into a knowledge graph as a modifiable node having a time-variant attribute.

78. The non-transitory computer-readable medium of claim 77, wherein the generating the logical structure comprises: generating a semantic triple based on the structured output of the one or more second LLMs, wherein the semantic triple comprises a subject, an object, and a predicate that relates the subject to the object; generating a modifier based on the structured output of the one or more second LLMs, wherein the modifier identifies the one or more contextual attributes for the semantic triple; and generating a symbolic triple as the logical structure by adding the modifier to the semantic triple.

79. The non-transitory computer-readable medium of claim 78, wherein the modifier comprises each of a subject modifier, a predicate modifier, and an object modifier for the semantic triple.

80. The non-transitory computer-readable medium of claim 77, wherein the text comprises academic literature and the particular category of information is at least one of a hypothesis, a research question, or a research finding.

81. The non-transitory computer-readable medium of claim 77, wherein the operations further comprise: modifying the modifiable node by adding to the logical structure a connection to another logical structure in the knowledge graph, wherein the connection has a new timevariant attribute while the time-variant attribute of the modifiable node is maintained.

82. The non-transitory computer-readable medium of claim 77, wherein the time-variant attribute comprises a timestamp.

83. The non-transitory computer-readable medium of claim 77, wherein the operations further comprise: generating a connection between the atom and at least one atom generated from a different text based on a similarity relationship between the atom and at least one atom generated from a different text; updating the atom based on feedback from the user to generate an updated atom; and generating an updated connection between the updated atom and the at least one atom generated from a different text.

84. The non-transitory computer-readable medium of claim 83, wherein the operations further comprise: receiving one or more evaluations from other users, wherein the one or more evaluations comprise values assigned by the other users to at least one of the text, the atom, the updated atom, the connection, or the updated connection.

85. The non-transitory computer-readable medium of claim 84, wherein the operations further comprising: measuring a rating of the user based on the received one or more evaluations of the other users.

86. The non-transitory computer-readable medium of claim 77, wherein the inputting the text comprises: tokenizing the text to obtain data chunks; performing keyword searching within the data chunks to obtain a ranked list of the data chunks associated with each of a plurality of keywords known to be indicative of the particular category of information; fusing a set of ranked lists associated with the each of the plurality of keywords to obtain a fused ranked list; extracting top-k data chunks from the data chunks based on the fused ranked list; and inputting a structured prompt containing the top-k data chunks into the one or more first LLMs.

87. The non-transitory computer-readable medium of claim 86, wherein one data chunk comprises three sentences.

88. The non-transitory computer-readable medium of claim 87, wherein a first sentence and a third sentence of the data chunk are overlapped with an adjacent data chunk obtained in tokenization.

89. The non-transitory computer-readable medium of claim 86, wherein an element of the ranked list comprises a relevance score between a keyword and a data chunk.

90. The non-transitory computer-readable medium of claim 86, wherein the fusing comprises combining a set of lists of relevance scores associated with the set of ranked lists.

91. The non-transitory computer-readable medium of claim 86, wherein the one or more first LLMs comprises a fine-tuned Mistral Instruct model.

92. The non-transitory computer-readable medium of claim 86, wherein the operations further comprise generating relevance metrics for the top-k data chunks.

93. The non-transitory computer-readable medium of claim 77, wherein the generating an atom comprises: embedding the atom generated from the text into a vector space, wherein the atom comprises a hypothesis; performing clustering within the vector space to generate one or more preliminary clusters of the atom; inputting a structured prompt containing the one or more preliminary clusters into the one or more first LLMs; generating one or more refined clusters of the atom based on an output of the one or more first LLMs; and generating another atom from the atom based on the output of the one or more first LLMs, wherein the other atom comprises a research question associated with the one or more refined clusters.

94. The non-transitory computer-readable medium of claim 93, wherein the embedding is performed by using at least a pre-trained BAAI / bge model.

95. The non-transitory computer-readable medium of claim 93, wherein the clustering comprises performing a k-means clustering.

96. The non-transitory computer-readable medium of claim 95, wherein a number of clusters in the k-means clustering is determined by at least an elbow method, a silhouette score, or a LLM capability.

97. The non-transitory computer-readable medium of claim 93, wherein the one or more first LLMs comprises a fine-tuned Llama-3 model.

98. The non-transitory computer-readable medium of claim 77, wherein the inputting the atom comprises: parsing the atom into a plurality of phrases, wherein the atom comprises one sentence from the received text; extracting a plurality of text entities from the plurality of phrases; and inputting a structured prompt containing the plurality of text entities into the one or more second LLMs.

99. The non-transitory computer-readable medium of claim 98, wherein the inputting a structured prompt comprises inputting the plurality of text entities as prompt examples into a few-shot learning model.

100. The non-transitory computer-readable medium of claim 98, wherein the one or more second LLMs comprises a fine-tuned Mistral Instruct model.

101. The non-transitory computer-readable medium of claim 77, wherein the generating an atom comprises generating a plurality of atoms from the text, and the method further comprises: identifying a relationship between at least two atoms from the plurality of atoms;generating a connection between the at least two atoms within the plurality of atoms.

102. The non-transitory computer-readable medium of claim 101, wherein the identified relationship from the plurality of atoms comprise a relationship at least between: a hypothesis and a variant, equivalent, contradictory, or related hypothesis; the hypothesis and a supporting, refuting, inconclusive, or unrelated research finding; or the hypothesis and a related research question.

103. The non-transitory computer-readable medium of claim 101, wherein the generated connection within the plurality of atoms comprises at least one of: a first connection between a hypothesis and a variant, equivalent, contradictory, or related hypothesis; a second connection between the hypothesis and a supporting, refuting, inconclusive, or unrelated research finding; or a third connection between the hypothesis and a related research question.

104. The non-transitory computer-readable medium of claim 83, wherein the atom comprises a hypothesis, and the updating comprises performing at least one of: inserting a first connection between the hypothesis and a variant, equivalent, contradictory, or related hypothesis; updating the first connection between the hypothesis and the variant, equivalent, contradictory, or related hypothesis; inserting a second connection between the hypothesis and a supporting, refuting, inconclusive, or unrelated research finding; updating the second connection between the hypothesis and the supporting, refuting, inconclusive, or unrelated research finding; inserting a third connection between the hypothesis and a related research question; or updating the third connection between the hypothesis and the related research question.

105. The non-transitory computer-readable medium of claim 85, wherein the measuring comprises: determining a user contribution metric from the evaluation; determining a user performance metric from the evaluation; determining a user collaborative metric from the evaluation; and determining a user score based on at least the user contribution metric, the user performance metric, or the user collaborative metric.

106. The non-transitory computer-readable medium of claim 105, wherein the user contribution metric comprises a number of contributions made by the user to at least one of the text, the atom, the updated atom, the connection, or the updated connection.

107. The non-transitory computer-readable medium of claim 105, wherein the user performance metric comprises at least one of a number of citations, an h-index, a number of publications, a number of patents issued, a number of clinical trials performed, a number of grants received, a number of comments, a number of times a particular term is used by the user as compared to published literature, or a number of peer reviews made by the user.

108. The non-transitory computer-readable medium of claim 105, wherein the user collaborative metric comprises at least one of a number of collaborations, a number of cocontributions, or a number of co-citations between the user and peers of the user.

109. The non-transitory computer-readable medium of claim 103, wherein the operations further comprise: generating a first output visualizing the first connection; generating a second output visualizing the second connection; and generating a third output visualizing the third connection.

110. The non-transitory computer-readable medium of claim 85, wherein the measuring comprises generating a user performance metric based on an aggregation of a plurality of contributions by the user.

111. The non-transitory computer-readable medium of claim 85, wherein the operations further comprise: generating a first output visualizing the measured rating of the user; and generating a second output visualizing a comparison of the measured rating between the user and peers of the user.

112. The non-transitory computer-readable medium of claim 111, wherein the first output comprises: a first chart to display a user contribution metric for the user over a time period; a second chart to display a user performance metric for the user over the time period; and a third chart to display a rating score for the user over the time period.

113. The non-transitory computer-readable medium of claim 111, wherein the second output comprises: a first chart to display a comparison of a user contribution metric between the user and the peers of the user over a time period; a second chart to display a comparison of a user performance metric between the user and the peers of the user over the time period; and a third chart to display a comparison of a rating score between the user and the peers of the user over the time period.

114. The non-transitory computer-readable medium of claim 109, wherein the operations further comprise providing a recommendation to peers of the user or third-party agencies based on at least the measured rating of the user or the comparison of the measured rating between the user and the peers of the user.

115. A computer-implemented method compri sing : generating, by a computer processor based on a distance metric, one or more clusters from one or more texts, wherein the distance metric is applied to determine a similarity between a first text and a second text within a cluster; generating, by the computer processor, an embedding of the one or more clusters into a first vector space;generating, by the computer processor based on the embedding, a reduced embedding of the one or more clusters into a second vector space, wherein the second vector space has a lower dimensionality than the first vector space; generating, by the computer processor, a ranking of the reduced embedding of the one or more clusters, wherein a ranking score is assigned to the first text based on the similarity between the first text and the second text within the cluster; generating, by the computer processor based on the ranking score, a mapping between the first text and the reduced embedding of the first text; querying, by the computer processor based on the mapping, one or more LLMs to generate a topic label summarizing the first text within the cluster; and in response to the querying, generating, by the computer processor, a similarity map associated with the one or more clusters from the one or more texts, wherein the generated topic label is inserted into the similarity map to visualize the similarity between the first text and the second text within the cluster.

116. The computer-implemented method of claim 115, wherein the one or more clusters are generated from the one or more texts based on a nearest neighbor clustering.

117. The computer-implemented method of claim 115, wherein the distance metric comprises a cosine similarity, a Euclidean distance, a Manhattan distance, or a Jaccard similarity.

118. The computer-implemented method of claim 115, wherein the embedding of the one or more clusters is generated based on at least a Word2Vec model, a GloVe model, or a pretrained BAAI / bge model.

119. The computer-implemented method of claim 115, wherein the reduced embedding of the one or more clusters is generated based on at least a linear dimension reduction model or a non-linear dimension reduction model.

120. The computer-implemented method of claim 115, further comprising: storing, by the computer processor into a database using a structured format, the mapping between the first text and the reduced embedding of the first text, wherein the first text is retrievable based on the reduced embedding.

121. The computer-implemented method of claim 115, further comprising: determining, by the computer processor, whether the generated topic label achieves a pre-defined quality threshold; in response to the topic label not achieving the pre-defined quality threshold, receiving, by the computer processor from a user, a user feedback associated with the topic label; and updating, by the computer processor using the one or more LLMs, the topic label based on the user feedback.

122. A computer-implemented method comprising: extracting, by a computer processor, one or more keywords from a text; retrieving, by the computer processor from a database, one or more documents relevant to the text based on the extracted one or more keywords; ranking, by the computer processor based on a distance metric, the one or more documents from most to least relevant to the text, wherein the distance metric is applied to determine the relevance between a document and the text; in response to the ranking, extracting, by the computer processor, evidence from the ranked one or more documents, wherein the evidence comprises at least a part of the document relevant to the text; generating, by the computer processor querying the one or more LLMs based on the evidence, an evidence set associated with the text, wherein the evidence set incorporates at least one of the part of the document or another part of another document relevant to the text; and generating, by the computer processor querying the one or more LLMs based on the evidence set, a summary of the evidence set.

123. The computer-implemented method of claim 122, wherein the text comprises a positive hypothesis, a negative hypothesis, an opposite hypothesis, or a null hypothesis.

124. The computer-implemented method of claim 122, wherein the one or more keywords are extracted based on at least an unsupervised model, a graph-based model, a supervised model, or the one or more LLMs.

125. The computer-implemented method of claim 122, wherein retrieving the one or more documents further comprises: generating a search query comprising a matching condition and the extracted one or more keywords from the text, wherein the matching condition indicates how closely the extracted one or more keywords align with the one or more documents to trigger a match of the document in the retrieving; performing, based on the search query, a semantic search of the one or more documents within the database; and in response to performing the semantic search, identifying the one or more documents within the database based on the document achieving the matching condition of the search query.

126. The computer-implemented method of claim 122, wherein the evidence comprises supporting evidence and refuting evidence.

127. The computer-implemented method of claim 122, wherein summary of the evidence set summarizes the evidence over a period of time, and wherein at least a starting time stamp or an ending time stamp associated with the period of time is provided by a user.

128. A computer-implemented method comprising: receiving, by an agent platform using a computer processor, a search query from a user for locating logical structures in a database; identifying, by the agent platform, a context associated with the search query, wherein the context is relevant to the logical structures in the database; generating, by the agent platform, another database storing historical search query relevant to the search query and historical logical structures having been located based on the historical search query; retrieving, by the agent platform querying one or more LLMs based on a prompt, the logical structures from the database, wherein the prompt comprises the identified context, the historical search query, and the historical logical structures; refining, by the agent platform, the retrieved logical structures based on a user feedback; andstoring, by the agent platform into the other database, the search query and the logical structures, wherein the other database is searchable by another search query from the user for locating other logical structures in the database.

129. The computer-implemented method of claim 128, wherein the search query comprises user utterance, a natural language phrase, or one or more search keywords.

130. The computer-implemented method of claim 128, wherein each logical structure comprises a hypothesis, evidence associated with the hypothesis, and a linking between the hypothesis and the evidence, and wherein the linking provides a shortcut connection for retrieving one of the hypothesis or the evidence based on another one of the hypothesis or the evidence.

131. The computer-implemented method of claim 128, further comprising: identifying, by the agent platform, a user preference from the historical search query relevant to the search query; and refining, by the agent platform, the search query based on the identified user preference.

132. The computer-implemented method of claim 128, further comprising: determining, by the agent platform, whether the retrieved logical structures from the database achieve a pre-defined metric; and in response to the retrieved logical structures not achieving the pre-defined metric, refining, by the agent platform, the retrieved logical structures based on internal logics or parameters pre-defined in the agent platform, wherein the internal logics or parameters are updated over a period of time.

133. A system comprising: one or more memories; at least one processor each coupled to at least one of the memories and configured to perform operations comprising:generating, based on a distance metric, one or more clusters from one or more texts, wherein the distance metric is applied to determine a similarity between a first text and a second text within a cluster; generating an embedding of the one or more clusters into a first vector space; generating, based on the embedding, a reduced embedding of the one or more clusters into a second vector space, wherein the second vector space has a lower dimensionality than the first vector space; generating a ranking of the reduced embedding of the one or more clusters, wherein a ranking score is assigned to the first text based on the similarity between the first text and the second text within the cluster; generating, based on the ranking score, a mapping between the first text and the reduced embedding of the first text; querying, based on the mapping, one or more LLMs to generate a topic label summarizing the first text within the cluster; and in response to the querying, generating a similarity map associated with the one or more clusters from the one or more texts, wherein the generated topic label is inserted into the similarity map to visualize the similarity between the first text and the second text within the cluster.

134. The system of claim 133, wherein the one or more clusters are generated from the one or more texts based on a nearest neighbor clustering.

135. The system of claim 133, wherein the distance metric comprises a cosine similarity, a Euclidean distance, a Manhattan distance, or a Jaccard similarity.

136. The system of claim 133, wherein the embedding of the one or more clusters is generated based on at least a Word2Vec model, a GloVe model, or a pre-trained BAAI / bge model.

137. The system of claim 133, wherein the reduced embedding of the one or more clusters is generated based on at least a linear dimension reduction model or a non-linear dimension reduction model.

138. The system of claim 133, wherein the operations further comprise: storing, into a database using a structured format, the mapping between the first text and the reduced embedding of the first text, wherein the first text is retrievable based on the reduced embedding.

139. The system of claim 133, wherein the operations further comprise: determining whether the generated topic label achieves a pre-defined quality threshold; in response to the topic label not achieving the pre-defined quality threshold, receiving, from a user, a user feedback associated with the topic label; and updating, using the one or more LLMs, the topic label based on the user feedback.

140. A system comprising: one or more memories; at least one processor each coupled to at least one of the memories and configured to perform operations comprising: extracting one or more keywords from a text; retrieving, from a database, one or more documents relevant to the text based on the extracted one or more keywords; ranking, based on a distance metric, the one or more documents from most to least relevant to the text, wherein the distance metric is applied to determine the relevance between a document and the text; in response to the ranking, extracting evidence from the ranked one or more documents, wherein the evidence comprises at least a part of the document relevant to the text; generating, by querying the one or more LLMs based on the evidence, an evidence set associated with the text, wherein the evidence set incorporates at least one of the part of the document or another part of another document relevant to the text; and generating, by querying the one or more LLMs based on the evidence set, a summary of the evidence set.

141. The system of claim 140, wherein the text comprises a positive hypothesis, a negative hypothesis, an opposite hypothesis, or a null hypothesis.

142. The system of claim 140, wherein the one or more keywords are extracted based on at least an unsupervised model, a graph-based model, a supervised model, or the one or more LLMs.

143. The system of claim 140, wherein retrieving the one or more documents further comprises: generating a search query comprising a matching condition and the extracted one or more keywords from the text, wherein the matching condition indicates how closely the extracted one or more keywords align with the one or more documents to trigger a match of the document in the retrieving; performing, based on the search query, a semantic search of the one or more documents within the database; and in response to performing the semantic search, identifying the one or more documents within the database based on the document achieving the matching condition of the search query.

144. The system of claim 140, wherein the evidence comprises supporting evidence and refuting evidence.

145. The system of claim 140, wherein summary of the evidence set summarizes the evidence over a period of time, and wherein at least a starting time stamp or an ending time stamp associated with the period of time is provided by a user.

146. A system comprising: one or more memories; at least one processor each coupled to at least one of the memories and configured to perform operations comprising: receiving, by an agent platform, a search query from a user for locating logical structures in a database; identifying, by the agent platform, a context associated with the search query, wherein the context is relevant to the logical structures in the database;generating, by the agent platform, another database storing historical search query relevant to the search query and historical logical structures having been located based on the historical search query; retrieving, by the agent platform querying one or more LLMs based on a prompt, the logical structures from the database, wherein the prompt comprises the identified context, the historical search query, and the historical logical structures; refining, by the agent platform, the retrieved logical structures based on a user feedback; and storing, by the agent platform into the other database, the search query and the logical structures, wherein the other database is searchable by another search query from the user for locating other logical structures in the database.

147. The system of claim 146, wherein the search query comprises user utterance, a natural language phrase, or one or more search keywords.

148. The system of claim 146, wherein each logical structure comprises a hypothesis, evidence associated with the hypothesis, and a linking between the hypothesis and the evidence, and wherein the linking provides a shortcut connection for retrieving one of the hypothesis or the evidence based on another one of the hypothesis or the evidence.

149. The system of claim 146, wherein the operations further comprise: identifying, by the agent platform, a user preference from the historical search query relevant to the search query; and refining, by the agent platform, the search query based on the identified user preference.

150. The system of claim 146, wherein the operations further comprise: determining, by the agent platform, whether the retrieved logical structures from the database achieve a pre-defined metric; and in response to the retrieved logical structures not achieving the pre-defined metric, refining, by the agent platform, the retrieved logical structures based on internal logics or parameters pre-defined in the agent platform, wherein the internal logics or parameters are updated over a period of time.

151. A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising: one or more memories; at least one processor each coupled to at least one of the memories and configured to perform operations comprising: generating, based on a distance metric, one or more clusters from one or more texts, wherein the distance metric is applied to determine a similarity between a first text and a second text within a cluster; generating an embedding of the one or more clusters into a first vector space; generating, based on the embedding, a reduced embedding of the one or more clusters into a second vector space, wherein the second vector space has a lower dimensionality than the first vector space; generating a ranking of the reduced embedding of the one or more clusters, wherein a ranking score is assigned to the first text based on the similarity between the first text and the second text within the cluster; generating, based on the ranking score, a mapping between the first text and the reduced embedding of the first text; querying, based on the mapping, one or more LLMs to generate a topic label summarizing the first text within the cluster; and in response to the querying, generating a similarity map associated with the one or more clusters from the one or more texts, wherein the generated topic label is inserted into the similarity map to visualize the similarity between the first text and the second text within the cluster.

152. The non-transitory computer-readable medium of claim 151, wherein the one or more clusters are generated from the one or more texts based on a nearest neighbor clustering.

153. The non-transitory computer-readable medium of claim 151, wherein the distance metric comprises a cosine similarity, a Euclidean distance, a Manhattan distance, or a Jaccard similarity.

154. The non-transitory computer-readable medium of claim 151, wherein the embedding of the one or more clusters is generated based on at least a Word2Vec model, a GloVe model, or a pre-trained BAAI / bge model.

155. The non-transitory computer-readable medium of claim 151, wherein the reduced embedding of the one or more clusters is generated based on at least a linear dimension reduction model or a non-linear dimension reduction model.

156. The non-transitory computer-readable medium of claim 151, wherein the operations further comprise: storing, into a database using a structured format, the mapping between the first text and the reduced embedding of the first text, wherein the first text is retrievable based on the reduced embedding.

157. The non-transitory computer-readable medium of claim 151, wherein the operations further comprise: determining whether the generated topic label achieves a pre-defined quality threshold; in response to the topic label not achieving the pre-defined quality threshold, receiving, from a user, a user feedback associated with the topic label; and updating, using the one or more LLMs, the topic label based on the user feedback.

158. A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising: one or more memories; at least one processor each coupled to at least one of the memories and configured to perform operations comprising: extracting one or more keywords from a text; retrieving, from a database, one or more documents relevant to the text based on the extracted one or more keywords;ranking, based on a distance metric, the one or more documents from most to least relevant to the text, wherein the distance metric is applied to determine the relevance between a document and the text; in response to the ranking, extracting evidence from the ranked one or more documents, wherein the evidence comprises at least a part of the document relevant to the text; generating, by querying the one or more LLMs based on the evidence, an evidence set associated with the text, wherein the evidence set incorporates at least one of the part of the document or another part of another document relevant to the text; and generating, by querying the one or more LLMs based on the evidence set, a summary of the evidence set.

159. The non-transitory computer-readable medium of claim 158, wherein the text comprises a positive hypothesis, a negative hypothesis, an opposite hypothesis, or a null hypothesis.

160. The non-transitory computer-readable medium of claim 158, wherein the one or more keywords are extracted based on at least an unsupervised model, a graph-based model, a supervised model, or the one or more LLMs.

161. The non-transitory computer-readable medium of claim 158, wherein retrieving the one or more documents further comprises: generating a search query comprising a matching condition and the extracted one or more keywords from the text, wherein the matching condition indicates how closely the extracted one or more keywords align with the one or more documents to trigger a match of the document in the retrieving; performing, based on the search query, a semantic search of the one or more documents within the database; and in response to performing the semantic search, identifying the one or more documents within the database based on the document achieving the matching condition of the search query.

162. The non-transitory computer-readable medium of claim 158, wherein the evidence comprises supporting evidence and refuting evidence.

163. The non-transitory computer-readable medium of claim 158, wherein summary of the evidence set summarizes the evidence over a period of time, and wherein at least a starting time stamp or an ending time stamp associated with the period of time is provided by a user.

164. A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising: one or more memories; at least one processor each coupled to at least one of the memories and configured to perform operations comprising: receiving, by an agent platform, a search query from a user for locating logical structures in a database; identifying, by the agent platform, a context associated with the search query, wherein the context is relevant to the logical structures in the database; generating, by the agent platform, another database storing historical search query relevant to the search query and historical logical structures having been located based on the historical search query; retrieving, by the agent platform querying one or more LLMs based on a prompt, the logical structures from the database, wherein the prompt comprises the identified context, the historical search query, and the historical logical structures; refining, by the agent platform, the retrieved logical structures based on a user feedback; and storing, by the agent platform into the other database, the search query and the logical structures, wherein the other database is searchable by another search query from the user for locating other logical structures in the database.

165. The non-transitory computer-readable medium of claim 164, wherein the search query comprises user utterance, a natural language phrase, or one or more search keywords.

166. The non-transitory computer-readable medium of claim 164, wherein each logical structure comprises a hypothesis, evidence associated with the hypothesis, and a linking between the hypothesis and the evidence, and wherein the linking provides a shortcutconnection for retrieving one of the hypothesis or the evidence based on another one of the hypothesis or the evidence.

167. The non-transitory computer-readable medium of claim 164, wherein the operations further comprise: identifying, by the agent platform, a user preference from the historical search query relevant to the search query; and refining, by the agent platform, the search query based on the identified user preference.

168. The non-transitory computer-readable medium of claim 164, wherein the operations further comprise: determining, by the agent platform, whether the retrieved logical structures from the database achieve a pre-defined metric; and in response to the retrieved logical structures not achieving the pre-defined metric, refining, by the agent platform, the retrieved logical structures based on internal logics or parameters pre-defined in the agent platform, wherein the internal logics or parameters are updated over a period of time.

Citation Information

Patent Citations

  • Utilizing a large language model to perform a query

    US12019663B1

  • System and method for interoperable cloud DSL to orchestrate multiple cloud platforms and services

    US20180276060A1

  • Artificial intelligence-based question-answer natural language processing traces

    US20220300712A1

  • Methods for automated therapy and bioactive discovery and for automated therapy and bioactive delivery

    US20230154585A1

  • Systems and methods for intent discovery

    US20230315999A1

Cited By

  • Incremental knowledge evolution and multi-dimensional expert constraint driven bidding document generation method and system

    CN121936432A