Graph language models: distilling reliable knowledge graphs from high-quality text

Graph language models address the issue of hallucinations in LLMs by distilling knowledge graphs from high-quality text, enhancing interpretability and trustworthiness through the use of neural inversion and credible sources.

WO2025106768A1PCT designated stage expired Publication Date: 2025-05-22THE TRUSTEES OF PRINCETON UNIV

Patent Information

Application Number
PCT/US2024/056058
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-16
Filing Date
2024-11-15
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Current large language models (LLMs) suffer from hallucinations, generating false or misleading text due to a lack of true understanding and reasoning, and are limited by insufficient causal modeling objectives.

Method used

The development of graph language models (GLMs) that distill reliable knowledge graphs from high-quality text using neural inversion, combining syntactic and semantic information to mitigate hallucinations and improve interpretability.

Benefits of technology

GLMs effectively reduce hallucinations by leveraging credible sources and producing responses that can be mapped to paths in the knowledge graph, providing causal understanding and improving the trustworthiness of AI models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024056058_22052025_PF_FP_ABST
    Figure US2024056058_22052025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are techniques for training a graph language model (GEM). The techniques generally include training the GEM on a training corpus, expanding a current training knowledge graph into an expanded knowledge graph by distilling semantic information from the training corpus using neural inversion. The knowledge graph is only expanded when there is new data, and if convergence has not been reached, the method may include updating the training corpus by injecting the expanded knowledge graph into the training corpus, thereby making the expanded knowledge graph the current training knowledge graph. The training, expanding, and updating steps are repeated until convergence. The training corpus may be formed by parsing a dataset into chain graphs, where syntactic language information is represented as a chain sequence of one or more root nodes, and where semantic information from the seed knowledge graph is added as leaf nodes around each root node.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] GRAPH LANGUAGE MODELS: DISTILLING RELIABLE KNOWLEDGE GRAPHS FROM HIGH-QUALITY TEXT

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] This present disclosure claims priority to U.S. Provisional Patent Application No. 63 / 599.846, filed November 16, 2023, the contents of which are incorporated by reference herein in its entirety.

[0004] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0005] This invention was made with government support under Grant No. CNS-2216746 awarded by the National Science Foundation. The government has certain rights in the invention.

[0006] TECHNICAL FIELD

[0007] Disclosed are graph language models that allow for distilling reliable knowledge graphs from high-quality text.

[0008] BACKGROUND

[0009] Large language models (LLM) are being adopted in more and more industries, including, e g., chatbots and / or search tools. An LLM is an artificial intelligence system designed to understand and generate human-like text. It is trained on vast amounts of text data and uses complex algorithms to predict and produce language based on context. LLMs can perform a variety' of tasks, including answering questions, summarizing information, translating languages, and engaging in conversation. Their strength lies in their ability to generate coherent and contextually relevant responses, although they lack true understanding and reasoning.

[0010] However, LLMs may "‘hallucinate”. Hallucinations refer to instances when an LLM generates text that is false, misleading, or nonsensical but is presented confidently. This can happen for various reasons, because the model doesn't have a true understanding of facts or concepts; it simply generates responses based on patterns in the data it was trained on. BRIEF SUMMARY

[0011] Disclosed is a method for training a graph language model (GLM). The method includes training the GLM on a training corpus. The method includes expanding a current training knowledge graph into an expanded knowledge graph by distilling semantic information from the training corpus using neural inversion. The training knowledge graph will be expanded only when there is new information / data (e.g., when the KG can be improved). The method may include updating the training corpus by injecting the expanded knowledge graph into the training corpus, thereby making the expanded knowledge graph the current training knowledge graph. The training, expanding, and updating steps are repeated until convergence.

[0012] The method may include receiving a dataset. The method may include forming a training corpus by injecting a current knowledge graph into the dataset, where the cunent knowledge graph is a seed knowledge graph, and where the training corpus includes syntactic information from an initial dataset (a textual or “syntactic” dataset) and semantic information from the seed knowledge graph. The method may include extracting texts and / or images from a source (e.g., processing a source such as PubMed to form the dataset). Forming the training corpus may include parsing the dataset into chain graphs, where syntactic language information is represented as a chain sequence of one or more root nodes, and where semantic information from the seed knowledge graph is added as leaf nodes around each root node. There may be a plurality of leaf nodes around each root node.

[0013] Training the GLM may include utilizing a masked language modeling (MLM) objective, forcing the GLM to predict masked tokens or nodes using neighboring information. Both a root node and / or a leaf node may be masked. A masking probability for root nodes and leaf nodes may be, independently, any appropriate value, such as 10%-25%.

[0014] Expanding the current training knowledge graph may include freezing weights and backpropagating gradients toward an input. Expanding the current training knowledge graph may include distilling out semantic learned representations of the GLM to add leaf nodes to the input, thus making the learned representation “explicit’. Expanding the current training knowledge graph may include extracting relevant leaf nodes after post-processing to expand the current training knowledge graph.

[0015] In various aspects, a non-transitory computer-readable storage device may be provided. The storage device may include instructions, that, when executed by one or more processing units, causes the one or more processing units to, collectively, perform a method as disclosed herein. In various aspects, a system may be provided. The system may include one or more processing units operably coupled to a non-transitory computer-readable storage device. The storage device may include instructions, that, when executed by the one or more processing units, causes the one or more processing units to, collectively, perform a method as disclosed herein. The one or more processing units may be configured to communicate with one or more remote devices. At least one of the one or more remote devices may be configured to receive input from a user.

[0016] BRIEF DESCRIPTION OF FIGURES

[0017] Figure 1 is a schematic illustration of a framework for training and using graph language models.

[0018] Figure 2 is a schematic illustration of a system.

[0019] Figure 3 is a flowchart for a method as disclosed herein.

[0020] Figure 4 is a flowchart for providing a training corpus.

[0021] Figure 5 is a schematic illustration of a knowledge graph (KG) injection process, where definitions in the KG are injected into the chain graph. Syntactic and semantic information are represented with rounded squares and oval shapes, respectively.

[0022] Figure 6 is a schematic illustration of seed KGs and chain graphs leading to fully- connected leaves.

[0023] Figure 7 is a schematic illustration of KG expansion based on neural inversion.

[0024] DETAILED DESCRIPTION

[0025] Modem interpretability methods for neural networks cannot precisely analyze complex artificial intelligence (Al) models. Moreover, to mitigate hallucination in LLMs, researchers attempt to improve training dataset fidelity and induce modeling biases. However, the causal modeling objective in current LLMs is insufficient for addressing these challenges. Instead, the present disclosure proposes fusing knowledge graph (KG) generation with the LLM objective, resulting in graph language models (GLMs). KGs are inexpensive (for inference), fast, explainable, editable, and auditable. The proposed graph / language modeling framework addresses, e.g., the lack of interpretability and hallucination in current LLMs. The proposed GLM supports connections in the semantic space, going beyond purely syntactic representations modeled by existing methods, and distills knowledge graphs from high-quality corpora. Any response by the disclosed GLM can be mapped to a path in the KG. thus providing causal understanding of the model response. It is hypothesized that prioritizing data quantity over quality' (the prevailing paradigm in the Al community) is one of the main reasons for fallacious text generations. Therefore, to eliminate hallucination, the disclosed framework relies on a training dataset that leverages only credible sources. Further, since the KG represents the neural GLM in symbolic form, one can audit the KG produced by the trained GLM. This is not possible with existing LLMs. Hence, through this work, introduced is a novel regime of training neuro-symbolic models that leverage self-supervision at scale, thus opening new avenues for building responsible and trustworthy artificial general intelligence (AGI) models of the future.

[0026] FIG. 1 shows a high-level overview of the proposed framework, which combines syntactic and semantic information in a unified model. The framework involves pre-processing information (such as books, articles, etc.) from credible sources. This may involve, e.g., extracting information having one or more modalities (including multiple modalities) from these sources, such as extracting text, figures, etc. A seed KG is then injected into the extracted information (e.g., the extracted text) to form an updated corpus, (with syntactic information from the text and semantic information from the seed KG) for GLM training. The framework leverages neural inversion to distill semantic information from the text in order to expand the seed KG. When there is new data / information, the KG is expanded. The expanded KG is then used to update the training corpus. This cycle continues until convergence (as is known, neural inversion generally always runs until convergence). This may be, e.g., when additional nodes do not result in information gain in the resultant KG. Hence, using interleaved training and neural inversion, one can generate a symbolic KG from text in the target domain. Finally, one can leverage the GLM and its KG counterpart to answ er user requests, employing the proposed question-answering system.

[0027] As will be understood, these approaches may be implemented on almost any system utilizing processing unit(s). As used herein, the term '‘processing unit" refers to, is part of, or includes hardw are components such as an electronic circuit, a logic circuit, a processor (shared, dedicated, or group) and / or memory7(shared, dedicated, or group), an Application Specific Integrated Circuit (ASIC), a field-programmable device (FPD) (e.g.. a field-programmable gate array (FPGA), a programmable logic device (PLD). a complex PLD (CPLD), a high-capacity PLD (HCPLD), a structured ASIC, or a programmable SoC), digital signal processors (DSPs), etc., that are configured to provide the described functionality. In some embodiments, the processing unit may execute one or more softw are or firmware programs to provide at least some of the described functionality. The term “processing unit7’ may also refer to a combination of one or more hardware elements (such as a combination of circuits used in an electrical or electronic system) with the program code used to carry out the functionality of that program code. The term "‘processing unif ’ may also refer to one or more application processors, one or more baseband processors, a physical central processing unit (CPU), a single-core processor, a dual-core processor, a triple-core processor, a quad-core processor, and / or any other device, or combination of devices, capable of executing or otherwise operating (either individually or as a combined unit of processing components), collectively, computer-executable instructions such as program code, software modules, and / or functional processes.

[0028] Here, an example, of a system is shown in FIG. 2. The system (200) may include one or more processing units (212). The processing unit(s) (212) may be one a single device (e.g,. device (210)), or may be on multiple devices (e.g., operably coupled devices (210, 211)).

[0029] The system may include one or more devices (210, 211). Each device (210, 211) may include one or more processing units (212). The processing unit(s) (212) may be operably coupled to memoiy (214). The processing unit(s) (212) may be operably coupled to a transceiver (216), such a wireless communications transceiver, an ethemet transceiver, etc.. Such a transceiver may allow a device (210) to communicate, e.g., with one or more remote processing unit(s) (222). and / or with one or more additional devices (211). The processing unit(s) may be coupled to a non-transitory computer-readable storage device, which may be an on-device storage device (e.g., storage device (218)) and / or remote storage (e.g., remote storage device (230)) operably coupled to the device(s). The storage device(s) may contain instructions that, when executed by the processing unit(s), causes the processing unit(s) to, collectively, perform various steps of a method as disclosed herein.

[0030] In some embodiments, the device(s) (210, 211) may be in communication with the remote processing unit(s) (222) of a remote device (220). such as a remote PC, laptop, smartphone, etc. The remote processing unit(s)(222) may be configured to receive input from a user (224) (e.g., via a keyboard, mouse, etc.).

[0031] Referring to FIG. 3, a method for training a GLM is shown. The method (300) may include providing (310) a training corpus. The training corpus may be provided in any appropriate manner.

[0032] For example, referring to FIG. 4, providing (310) the training corpus may include receiving (410) a dataset, such as a dataset from a data repository. The dataset may include one or more sources of text and / or images (e.g., a publication, such as a journal article). The method may include extracting (420) texts and / or images from a source.

[0033] The method may include forming (430) a training corpus by injecting a cunent knowledge graph into the dataset. The current knowledge graph may be a seed knowledge graph. The training corpus may include syntactic information from the initial dataset (which is a textual dataset, sometimes referred to as a “syntactic” dataset) and semantic information from the seed knowledge graph. Forming the training corpus may be done in any appropriate manner. For example, referring to FIG. 5, one approach includes parsing the received dataset into chain graphs (520), where syntactic language information is represented as a chain sequence (e.g., leafy chain graph (530)) of one or more root nodes (532), and where semantic information from the seed knowledge graph (510) is added as leaf nodes (534) around each root node (532). In some embodiments, there may be a plurality of leaf nodes (534) around each root node (532).

[0034] Referring to FIG. 4, the method may include sending (440) (or receiving) the training corpus from one device to another (e.g., if a first server’s processing unit(s) are used to form the training corpus, while a second server’s processing unit(s) are used to train the GLM, the training corpus may be sent from one device to another.

[0035] Referring to FIG. 3, the method may include training (320) the GLM on a training corpus. The GLM may be trained in this step in any appropriate manner. In some embodiments, this may include utilizing a masked language modeling (MLM) objective, forcing the GLM to predict masked tokens or nodes using neighboring information.

[0036] Some or all of the root node(s) may be masked. Some or all of the root node(s) may be unmasked. Some or all of the leaf node(s) may be masked. Some or all of the leaf node(s) may be unmasked. Both root node(s) and leaf node(s) may be masked. Any appropriate masking probability may be utilized. The masking probability for root and leaf nodes may be at least 10%. The masking probability for root and leaf nodes may be at no more than 25%. In various aspects, a masking probability for root nodes and leaf nodes may be, independently, 10%-25%. In various aspects, a masking probability for each root node and leaf node may be, independently, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10% up to 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50%, including all ranges and subranges thereof.

[0037] The method may include expanding (330) the current knowledge graph into an expanded knowledge graph by distilling semantic information from the training corpus using neural inversion. Expanding the seed knowledge graph may include freezing weights and backpropagating gradients toward an input. Expanding the seed knowledge graph may include expanding the seed knowledge graph includes distilling out semantic learned representations of the GLM to add leaf nodes to the input, and thus making the learned representations ‘explicit’. Expanding the seed knowledge graph may include extracting relevant leaf nodes after post-processing to expand the current knowledge graph. The process may include updating (350) the training corpus by injecting the expanded knowledge graph into the training corpus, thereby making the expanded knowledge graph the current training graph. The method may include repeating the training, expanding, and updating steps until convergence is determined (340).

[0038] The training corpus may be provided in any appropriate manner. Referring to FIG. 4, for example, providing (310) the training corpus may include receiving (410) a dataset. It may include extracting (420) texts and / or images from a source within the dataset. providing the training corpus may include forming (430) a training corpus by injecting a current knowledge graph into the dataset, where the current knowledge graph is a seed knowledge graph, and where the training corpus includes syntactic information from the initial dataset and semantic information from the seed knowledge graph. Forming (430) the training corpus may include parsing (432) the dataset into chain graphs, where syntactic language information is represented as a chain sequence of one or more root nodes, and where semantic information from the seed knowledge graph is added as leaf nodes around one or more root nodes (and preferably, around each root node).

[0039] This concept can be seen schematically in FIG. 5, where a seed KG (510) may be combined with the chain graph (520), resulting in a leafy chain graph (530), where each root node (532) may include leaf nodes (534) from the seed KG (510) around the root node(s). As will be understood, for simplicity, a simple seed KG, chain graph, and leafy chain graph are shown here. However, more complex seed KGs, chain graphs, and leafy chain graphs (such as having multiple leaf nodes around every node) may readily be generated.

[0040] It will be understood that a device used to generate the training corpus may be the same device that trains the GLM on the training corpus, or, it may be a different device. For example, referring to FIG. 2. every thing may be done on device (210), or the preparation of the training corpus may be done on device (211) and the training of the GLM may be done on device (210). Thus, referring to FIG. 4, providing the training corpus may include sending (440) the training corpus, e.g., to a device training the GLM on the training corpus.

[0041] Example

[0042] Knowledge Graph Extraction

[0043] A KG structures conceptual knowledge in a very intuitive way. Nodes (also called entities) in a KG are concepts connected by edges (also called relations). Compared to internal representations learned in an LLM. a KG is interpretable, resulting in solutions more acceptable to physicians and domain experts. To standardize medical KG development, conventional approaches suggest hierarchical, interconnected classes and relationships for biomedical entities.

[0044] The conventional method for building a medical KG is to extract concepts and relations from the Semantic MEDLINE Database (SemMedDB) or parse medical texts using the SemRep natural language processing (NLP) tool. For instance, some efforts have retrieved information in the breast cancer domain using SemMedDB to discover new connections between paper abstracts in the PubMed library. However. SemRep narrows down the scope of extracted information. For example, from the diabetes definition ("Diabetes is a group of metabolic diseases characterized by hyperglycemia resulting from defects in insulin secretion, insulin action, or both"’), the interactive SemRep extracts the following relations (with request parameters of knowledge source 2018, lexicon year 2018, in the relaxed model setting. Metainformation has been removed from the output):

[0045] IS A

[0046] Diabetes — > Metabolic Diseases

[0047] COEXISTS WITH

[0048] Hyperglycemia - > Metabolic Diseases

[0049] COEXISTS WITH (SPEC)

[0050] Hyperglycemia - * Diabetes

[0051] On the other hand, for a less general sentence (“The chronic hyperglycemia of diabetes is associated with long-term damage, dysfunction, and failure of different organs, especially the eyes, kidneys, nerves, heart, and blood vessels”), SemRep yields an empty result. This example demonstrates the limitations of SemRep. It is unpredictable and often incorrect. Methodologies that employ SemRep NLP tools inherit all its limitations.

[0052] Other conventional approaches combine data from multiple medical sources and ontologies into a KG. Their question-answering system parses queries and ranks responses. They characterize responses as graphs that match the query in terms of topology and entity types. However, their aggregation system does not integrate the retrieved information intelligently.

[0053] In contrast, the presently disclosed approach extract concepts from original text-based sources rather than rely on curated databases from third parties.

[0054] Hallucination in LLMs

[0055] LLMs are infamous for hallucinating, which deters reliance on their truthfulness and faithfulness. Hallucination refers to generated content that is nonsensical or unfaithful to the provided source content. For LLMs, memorization along with biases in the training corpus (e.g., reliance on term frequency) emerge as the most significant factor in hallucination. For instance, LLMs often learn spurious correlations, such as proximal or highly co-occurring associations, as factual knowledge. Another primary' factor is the source-reference (target) divergence in the training dataset. In this case, the mismatch between the source and the reference data hampers factual consistency. Therefore, the model does not leam the correct correlation.

[0056] Hallucination hinders Al system performance in multiple ways. Incorrect but deceptively coherent output jeopardizes safety in areas with high requirements for truthfulness, such as medicine, national security, and cybersecurity. In less sensitive fields, such as customer support and recommendation systems, the erroneous content undermines the provider’s credibility, deceives users, and spreads mystification. Current natural language generation (NLG) systems do not guarantee output factuality or semantic coherence. In practical applications, the user is required to check the factual correctness of the generated response. This problem is yet to be resolved.

[0057] The concept of data quality in the NLG domain needs to be defined. Poor-quality data emanate from non-verified data sources (such as Twitter), skewed datasets, and scarce data for isolated and rare cases. Research demonstrates that biases in the training data hinder the reasoning processes in current LLMs. This amplifies human biases present in the training corpora. However, one can estimate data quality heuristically. One expects journal papers to be high-quality sources - papers and academic books represent the quintessence of humanity’s knowledge, free of nonsense. Due to the high quality of academic knowledge, students who leam from books do so quickly and avoid the most frequent errors. Further, one can hypothesize that a neural network trained on academic knowledge would generally exhibit robust performance characteristics (in terms of truthfulness given any input prompt).

[0058] The current trend in the deep learning community is to upscale NLG models and datasets. The notion that larger neural models capture more sophisticated patterns in the training corpora drives this trend. For instance, OpenLLaMA is trained on the RedPajama dataset, much larger than previous datasets. Research shows that LLMs leam more abstract concepts than their smaller counterparts. However, scaling model sizes comes with a cost, i.e., LLM training and inference demand higher computational resources. Further, creating a substantially large, diverse, and unbiased training corpus poses a challenge. Finally, verifying the factuality of large datasets is intractable.

[0059] Studies also show' that solely scaling up models negatively impacts truthfulness, requiring finetuning and prompt engineering.

[0060] Existing Medical Language Models With the advent of generative Al, the healthcare industry has rapidly incorporated LLMs tailored to clinical needs. Specifically. Med-PaLM 2, a PaLM model finetuned in the medical domain, is the first LLM to exceed the "passing'' score on the US Medical Licensing Examination (USMLE) and achieves state-of-the-art scores on MedMCQA, PubMedQA, and MMLU clinical dataset.

[0061] Graph Transformer Architecture

[0062] The transformer architecture has established itself as the de facto standard for NLP tasks. Since its first release in 2017, the model has been applied to other domains, including computer vision, audio, and even multimodal domains. At the heart of the transformer architecture lies the “self-attention” operation.

[0063] Some prior efforts modified the self-attention module for training transformers on graphical input. The proposed Graphormer model encodes node centrality as a learnable vector. For each node pair, it adds the shortest path distance between them and formulates edge types to capture the extra information. Graphormer outperforms previously-proposed graph neural networks (GNN). The presently disclosed envisions leveraging the Graphormer architecture in the disclosed GLM.

[0064] In this example, an approach using modules shown in FIG. 1 is described.

[0065] Data Pre-processing

[0066] Data quality is central to the methodology. Here, the training dataset was sourced from MEDLINE journals and the seed KG from UMLS Metathesaurus. The MEDLINE library was selected based on the National Institute of Health (NTH) charted advisory committee of expert recommendations (NIH maintains UMLS). To create the training dataset, we retrieved MEDLINE papers on diabetes from PubMed Central using the following query:

[0067] (

[0068] "diabetes mellitus"[MeSH Terms]

[0069] OR "diabetes insipidus" [MeSH Terms])

[0070] AND (medline[sb]

[0071] AND "2018 / 06 / 1 l"[PubDate] : "2023 / 06 / 09" [PubDate]

[0072] )

[0073] This yields 43,087 papers. Next, non-English papers are removed, and the rest are parsed into sentences using the PubMed parser.

[0074] The seed KG represents a small set of the most frequent medical terms in the dataset with simple definitions retrieved from UMLS. See Table 1, below. After tokenization, these definitions must fit into three tokens (details below). For instance, the UMLS Metathesaurus gives the following definition of Apoptosis: A regulated cell death mechanism characterized by distinctive morphologic changes in the nucleus and cytoplasm, including the endonucleolytic cleavage of genomic DNA, at regularly spaced, internucleosomal sites, i.e., DNA fragmentation. It is genetically programmed and serves as a balance to mitosis in regulating the size of animal tissues and in mediating pathologic processes associated with tumor growth. Here, the definition is too long and would thus consume several tokens to represent. To limit the number of tokens (and thus, nodes) used to represent semantic information, the definitions are reduced to two or three tokens. For the above example, cell death mechanism is set in the seed KG; hence, the term definition fits within the three-token limit.

[0075] Table 1 (Example of a seed KG. A table of terms and definitions.)

[0076] One can represent any sentence input to the GLM as a chain graph. In other words, one can represent syntactic language information as a chain sequence. To represent semantic information around the nodes in that chain, leaf nodes were added around the nodes in this chain graph (such nodes in the chain may be referred to as root nodes). The chain graphs were initialized with three leaf nodes (per root node in the chain graph), where the leaf nodes were set as <pad> tokens. We refer to the parameter that determines the number of leaf nodes as ny

[0077] During the KG injection step, the <pad> tokens (in the leaf nodes) were replaced with the tokenized definitions from the seed KG. For terms with several definitions, the one that is most commonly used was choose. As a result, root nodes in the chain graph carry syntactic information, while leaf nodes represent semantic knowledge. FIG. 5 shows the dataset preprocessing step. The training dataset was parsed into chain graphs with only <pad> tokens in leaves first. Then, the seed KG was injected. Note that, here, the leaf nodes are shown as being added around the first token (W) of the term (WFS1). However, as will be understood, the tokenizer can be configured to represent standard medical terms as singular tokens.

[0078] GLM Architecture

[0079] The proposed GLM in this example is a transformer-based GNN, (x, 0), with 6.9M trainable parameters. It is based on the Graphormer model. The length of the chain graphs were set to 128. To implement this, the input text was tokenized and chains were formulated with 128 tokens. Here, w was set equal to 3, i.e., each root node in the chain graphs supports up to three leaf nodes. A root node is always a non-empty token. A leaf node is either a <pad> token or represents semantic information. If the chain is shorter than the target length of the chain graphs (128 tokens), the rest of the chain may be padded using the <pad> token. In a fully- connected setting, every root node connects to a non-<pad>-token leaf node, as is shown in FIG. 6, where a seed KG (610) is injected (640) into a dataset of chain graphs (620), resulting in fully connected leaves (630).

[0080] Neural Inversion and KG Expansion

[0081] Neural inversion is leveraged to expand the KG, distilling the internal representations of the trained GLM to add semantic information to the syntactic chain graphs. Neural inversion refers to the process of backpropagating the gradients toward the input. For a given loss function £. this refers to Xj+1= X, — 7]VX£(1F) for the i-th step. Here, / is a hyperparameter.

[0082] FIG. 7 shows the KG expansion step that leverages neural inversion. First, the GLM is trained on the KG-injected data. For this, the masked language modeling (MLM) objective was used. MLM training is typical of most encoder-based models. It forces the model to predict masked tokens (or nodes in the graphical formulation) using neighboring information. The proposed training procedure involves masking not only the nodes in the chain graph but also the leaf nodes (if they are not <pad> tokens).

[0083] Any masking probabilities for both syntactic and semantic nodes may be used. Here, a 15% masking probability was used for both syntactic and semantic nodes, although different masking probabilities for both modalities was also explored. Next, the trained GLM was leveraged for KG expansion. For this, the weights were frozen and the gradients backpropagated toward the input. This enables one to distill out the internal syntactic and semantic learned representations of the GLM to add leaf nodes to the input. Relevant leaves are then extracted (after post-processing) to expand the current KG. The expanded KG is then used to inject semantic information into the source dataset and retrain the GLM. This cycle of GLM training and KG expansion is repeated until convergence is reached (i.e., no novel information is added to the KG). Question-answering System

[0084] The proposed question-answering system consists of two components. The first runs inference with the trained GLM for a neural response to the user request. The second uses the formulated KG for a symbolic response to the question. A response generated using the KG is a logical path in the KG that answers the question. It is interpretable and shows the causality behind the output response. V arious inference methods employing the generated KG may be used.

[0085] As seen, this is a novel framework for extracting neural and symbolic representations from high-quality text. The trained GLM (using syntactic and semantic information) results in robust internal representations that mitigate hallucination in traditional LLMs. Using interleaved training and neural inversion, the proposed GLM is employed to distill an interpretable KG from the source text. As a prospective avenue for further investigations, it is anticipated that the GLM framework will demonstrate its potential to unveil novel insights in the medical domain through previously undiscovered conceptual connections. While the case study was conducted within the medical domain, this approach has broader applicability' across a large number of domains. Indeed, at a minimum, it will be understood that any field where high-quality journals exist for that field would be readily applicable here. Scientific fields, such as molecular biology7, chemistry7(including organic, inorganic, physical, and / or analytical chemistry), physics (including theoretical physics, astrophysics, and condensed matter physics), environmental science, engineering (including mechanical, electrical, chemical, etc.), social sciences (including psychology, sociology, and anthropology ), etc.

Claims

What is claimed is:

1. A method for training a graph language model (GLM), comprising: training the GLM on a training corpus; expanding a current training knowledge graph into an expanded knowledge graph by distilling semantic information from the training corpus using neural inversion; determining if convergence has been reached; updating the training corpus by injecting the expanded knowledge graph into the training corpus, thereby making the expanded knowledge graph the current training knowledge graph; and repeating the training, expanding, and updating steps until convergence.

2. The method of claim 1, further comprising: receiving a dataset; and forming a training corpus by injecting a current knowledge graph into the dataset, where the current knowledge graph is a seed knowledge graph, and where the training corpus includes syntactic information from an initial dataset and semantic information from the seed knowledge graph.

3. The method of claim 2, further comprising extracting texts and / or images from a source.

4. The method of claim 2 or 3, wherein forming the training corpus includes parsing the dataset into chain graphs, where syntactic language information is represented as a chain sequence of one or more root nodes, and where semantic information from the seed knowledge graph is added as leaf nodes around each root node.

5. The method of claim 4, wherein there are a plurality of leaf nodes around each root node.

6. The method of any one of claims 1-5. wherein training the GLM includes utilize a masked language modeling (MLM) objective, forcing the GLM to predict masked tokens or nodes using neighboring information.

7. The method of claim 6, wherein both a root node and a leaf node are masked.

8. The method of claim 6 or 7, wherein a masking probability for root nodes and leaf nodes are, independently. 10%-25%.

9. The method of any of claims 1-8, wherein expanding the current training knowledge graph includes freezing weights and backpropagating gradients toward an input.

10. The method of claim 9, wherein expanding the current training knowledge graph includes distilling out semantic learned representations of the GLM to add leaf nodes to the input.

11. The method of claim 9 or 10, wherein expanding the current training knowledge graph includes extracting relevant leaf nodes after post-processing to expand the current training knowledge graph.

12. A non-transitory computer-readable storage device, comprising instructions, that, when executed by one or more processing units, causes the one or more processing units to, collectively, perform a method of any one of claims 1-1 1.

13. A system, comprising one or more processing units operably coupled to a non-transitory computer-readable storage device, comprising instructions, that, when executed by one or more processing units, causes the one or more processing units to, collectively, perform a method of any one of claims 1 -1 1.

14. The system of claim 13, wherein the one or more processing units are configured to communicate with one or more remote devices.

15. The system of claim 14, wherein at least one of the one or more remote devices are configured to receive input from a user.

Citation Information

Patent Citations

  • Techniques for building a knowledge graph in limited knowledge domains

    US20200057946A1

  • Systems and methods for machine learning-based site-specific threat modeling and threat detection

    US20200202184A1

  • Differentially private dataset generation and modeling for knowledge graphs

    US20210374279A1

  • Grouping nodes in a system

    US20230259744A1

Cited By

  • Data quality detection and management method based on large model

    CN121833692A

  • Multimodal data based generation of quality assured knowledge graph and contexts for user queries

    US20260245681A1