Aging biomarkers
By integrating knowledge graphs and machine learning to analyze biomarker data, the method addresses the challenge of organizing and identifying key biomarkers, enhancing the understanding of aging processes and health outcomes.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- KANSAS STATE UNIV RES FOUND
- Filing Date
- 2025-11-03
- Publication Date
- 2026-05-07
AI Technical Summary
Existing biomarker data is not well organized or classified, leading to difficulties in identifying meaningful relationships and connections, particularly in the field of aging research, where different fields focus on different types of biomarkers without a common system or language, resulting in missed connections that could be important for understanding diseases and health outcomes.
A method and system utilizing knowledge graphs and machine learning to integrate and analyze heterogeneous datasets, including genomics, transcriptomics, and epigenomics, to generate an ontology for aging biomarkers, capturing relationships and correlations through vector embeddings and graph embeddings, and using LLMs and classical machine learning to refine and visualize these connections.
Enables the identification of key biomarkers and predictions regarding health-related outcomes by organizing and analyzing vast biomarker data, uncovering correlations and relationships that traditional methods miss, facilitating a more holistic understanding of aging processes.
Smart Images

Figure US2025053785_07052026_PF_FP_ABST
Abstract
Description
AGING BIOMARKERSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] The current patent application is a non-provisional utility patent application which claims priority benefit of earlier-filed U.S. Provisional Application Ser. No. 63 / 714,957; titled “AGING BIOMARKERS”; and filed November 1, 2024. The Provisional Application is hereby incorporated by reference, in its entirety, into the current patent application.BACKGROUND OF THE INVENTION
[0002] Aging is a complex biological process influenced by various biomarkers, such as cholesterol and blood sugar levels, which serve as measurable indicators of health and disease. However, the vast amount of biomarker data presents challenges in identifying meaningful relationships. One major problem is that these biomarkers are not well organized or classified. This means that even though researchers have a lot of information about them, it is very difficult to bring all of the data together and make sense of it. Different fields of study may focus on different types of biomarkers without a common system or language to link them. As a result, researchers and healthcare professionals may miss connections that could be very important for understanding diseases, evaluating treatments, or preventing health problems before they begin. This problem becomes even clearer in the area of aging research. Aging is a very complex process that occurs at many levels, from genes and cells to entire bodies. Scientists use different “aging clocks”, such as epigenetic clocks (focusing on DNA methylation), transcriptomic clocks (focusing on RNA), and proteomic clocks (focusing on proteins) to measure the biological age of a person and predict how their body might change over time. These clocks are helpful, but they typically look at only one type of data and can miss important information.
[0003] Thus, there is a need for an improved means for organizing biomarker data. This background discussion is intended to provide information related to the present invention which is not necessarily prior art.SUMMARY OF THE INVENTION
[0004] Embodiments of the current invention address one or more of the above-mentioned problems and provide a distinct advance in the art of methods and systems of processing biomarker data.
[0005] One embodiment of the invention is a method including retrieving, at a computing device, from a database biomarker files, each of the biomarker files including a patient identification, an age, and a biomarker value; receiving, at the computing device, two selections of biomarkers for comparison; receiving, at the computing device, two ranges of values of a first selected biomarker to generate two groups of the biomarker files; determining, at the computing device, a first interquartile range of a second selected biomarker of a first group; determining, at the computing device, a second interquartile range of the second selected biomarker of a second group; determining, at the computing device, the first interquartile range is outside the second interquartile range; and recording, at the computing device, an indication of a relationship between the two selections of biomarkers.
[0006] Another embodiment of the invention is a system including a database and a processor. The database stores an ontology having biomarkers, properties, and connections between the biomarkers. The processor is configured to retrieve from the database biomarker files, each of the biomarker files including a patient identification, an age, and a biomarker value; receive a first biomarker selection and a second biomarker selection for comparison; receive two ranges for values of the first biomarker selection to generate a first group and a second group of the biomarker files; determine a first interquartile range of the second biomarker selection of the first group; determine a second interquartile range of the second biomarker selection of the second group; determine the first interquartile range is outside the second interquartile range; and update the ontology by creating and storing, in the ontology, a connection between the first and second biomarker selections.
[0007] Another embodiment of the invention is a system including a database, a user interface, and a processor. The database stores an ontology including biomarkers, biomarker categories, biomarker identification methods, condition indications, and connections between the biomarkers, the biomarker categories, the biomarker identification methods, and the condition indications. The connections between two biomarkers are representative of statistically significant correlations derived from structured data. The connections between the biomarkers and thecondition indications are representative of published causal data derived from unstructured data. The processor is in communication with the database and the user interface and is configured to display on the user interface the ontology depicted as a graph with nodes representative of the biomarkers, the biomarker categories, the biomarker identification methods, and the condition indications. The processor is configured to display lines on the graph representative of the connections.
[0008] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Other aspects and advantages of the current invention will be apparent from the following detailed description of the embodiments and the accompanying drawing figures.BRIEF DESCRIPTION OF DRAWINGS
[0009] Embodiments of the current invention are described in detail below with reference to the attached drawing figures, wherein:
[0010] FIG. 1 is a block diagram depicting selected components of an exemplary environment in which embodiments of the present invention may be implemented;
[0011] FIGS. 2A and 2B is a flowchart depicting exemplary steps of a method according to an embodiment of the present invention; and
[0012] FIG. 3 is an exemplary ontology depicted as a knowledge graph according to embodiments of the present invention.
[0013] The drawing figures do not limit the current invention to the specific embodiments disclosed and described herein. The drawings are not necessarily to scale, emphasis instead being placed upon clearly illustrating the principles of the invention.DETAILED DESCRIPTION OF THE INVENTION
[0014] The following detailed description of the technology references the accompanying drawings that illustrate specific embodiments in which the technology can be practiced. The embodiments are intended to describe aspects of the technology in sufficient detail to enable those skilled in the art to practice the technology. Other embodiments can be utilized and changes canbe made without departing from the scope of the current invention. The following detailed description is, therefore, not to be taken in a limiting sense. The scope of the current invention is defined only by the appended claims, along with the full scope of equivalents to which such claims are entitled.
[0015] Aging biomarkers are essential for understanding the biological mechanisms of aging and developing interventions to promote healthy aging. However, identifying reliable biomarkers for aging is challenging due to the complex, multifactorial nature of the aging process. A biomarker includes any feature that researchers measure to show normal biological processes, harmful processes, or the response of the body to various factors. These factors can include environmental exposures or medical treatments. Biomarkers are different from clinical outcome assessments (COAs), which measure how a patient feels, functions, or survives in daily life.
[0016] Several key categories of biomarkers exist, including diagnostic, monitoring, safety, susceptibility, prognostic, or the like. Diagnostic biomarkers help detect or confirm if a disease is present. They can also identify different types (subtypes) of a disease. Monitoring biomarkers are used to make repeated measurements over time so that researchers can see how a disease changes or how well a treatment is working. Safety biomarkers show whether someone is likely to experience toxic or harmful effects from a treatment or if they are already experiencing toxicity. Susceptibility or risk biomarkers tell us if a person is likely to develop a disease in the future. Prognostic biomarkers help predict how a disease might progress in people who already have a condition, including the chances of future clinical problems or events. It is important to note that a biomarker can belong to more than one category, but evidence must be collected to prove its usefulness in each category. It is foreseeable that embodiments of the invention may include other categories of biomarkers without departing from the scope of the present invention.
[0017] Knowledge graphs combined with Artificial Intelligence (Al) address these challenges by integrating and analyzing vast, heterogeneous datasets, including genomics, transcriptomics, proteomics, and epigenomics, along with clinical and phenotypic data. In one or more embodiments, a KNowledge Acquisition and Representation Methodology (KNARM or KARMA) and Ontology Learning with Integrated Vector Embeddings (OLIVE) workflow is used to generate an ontology for formal descriptions of aging biomarkers and the relationships among them. The biomarker ontology may comprise comprising a plurality of biomarkers, a plurality of properties, and a plurality of connections between the plurality of biomarkers and / or the properties.The biomarker properties may include a biomarker category, a biomarker identification method, a disease indication, a treatment method, a condition, or the like. Furthermore, Large Language Models (LLMs) and classical machine learning approaches such as clustering, linear and nonlinear regression in addition statistical approaches may be used to help generate the knowledge graph of aging-biomarkers. In one or more embodiments, these rely on or draw from existing literature in addition to manually- and / or automatically-curated structured data.
[0018] Various examples of the present disclosure relate to systems and methods for generating and implementing knowledge graphs to correlate and make predictions based on aging biomarkers. Relationships between actions and entities or agents can be captured and modeled by computers using vector embeddings. The vector embeddings can, in turn, be in the format of graph embeddings in a semantic web (i.e., knowledge graph or the like). For example, nodes in such a knowledge graph may represent entities or agents (or concepts), and edges may represent relationships between the nodes and / or actions of the entities or agents represented by the nodes.
[0019] The knowledge graphs and, more particularly, the vector embeddings in the knowledge graphs, are machine-readable and manipulable representations of how the entities and agents behave and / or are constituted. The knowledge graphs permit machine-learning and other types of programs to reason, for example via deductive reasoning around the behaviors and traits embodied in the graph. In one or more embodiments, such reasoning or other analyses of the knowledge graphs uncovers correlations or relationships between the nodes and / or edges of the graph, and / or permits predictions for the nodes and / or edges and / or for external factors or issues (such as where a biomarker knowledge graph may be used to predict mortality risks). The machinelearning algorithms and techniques may also infer relationships between nodes and / or edges and, correspondingly, revise the knowledge.
[0020] The outputs, conclusions and other observations (e g., inferences or deductions) which are generated by a machine learning algorithm or program have variable quality, depending at least in part on the accuracy with which the entities and agents and their respective relationships and actions are modeled via the vector embeddings of the knowledge graph, and on the “reasoning” modalities implemented by the program which interpret and deductively reason from those embeddings.
[0021] In various examples, the present technology seeks to improve identification of important biomarkers of aging. In one or more embodiments, traditionally weighted clocks are notrequired or used for such identification. The biomarkers may be referred to as "lynchpin biomarkers," and may be pinpointed through mapping and clustering of interconnected biomarkers in a knowledge graph. Associations among or relationships between biomarkers may be derived from age-related shifts using a combination of extensive biobank databases (a data-driven approach), Al techniques, and classical statistical methods, all supported by a graph database and knowledge graphs developed through KNARM / KARMA and OLIVE methodologies. Examples of KNARM / KARMA methodologies, and other methodologies which may be used in connection with examples of the present disclosure.
[0022] In one or more embodiments, LLMs and classical machine learning approaches such as clustering, linear and non-linear regression in addition statistical approaches are utilized to generate a knowledge graph of aging-biomarkers. The knowledge graph or biomarker connectome may accordingly be built using unique statistical methods. Lynchpin biomarkers may be identified using a combination of machine learning, Al, and statistical approaches for filtering and analysis of the data. The workflow permits graph database / knowledge graph generation and reuse of the knowledge graph via LLMs to explore biomarkers.
[0023] Analysis applications that use the knowledge graph(s) and / or workflows may include machine learning, Al, statistical and / or LLM applications further enhance the knowledge graph(s) and biomarker identification as well as physician communication between the analysis and data. The biomarker(s) may be dynamically visualized based on other biomarker data using the graph database / knowledge graph as well as an LLM interface.
[0024] For example, the resultant backend framework may support a physician-oriented tool, such as a software tool, client application, mobile application or the like, enabling a medical professional to dynamically view a patient’s biomarker history and observe changes over time. Additionally, the tool may offer Al-driven predictions on biomarker evolution based on parameters specified by the physician. In one or more embodiments, the tool may access or embody a knowledge graph accessible by and / or translated through an LLM.
[0025] FIG. 1 illustrates an example computing device 10 for generating and implementing knowledge graphs to correlate and make predictions based on aging biomarkers. Generally, the computing device 10 may include a tablet computer, laptop computer, desktop computer, workstation computer, smart phone, smart watch, and the like. In addition or alternatively, the computing device 10 may include cloud servers, domain controllers, application servers, databaseservers, database web servers, file servers, mail servers, catalog servers or the like, or combinations thereof.
[0026] The computing device 10 may include circuitry capable of wired and / or wireless communication with other computing devices and other electronic devices (such as a communication element 12), a memory element 14, a processing element 16, a software program 18 and a user interface 20.
[0027] The user interface 20 may include video devices of any of the following types: plasma, standard or ultra-high-definition light-emitting diode (LED), organic LED (OLED), quantum dot LED (QLED), Light Emitting Polymer (LEP) or Polymer LED (PLED), liquid crystal display (LCD), thin fdm transistor (TFT) LCD, LED side-lit or back-lit LCD, or the like, or combinations thereof. The user interface 20 may possess a square or a rectangular aspect ratio and may be viewed in either a landscape or a portrait mode. In various embodiments, the user interface 20 may also include a touch screen occupying all or part of the screen. The user interface 20 may provide visualizations of the semantic web(s), knowledge graph(s) and / or LLM interface(s), including for prompting and receiving data visualizations of portions of and / or predictions or reasoning from the knowledge graph(s), as well as receiving input from a user for generating the knowledge graph(s). More particularly, the user interface 20 may participate in receiving user input for constructing components of such web(s) or graph(s), output(s), and / or otherwise for manipulating data or providing instructions within the scope of the present invention.
[0028] Further, the software application or program 18 may be configured with instructions for performing and / or enabling performance of at least some of the steps set forth herein. In various examples, the software program 18 comprises instructions stored on computer- readable media of the memory element 14. The software application 18 may correspond to the operations described herein. In particular, the software application 18 may perform operations described in more detail below for generating and implementing knowledge graphs to correlate and make predictions based on aging biomarkers, including via a mobile app configured to display representations of the knowledge graph(s) to, to reason over such knowledge graph(s) responsive to input from, and / or to make predictions from such knowledge graph(s) responsive to input from, medical professionals.
[0029] The memory element 14 may include electronic hardware data storage components such as read-only memory (ROM), programmable ROM, erasable programmable ROM, random-access memory (RAM) such as static RAM (SRAM) or dynamic RAM (DRAM), cache memory, hard disks, floppy disks, optical disks, flash memory, thumb drives, universal serial bus (USB) drives, or the like, or combinations thereof. In some embodiments, the memory element 14 may be embedded in, or packaged in the same package as, the processing element 16. The memory element 14 may include, or may constitute, a “computer-readable medium.” The memory element 14 may store the instructions, code, code segments, software, firmware, programs, applications, apps, services, daemons, or the like that are executed by the processing element 16. In an embodiment, the memory element 14 stores the software applications / program 18. The memory element 14 may also store settings, data, documents, sound files, photographs, movies, images, databases, and the like.
[0030] The processing element 16 may include electronic hardware components such as processors. The processing element 16 may include digital processing unit(s). The processing element 16 may include one or more microprocessors (single-core and multi-core), microcontrollers, digital signal processors (DSPs), field-programmable gate arrays (FPGAs), analog, or digital application-specific integrated circuits (ASICs), or the like, or combinations thereof, without limitation. The processing element 16 may generally execute, process, or run instructions, code, code segments, software, firmware, programs, applications, apps, processes, services, daemons, or the like. For instance, the processing element 16 may execute the software applications / program 18. The processing element 16 may also include hardware components such as finite-state machines, sequential and combinational logic, and other electronic circuits that can perform the functions necessary for the operation of embodiments of the current invention. The processing element 16 may be in communication with the other electronic components through serial or parallel links that include universal busses, address busses, data busses and the like.
[0031] Through hardware, software, firmware, or various combinations thereof, the processing element 16 may - alone or in combination with other processing elements - be configured to perform the operations of embodiments of the present invention. Specific embodiments of the technology will now be described in connection with the attached drawing figures. The embodiments are intended to describe aspects of the invention in sufficient detail to enable those skilled in the art to practice the invention. Other embodiments can be utilized and changes can be made without departing from the scope of the present invention. The system may include additional, less, or alternate functionality and / or device(s), including those discussedelsewhere herein. The following detailed description is, therefore, not to be taken in a limiting sense. The scope of the present invention is defined only by the appended claims, along with the full scope of equivalents to which such claims are entitled, unless otherwise expressly stated and / or readily apparent to those skilled in the art from the description.
[0032] As introduced above, embodiments of the present disclosure utilize a computing device (e.g., computing device 10) to generate and / or update a knowledge graph representative of the biomarker ontology stored on the memory element 14. Embodiments of the present disclosure enable one or more users to utilize the computing device 10 to use the knowledge graph to derive relationships among aging biomarkers and / or make predictions regarding health-related outcomes or events. In one or more embodiments, the knowledge graph is utilized by medical professionals, including to perform the tasks described in more detail above based on the knowledge graph(s), via such a computing device (e g., computing device 10) and one or more LLMs interfacing between the medical professionals and the knowledge graph(s) 300 (depicted in FIG. 3).
[0033] Generating the knowledge graph may include retrieving biomarker datasets from one or more databases 24a, 24n over the communication network 22. The biomarker datasets include biomarker values for patients of various ages and both sexes along with patient identifications, which may be referred to as structured data. The biomarker datasets may be analyzed according to a plurality of operations. One or more data processing operations may include using one or more libraries, e.g., those offered under the trademarks PYTHON (a registered trademark of Python Software Foundation) and / or PANDAS™ (a trademark of or licensed to NumFOCUS, Inc.). Such library(ies) may be used to process the biomarker datasets to calculate key statistical metrics for each biomarker across different age groups. The statistical metrics across age groups may reveal relationships regarding how biomarker levels shift with age. In one or more embodiments, one or more of the databases 24a, 24n comprise unstructured data, such as literature (e.g., academic and / or research papers), related biomarkers and their connections to certain properties, such as categories, conditions, invasiveness levels, treatments, etc. and connections to other biomarkers.
[0034] The analysis operations may include biomarker data processing and filtering. The analysis operations may include using natural language processing to parse the unstructured data to glean biomarker connections and properties. A filtering strategy may be implemented based at least in part on isolating and analyzing specific biomarkers using other biomarkers’ value changesover time as one or more filter(s). For example, patients associated with the biomarkers may be grouped based on body mass index (BMI) ranges, and corresponding biomarker levels may be compared across the BMI-defined cohorts. Such biomarker filtering may reveal potential associations between one biomarker and other biomarkers, helping to clarify whether the two biomarkers significantly influence the levels of changing biomarkers and how statistically significant each change is. Change patterns of biomarkers over time may also be identified (i.e., biomarker increase or decrease patterns over time). Such time-based change patterns may be evaluated after a certain point in life, and may be evaluated using optimization algorithms and concepts such as local minima, local maxima, global minima-maxima values, or the like, in each case as applied to a certain biomarker’s graph using the filtering approach.
[0035] The analysis operations may include significance testing. In one or more embodiments, a large number of biomarkers and possible comparisons renders manual visualization of each graph (to detect relationships) inefficient or impossible. An automated approach to detect significant changes in biomarker values may consequently be applied by assessing statistical significance to determine whether changes in one biomarker are linked or correlated to substantial variations in others. This operation may efficiently identify meaningful biomarker interactions. It should also be noted that manual review may be implemented to revise such automatically-derived relationships or correlations, e.g., to remove false positive correlations. Also or alternatively, one or more machine learning methods may be implemented to rapidly explore the datasets for new potential connections or correlations among biomarkers. For example, a Correlation Matrix Calculation and / or Correlation Threshold Filtering may be used to identify simple correlations and understand high-correlation pairs using varying cut-off thresholds.
[0036] Existing solutions for determining statistically significant differences between values include the t-test and the Mann-Whitney U-test. However, both of these methods yield unexpectedly poor results due to gaps in some of the structured data. In one or more embodiments, the analysis operations include determining medians and interquartile ranges of a first biometric value for two or more groups of biomarker files filtered according to ranges of a second biometric value. The analysis may include determining that the medians of one or both of the groups fall within the interquartile range of the other group, which is indicative of a lack of connection between the biomarkers. However, if the interquartile ranges of the two groups do not overlap,then the computing device 10 is configured to indicate and store this as a potential relationship between the two biomarkers.
[0037] The analysis operations may include ML and Al Methods Integration. More particularly, in one or more embodiments diverse analytical methods may be integrated, including survival analysis, neural networks, and other ML techniques. This multi-modal approach allows for a more holistic understanding of mortality risk by combining insights from time-to-event data (via Cox Proportional Hazards modeling) with complex feature interactions discovered through machine learning. By unifying these methods, the analysis can simultaneously consider the temporal aspects of mortality and the intricate relationships among biomarkers, demographic factors, and other health indicators.
[0038] The analysis operations may include knowledge graph generation. Relying at least in part on relationships and correlations between biomarkers and / or associated data revealed by operations described in more detail above, the KNARM / KARMA and OLIVE methodologies may be used (e.g., in conjunction with a database backend, such as that offered under the mark NEO4J (a registered trademark of NEO4I, Inc.)) to embody the enriched and related biomarker data in a knowledge graph. Such database backend(s) and methodologies may permit visualization and structural representation of the relationships as a knowledge graph, thereby enhancing interpretation of biomarker interactions. By integrating a database backend, a semi -automated ontology-building approach may be applied to further refine the creation of the knowledge graph.
[0039] In one or more embodiments, the LLM-powered OLIVE workflow may also be used to enhance and automate the process of knowledge graph creation. This workflow permits integration of knowledge from existing research, e.g., by integrating correlations and relationships derived from written portions of literature (e.g., abstract sections of published papers) into the knowledge graph. More particularly, OLIVE may use or prompt an LLM to search scientific literature for abstracts related to specific biomarker keywords. Based on a biomarkers vocabulary derived from the biomarkers present in initial biomarker datasets, such identified literature may be analyzed to yield additional correlations and relationships between biomarkers and associated data for automated integration into the knowledge graph. The LLM may analyze the identified publications (e.g., abstracts) to extract relevant information (e.g., relationships or correlations between biomarkers) and automatically convert it into a knowledge graph or revisions to the knowledge graph based on the extraction, and the generated knowledge graph can be visualizedusing a user interface. Such a user interface may be enhanced and included in a mobile app or the like for use by medical professionals (discussed in more detail above).
[0040] Accordingly, embodiments of the present invention generate a knowledge graph that captures relationships between various biomarkers, including based on information extracted from literature and curated according to analysis operations discussed in more detail above, and utilize the knowledge graph and associated algorithms to implement a tool for use by medical professionals. An exemplary knowledge graph 300 representative of a biomarker ontology is depicted in FIG. 3. The knowledge graph 300 includes a number of nodes, including biomarkers, categories, invasiveness levels, identification methods, sources of samples for identification, and arrows indicating connections between the various nodes. The knowledge graph 300 may include any number of nodes and node types, such as sub-categories or sub-classes, without departing from the scope of the present invention.
[0041] Turning to FIGS. 2A and 2B, the flow chart depicted therein includes the steps of an exemplary method 200 of processing biomarker data. In some alternative implementations, the functions noted in the various blocks may occur out of the order depicted in FIGS. 2A and 2B. For example, two blocks shown in succession in FIGS. 2A and 2B may in fact be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order depending upon the functionality involved. In addition, some steps may be optional.
[0042] The method 200 is described below, for ease of reference, as being executed by exemplary devices and components introduced with the embodiments illustrated in FIG. 1. The steps of the method 200 may be performed by the control system through the utilization of processors, transceivers, hardware, software, firmware, or combinations thereof. However, some of such actions may be distributed differently among such devices or other devices without departing from the spirit of the present invention. Control of the system may also be partially implemented with computer programs stored on one or more non-transient computer-readable medium(s). The computer-readable medium(s) may include one or more executable programs stored thereon, wherein the program(s) instruct one or more processing elements to perform all or certain of the steps outlined herein. The program(s) stored on the computer-readable medium(s) may instruct processing element(s) to perform additional, fewer, or alternative actions, including those discussed elsewhere herein.
[0043] Referring to step 202, biomarker data is collected, via the computing device, from one or more databases. In one or more embodiments, the computing device receives from one or more databases structured data, such as a plurality of biomarker files including patient identification, ages, and biomarker values. In one or more embodiments, this step includes obtaining, at the one or more computing devices, unstructured data comprising published connections between two or more biomarkers from one or more databases.
[0044] Referring to step 204, the biomarker data is integrated via the computing device. The computing device may be configured to merge the biomarker data according to a property, such as by patient identification, so that all data associated with any particular patient identification is stored in association with that that patient identification. This step may also include parsing, at the one or more computing devices, the unstructured data using one or more natural language processing algorithms to generate validation data. This step may also include using or prompting, via the computing device, an LLM to search scientific literature for abstracts or other excerpts related to specific biomarker keywords. Based on a biomarkers vocabulary derived from the biomarkers present in initial biomarker datasets, such identified literature may be analyzed to yield additional correlations and relationships between biomarkers and associated data for automated integration into the biomarker database stored on the memory element. The computing device may be configured to use the LLM to analyze the identified publications (e.g., abstracts) to extract relevant information (e.g., relationships or correlations between biomarkers).
[0045] Referring to step 206, selections of two or more biomarkers for comparison may be received at the computing device. This step may include receiving, at the computing device, two or more ranges of values of the first selected biomarker to generate two or more groups of the plurality of biomarker files. The selections may be received from the user interface and may be in response to use inputs provided through the user interface.
[0046] Referring to step 208, the two groups are displayed, on the user interface, for visual comparison. The groups may be displayed side by side as scatter plots with the selected biomarker value being shown as a function of another variable, such as age or BMI.
[0047] Referring to step 210, statistical analysis is performed on the two groups to identify statistically significant differences between their biomarker values. Existing solutions for determining statistically significant differences between values include the t-test and the Mann- Whitney U-test. However, both of these methods yielded unexpectedly poor results. Accordingly,embodiments of the invention include determining interquartile ranges (IQR) of biomarkers and median values. When comparing values of the first biomarker of two groups filtered according to the second biomarker, if their IQRs do not overlap, then embodiments of the invention indicate that there is a significant difference between the values of the first biomarker between the two groups. This indicates a relationship between the first and second biomarkers. Accordingly this step may include determining, at the one or more computing devices, a first interquartile range of the second selected biomarker of the two or more selections of a first group of the two or more groups. This step may further include determining, at the one or more computing devices, a second interquartile range of the second selected biomarker of the two or more selections of a second group of the two or more groups. This step may further include determining, at the one or more computing devices, the first interquartile range is outside the second interquartile range. This step may include recording, at the one or more computing devices, an indication of a relationship between the two or more selected biomarkers based on the detected statistically significant difference.
[0048] Referring to step 212, in one or more embodiments, automated statistical analysis is performed, via the computing device, on a plurality of the biomarkers for detecting significant differences. This step may include performing the statistical analysis discussed in step 210 for several combinations of the biomarkers and / or for each combination of the biomarkers. However, the computing device may perform statistical analysis on any number of combinations of the biomarkers without departing from the scope of the present invention. This step may include recording, at the one or more computing devices, indications of relationships between the biomarkers based on the statistical analysis. This step may also include updating the ontology comprising the biomarkers, associated properties, and connections between the biomarkers and / or other properties. This step may also include comparing the plurality of connections between biomarkers derived by the statistical analysis with the validation data.
[0049] Referring to step 214, one or more biomarkers with groups demonstrating statistically significant differences are displayed on the user interface. The biomarkers may be displayed as a knowledge graph (similar to the knowledge graph 300 of FIG. 3) with the nodes comprising biomarkers and lines depicting the connections inferred from the data analysis. A user may be able to cycle through the various biomarkers and view the relationships based on statistically significant differences.
[0050] Referring to step 216, an adjacency matrix is generated, at the computing device, based on the relationships between biomarkers inferred from the statistically significant differences. The matrix may include biomarkers on the heading row and first column with values in the matrix representative of connections between two biomarkers (one in the row and one in the column).
[0051] Turning to FIG. 2B, and referring to step 218, the adjacency matrix is reduced, via the computing device, using a singular value decomposition algorithm. The adjacency matrix of a large biomarker graph can be very high-dimensional, meaning it has a lot of information but is difficult to process efficiently. To make this data more manageable, embodiments of the invention use the dimensionality reduction technique of singular value decomposition. This technique reduces the number of columns and rows but retains the essential relationships between biomarkers. This enables generation of graphs to visualize biomarker relationships in a 2D or 3D space, making it easier to detect patterns. Accordingly, this step may also include displaying the biomarkers in relation to their interconnections in a two-dimensional or three-dimensional scatterplot graph.
[0052] Referring to step 220, a clustering algorithm is performed to detect additional relationships between the biomarkers. The clustering algorithm may be a k-means clustering algorithm or a density-based spatial clustering of applications with noise clustering algorithm. This step may include updating, via the computing device, the ontology and / or the knowledge graph (depicted in FIG. 3) with the additional relationships gleaned from the clustering algorithm.
[0053] The method 200 may include additional, less, or alternate steps and / or device(s), including those discussed elsewhere herein. For example, in one or more embodiments, the method may include depicting the ontology as the knowledge graph having the plurality of nodes representing biomarker properties, including biomarker categories, biomarker identification methods, disease indications, treatment methods, conditions, and / or the like. Additionally, each property or node may include any number of subclasses without departing from the scope of the present invention.
[0054] Throughout this specification, references to “one embodiment”, “an embodiment”, or “embodiments” mean that the feature or features being referred to are included in at least one embodiment of the technology. Separate references to “one embodiment”, “an embodiment”, or “embodiments” in this description do not necessarily refer to the same embodiment and are alsonot mutually exclusive unless so stated and / or except as will be readily apparent to those skilled in the art from the description. For example, a feature, structure, act, etc. described in one embodiment may also be included in other embodiments, but is not necessarily included. Thus, the current invention can include a variety of combinations and / or integrations of the embodiments described herein.
[0055] Although the present application sets forth a detailed description of numerous different embodiments, it should be understood that the legal scope of the description is defined by the words of the claims set forth at the end of this patent and equivalents. The detailed description is to be construed as exemplary only and does not describe every possible embodiment since describing every possible embodiment would be impractical. Numerous alternative embodiments may be implemented, using either current technology or technology developed after the filing date of this patent, which would still fall within the scope of the claims.
[0056] As used herein, the phrase “and / or,” when used in a list of two or more items, means that any one of the listed items can be employed by itself or any combination of two or more of the listed items can be employed. For example, if a composition is described as containing or excluding components A, B, and / or C, the composition can contain or exclude A alone; B alone; C alone; A and B in combination; A and C in combination; B and C in combination; or A, B, and C in combination.
[0057] The present description also uses numerical ranges to quantify certain parameters relating to various embodiments of the invention. It should be understood that when numerical ranges are provided, such ranges are to be construed as providing literal support for claim limitations that only recite the lower value of the range as well as claim limitations that only recite the upper value of the range. For example, a disclosed numerical range of about 10 to about 100 provides literal support for a claim reciting “greater than or equal to about 10” (with no upper bounds) and a claim reciting “less than or equal to about 100” (with no lower bounds).
[0058] Furthermore, unless otherwise specified, any directional references (e.g., upper, lower, above, below, etc.) are used herein solely for the sake of convenience and should be understood only in relation to each other. For instance, a component might in practice be oriented such that faces referred to as “upper” and “lower” are sideways, angled, inverted, etc. relative to the chosen frame of reference.
[0059] Throughout this specification, plural instances may implement components, operations, or structures described as a single instance. Although individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed concurrently, and nothing requires that the operations be performed in the order illustrated. Structures and functionality presented as separate components in example configurations may be implemented as a combined structure or component. Similarly, structures and functionality presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter herein.
[0060] Certain embodiments are described herein as including logic or a number of routines, subroutines, applications, or instructions. These may constitute either software (e.g., code embodied on a machine-readable medium or in a transmission signal) or hardware. In hardware, the routines, etc., are tangible units capable of performing certain operations and may be configured or arranged in a certain manner. In example embodiments, one or more computer systems (e.g., a standalone, client or server computer system) or one or more hardware modules of a computer system (e.g., a processor or a group of processors) may be configured by software (e.g., an application or application portion) as computer hardware that operates to perform certain operations as described herein.
[0061] In various embodiments, computer hardware, such as a processing element, may be implemented as special purpose or as general purpose. For example, the processing element may comprise dedicated circuitry or logic that is permanently configured, such as an applicationspecific integrated circuit (ASIC), or indefinitely configured, such as an FPGA, to perform certain operations. The processing element may also comprise programmable logic or circuitry (e.g., as encompassed within a general-purpose processor or other programmable processor) that is temporarily configured by software to perform certain operations. It will be appreciated that the decision to implement the processing element as special purpose, in dedicated and permanently configured circuitry, or as general purpose (e.g., configured by software) may be driven by cost and time considerations.
[0062] Accordingly, the term “processing element” or equivalents should be understood to encompass a tangible entity, be that an entity that is physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a certainmanner or to perform certain operations described herein. Considering embodiments in which the processing element is temporarily configured (e g., programmed), each of the processing elements need not be configured or instantiated at any one instance in time. For example, where the processing element comprises a general-purpose processor configured using software, the general- purpose processor may be configured as respective different processing elements at different times. Software may accordingly configure the processing element to constitute a particular hardware configuration at one instance of time and to constitute a different hardware configuration at a different instance of time.
[0063] The processing element may include processors, microprocessors (single-core and multi-core), microcontrollers, DSPs, field-programmable gate arrays (FPGAs), analog and / or digital application-specific integrated circuits (ASICs), or the like, or combinations thereof. The processing element may generally execute, process, or run instructions, code, code segments, software, firmware, programs, applications, apps, processes, services, daemons, or the like. The processing element may also include hardware components such as finite-state machines, sequential and combinational logic, and other electronic circuits that can perform the functions necessary for the operation of the current invention. The processing element may be in communication with the other electronic components through serial or parallel links that include address busses, data busses, control lines, and the like.
[0064] Computer hardware components, such as communication elements, memory elements, processing elements, and the like, may provide information to, and receive information from, other computer hardware components. Accordingly, the described computer hardware components may be regarded as being communicatively coupled. Where multiple of such computer hardware components exist contemporaneously, communications may be achieved through signal transmission (e.g., over appropriate circuits and buses) that connect the computer hardware components. In embodiments in which multiple computer hardware components are configured or instantiated at different times, communications between such computer hardware components may be achieved, for example, through the storage and retrieval of information in memory structures to which the multiple computer hardware components have access. For example, one computer hardware component may perform an operation and store the output of that operation in a memory device to which it is communicatively coupled. A further computer hardware component may then, at a later time, access the memory device to retrieve and processthe stored output. Computer hardware components may also initiate communications with input or output devices, and may operate on a resource (e.g., a collection of information).
[0065] The memory device or element may include data storage components, such as readonly memory (ROM), programmable ROM, erasable programmable ROM, random-access memory (RAM) such as static RAM (SRAM) or dynamic RAM (DRAM), cache memory, hard disks, floppy disks, optical disks, flash memory, thumb drives, universal serial bus (USB) drives, or the like, or combinations thereof. In some embodiments, the memory element may be embedded in, or packaged in the same package as, the processing element. The memory element may include, or may constitute, a “computer-readable medium”. The memory element may store the instructions, code, code segments, software, firmware, programs, applications, apps, services, daemons, or the like that are executed by the processing element.
[0066] The communication element may generally allow communication with systems and / or external devices. The communication element may include signal or data transmitting and receiving circuits, such as antennas, amplifiers, filters, mixers, oscillators, digital signal processors (DSPs), and the like. The communication element may establish communication wirelessly by utilizing RF signals and / or data that comply with communication standards such as cellular 2G, 3G, 4G, 5G, or LTE, WiFi, WiMAX, Bluetooth®, BLE, or combinations thereof. The communication element may be in communication with the processing element and the memory element.
[0067] The user interface generally allows the user to utilize inputs and outputs to interact with the device and is in communication with the one or more processing element. Inputs may include buttons, pushbuttons, knobs, jog dials, shuttle dials, directional pads, multidirectional buttons, switches, keypads, keyboards, mice, joysticks, microphones, or the like, or combinations thereof. The outputs of the present invention may include a display and / or any number of additional outputs, such as audio speakers, lights, dials, meters, printers, or the like, or combinations thereof, without departing from the scope of the present invention.
[0068] The various operations of example methods described herein may be performed, at least partially, by one or more processing elements that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processing elements may constitute processing element- implemented modules that operate to perform one or more operations or functions. The modulesreferred to herein may, in some example embodiments, comprise processing element-implemented modules.
[0069] Similarly, the methods or routines described herein may be at least partially processing element-implemented. For example, at least some of the operations of a method may be performed by one or more processing elements or processing element-implemented hardware modules. The performance of certain of the operations may be distributed among the one or more processing elements, not only residing within a single machine, but deployed across a number of machines. In some example embodiments, the processing elements may be located in a single location (e.g., within a home environment, an office environment or as a server farm), while in other embodiments the processing elements may be distributed across a number of locations.
[0070] Unless specifically stated otherwise, discussions herein using words such as “processing,” “computing,” “calculating,” “determining,” “presenting,” “displaying,” or the like may refer to actions or processes of a machine (e.g., a computer with a processing element and other computer hardware components) that manipulates or transforms data represented as physical (e.g., electronic, magnetic, or optical) quantities within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.
[0071] As used herein, the terms “comprises,” “comprising,” “includes,” “including,” “has,” “having” or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0072] The patent claims at the end of this patent application are not intended to be construed under 35 U.S.C. § 112(f) unless traditional means-plus-function language is expressly recited, such as “means for” or “step for” language being explicitly recited in the claim(s).
[0073] Although the technology has been described with reference to the embodiments illustrated in the attached drawing figures, it is noted that equivalents may be employed and substitutions made herein without departing from the scope of the technology as recited in the claims.
[0074] Having thus described various embodiments of the technology, what is claimed as new and desired to be protected by Letters Patent includes the following:
Claims
CLAIMS1. A method comprising:(a) retrieving, at one or more computing devices, from one or more databases a plurality of biomarker files, each of the plurality of biomarker files including a patient identification, an age, and one or more biomarker values;(b) receiving, at the one or more computing devices, two or more selections of biomarkers for comparison;(c) receiving, at the one or more computing devices, two or more ranges of values of a first selected biomarker of the two or more selections of the biomarkers to generate two or more groups of the plurality of biomarker files;(d) determining, at the one or more computing devices, a first interquartile range of a second selected biomarker of the two or more selections of a first group of the two or more groups;(e) determining, at the one or more computing devices, a second interquartile range of the second selected biomarker of the two or more selections of a second group of the two or more groups;(f) determining, at the one or more computing devices, the first interquartile range is outside the second interquartile range; and(g) recording, at the one or more computing devices, an indication of a relationship between the two or more selections of biomarkers.
2. The method of claim 1, further comprising displaying, at the one or more computing devices, scatter plots comprising the age and values of the second selected biomarker of the two or more groups.
3. The method of claim 1, further comprising determining, at the one or more computer devices, median values of the second selected biomarker for the two or more groups.
4. The method of claim 1, further comprising repeating steps (b) through (g) for a plurality of additional biomarker selections to generate a plurality of additional indications of relationships between the plurality of additional biomarker selections.
5. The method of claim 4, further comprising storing, on the one or more computing devices, an ontology comprising a plurality of biomarkers, a plurality of properties, and a plurality of connections between the plurality of biomarkers.
6. The method of claim 5, further comprising updating, at the one or more computing devices, the ontology by creating and storing, in the ontology, the indication of the relationship and the additional indications of relationships as the plurality of connections between the plurality of biomarkers.
7. The method of claim 6, further comprising: generating, at the one or more computing devices, an adjacency matrix for the plurality of biomarkers; reducing, at the one or more computing devices, the adjacency matrix via singular value decomposition; and dividing, at the one or more computing devices, the plurality of biomarkers into one or more clusters using one or more clustering algorithms.
8. The method of claim 7, wherein the one or more clustering algorithms includes at least one of a k-means clustering algorithm or a density-based spatial clustering of applications with noise clustering algorithm.
9. The method of claim 7, further comprising updating, at the one or more computing devices, the ontology by creating and storing, in the ontology, a plurality of connections between the plurality of biomarkers in the one or more clusters.
10. The method of claim 7, further comprising displaying, at the one or more computing devices, the singular value decomposition of the adjacency matrix in a graph.11 . The method of claim 6, further comprising: obtaining, at the one or more computing devices, unstructured data comprising published connections between two or more biomarkers from the one or more databases; parsing, at the one or more computing devices, the unstructured data using one or more natural language processing algorithms to generate validation data; and comparing the plurality of connections between the plurality of biomarkers based on the validation data.
12. The method of claim 5, wherein the plurality of properties include a biomarker category, a biomarker identification method, a disease indication, a treatment method, and a condition.
13. A system comprising: one or more databases storing an ontology comprising a plurality of biomarkers, a plurality of properties, and a plurality of connections between the plurality of biomarkers; one or more processors in communication with the one or more databases and configured to: retrieve from the one or more databases a plurality of biomarker files, each of the plurality of biomarker files including a patient identification, an age, and one or more biomarker values; receive a first biomarker selection and a second biomarker selection for comparison; receive two ranges for values of the first biomarker selection to generate a first group and a second group of the plurality of biomarker files; determine a first interquartile range of the second biomarker selection of the first group; determine a second interquartile range of the second biomarker selection of the second group; determine the first interquartile range is outside the second interquartile range; and update the ontology by creating and storing, in the ontology, a connection between the first and second biomarker selections.
14. The system of claim 13, wherein the one or more processors are configured to: generate an adjacency matrix for the plurality of biomarkers; reduce the adjacency matrix via singular value decomposition; divide the plurality of biomarkers into one or more clusters using one or more clustering algorithms; and update the ontology by creating and storing, in the ontology, a plurality of connections between the plurality of biomarkers in the one or more clusters.1 . The system of claim 14, wherein the one or more processors are configured to: obtain unstructured data comprising published connections between one or more biomarkers from the one or more databases; parse the unstructured data using one or more natural language processing algorithms to generate validation data; and compare the plurality of connections between the plurality of biomarkers based on the validation data.
16. The system of claim 13, wherein the one or more processors are configured to display on a user interface a graph comprising a plurality of nodes representative of the plurality of biomarkers and lines connected to the plurality of nodes representative of the plurality of connections between the plurality of biomarkers.
17. The system of claim 16, wherein the plurality of properties include a biomarker category, a biomarker identification method, a disease indication, a treatment method, and a condition.
18. The system of claim 17, wherein the plurality of nodes include nodes representative of the biomarker category, the biomarker identification method, the disease indication, the treatment method, and the condition.
19. A system comprising: one or more databases storing an ontology comprising biomarkers, biomarker categories, biomarker identification methods, condition indications, and connections between the biomarkers, the biomarker categories, the biomarker identification methods, and the condition indications, wherein: the connections between two biomarkers are representative of statistically significant correlations derived from structured data, and the connections between the biomarkers and the condition indications are representative of published causal data derived from unstructured data; a user interface; and one or more processors in communication with the one or more databases and the user interface and configured to display on the user interface the ontology depicted as a graph comprising nodes representative of the biomarkers, the biomarker categories, the biomarker identification methods, and the condition indications and lines representative of the connections.
20. The system of claim 19, wherein the one or more processors is configured to receive a search inquiry and execute a depth-limited search algorithm based at least in part on the search inquiry.
Citation Information
Patent Citations
System and method for biomarker-outcome prediction and medical literature exploration
US11257594B1
System and Method for the Visualization of Medical Data
US20140022255A1
Prostate cancer biomarkers
US20230003731A1
Methylation biomarker selection apparatuses and methods
WO2023052917A1
System and method for non-invasive quantification of blood biomarkers
WO2024052898A1