Molecular community networking for metabolomics
The system addresses the challenges in molecular network analysis by identifying densely connected molecular communities and accurately annotating molecules, enhancing understanding for diagnostics and therapeutics.
Patent Information
- Application Number
- PCT/US2025/023517
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-05
- Filing Date
- 2025-04-07
- Publication Date
- 2025-10-09
AI Technical Summary
Current methods for understanding molecular networks, particularly in microbiomes, face challenges such as ambiguity in molecular origin, low resolution in mass spectrometry, incomplete data annotation, and reliance on indirect correlations, leading to incomplete interaction data and missed transient interactions.
A system and method for generating an unpruned molecular network, applying community detection algorithms to identify densely connected molecular communities, and using reference spectra or molecular structure prediction tools to accurately annotate molecules within these communities.
Enables comprehensive analysis of molecular networks, revealing dynamic biochemical interactions and providing context-specific insights for medical diagnostics and therapeutics.
Smart Images

Figure US2025023517_09102025_PF_FP_ABST
Abstract
Description
MOLECULAR COMMUNITY NETWORKING FOR METABOLOMICSCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63 / 574,916, filed on April 05, 2024, and the entire contents of which are incorporated herein by reference for all purposes.BACKGROUND OF THE INVENTION1. Field of the Invention
[0002] The invention disclosed herein relates to identification of molecules, and in particular to techniques for discovery, characterization, and quantification of molecules associated with biomolecular structures and biochemical processes. In particular, the disclosed invention facilitates an understanding of molecular connectivity, revealing previously hidden relationships, organizing complex molecular datasets, and enabling the discovery of novel metabolites.2. Description of the Related Art
[0003] Biochemical processes involve complex interactions among diverse molecules, including proteins, nucleic acids, lipids, and small molecules. These interactions form structured systems referred to as “molecular networks”. A molecular- network is a collection of molecules connected by specific chemical, physical, physicochemical, biochemical or biophysical relationships that contribute to a particular- cellular or physiological function. The nodes in the network represent a group or cluster of molecules or biomolecules, while the edges denote interactions such as binding events, metabolic conversions, regulatory influences, or signal transductions. The edges may also denote similarity metrics between nodes.
[0004] Molecular networks provide insights into the organization and regulation of cellular functions in both single organisms and complex ecosystems. In an individual organism, a molecular- network may define the pathways for signal transduction, gene regulation, or metabolic activity and a myriad of other processes. In contrast, molecular- networks associated with a microbiome represent interspecies molecular interactions, including symbiotic, antagonistic, or cooperative exchanges. It is desirable to develop better methodologies for characterizing andunderstanding of molecular networks (in particular, microbiome molecular networks), as these are complex and not well understood due to the challenges of simultaneously studying interactions between many different molecules and species in a molecular network.
[0005] Presently, the ability to understand interactions between different molecules and species is limited for a variety of reasons. These include ambiguity in molecular origin, where for example, in microbiome-rich samples, distinguishing host-derived molecules from microbial or dietary molecules remains challenging without organism- specific labels. Mass spectrometry and sequencing technologies sometimes suffer from low resolution and do not detect low-abundance molecules or distinguish between structural isomers.
[0006] Many inferred networks rely on indirect correlations, leading to false positives or omission of weak yet functionally relevant interactions. In other words, data that details interaction between certain molecules is often incomplete. Methods and systems used to gather this data often detail only static snapshots, capturing only a single time point and missing the dynamic context that is often desirable to observe transient or condition-dependent interactions. Sometimes biases in data annotation present a problem, as functional databases are often incomplete and skewed toward well-studied model organisms, leaving gaps in the understanding of non-model systems.
[0007] Improved methods for characterizing a molecular network are therefore needed so that information derived from molecular networks contribute to advancements in medical diagnostics and therapeutics by enabling a comprehensive analysis of biological states, disease mechanisms, and treatment responses. Well-characterized molecular networks may reveal the structure and dynamics of biochemical interactions within a cell or tissue, providing context- specific insights that single-molecule measurements cannot deliver.SUMMARY OF THE INVENTION
[0008] In one embodiment, disclosed herein is a system for detection of a molecular community in a molecular network, the system comprising an information handling system that comprises at least one processor; and a non-transitory storage medium storing instructions readable and executable by the processor to perform a method comprising receiving molecular data from an analytical device and / or a data repository; generating an unpruned molecular network from the molecular- data; treating the unpruned network with a community detection algorithm to form one or more molecular communities; where each molecular community comprises nodes and edges;where all network connections are included in the molecular community and the molecular community has a higher density of connections internally as compared with the density of connections outside the molecular community; parsing the molecular communities to prevent the formation of singletons; and identifying at least one molecule present in the molecular community by comparing the molecular community with reference spectra from a library of spectra or by using a molecular structure prediction tool; where the molecular structure prediction tool generates an annotation with an associated metric that indicates the accuracy of this annotation.
[0009] In another embodiment, disclosed herein too is a computer program product stored on non- transitory machine readable media, the computer program product comprising machine executable instructions configured for detecting a molecular community, comprising generating an unpruned molecular network from molecular data received from an analytical device or a data repository; treating the unpruned network with a community detection algorithm to form one or more molecular communities; where each molecular community comprises nodes and edges; where all network connections are included in the molecular community and the molecular community has a higher density of connections internally as compared with the density of connections outside the community; parsing the molecular communities to prevent the formation of singletons; and identifying at least one molecule present in the molecular community by comparing the molecular community with reference spectra from a library of spectra or by using a molecular structure prediction tool; where the molecular structure prediction tool generates a likely annotation with a metric that gives the accuracy of this annotation.BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The features and advantages of the invention are apparent from the following description taken in conjunction with the accompanying drawings in which:
[0011] FIG. 1A is a schematic depiction of an exemplary information handling system (IHS) that receives molecular data from an analytical device, processes the data and transports the output to a terminal, where it can be viewed;
[0012] FIG. IB is a schematic depiction of the analytical device, the information handling system and the terminal, where it can be viewed;
[0013] FIG. 2 depicts a block diagram of components of a machine learning training and inference system;
[0014] FIG. 3 is a depiction of a method used for identifying molecular communities in a molecular network; the molecular network may be a biological organism;
[0015] FIG. 4 is a depiction of the generation of the El GC-MS spectra network for the data generated from a kimchi volatilome dataset;
[0016] FIG. 5 is a depiction of a MCN of the reference compounds for El GC-MS spectra, generated for the subset of 1000 compounds parsed from a GNPS library;
[0017] FIG. 6 depicts an example of MCN connectivity: a close up portion of a molecular community of the MCN shown in FIG. 3;
[0018] FIG. 7 depicts an example of MCN connectivity: a close up portion of a molecular community of the MCN shown in FIG. 3; and
[0019] FIG. 8 depicts an example of MCN connectivity: a close up portion of a molecular community of the MCN shown in FIG. 3.DETAILED DESCRIPTION OF THE INVENTION
[0020] Disclosed herein is a system and a method for discovery, characterization, and quantification of diverse molecules that are present in a molecular network. The system includes one or more data sources in operative communication with an “information handling system”. The information handling system includes a network interface device (NID) that performs network data processing, traffic analysis, routing decisions, real-time learning from network behavior, and so on. The data sources include analytical devices that are operative to probe a sample for molecular information. The information handling system is configured to receive the input from the one or more data sources, perform analyses, and provide output information related to molecular identity of one or more molecular communities present in a molecular network. The input may come in diverse form and may be standardized for “networking” processes. In embodiments disclosed herein, the analyses performed by the information handling system are implemented through, at least in part, artificial intelligence (Al). The output may be used in a variety of ways, including, for example development of patient specific diagnostics and / or therapeutics.
[0021] The method includes generating molecular data from the one or more data sources and transmitting this molecular data to the information handling system, which generates an unpruned molecular network from the molecular data. The unpruned molecular network contains all nodes and edges present in the molecular data and may optionally be deployed in graphical form. Eachnode represents a molecular community within the molecular network. The edges reflect the degree of similarity or functional rclatcdncss between the nodes in the molecular' network. The method further includes treating the unpruned network with a community detection algorithm to find groups of nodes (molecular communities) in the network that are more densely connected internally than with the rest of the network. The method further includes pruning the clusters to remove nodes or edges that are not present in large quantities or that are not relevant. It may include removing weak, redundant, or non-informative edges or nodes to clarify a true signal (e.g., a biological signal, chemical structure, and the like). The pruning leaves behind a molecular community network with statistical features (e.g., in the form of graphical information or statistical information) that facilitates an identification of at least one molecule present in the molecular data. The molecular community network includes one or more molecular communities within the molecular network. In one embodiment, the identification of the molecular communities is conducted by comparing spectra of the molecular communities with a reference spectra present in a library.
[0022] In an embodiment, a system for detection of a molecular community comprises an information handling system that comprises at least one processor; and a non-transitory storage medium storing instructions readable and executable by the processor. The information handling system performs a method comprising receiving molecular data from an analytical device and / or a data repository. It generates an unpruned molecular network from the molecular data and treats the unpruned network with a community detection algorithm to form one or more molecular communities. Each molecular community comprises nodes and edges. All network connections are included in the molecular community and the molecular community has a higher density of connections internally as compared with the density of connections outside the molecular community. The molecular communities are parsed to prevent the formation of singletons. At least one molecule present in the molecular community is identified by comparing the molecular community with reference spectra from a library of spectra or by using a molecular structure prediction tool; where the molecular structure prediction tool generates an annotation (e.g., an atom or molecular moiety) with an associated metric (e.g., a probability) that indicates the accuracy of this annotation.
[0023] Disclosed herein too is a computer program product stored on non-transitory machine readable media, the computer program product comprising machine executable instructionsconfigured for detecting a molecular community. The detection of the molecular community includes generating an unpruned molecular network from molecular data received from an analytical device or a data repository. The unpruned network is treated with a community detection algorithm to form one or more molecular communities, where each molecular community comprises nodes and edges. All network connections are included in the molecular community and the molecular community has a higher density of connections internally as compared with the density of connections outside the community. The molecular communities are parsed to prevent the formation of singletons. At least one molecule present in the molecular community is identified by comparing the molecular community with reference spectra from a library of spectra or by using a molecular structure prediction tool; where the molecular structure prediction tool generates a likely annotation (e.g., an atom or a molecular moiety) with a metric (a probability) that gives the accuracy of this annotation. An example of a molecular moiety is a molecular group such as a hydroxyl group, an alkyl group, an alkylene group, a thiol group, and so on.
[0024] In another embodiment, a molecular structure prediction tool may be used as a foundational component for identifying molecular communities within a molecular network by enabling high- resolution structural characterization of molecular' entities, such as proteins or peptides. The molecular- structure prediction tool may utilize computational techniques such as energy minimization, homology modeling, fragment assembly, or physics-based force field simulations to predict the spatial arrangement of atoms within the molecule. In some implementations, the tool further incorporates rule-based algorithms, structural templates, or experimentally derived constraints (e.g., from nuclear’ magnetic resonance or cryo-EM data) to refine the predicted conformation. Alternatively or additionally, deep learning or neural network-based models may be used to infer structural coordinates from learned patterns in large databases of known molecular structures. The resulting predicted structure may be used for downstream applications including interaction modeling, drug design, functional annotation, or structural comparison.
[0025] In order to provide some greater context for the technology disclosed herein, some terminology is now introduced.
[0026] As used herein, the term “algorithm” generally refers a set of instructions (such as a set of machine executable instructions stored on non-transitory machine-readable media) that provide for functionality as disclosed herein. When implemented in a computing environment,embodiments of the algorithm may include libraries, separate procedures, compiled code and other similar structures.
[0027] As used herein, a “community detection algorithm” is used to identify groups of nodes in a molecular network that are more densely connected to each other than to the rest of the molecular network. These groups are called molecular' communities, clusters, or modules.
[0028] An information handling system (IHS) collects, processes, stores, and transmits information. Examples may include desktops, laptops, servers, embedded devices, and the like. Its functions may include data input (receiving data from the analytical devices) / output (providing information about the identity of at least one molecular in the molecular data), computation, storage, running software / firmware, or the like. The information handling system that is designed to collect, analyze and manage connections between molecules based on large-scale data-often enhanced with machine learning (ML) or Al capabilities.
[0029] As used herein, machine learning training is the process by which a machine learning (ML) model learns patterns from data so it can make predictions or decisions on new, unseen data. As used herein, “artificial intelligence” is a field of computer science focused on creating systems or machines that can perform tasks that normally use human intelligence. These tasks include things like learning, problem- solving, pattern recognition, decision-making, language understanding, and even visual perception.
[0030] The term “operative communication” may involve one or more of electrical communication, wireless communication or optical communication.
[0031] As used herein, the term “molecular network” generally refers to a collection of molecules connected by specific chemical, physical, physicochemical, biochemical or biophysical relationships that contribute to a particular cellular or physiological function. The molecular network may be a graphical model that represents relationships between molecules. Molecular networks are often derived from experimental data and computational predictions. They provide a framework for analyzing the role of individual molecules and their interactions in cellular or physiological functions. The molecular' network may include nodes and edges. Each node represents a community of molecules (also called a “molecular community”) and edges that reflect the degree of similarity or functional relatedness between the community of molecules.
[0031] As used herein, the terms “molecular system” “molecular network”, “molecular community” or “molecular network community” are respectively inclusive of “biomolecularsystem”, “biomolecular network”, “biomolecular community” or “biomolecular network community”.
[0032] The term ‘‘chemical” is inclusive of the term “biochemical” and the term “physical” is inclusive of the term “biophysical”.
[0033] As used herein, the terms “molecular community” and “molecular network” describe related but organizational levels of molecular systems. A molecular network includes a set of molecules (also referred to as a group of molecules or a cluster of molecules) (referred to herein as a molecular community) connected by defined interactions that collectively represent a chemical, physical or biological process, pathway, or system function. The network structure may be represented as a graph, where nodes represent molecules (e.g., polymers, proteins, nucleic acids, metabolites) and edges represent specific interactions (e.g., binding, catalysis, regulation, or transformation).
[0034] A singleton in a molecular network is a node with no connections to any other.
[0035] An annotated neighbor is a node connected to a node of interest that has known functional, structural, chemical, physical, or physicochemical information (annotations) associated with it.
[0036] Generally, the term “molecular’ community” refers to a group of molecules within a larger molecular- network that displays a high degree of connectivity and functional coherence. Functional coherence alludes to the fact that some groups or clusters of molecules have characteristics (e.g., chemical structure, biological function, biosynthetic origin, co-occurrence patterns, coordinated activity, or the like) that are more closely associated with each other than with molecules outside the group. These communities may represent metabolic modules, signaling pathways, functional classes (e.g., lipids, peptides, secondary metabolites, or the like), or coregulated molecular ensembles.
[0037] In graph theory terms, a molecular community is a subnetwork or module composed of nodes that are more densely connected to each other than to nodes outside the group. Molecular communities frequently correspond to functional units such as signaling cascades, protein complexes, or metabolic sub-pathways; co-regulated entities such as genes or proteins under control of shared regulatory elements; spatiotemporal clusters such as molecules co-localized in specific subcellular regions or active under similar conditions.
[0038] The generation of one or more molecular communities in the molecular network is sometimes referred to herein as “partitioning” of the molecular network.
[0039] As used herein, a “molecular community network” refers to a plurality of molecular communities present in an unpruned or a pruned molecular’ network, where the molecular communities have some similarities to each other based on relationships like chemical similarity, similar behavior, similar physical structure, known biology, amongst other factors. In other words, the molecular network can also be referred to as a molecular’ community network once molecular communities having some shared identity are identified within the molecular network. There may be a plurality of molecular- communities in the molecular community network with each molecular community (of the plurality of molecular- communities) having a different similarity or set of similarities.
[0040] As used herein, “molecular annotation” refers to the assignment of descriptive information or functional labels to molecular entities, such as polymers, genes, proteins, or metabolites, based on known characteristics. Such annotations may include molecular function, cellular- localization, biological pathway participation, structural classification, disease associations, or other relevant biological attributes derived from curated databases, computational predictions, or experimental data, amongst other things. Molecular annotation facilitates interpretation, classification, and analysis of molecular data within biological systems or computational models.
[0041] As used herein, modularity (Q) is a measure of how well a network is divided into communities. It compares the actual number of edges within communities to the expected number if the edges were randomly distributed. A high modularity indicates dense connections within communities and sparse connections between them. It ranges from about -0.5 to 1, but is typically between 0 and 1 in real networks.
[0042] A bridge is an edge (or sometimes a node) that connects two otherwise separate parts of a network.
[0043] A bottleneck is a node or edge that a lot of traffic or flow must pass through, creating a critical point of control or potential congestion. A bottleneck node has high betweenness centrality, meaning - it sits on many shortest paths between other nodes; it plays a major role in information flow or network communication; if it fails or is removed - it can disrupt connectivity or slow things down.
[0044] Betweenness centrality measures how often a node lies on the shortest path between other pairs of nodes in a network.
[0045] As used herein, “network topology” refers to the structural configuration and arrangement of nodes and edges within a network, describing how molecular entities (c.g., polymers, genes, proteins, or metabolites) are interconnected. The topology of a molecular network encompasses quantitative and qualitative features such as the number and distribution of node connections (degree distribution), the presence of highly connected nodes (hubs), the existence of modular substructures or communities, clustering coefficients, path lengths between nodes, and overall network connectivity or sparsity.
[0046] As used herein, “embedding” is a mathematical transformation that maps each node (e.g., a molecular community such as a gene, protein, or metabolite) to a vector in a lower-dimensional numerical space, wherein the proximity or orientation of vectors reflects the structural or functional relationships between corresponding nodes in the original network. These embeddings may be generated using graph-based techniques such as random walk algorithms (e.g., Node2Vec, DeepWalk), matrix factorization, or neural network-based models such as graph neural networks (GNNs), which aggregate and encode topological and contextual information from neighboring nodes. The resulting vector representations preserve key network features and enable downstream computational tasks such as clustering, classification, similarity analysis, or community detection. Embedding techniques thus provide a scalable and flexible means to capture complex molecular relationships in a format amenable to machine learning and statistical analysis.
[0047] Clustering is a fundamental technique in unsupervised learning aimed at partitioning a dataset into groups, or clusters, where data points within the same cluster are more similar to each other than to those in other clusters. The objective is to discover inherent structures or patterns within the data without any prior knowledge of class labels.
[0048] Examples of clustering techniques include centroid-based clustering, density-based clustering, and hierarchical clustering. In centroid-based clustering algorithms such as k-means, clusters are represented by a central point, or centroid, which serves as the prototype for the cluster. Data points are iteratively assigned to the nearest centroid, and centroids are updated based on the mean of the data points within each cluster. In density-based clustering algorithms, like DBSCAN (Density-Based Spatial Clustering of Applications with Noise), identify clusters based on regions of high data density. Data points are classified as core points, border points, or noise points, and clusters are formed around core points within a specified neighborhood density threshold. In hierarchical clustering methods, a hierarchy of clusters is created by iteratively merging or splittingclusters based on proximity measures. Agglomerative hierarchical clustering starts with individual data points as clusters and progressively merges them based on proximity, while divisive hierarchical clustering begins with one cluster containing all data points and splits them into smaller clusters recursively.
[0049] Some key elements clustering techniques include principal component analysis (PCA), t- Distributed Stochastic Neighbor Embedding (t-SNE), and autoencoders. With regard to Principal Component Analysis (PCA): PCA is a widely used dimensionality reduction technique that transforms the original features into a new set of orthogonal (uncorrelated) features called principal components. These components capture the maximum variance in the data, allowing for dimensionality reduction while retaining most of the information. With regard to t-Distributed Stochastic Neighbor Embedding (t-SNE): t-SNE is a nonlinear dimensionality reduction technique particularly effective for visualizing high-dimensional data in lower-dimensional space. It preserves the local structure of the data, making it suitable for exploratory data analysis and visualization. With regard to autoencoders: Autoencoders are neural network architectures trained to reconstruct input data from a compressed, lower-dimensional representation called the latent space. By encoding and decoding data, autoencoders learn a compact representation that captures the essential features of the original data, enabling dimensionality reduction.
[0050] Dimensionality reduction is a technique used to reduce the number of features, or dimensions, in a dataset while preserving its essential information and structure. By reducing the dimensionality of the data, it becomes more manageable, less prone to overfitting, and computationally less intensive.
[0051] As used herein, a “lower-dimensional numerical space” refers to a vector space comprising fewer dimensions than the original data representation of molecular entities, wherein each dimension corresponds to a numerical feature. For example, a molecular' entity such as a gene may initially be characterized by a high-dimensional feature vector comprising hundreds or thousands of attributes (e.g., greater than 1000 dimensions), including but not limited to interaction profiles, gene expression levels, or regulatory relationships. Through the application of dimensionality reduction techniques, such as embeddings, the high-dimensional representation may be transformed into a lower-dimensional vector comprising substantially fewer features (e.g., 10 to 100 dimensions), wherein the resulting representation captures essential aspects of the gene’s behavior or function within a molecular network. This transformation retains the salient structuralor functional characteristics of the molecular entities while enabling efficient computational processing, storage, and analysis, including clustering, classification, and pattern recognition within the network.
[0053] As used herein, a “molecular sample” or “sample” includes a tissue sample from a living being, a chemical or pharmaceutical sample that is a residue from a sample collection process or that is manufactured in a manufacturing process. Molecular data pertaining to “molecular communities”, “molecular networks” and “molecular community networks” may be obtained from the molecular sample.
[0054] Molecular communities support interpretation of large-scale molecular networks by breaking down complex systems into smaller, manageable components with shared functions or behaviors. While molecular networks provide a holistic view, molecular communities offer mechanistic granularity and facilitate targeted investigations in diagnostics, drug development, and functional genomics. The Table 1 below provides a convenient summary of the definitions and scope of the terms “molecular network” and “molecular community” amongst other things.Table 10055] As sued herein, “metabolomics” is the study of metabolites. The metabolites are the intermediates and end products of metabolic pathways, and they reflect the biochemical activity and state of a biological system. A metabolite is any molecule involved in metabolism — either as a substrate, intermediate, or product of a biochemical reaction. They are typically divided intoprimary metabolites, which drive basic cellular function (e.g., amino acids, sugars, nucleotides) or secondary metabolites - not directly involved in growth or survival of a bio-organism, but which are nevertheless important for interactions (e.g., antibiotics, pigments, and the like).
[0056] The term “metabolite community” refers to a group of interrelated metabolites that are functionally, structurally, chemically, physically, or physicochemically (e.g., hydrogen bonding or ionic interactions) connected within a biological system. These metabolites interact within shared metabolic pathways and exhibit relationships that can be analyzed through molecular networking to reveal biochemical activities, metabolic states, and functional associations in metabolomics studies. In a “metabolite network”, the term “metabolite community” is analogous to a node. A metabolite community is analogous to a network community which is also referred to herein a node.
[0057] With reference now to FIGS. 1A and IB, an exemplary information handling system (IHS) 100 is in operative communication with an analytical device or a data repository 150 for receiving input (in the form of molecular data) and providing an output 155 that includes the identification of at least one molecule or molecular community from the molecular’ data. FIG. 1 A is a schematic depiction of an exemplary information handling system (IHS), while FIG. IB is a schematic depiction of a networked system 200 that includes the information handling system (IHS) 100, an analytical device or a data repository 150 and a device 155 for depicting output information. The analytical device may probe a molecular sample such as a tissue sample (e.g., blood, urine, from a living organism), a chemical or pharmaceutical sample from a reactor (e.g., a reaction product from a chemical or pharmaceutical reaction), and so on. The tissue sample may be examined in- vivo (outside a living being) or in-vitro (inside a living being via laproscopic methods). The analytical equipment will be discussed in detail later. The data repository 150 may be an independent storage device that may be in direct or indirect communication with an analytical device and receives molecular’ date from it. The IHS 100 performs a methodology (a sequence of steps detailed in the FIG. 3) on the molecular data and provides the result to an output device 155. The result includes an identification of at least one molecule or one molecular community in the molecular- sample.
[0058] As shown in FIG. 1 A, the IHS 100 includes one or more processor(s) 105 coupled to system memory 110 via system interconnect 115. The one or more processor(s) 105 may also be referred to as a “central processing unit (CPU) 105.” System interconnect 115 can be interchangeablyreferred to as a “system bus 1 15,” in one or more embodiments. Also coupled to system interconnect 115 is storage 120 within which can be stored one or more software and / or firmware modules and / or data (not specifically shown). In one embodiment, storage 120 can be a hard drive or a solid state drive. The one or more software and / or firmware modules within storage 120 can be loaded into system memory 110 during operation of IHS 100. As shown, system memory 110 can include therein a plurality of software and / or firmware modules including firmware (F / W) 112, basic input / output system / unified extensible firmware interface (BIOS / UEFI) 114, operating system (O / S) 116 and application(s) 118. The various software and / or firmware modules have varying functionality when their corresponding program code is executed by one or more processors 105 or other processing devices within IHS 100. During boot-up or booting operations of IHS 100, processor 105 selectively loads at least BIOS / UEFI driver or image from non-volatile random access memory (NVRAM) to system memory 110 for storage in BIOS / UEFI 114. In one or more embodiments, BIOS / UEFI image comprises the additional functionality associated with unified extensible firmware interface and can include UEFI images and drivers.
[0059] IHS 100 further includes one or more input / output (I / O) controllers 130 which support connection by, and processing of signals from, one or more connected input device(s) 132, such as a keyboard, mouse, touch screen, or microphone. I / O controllers 130 also support connection to and forwarding of output signals to one or more connected output device(s) 134, such as a monitor or display device or audio speaker(s).
[0060] IHS 100 further includes a network interface device (NID) 180 that enables IHS 100 to communicate and / or interface with other devices, services, and components that are located external to IHS 100. These devices, services, and components can interface with IHS 100 via an external network, such as exemplary network 190, using one or more communication protocols. In one embodiment, a customer provisioned system / platform can comprise multiple devices located across a distributed network, and NID 180 enables IHS 100 to be connected to these other devices. Network 190 can be a local area network, wide area network, personal area network, and the like, and the connection to and / or between network and IHS 100 can be wired or wireless or a combination thereof. For purposes of discussion, network 190 is indicated as a single collective component for simplicity. However, it is appreciated that network 190 can include one or more direct connections to other devices as well as a more complex set of interconnections as can exist within a wide area network, such as the Internet.
[0061] As discussed herein, and for purposes of clarity, the IHS 100 includes a plurality of “computing resources.” Generally, the computing resources provide system functionality needed for computing functions. Exemplary computing resources include, without limitation, processor(s) 105, system memory 110, storage 120, and the input / output controller(s) 130, and other such components.
[0062] In an embodiment, the IHS of the FIG. 1A can incorporate components of machine learning training via network 190 (i.e., in the cloud) or via storage 120. FIG. 2 depicts a block diagram 400 of components of a machine learning training and inference system according to one or more embodiments described herein. One or more embodiments described herein can utilize machine learning techniques to perform tasks, such as classifying molecular features, predicting molecular interactions, or annotating unknown compounds based on spectral similarity of the molecular data with molecular structures obtained from the reference spectra. More specifically, one or more embodiments described herein can incorporate and utilize rule-based decision making and artificial intelligence (Al) reasoning to accomplish the various operations described herein, namely automated feature annotation, network topology inference, and compound identification in complex biological samples.
[0063] In an embodiment, the information handling system 100 may include a personal computer. The personal computer may be connected to or in communication with a network, and other computing systems such as a server. The server may be local or remote. The personal computer and / or server may contain a library containing information regarding known chemical compounds or metabolites. Information may be associated with the chemical compounds or metabolites includes, for example, spectral data. Other components suited for practice of the teachings herein include any type of device that will yield spectral data suited for characterizing a molecular structure. For example, mass spectroscopy systems, chromatography systems, nuclear magnetic resonance (NMR) and a variety of other similarly purposed devices may be provided as analytical equipment.
[0064] In some embodiments, a computer program product stored on non-transitory machine- readable media is provided. The computer program product may contain machine executable instructions that are configured to operate and control analysis equipment. The computer program product may be configured as the engine that makes classifications, associations, notations, and other types of characterizations of the unknown molecular structures. The characterizations maybe based on the analysis data, libraries of data for known metabolites and with input from a user. For example, a user may set tolerances or thresholds for making characterizations. Generally, embodiments of computer program products as disclosed herein may also be referred to as “software.”
[0065] The computer program product may be configured to update a central reference library or a machine learning model with additional known chemical compounds or metabolites upon satisfactory verification of identification.
[0066] The computer program product may be trained using appropriate data. For example, the computer program product may be trained using training systems implementing artificial intelligence and machine learning. The computer program product may be built using programming tools such as Python, C++ and other such tools. The computer program product may incorporate aspects of commercially available software such as Excel. The computer program product may communicate with, cooperate with, or replace elements of related software such as any one or more of: AlpsNMR; SigMa; NMRfilter; MSHub + EI-GNPS; RGCxGC toolbox; CROP; ncGTW; TidyMS; AutoTuner; hRUV; MetumpX; MetaQuac; dbnorm; MetaClean; NeatMS; MESSAR; SMART 2.0; MetFID; CPVA; NRPro; MetENP / MetENPWeb; CANOPUS; MolDiscovery; MetIDfyR; Qemistree; and IIMN.
[0067] Generally, the techniques may make use of machine learning technologies. Machine learning (ML) is a subset of artificial intelligence (Al) focused on developing algorithms that allow computers to learn from data. Types of machine learning include, without limitation: supervised learning; unsupervised learning; reinforcement learning; semi-supervised learning; deep learning and others. The phrase “machine learning” broadly describes a function of electronic systems that learn from data. A machine learning system, engine, or module can include a trainable machine learning algorithm that can be trained, such as in an external cloud environment, to learn functional relationships between inputs and outputs, and the resulting model (sometimes referred to as a “trained neural network,” “trained model,” and / or “trained machine learning model”) can be used for predicting the biological activity or structural class of novel molecular' features, for example.
[0068] In one or more embodiments, machine learning functionality can be implemented using an artificial neural network (ANN) having the capability to be trained to perform a function. In machine learning and cognitive science, ANNs are a family of statistical learning models inspired by the biological neural networks in nature. ANNs can be used to estimate or approximate systemsand functions that depend on a large number of inputs. Convolutional neural networks (CNN) are a class of deep, feed-forward ANNs that arc particularly useful at tasks such as analyzing visual imagery and natural language processing (NLP). Recurrent neural networks (RNN) are another class of deep, feed-forward ANNs and are particularly useful at tasks such as, but not limited to, unsegmented connected handwriting recognition and speech recognition. Other types of neural networks are also known and can be used in accordance with one or more embodiments described herein.
[0069] ANNs can be embodied as so-called “neuromorphic” systems of interconnected processor elements that act as simulated “neurons” and exchange “messages” between each other in the form of electronic signals. Similar to the so-called “plasticity” of synaptic neurotransmitter connections that cany messages between biological neurons, the connections in ANNs that carry electronic messages between simulated neurons are provided with numeric weights that correspond to the strength or weakness of a given connection. The weights can be adjusted and tuned based on experience, making ANNs adaptive to inputs and capable of learning. For example, an ANN for handwriting recognition is defined by a set of input neurons that can be activated by the pixels of an input image. After being weighted and transformed by a function determined by the network’s designer, the activation of these input neurons are then passed to other downstream neurons, which are often referred to as “hidden” neurons. This process is repeated until an output neuron is activated. The activated output neuron determines which character was input. It should be appreciated that these same techniques can be applied in the case of inferring molecular relationships and network topologies based on spectral or structural data as described herein.
[0070] Generally, supervised learning involves training a model on labeled data, where the algorithm learns to make predictions based on input-output pairs. Common applications include regression and classification tasks. Unsupcrviscd learning involves training a model on unlabeled data, where the algorithm learns to identify patterns and structures within the data without explicit guidance. Clustering and dimensionality reduction are typical unsupervised learning techniques. Reinforcement learning involves training an agent to interact with an environment in order to achieve a specific goal. The agent learns through trial and error, receiving feedback in the form of rewards or penalties based on its actions. Semi-supervised learning combines elements of both supervised and unsupervised learning, leveraging a small amount of labeled data along with a larger pool of unlabeled data to improve model performance. Deep learning is a subset of ML thatutilizes artificial neural networks with multiple layers (deep neural networks) to extract high-level features from raw data. It has shown remarkable success in various domains, including image recognition, natural language processing, and speech recognition.
[0071] Systems for training and using a machine learning model are now described in more detail with reference to FIG. 2. In an embodiment, the information handling system 100 of the FIG. 1 may include the machine learning training and inference system 400 of the FIG. 2. Particularly, FIG. 2 depicts a block diagram of components of a machine learning training and inference system 400 according to one or more embodiments described herein. The system 400 performs training 402 and inference 404. During training 402, a training engine 416 trains a model (e.g., the trained model 418) to perform a task, such as to predict molecular relationships, annotate unknown compounds, or classify spectral features based on structural similarity. Inference 404 is the process of implementing the trained model 418 to perform the task, such as to generate molecular networks, assign compound annotations, or identify clusters of structurally related molecules, in the context of a larger system (e.g., a system 426).
[0072] The training 402 begins with training data 412, which may be structured or unstructured data. According to one or more embodiments described herein, the training data 412 includes mass spectrometry (MS / MS) spectral data, molecular structure data, known compound annotations, or molecular fingerprints. The training engine 416 receives the training data 412 and a model form 414. According to one or more embodiments described herein, the model form 414 represents a base model that is untrained. The model form 414 can have preset weights and biases, which can be adjusted during training. It should be appreciated that the model form 414 can be selected from many different model forms depending on the task to be performed. For example, where the training 402 is to train a model to perform image classification, the model form 414 may be a model form of a CNN, although other types of model forms and / or algorithms can be implemented.
[0073] According to one or more embodiments described herein, the model form 414 represents an algorithm that can be trained to perform a particular task. In some embodiments, the model form 414 is an algorithm that can include, for example, supervised learning algorithms, unsupervised learning algorithm, artificial neural network algorithms, association rule learning algorithms, hierarchical clustering algorithms, cluster analysis algorithms, outlier detection algorithms, semi-supervised learning algorithms, reinforcement learning algorithms and / or deeplearning algorithms. Examples of supervised learning algorithms can include, for example, AODE; Artificial neural network, such as Backpropagation, Autocncodcrs, Hopficld networks, Boltzmann machines, Restricted Boltzmann Machines, and / or Spiking neural networks; Bayesian statistics, such as Bayesian network and / or Bayesian knowledge base; Case-based reasoning; Gaussian process regression; Gene expression programming; Group method of data handling (GMDH); Inductive logic programming; Instance-based learning; Lazy learning; Learning Automata; Learning Vector Quantization; Logistic Model Tree; Minimum message length (decision trees, decision graphs, etc.), such as Nearest Neighbor algorithms and / or Analogical modeling; Probably approximately correct learning (PAC) learning; Ripple down rules, a knowledge acquisition methodology; Symbolic machine learning algorithms; Support vector machines; Random Forests; ensembles of classifiers, such as Bootstrap aggregating (bagging) and / or Boosting (meta- algorithm); Ordinal classification; Information fuzzy networks (IFN); Conditional Random Field; ANOVA; Linear classifiers, such as Fisher's linear discriminant, Lineai- regression, Logistic regression, Multinomial logistic regression, Naive Bayes classifier, Perceptron, and / or Support vector machines; Quadratic classifiers; k-nearest neighbor; Boosting; Decision trees, such as C4.5, Random forests, ID3, CART, SLIQ, and / or SPRINT; Bayesian networks, such as Naive Bayes; and / or Hidden Markov models. Examples of unsupervised learning algorithms can include Expectation-maximization algorithm; Vector Quantization; Generative topographic map; and / or Information bottleneck method. Examples of artificial neural network can include Self-organizing maps. Examples of association rule learning algorithms can include Apriori algorithm; Eclat algorithm; and / or FP-growth algorithm. Examples of hierarchical clustering can include Single-linkage clustering and / or Conceptual clustering. Examples of cluster analysis can include K-means algorithm; Fuzzy clustering; DBSCAN; and / or OPTICS algorithm. Examples of outlier detection can include Local Outlier Factors. Examples of semi- supervised learning algorithms can include Generative models; Low-density separation; Graph-based methods; and / or Co-training. Examples of reinforcement learning algorithms can include Temporal difference learning; Q-leaming; Learning Automata; and / or SARSA. Examples of deep learning algorithms can include Deep belief networks; Deep Boltzmann machines; Deep Convolutional neural networks; Deep Recurrent neural networks; and / or Hierarchical temporal memory.
[0074] According to one or more embodiments described herein, the model form 414 is a foundational model that is trained on a wide variety of generalized, unlabclcd training data to perform one or more different general tasks, such as generating content (text, images, etc.), performing natural language processing, and / or the like including combinations and / or multiples thereof. In the case of the model form 414 being a foundational model, the training 402 can include tuning the foundational model (e.g., the model form 414) using the training data 412. Tuning the foundational model provides the benefits of the broad capabilities of the foundational model while enabling the foundational model to be customized using training data (e.g., the training data 412) related to a particular task or environment to which the foundational modal is then applied. In this way, the training 402 need not train a new model from scratch, which is time consuming and resource intensive.
[0075] The training 402 can be supervised learning, semi-supervised learning, unsupervised learning, reinforcement learning, and / or the like, including combinations and / or multiples thereof. For example, supervised learning can be used to train a machine learning model to classify an object of interest in an image. To do this, the training data 412 includes labeled images, including images of the object of interest with associated labels (ground truth) and other images that do not include the object of interest with associated labels. In this example, the training engine 416 takes as input a training image from the training data 412, makes a prediction for classifying the image, and compares the prediction to the known label. The training engine 416 then adjusts weights and / or biases of the model based on results of the comparison, such as by using backpropagation. The training 402 may be performed multiple times until a suitable model is trained (e.g., the trained model 418).
[0076] Once trained, the trained model 418 can be used to perform inference 404 to perform a task, such as to predict structural relationships between molecules, identify unknown compounds, or assign molecular families within a molecular network. The inference engine 420 applies the trained model 418 to new data 422 (e.g., real-world, non-training data). For example, if the trained model 418 is trained to classify images of a particular object, such as a chair, the new data 422 can be an image of a chair that was not part of the training data 412. In this way, the new data 422 represents data to which the model 418 has not been exposed. The inference engine 420 makes a prediction 424 (e.g., a classification of an object in an image of the new data 422) and passes the prediction 424 to the system 426. The system 426 can, based on the prediction 424, take an action,perform an operation, perform an analysis, and / or the like, including combinations and / or multiples thereof. In some embodiments, the system 426 can add to and / or modify the new data 422 based on the prediction 424.
[0077] In accordance with one or more embodiments, the predictions 424 generated by the inference engine 420 are periodically monitored and verified to ensure that the inference engine 420 is operating as expected. Based on the verification, additional training 402 may occur using the trained model 418 as the starting point. The additional training 402 may include all or a subset of the original training data 412 and / or new training data 412. In accordance with one or more embodiments, the training 402 includes updating the trained model 418 to account for changes in expected input data.
[0078] This disclosure will now detail the methodology performed by the information handling system 100 that results in the eventual identification of at least one molecule or at least one molecular community in the molecular data obtained from an analytic device or a data repository.
[0079] FIG. 3 depicts a method 500 that includes generating an unpruned network of molecules from molecular data derived from a biological organism, a chemical composition or a pharmaceutical composition. The unpruned molecular network may be derived from experimental or computational molecular data (see step 502). One or more molecular communities are generated in graphical form from the unpruned molecular network based on relationships between molecules such as chemical similarity, co-behavior, physical structure, or known biology (see step 504). The one or more molecular communities are then pruned to identify meaningful, biologically relevant groups of molecules (or metabolites) (see step 506). Pruning facilitates community detection by filtering out weak or noisy connections, making the resulting molecular network more structured and easier to partition into biologically and chemically meaningful groups. Pruning is generally conducted prior to or during the generation of molecular communities. Pruning may be followed by parsing. (Step 508) Parsing is optional and is generally conducted after the formation of molecular communities and may involve extracting and cataloging the individual members of the community, assessing their functional annotations, quantifying intra-community connectivity, and evaluating biological relevance through pathway enrichment, disease association, or regulatory relationship analysis. Following parsing, the molecular communities may be identified by comparing respective molecular communities with reference molecules generated in a reference library. (Step 510) In an embodiment, a molecular community may be identified by a molecularstructure prediction tool that is configured to generate a three-dimensional structural model of a molecule.
[0080] The method detailed in the FIG. 3 may be implemented by generating molecular data in quantitative analytical devices (e.g., mass spectrometers, nuclear magnetic resonance equipment, liquid or gas chromatographs, and so on) and downloading this data to the information handling system detailed in FIG. 1A, which then processes the data according to the steps listed above (and which are described in further detail) to facilitate an accurate identification of at least one molecule or metabolite or at least one molecular community.
[0081] With regard to step 502, experimental molecular data may be used to generate the unpruned network. The experimental data may be generated from a variety of different quantitative analytical devices and techniques that include probing living organisms or molecular structures (from chemical or pharmaceutical compositions) with stimuli such as heat, electromagnetic radiation (e.g., visible light, ultraviolet light, radio-frequency and microwave radiation, infrared radiation, xrays, and the like), acoustic radiation or mechanical energy (e.g., mechanical waves traveling through a medium), plasma (e.g., high-energy, ionized gas containing electrons, ions, and neutral atoms), electricity, magnetism, particle or ion beams, and the like, or a combination thereof.
[0082] The analytical techniques that provide the molecular data may be destructive or nondestructive techniques. They may include determining molecular structural sizes on the order of Angstroms to several millimeters. They may also include determining properties of the biological organisms, chemical and / or pharmaceutical compositions and correlating these properties with the molecular structures to derive the unpruned network. Examples of such techniques include mass spectrometry (MS), gas chromatography (GC), liquid chromatography (LC), electric ionization mass spectrometry (EI-MS), small angle xray scattering (SAXS), wide angle xray scattering (WAXS), acoustic and mechanical vibration (e.g., low-frequency sound (Hz-kHz) - typically used in sonar- or medical ultrasound, high-frequency ultrasound (MHz) - used in medical imaging and soft tissue analysis, hypersound or gigahertz acoustics (GHz) -molecular’ scales), or a combination thereof.
[0083] In an embodiment, the molecular data is generated by mass spectrometry (MS), tandem mass spectrometry (MS / MS), liquid chromatography-mass spectrometry (LC-MS), hydrophilic interaction chromatography (HILIC), reverse-phase chromatography (RPC), gas chromatography-mass spectrometry (GC-MS), electrospray ionization mass spectrometry (ESI-MS), matrix- assisted laser dcsorption / ionization, timc-of-flight mass spectrometry (MALDI-TOF MS), quadrupole time-of-flight mass spectrometry (QTOF-MS), fourier transform ion cyclotron resonance mass spectrometry (FT-ICR MS), inductively coupled plasma mass spectrometry (ICP- MS), atmospheric pressure chemical ionization mass spectrometry (APCI-MS), proton transfer reaction mass spectrometry (PTR-MS), thermal ionization mass spectrometry (TIMS), desorption electrospray ionization mass spectrometry (DESI-MS), dynamic mechanical analysis (DMA) or a combination thereof.
[0084] Table 2 lists electromagnetic radiation-based techniques that may be used for probing structures in biological organisms. Table 3 lists acoustic techniques, Table 4 lists electrical and electric field-based techniques, Table 5 lists magnetic techniques, Table 6 lists heat based techniques, Table 7 lists techniques that deploy particle or ion beams, Table 8 lists plasma based techniques and Table 9 lists techniques that use chemical interactions to probe for chemical structures.Table 2Table 3Table 4Table 5Table 6Table 7Table 8Table 9from a data depository is used to generate an unpruned molecular network. In an embodiment, the network may be represented in a graphical fashion. The unpruned molecular network obtained from the aforementioned techniques includes all original connections between nodes, even if some edges (which are used to connect nodes that bear some similarity to each other) have minimalsignificance for community structure. The unpruned network includes all pairwise relationships between molecules, without filtering weak or redundant connections. The unpruned network generally refers to a full connectivity map — a network that includes all possible associations (e.g., interactions, correlations, or similarities) between molecular entities (genes, proteins, compounds, and the like), before any filtering or pruning is applied to remove weak or less meaningful edges.
[0086] The unpruned molecular network therefore serves as the raw input for community detection algorithms before any preprocessing or refinement. No techniques such as degree -based filtering, backbone extraction, or sparsification are applied to the unpruned molecular network. The unpruned molecular network is generally not treated to remove weaker connections.
[0087] Step 504 is directed to detecting and generating molecular communities (to produce a molecular community network) from the unpruned molecular network based on relationships like chemical similarity, similar behavior, similar physical structure, known biology, amongst other factors. For example, groups or clusters of molecules that bear some of the aforementioned relationships (e.g., chemical similarity) may be designated as nodes. The nodes therefore represent molecular communities that bear some similarity to each other. The nodes may comprise singletons (nodes that are not connected to each other in the molecular network) or annotated neighbors (an annotated neighbor is a node connected to a node of interest that has information of interest (annotations) associated with it). The number of singletons may be greater or less than the number of annotated nodes. Edges are then added between nodes based on pairwise similarity scores. In other words, edges are used to connect nodes that have certain similar features. The edge may be designated by a weight (a numerical value) that represents the similarity between the nodes.
[0088] In an embodiment, the number of original nodes that have neighbors in an unpruned molecular network may be greater than 1000, preferably greater than 10,000, more preferably greater than 100,000, preferably greater than 1,000,000, preferably greater than 2,000,000, preferably greater than 4,000,000, preferably greater than 6,000,000, and more preferably greater than 8,000,000.
[0089] In an embodiment, the number of singletons (a node that is not connected to other nodes in the original network) is less than 10%, preferably less 20%, preferably less than 30%, preferably less than 40% and preferably less than 50% of number of original nodes that have neighbors in the network. In an embodiment, the number of singletons in the original network is greater than 300,preferably greater than 500, preferably greater than 1000, preferably greater than 500,000, preferably greater than 1,000,000, upto 2,000,000 nodes.
[0090] In an embodiment, a node in the molecular community network (which is formed after partitioning of the molecular network into molecular communities) may be connected to at least one or more other nodes and preferably connected to at least two other nodes by weighted edges with a cosine similarity greater than 0.70, preferably greater than 0.75, preferably greater than 0.8, preferably greater than 0.85, preferably greater than 0.9, and more preferably greater than 0.91, and preferably greater than 0.915. In an embodiment, each node in the molecular community network represents a different chemical compound.
[0091] The method disclosed herein detects and connects at least 85% of original nodes existing in the molecular network, preferably at least 90% of original nodes existing in the molecular network, and more preferably at least 95% of original nodes existing in the molecular network.
[0092] The step 504 may include network construction (creating a graphical representation of the network or creating a virtual network), relationship modeling, and the use of community detection algorithms to determine molecular communities, amongst other things. This results in the determination of a plurality of molecular communities in the unpruned molecular network. After the detection of the plurality of molecular communities, the unpruned molecular network may also be referred to as a molecular community network because it now includes a network of a plurality of molecular communities.
[0093] The nodes in the molecular network may be determined by finding approximate similarities between groups or clusters of molecules using spectroscopic techniques. Each node comprises structurally or functionally related molecules. Spectroscopic techniques and analytical techniques are listed above and will not be repeated again. An exemplary technique for determining nodes is mass spectroscopy. For example, in tandem mass spectroscopy (MS / MS), a node is determined by identifying a distinct molecular feature, typically represented by a high-quality MS / MS spectrum, which is then used to represent a potential group or cluster of molecules (also referred to as a molecular community).
[0094] When mass spectroscopy is performed on a sample (containing a metabolite or a chemical or a pharmaceutical composition), a distinct fragmentation pattern obtained from a specific ion (typically a precursor ion). This specific ion serves as a kind of “fingerprint” for a molecule or molecular feature. By subjecting the sample (containing a metabolite or a chemical or apharmaceutical composition) to multiple fragmentations across different across different scans, a consensus spectrum may be obtained. The consensus spectrum clarifies that there is a group of molecules (a molecular community) that have similar features and this group of molecules can be represented by a single node. The consensus spectrum is a unique spectrum that helps avoid redundancy in the network. It generally ensures that each node represents a distinct molecular feature. The term “scan” refers to a single data acquisition cycle where the MS instrument measures ion signals across a range of mass-to-charge (m / z) values.
[0095] Following the determination of nodes, edges are created by determining whether respective nodes are similar enough to be grouped together. In an embodiment, edges are created if the nodes meet certain clustering thresholds. A clustering threshold is a cutoff value used to determine whether two data points (or spectra, molecules, and the like) are similar enough to be grouped (clustered) together. A clustering threshold is generally applied to pairwise similarity scores to determine the edges.
[0096] In an embodiment, this may optionally be accomplished by constructing a graph of the network and using similarity scoring techniques to identify groups or clusters of molecules that bear' some similarity to each other. Similarity scoring includes mathematical techniques such as the Tanimoto similarity, cosine scores, the Dice coefficient and / or the Soergel distance to determine similarities between groups or clusters of molecules (molecular communities) in the molecular network. This similarity scoring process helps normalize intensity variations that can arise from different instrument conditions and matrix effects.
[0097] The Tanimoto similarity, also known as the Jaccard index, may be used to assess the similarity between two chemical compounds based on their binary fingerprints, with values ranging from 0 to 1, where higher values (closer to 1) indicate greater similarity.
[0098] The Tanimoto coefficient calculates the ratio of the number of common features in two molecular community to the total number of features present in either or both molecular communities. A Tanimoto coefficient of 1 indicates that the two molecular communities have identical structures. A Tanimoto coefficient of 0 indicates that the two molecula ’ communities have no common features (and therefore, no structural similarity). A Tanimoto coefficient of 0.85 or greater is often considered a high degree of similarity and may be used to connect nodes (i.e., create edges).
[0099] A cosine score (or cosine similarity) is a measure used to compare how similar two vectors arc, regardless of their magnitude. In molecular communities, cosine similarity measures how similarly two molecules behave across samples and helps define or evaluate connections and clusters in the molecular network. In an embodiment, a cosine score of greater than 0.7, preferably greater than 0.80 and more preferably greater than 0.85 is desirable for creating a similarity graph from which molecular communities can be generated. For example, molecular communities (nodes) having a similarity score of 0.85 to 0.90 may be connected to form a first set of edges. Molecular' communities (nodes) having a similarity score of 0.91 to 0.95 may be connected to form a second set of edges.
[0100] Edges having different cosine values may be allocated a numerical weighting. The “numerical weighting of an edge” refers to a quantitative value assigned to a connection between two nodes within a network, wherein the weight represents a measure of the strength, reliability, similarity, or significance of the relationship between the corresponding molecular entities. In various embodiments, edge weights may encode biological metrics such as gene co-expression correlation coefficients, protein-protein interaction confidence scores, chemical similarity indices, or pathway co-membership probabilities. These numerical weights facilitate weighted network analyses, allowing for the prioritization, filtering, or differential evaluation of connections during computational tasks such as clustering, community detection, network traversal, or functional inference. In some implementations, higher numerical weights may indicate greater biological relevance or stronger molecular associations, while in others, inverse weighting schemes may be employed to represent distance or dissimilarity.
[0101] In an embodiment, in order to determine the connectivity in the network, the edges may be classified into three categories (i) “high-quality”: edges with cosine similarities in the range of 0.9 to 1.0; these edges should have the highest impact on identifying the correct connected molecular communities; (ii) “medium-quality” or “gray area” cosine similarities in the range of 0.6 to 0.8; (iii) “low-quality” cosine similarities of approximately 0.5 and below; these edges are still preserved in the molecular community network to maintain connectivity; however, the impact of these “low-quality” edges on community partitioning would be minimized. This approach facilitates the preservation of overall network connectivity, while simultaneously ensuring that the impact of each edge on the molecular community network is commensurate with the “quality” of the respective edge.
[0102] This cosine similarity is useful when used in conjunction with mass spectrometry data. It provides a measure of spectral similarity by treating spectra as vectors and calculating their angular distance, which accommodates the intensity variations common in mass spectrometry data. Spectral similarity may be quantified by representing individual spectra as mathematical vectors and comparing their orientation in a multidimensional space. This involves calculating the angular distance between two spectral vectors, which reflects the similarity in their overall shape and distribution of intensity values, rather than relying solely on absolute intensity magnitudes. This is effective for mass spectrometry data, where variations in peak intensities are common due to experimental conditions or instrument sensitivity. By focusing on the angle between spectral vectors rather than their magnitude, the method provides a measure of similarity that accounts for structural or compositional resemblance between compounds while minimizing the impact of intensity fluctuations. This enables a reliable identification and clustering of chemically related molecules based on their spectral profiles.
[0103] The cosine similarity calculation incorporates fragment intensity information from mass spectrometry information by applying a square root transformation to the intensity values before computing the spectral similarity score. This transformation reduces the dominance of highly intense peaks and increases the relative influence of lower-intensity fragments, which can carry meaningful structural information. The cosine similarity score is then computed using the transformed intensities, according to the following formula:Cosinewhere Ai and Bi represent the original intensity values of the i* matched fragment ion in spectra A and B, respectively, and n is the total number of matched peaks. The square root transformation balances the contribution of each peak, allowing for more reliable comparison of mass spectra, which is particularly useful in applications such as compound identification, molecular clustering, and spectral library searching.
[0104] The Dice coefficient is another similarity measure used to compare two sets — commonly used to evaluate how similar two molecular fingerprints are. It is defined as:Dicewhere c = number of bits set to 1 in both fingerprints (intersection); a = number of bits set in fingerprint A; and b = number of bits set in fingerprint B . The score ranges from 0 (no similarity) to 1 (identical). In an embodiment, a Dice coefficient of greater than 0.7, preferably greater than 0.80 and more preferably greater than 0.85 is desirable for creating a similarity graph.
[0105] The Soergel distance is a dissimilarity measure used to compare two sets or vectors, especially in cheminformatics to assess how different two molecular fingerprints are: a 4 & ---- &Soergel Distance ■■' - 7 -» - a "" c where a = number of bits set in fingerprint A, b = number of bits set in fingerprint B, and c =number of bits set in both A and B (the intersection). It measures the proportion of dissimilar bits between two clusters of molecules. A fingerprint is a digital representation of a molecule. It ranges from 0 (identical) to 1 (completely different). In an embodiment, a Soergel Distance of less than 0.25, preferably less than 0.20 and more preferably greater than 0.15 is desirable for creating a similarity graph.
[0106] In an embodiment, a similarity network generated by using the Tanimoto similarity, the Dice similarity, the cosine similarity or the Soergel distance may be used to construct a weighted graph. The weighted graph may be fed to a community detection algorithm to generate a molecular network community as detailed below.
[0107] As noted above, each edge is formed by linking nodes that have some similarity to each other. An additional (optional) fine-tuning procedure may be used to ensure that low similarity connections, even if numerous, do not overwhelm high similarity connections. In order to ensure that low cosine similarity connections (cosine similarity less than 0.5, preferably less than 0.3 and preferably less than 0.2) do not overwhelm high similarity connections (cosine similarity greater than 0.8, preferably greater than 0.9), a sigmoid weighting function may be used to facilitate a transformation of the initially calculated cosine similarities.
[0108] The sigmoid function is a function that maps any input “x” to an output value between 0 and 1, that matches the range or cosine similarity. Input x is the respective cosine similarity calculated from the original data. The formula for the sigmoid function is given by:
[0109] In an embodiment, the sigmoid function CJ(X, k, c) is parameterized by (steepness) and (center), as shown below:where k is the steepness and c is the center. This function is applied to rescale the weights of edges in the molecular community network. The steepness parameter k determines how sharply the function transitions between its asymptotic limits, such that larger values of & yield a steeper slope, while smaller values produce a more gradual curve. The center parameter c determines the position along the input axis at which the function output equals 0.5, thereby controlling the horizontal displacement of the curve.
[0109] The sigmoid function enables the modeling of threshold-like behavior in systems where outputs increase gradually with respect to input changes around a defined center. By doing so, very low weights corresponding to non-informative edges are further reduced, effectively minimizing the influence of weak connections on the molecular communities determined by the Louvain algorithm (which is detailed below). Higher weights are preserved or slightly adjusted, amplifying useful relationships that highlight significant edges in the graph. The sigmoid function is tuned to prioritize the dense connectivity regions such as, for example, in biological data while discriminating against connections that are likely random, especially those having a cosine similarity of less than 0.5.
[0110] The use of the sigmoid function enables the inclusion of “low-quality” (less than or equal to 0.5) and “medium-quality” (0.5 to 0.7) edges in the molecular community network, thus maintaining connectivity and preserving all the information contained in the original dataset. Even though the respective weights of “low-quality” (0.5) and “medium-quality” (0.5-0.7) links are decreased by using the sigmoid transformation, all of these weights are still positive (non-zero), and none of the edges are discarded, maintaining connectivity and preserving all the information contained in the original dataset.
[0111] Community detection algorithms like Louvain or Infomap are then applied to the graph or molecular network to identify densely connected groups of nodes, revealing clusters of structurally or functionally related molecules. Even without pruning, the algorithm may organize the molecular network into molecular communities based on intrinsic connectivity patterns.
[0112] A variety of community detection approaches exist and can be used. Examples of community detection algorithms include the Louvain community detection algorithm, the Girvan- Newman community detection algorithm, the label propagation algorithm, the spectral clustering algorithm, or a combination thereof.
[0113] Some of the foregoing community detection algorithms use a modularity maximization algorithm for community detection. A modularity maximization algorithm is a computational method used in network clustering and community detection that seeks to partition the molecular network into molecular communities by maximizing modularity, a measure of the strength of community structure. Modularity measures how well the molecular network is divided into molecular communities.
[0114] In some embodiments, a modularity maximization algorithm is an optimization-based approach that partitions a network into communities by maximizing the modularity function, which quantifies the difference between the actual density of edges within molecular clusters and the expected density in a randomly connected network. In some embodiments, a modularity maximization algorithm groups nodes into communities such that intra-community connections are maximized while inter-community connections are minimized.
[0115] The Louvain community detection algorithm is one such a modularity maximization algorithm that may be used to detect and generate a molecular community network. The Louvain community detection algorithm is useful for finding molecular communities in large molecular networks, especially in applications like social networks, biological networks, and molecular similarity networks. It is fast, scalable, and may operate unsupervised. It can identify communities via modularity optimization in very large networks (e.g., networks with approximately 100 million nodes).
[0116] The Louvain algorithm groups nodes into communities such that the modularity of the partition is maximized. Modularity is a measure of how densely connected nodes are within communities, compared to that in a random network and is expressed as follows:where:^4s: edge between node i and ' degree of node i m: -total number of edgesCs: community of node i d: 1 if nodes are in the same community, 0 otherwise
[0117] The Louvain community detection algorithm to determine molecular’ communities is conducted in two phases. The first phase (Phase I) includes the following steps: i) a node is first selected to represent a molecular community; ii) for each node, the gain in modularity is evaluated if it is moved to a molecular’ community of each of its neighbors; iii) the node is then moved to the molecular- community that gives the largest increase (or smallest decrease) in modularity; iv) the foregoing steps are repeated until no more improvements can be made. A gain in modularity means that when a node (or group of nodes) is moved to a different community, the overall modularity score increases. This results in small local molecular communities. After completing this phase, the first phase can be reapplied to create larger molecular communities with increased modularity.
[0118] In the second phase, a new molecular network is constructed using the molecular communities identified in the first phase as nodes. The weights of links between these new nodes are determined by the sum of weights of links between nodes in the corresponding molecular communities. The second phase (Phase II) is directed to community aggregation and includes the following steps: i) collapse each molecular community into a single super-node; ii) build a new network where the nodes are the molecular communities from Phase 1 ; the edge weights are equal to the sum of edges between the nodes in the original network. The process is repeated between Phase I and Phase II until there is no further increase in modularity. This results in a final hierarchy of molecular communities. In an embodiment, the modularity obtained from the Louvain community detection algorithm is 0.4 to 0.9.
[0119] Using the Louvain community detection algorithm facilitates clustering similar molecules (e.g., as obtained from a Tanimoto similarity computation), identifying scaffold or activity-based communities and revealing hidden chemical series or structure-activity relationship (SAR) patterns. The number of communities identified by the Louvain method is not fixed or predefined. Instead, the algorithm produces molecular communities that arise “naturally” from the input data, without any initial restrictions on the number or sizes of those communities.
[0120] Other community detection algorithms may be used to determine molecular community networks in an unpruned network. In some embodiments, Infomap is used for molecular community networking detection and optimization. Infomap employs a map equation to detect modules by tracing random walks through the network and is useful at discovering smaller, detailed clusters.
[0121] In some embodiments, “label propagation” is used for molecular community networking detection and optimization. Label propagation assigns community labels through iterative consensus among neighboring nodes and provides a fast, approximate solution for large-scale molecular networks. In some embodiments, the Girvan-Newman Algorithm is used for molecular community networking detection and optimization. The Girvan-Newman Algorithm progressively removes edges with high betweenness centrality to split the network into communities. Betweenness centrality measures how often a node lies on the shortest path between other pairs of nodes in a network.
[0122] In some embodiments, Walktrap is used for molecular community networking detection and optimization. Walktrap uses random walks to compute distances between nodes, then merges nodes with similar’ visitation patterns and produces a hierarchical decomposition of the molecular network. In some embodiments, Spinglass is used for molecular community networking detection and optimization. Spinglass treats community detection as a statistical physics problem involving spins on a graph and optimizes a modularity-like quality function.
[0123] In an embodiment, the community detection may be performed by first identifying a maximum weight spanning tree. In the context of graph theory, a tree is a fundamental structure that consists of nodes connected by edges without forming any cycles (i.e., there are no closed loops - i.e., the tree does not start and end at the same node). One specific type of tree associated with a graph is a spanning tree. A spanning tree of a graph is a subgraph that includes all the nodes of the original graph while preserving connectivity and remaining acyclic. In computer networking, a spanning tree protocol may be used to ensure a loop-free topology, enabling efficient and reliable communication between network devices.
[0124] A maximum weight spanning tree is a spanning tree where the sum of the weights assigned to its edges is maximized. In other words, it is a tree that spans all the nodes of the graph while emphasizing edges with the highest weights. The weight of an edge represents the strength, significance, or type of interaction between two nodes that the edge connects.
[0125] Constructing a maximum weight spanning tree involves employing algorithms designed for spanning trees, such as Prim's algorithm or Kruskal's algorithm, with a modification to prioritize edges with higher weights. The approach typically involves stalling with an arbitrary original node, selecting the edge with the maximum weight connected to it, and iteratively adding edges to the tree while avoiding cycles (avoiding returning to the original node). This process continues until all nodes are included in the tree, resulting in a spanning tree that maximizes the sum of edge weights.
[0126] In a preferred embodiment, the cosine similarity score is used to estimate the weight of each edge. The weight of each edge is equal to its respective cosine similarity score. The molecular network is then split into molecular communities by maximizing modularity.
[0127] Following the generation of a final hierarchy of molecular communities (also termed clusters), the one or more molecular communities are pruned to identify meaningful, relevant groups of molecules (or metabolites) (see step 506). Pruning refers to the process of simplifying or refining a network (or graph) by removing nodes or edges that are considered irrelevant, noisy, or less significant. Pruning molecular communities from an unpruned network of molecules involves filtering, clustering, and refinement to identify meaningful, biologically relevant groups of molecules (or metabolites). In an embodiment, pruning includes removing low-weight edges while retaining full connectivity by keeping at least one strongest connection for each node.
[0128] In molecular networks (like similarity networks, interaction networks, or communities of related compounds), pruning helps focus on the most biologically or chemically meaningful structures. In an embodiment, each node in the plurality of nodes represents a community of metabolites while the edges connect nodes that have similar characteristics (e.g., chemical similarity, co-behavior, physical structure, biological function correlation, mutual information, or distance metrics).
[0129] Pruning may be conducted by one of the following methods - edge pruning, node pruning, graph topology-based pruning, community-aware pruning, or a combination thereof.
[0130] Edge pruning, based on weights or similarities includes removing weak or low-confidence connections between nodes. For example, it includes removing connections between nodes that have a Tanimoto similarity of less than 0.3.
[0131] Node pruning includes removing nodes that are likely uninformative or noisy. Examples of node pruning include removing isolated nodes (nodes with degree - 0, i.e., nodes that are notconnected to other nodes), removing nodes with very low degree (e.g., less than 2; i.e., nodes that arc connected to only one other node), and / or removing nodes that lack biological relevance (e.g., failed assays or no known activity).
[0132] Graph topology-based pruning looks at the structural role of edges / nodes in the molecular network and includes, for example, removing bridges (a bridge is an edge (or sometimes a node) that connects two otherwise separate parts of a network) or bottlenecks (a bottleneck is a node or edge that a lot of traffic or flow must pass through, creating a critical point of control or potential congestion) that link otherwise unrelated pails by simplifying to a backbone graph (e.g., via maximum weight spanning tree) or retaining only core-periphery structure.
[0133] Community-aware pruning includes detecting communities (e.g., via Louvain or spectral clustering) and removing inter-community edges by focusing on inter-cluster similarity. In an embodiment, pruning of a community network may include, for example, computing all pairwise similarities, keeping only edges where the Tanimoto similarity is greater than or equal to 0.4, removing nodes with less than 2 connections, running a Louvain community detection on the network and discarding communities of size less than 3, followed by visualizing the pruned, cleaned-up network.
[0134] In mass spectrometry -based metabolomics where the mass-spectrometry data often includes a large number of features (many of which arise from noise, isotopes / adducts, redundant peaks, contaminants, and the like), pinning techniques may include one or more of intensity thresholding to remove low-abundance signals, frequency filtering to remove features not consistently found across replicates or conditions, blank subtraction to eliminate features also present in blank / control runs, relative standard deviation / variance filtering to remove features with high technical variability, and adduct deconvolution, which may be conducted to collapse related peaks into a single compound.
[0135] In one embodiment, in pruning metabolite networks, i.e., when building a metabolite correlation network from NMR or LC-MS data, the metabolites may be assumed to be nodes, edges may be determined using a Pcarson / Spcarman correlation (Pearson and Spearman correlations are two commonly used statistical measures to determine how strongly two variables are related) or a spectral similarity. Pruning here may include i) edge pruning where only strong correlations (e.g., the absolute cosine score being greater than 0.7) are retained and removing statistically insignificant edges (e g., using a false discover rate (FDR) of greater than 0.05); ii)node pruning where metabolites with low confidence (e.g., unknown IDs) are removed and removing nodes with a few connections.
[0136] Pruning reduces noise and weak connections and removes low-similarity edges. This ensures only strong, meaningful connections remain. It prevents the formation of false communities that lead to blended or inaccurate communities. It simplifies network structure. The presence of fewer edges and fewer paths in the community network leads to easier computation and better clustering resolution. This helps algorithms detect dense modules more precisely and also highlights core chemical similarities, which leads to a more accurate identification of molecules. The network becomes sparser and cleaner, highlighting strong relationships only, which results in an easier identification of the molecular community network. In short, after pruning, the nodes which represent molecular communities; the edges, which represent relationships that include correlation, structural similarity, co-expression, and the like between molecular communities; and edge weights, which represent strength or confidence between the respective correlations, structural similarities, co-expressions, are strengthened to identify the molecular communities with a greater degree of accuracy.
[0137] In an embodiment, after pruning, the molecular community may optionally be parsed. (See Step 508) Parsing a molecular community includes further refining it based on the molecular community’s characteristics, structure, function, how it is connected to other nodes, amongst other things. This may include examining the particular molecular' community in isolation and analyzing its internal components (removing weak nodes, splitting molecular communities into sub- molecular communities, and the like), to better understand its structure, members, and potential function. Parsing generally facilitates uncovering meaningful structure or biological insight.
[0138] A weak node refers to a node within a network that exhibits limited connectivity or influence relative to other nodes in the system. Numerically, this may be characterized by a low degree, defined as the number of edges or direct connections to other nodes, typically falling below a statistically derived threshold such as less than half the average degree of the molecular' network. In the context of a chemical molecular network, a weak node may represent molecular communities that shares minimal structural or functional similarity with other molecular communities, resulting in sparse or marginal linkages within the network topology. Such nodes often reside at the periphery of molecular communities and may exhibit low centrality measures,including betweenness and eigenvector centrality, indicating a reduced role in the transmission of information or functional relationships across the molecular network.
[0139] In an embodiment, a weak node may be defined as a node having a degree centrality of less than 3, indicating it is directly connected to fewer than three other nodes. Additionally, or alternatively, a weak node may possess a betweenness centrality value below 0.01, indicating minimal participation in the shortest paths between node pairs across the network. In weighted networks, a node may further be considered weak if the average weight of its connected edges is less than 0.2 on a normalized scale from 0 to 1.
[0140] Splitting a molecular community into sub-molecular communities may be achieved through a process of hierarchical refinement or recursive partitioning, wherein a previously identified molecular community is reanalyzed to reveal finer-grained structural or functional groupings. This may involve the reapplication of clustering algorithms, such as k-means, hierarchical clustering, or density-based methods (e.g., DBSCAN), restricted to the data points within the original molecular community. Alternatively, sub-molecular community formation may be guided by distinct similarity metrics, feature sets, or dimensionality reduction techniques, allowing for the identification of latent substructures not apparent under the initial clustering criteria. In an embodiment, community detection algorithms — such as the Louvain or Leiden methods — can be applied iteratively to molecular communities, enabling the delineation of nested or context-specific sub-communities. This multi-level clustering approach enables more granular classification and improved interpretability of complex datasets.
[0141] Table 10 lists some differences between the pruning and parsing.Table 1000142] After pruning and parsing, the molecular communities within the molecular community network may be identified. (See Step 510) There are several methods by which molecular communities may be identified. The molecular community network is compared against existing reference spectra (contained in a library of reference spectra) generated for known molecules via electron ionization (El)-gas chromatography-mass spectrometry (GC-MS) and liquidchromatography-mass spectrometry (LC-MS). A subset of up to 500, preferably up to 1000, and preferably up to 1500 reference spectra were randomly examined from the library. A number of reference spectra greater than 2000, preferably greater than 2500, preferably greater than 5000 reference spectra may be examined from the library. In an embodiment, the reference database is hosted by the Global Natural Products Social (GNPS) platform, which is a publicly available, community-curated library of mass spectrometry data. Each spectrum in the GNPS library represents the fragmentation pattern of a chemical compound as detected by tandem mass spectrometry (MS / MS). The random selection process ensures an unbiased subset of spectral data. The molecular community network may then be used to explore connectivity patterns of examined reference compounds. In other words, the reference compounds from the library may be used to explore connections (chemical bonds) between molecules in the molecular community and also explore connections between adjacent molecular communities in the molecular community network.
[0143] In an embodiment, the molecular community network and the reference compounds (from the library) may be viewed on a monitor (see output 155) in system 200 (see FIG. IB) Different reference compounds from the library may be viewed on the monitor till an appropriate fit for the molecular- community is arrived at. In one embodiment, the disclosed method utilizes a plurality of reference spectra as informational “seeds” to determine the composition and structure of a molecular- community within a sample. The reference spectra, which may include mass spectra, NMR spectra, Raman spectra, or any combination thereof, are compared to the experimental spectra acquired from the molecular community. In another embodiment, structures obtained from the reference spectra are statistically compared with structure obtained by molecular community networking. By leveraging pattern recognition algorithms and similarity scoring metrics, the method identifies spectral features in the sample that correlate with the known reference spectra. These correlations serve as anchor points or “seeds” from which molecular relationships are inferred, allowing for the propagation of identity or structural information across related, but initially unassigned, spectral signals. Through iterative association and network-based clustering, the method delineates a molecular community comprising structurally or functionally related molecules, even in cases where full structural elucidation is unavailable. This approach enhances molecular annotation in complex mixtures, facilitating discovery and characterization of biologically or chemically relevant molecular groups.
[0144] In order to generate a comprehensive and reliable reference dataset for liquid chromatography-mass spectrometry (LC-MS) analysis, data were acquired from a set of 800 authenticated metabolite standards using four distinct analytical methodologies. These methodologies included reverse-phase (RP) and hydrophilic interaction liquid chromatography (HILIC), each performed under both positive and negative ionization modes. The use of multiple chromatographic techniques and ionization polarities was designed to capture the inherent spectral variability that arises from differences in chemical properties and ionization behavior among metabolites. Simultaneously, by applying these methods in a controlled experimental setting, instrumental variation was minimized, thereby ensuring consistency and reproducibility across the resulting spectral profiles. This multi-modal acquisition strategy enhances the robustness of the reference library and improves the reliability of downstream applications such as compound identification, spectral matching, and molecular network construction.
[0145] In another embodiment, molecular structure prediction tools (such as for example a protein structure prediction tool) may be used as a foundational component for identifying molecular communities within a molecular network by enabling high-resolution structural characterization of molecular entities, such as proteins or peptides. In this approach, the molecular structure prediction tool may be used to generate three-dimensional structural models of target molecules from amino acid sequences. In one embodiment, the molecular structure prediction tool generates an annotation with an associated metric that indicates the accuracy of this annotation. The annotation referred to herein represents an atom or compound present in the molecular community. The accuracy is expressed as a probability that the identification is correct. A high probability indicates a high degree of confidence in the annotation. These predicted structures are then analyzed to extract relevant structural features — such as secondary structure elements, active site conformations, surface charge distributions, or protein-protein interaction (PPI) interfaces.
[0146] Once these structural features are extracted, they are converted into numerical representations, such as feature vectors, molecular fingerprints, or graph-based embeddings. These representations are then used to construct a molecular- network, where nodes represent individual proteins or molecules and edges encode structural similarity, predicted interaction strength, or functional compatibility derived from the protein structure prediction tool-generated models. Additional data sources — such as experimental interaction databases, sequencehomology, or known functional annotations — can be integrated to enhance the accuracy of detection and identification of the molecular networks.
[0147] Examples of molecules or molecular communities that may be detected by the method and system disclosed herein includes carbohydrates (e.g., glucose, fructose, galactose, ribose, lactose, or the like), amino acids (e.g., glycine, alanine, glutamine, tryptophan, histidine, or the like), lipids (e.g., palmitic acid, arachidonic acid, cholesterol, phosphatidylcholine, sphingosine, or the like), lipids (palmitic acid, arachidonic acid, cholesterol, phosphatidylcholine, sphingosine, or the like), nucleotides and derivatives (e.g., ATP, GTP, cAMP, NAD+, FAD, or the like), organic acids (e.g., pyruvate, citrate, lactic acid, succinic acid, acetyl-CoA, or the like), hormones (e.g., cortisol, estradiol, epinephrine, thyroxine, or the like), or a combination thereof.
[0148] In an exemplary embodiment, the method and system may be used to detect metabolites present in the molecular community networks. These include microbially-derived bile acids. Examples of these microbially-derived bile acids include allolithocholic acid (alloLCA), deoxycholic acid (DCA), glycodeoxycholic acid (GDCA), glycolithocholic acid (GLCA), hyodeoxycholic acid (HDCA), iso-deoxycholic acid (iso-DCA), iso-lithocholic acid (iso-LCA), lithocholic acid (LCA), murideoxycholic acid (MDCA), taurodeoxycholic acid (TUDCA), taurolithocholic acid (TLCA), ursodeoxycholic acid (UDCA), or a combination thereof.
[0149] While the foregoing discussion emphasizes the identification of metabolites, the application is not limited thereto. It may be used to identify other communities of biomolecules, molecules that are not necessarily connected to biological organisms such as, for example, polymers, drugs and their degradation products, pesticides and their degradation products, fluoropolymers (e.g., PFAS), and the like.
[0150] In an embodiment, the information handling system 100 in conjunction with the machine learning training and inference system 400 may employ iterative machine learning to determine or to facilitate the determination of molecular communities in a molecular network. In addition, the machine learning training and inference system 400 may be used to form one or more predictive models (that represent the particular chemical compositions or the biological organisms) that may effectively inform other researchers what parameters (e.g., cosine similarities, sigmoidal functions, the type of community detection algorithm, what particular pruning parameters to deploy, and so on) to choose when evaluating a particular type of chemical composition or biological tissue, thereby reducing the amount of effort expended on identifying the particular composition. Thepredictive model may use some form of initial spectra (e.g., MS / MS spectra) or molecular data to decide upon the best parameters to deploy to make a molecular determination. It may also reduce the amount of computational power expended on the identification thereby making the identification quick and efficient.
[0151] As the machine learning process (and the model) improves its understanding and knowledge of composition and / or tissue identification for the biological organism, it may be used as a rapid quality assurance tool for compositions and tissues obtained from approximately similar or equivalent processes.
[0152] For example, as the information handling system 100 in conjunction with the machine learning training and inference system 400 receives chemical information (in the form of spectra or molecular data) from a particular manufacturing process or from a particular piece of biological tissue (from a person or a group of persons having some similar identifying features (e.g., Italians), it may be able to rapidly pinpoint molecular deviations from the norm and even provide reasons for the deviation. It may be able to provide for corrective actions or prescribe medications or treatments that can correct for the deviation from the norm. In a chemical process, the corrective action may include replacing a batch of input or changing an operating temperature to fix the deviation.
[0153] Some examples of machine learning by the information handling system 100 are provided below. In an embodiment, machine learning and network analysis techniques can be employed to identify molecular communities within molecular networks, such as gene coexpression networks, protein-protein interaction networks, or metabolic networks. In one embodiment, in a machine learning operation, graph-based community detection algorithms — such as modularity optimization (e.g., Louvain or Leiden community detection algorithms), spectral clustering, or label propagation — may be used to identify densely connected subgroups of molecules that represent functional modules. Alternatively, or additionally, machine learning approaches, including unsupervised methods like graph neural networks (GNNs), may be used to learn low-dimensional representations (embeddings) of molecular entities based on the network topology. These embeddings can be clustered using standard algorithms (e.g., K-means, DBSCAN) to delineate molecular communities.
[0154] An embedding is a mathematical transformation that converts each molecular node into a vector in a lower-dimensional space (the vector represents fewer features than the original nodeor molecular community), preserving key structural or functional relationships for use in computational analysis such as clustering or community detection. A “lowcr-dimcnsional numerical space” refers to a vector space comprising fewer numerical features than the original high-dimensional representation of molecular entities, wherein dimensionality reduction techniques, such as embedding methods, are applied to preserve essential structural or functional characteristics while enabling efficient computational analysis, including clustering, classification, and pattern recognition.
[0155] In certain embodiments, supervised learning models may be trained on annotated biological data (from libraries such the reference database hosted by the Global Natural Products Social (GNPS) platform, which is discussed above) so as to predict molecular communities based on node features such as connectivity or expression profiles. Advanced implementations may incorporate multi-omics (multiple types of chemical or biological data — such as genomics, transcriptomics, proteomics, and metabolomics) data into heterogeneous molecular networks and apply multimodal graph learning to identify cross-modal communities.
[0156] Advanced implementations may combine multiple types of chemical or biological data — such as genomics, transcriptomics, proteomics, and metabolomics — into a single heterogeneous molecular network, where each data type represents a different layer or modality of molecular information. By applying multimodal graph learning techniques to this integrated network, the information handling system 100 can identify cross-modal communities — groups of related molecular entities that span different data types — thereby uncovering more comprehensive and chemically or biologically meaningful patterns that would not be apparent when analyzing each data type in isolation.
[0157] Dynamic versions of the machine learning model may be used to analyze changes in molecular communities over time or across conditions (e.g., disease states). This may include systematic identification of chemically, biologically or functionally relevant molecular modules that may be associated with physiological or pathological processes. In an embodiment, the information handling system 100 may be used to deliver newly discovered information (from a particular newly identified network) to the reference library (such as the (GNPS) platform) for additional annotation and inclusion.
[0158] In another embodiment, machine learning may be used to make repeated attempts to determine the identities of molecules or nodes in a molecular network that may not have a highcosine similarity initially. For example, during a first pass at determining a molecular communities in the network, nodes with cosine similarities greater than 0.9 may be considered for identification and subjected to further treatment (e.g., application of the community detection algorithm, pruning and parsing followed by identification via a reference library), while those with lower cosine similarities 0.7 to 0.9 are identified in a second pass. During the first pass and second pass, the nodes with cosine similarities lower than 0.5 are retained so that molecular connectivity withing the network is maintained. During the first and second passes, the molecular features of the nodes with low cosine similarities are not considered or given a low weighting. However, in certain instances it may be valuable to identify the nodes with low cosine similarities (less than or equal to 0.5). During the third pass, an attempt may be made to identify the nodes with low cosine similarities by iteratively using the reference library (e.g., GPNS) along with any available experimental (from the analytical techniques listed above).
[0159] Information generated during the third pass such as negatives, false positives and eventual successful outcomes may be fed back to the reference library (or to the machine learning model) to help future researchers identify molecules that are present in small populations (e.g., singletons) and that have few similarities with the remaining molecular communities in the network. The identification of a plurality of molecules that are not present in large numbers and / or that have a cosine similarity of less than or equal to 0.5 can facilitate a rapid identification of future compositions and tissues when such an identification is desirable because the reference library now has a repository of these randomly occurring but hitherto present chemical or biological moieties. Molecular species that have low resolution in certain techniques, low-abundance molecules or structural isomers may be rapidly identified.
[0160] In short, the information handling system 100 (See FIG. 1A) in conjunction with the machine learning training and inference system 400 may be used to update existing reference libraries or machine-learning models as well as other databases to facilitate a more rapid identification process of hitherto unknown molecules and molecular communities so that future searches may be conducted in an expedited function or be avoided all together (because knowledge of particular compositions or tissues becomes commonplace).
[0161] In certain embodiments, artificial intelligence (Al) may be utilized to evaluate molecular communities and molecular networks through rule-based systems, logical inference, or deterministic graph-theoretic approaches, without reliance on machine learning algorithms. Forexample, the information handling system 100 may be configured with predefined biological rules, such that molecular entities or communities arc classified based on the presence of known functional annotations, interaction patterns, or pathway associations. Logical reasoning engines may infer higher-level biological functions by applying structured rules to underlying molecular attributes, enabling the identification of chemically or biologically meaningful molecular communities based on curated knowledge rather than statistical learning.
[0162] In addition, molecular networks may be analyzed using algorithmic methods derived from graph theory, including but not limited to assessments of node centrality, clustering coefficients (e.g. similarities), modularity, and connectivity. These evaluations can be used to characterize global network architecture or identify sub-structures of interest (e.g., singletons and groups or clusters (molecular communities) that have one or more similarities), such as highly connected nodes or functionally cohesive clusters. Furthermore, symbolic Al techniques, such as logic programming or constraint satisfaction (e.g., only considering molecular communities with amino acids), may be applied to query knowledge graphs or enforce biological constraints during network partitioning, thereby enabling interpretable and knowledge-driven analysis of molecular systems without the use of data-dependent model training.
[0163] In an embodiment, the system 100 (see FIGS. 1A and IB) may implement method 500 (See FIG. 3) to determine metabolites involved in cellular metabolism, including energy production, biosynthesis of molecules such as proteins and lipids, and degradation of substances. In other words, the system and method described herein may be applied to the analysis of metabolomics data. For disease mechanisms, changes in metabolite levels can indicate alterations in cellular function associated with diseases such as cancer, metabolic disorders, and neurological conditions. Metabolomics can therefore help in disease diagnosis, monitoring, and understanding the underlying mechanisms. As to drug metabolism and toxicity, metabolomics can assess how drugs are metabolized within the body and identify potential metabolic byproducts or toxic effects. This information is valuable for drug development and safety assessment. Metabolomics may be used to analyze the metabolic response to dietary intake, identifying biomarkers of nutritional status and assessing the effects of dietary interventions on metabolism. It may be used to study environmental exposures. Metabolomics may be used detect changes in metabolite levels resulting from exposure to environmental pollutants, toxins, or other stressors, providing insights into their impact on biological systems. Overall, metabolomics plays a crucial role in systems biology byproviding a comprehensive view of the biochemical processes underlying physiological and pathological states, with applications in medicine, agriculture, environmental science, and biotechnology.
[0164] In an embodiment, the technology disclosed herein is directed to a computer- implemented method for metabolomics community network detection and optimization, where the method includes receiving metabolomics data from at least one data source. The metabolomics data obtained from the at least one data source is used to generate an unpruned network. The data source is at least one of the analytical devices listed above. A topology graph is generated by clustering the unpruned network thereby generating a tree structure, wherein the tree structure includes a plurality of nodes, wherein each node in the plurality represents a community of metabolites from the metabolomics data. The tree structure also includes edges connecting pairs of nodes wherein the edges have weights representing similarities between the nodes. The topology graph containing communities of metabolites may be viewed on a display. The tree structure may be pruned to reduce noise and weak connections and to remove low-similarity edges thus generating a plurality of modules. An algorithm may detect dense modules more precisely and also highlight core chemical similarities, which leads to a more accurate identification of molecules.
[0165] Another aspect of the technology disclosed herein is related to a system for metabolomics community network detection and optimization, where the system 100 includes at least one processor and a non-transitory storage medium storing instruction readable and executable by the processor to perform a metabolomics community network detection and optimization. This includes receiving metabolomics data from at least one data source. The data source is at least one of the analytical devices listed above. A topology graph is generated by clustering the unpruned network thereby generating a tree structure, wherein the tree structure includes a plurality of nodes, wherein each node in the plurality represents a community of metabolites from the metabolomics data. The tree structure also includes edges connecting pairs of nodes wherein the edges have weights representing similarities between the nodes. The topology graph containing communities of metabolites may be viewed on a display. The tree structure may be pruned to reduce noise and weak connections and to remove low- similarity edges thus generating a plurality of modules. An algorithm may detect dense modules more precisely and also highlight core chemical similarities, which leads to a more accurate identification of molecules.
[0166] The method disclosed herein is advantageous in that the use of molecular community networking optimizes connectivity patterns for each molecule, linking up to 95% of detectable molecules compared with approximately 40% when utilizing the methods without using the disclosed metabolomics community network detection and optimization, thereby revealing previously hidden molecular relationships. In some embodiments, the use of molecular community networking aids in organizing complex molecular space and tracking spectral variants of the same molecule, providing a clearer picture of true molecular diversity in complex samples.
[0167] This method is advantageous in that it can facilitate more accurate disease detection. In medical diagnostics, molecular networks facilitate identification of disease- specific biomarkers by highlighting disrupted pathways and abnormal molecular interactions. Comparative analysis between networks derived from healthy and diseased tissues can isolate molecules or interactions exhibiting differential regulation or expression. These differentials often serve as biomarkers for early disease detection, prognosis, or disease classification.
[0168] A key diagnostic application includes disease classification. Molecular network signatures may be used to distinguish among subtypes of complex diseases, such as cancers or autoimmune disorders, based on distinct pathway alterations. Another application is biomarker discovery. Network centrality and connectivity metrics may be used to identify hub molecules or key regulators associated with disease phenotypes. The methodology may be used to predict disease risk for individual patients. Genetic and epigenetic data integrated into regulatory networks allow prediction of disease susceptibility based on inherited or acquired network perturbations.
[0169] In therapeutics, molecular networks guide drug development and personalized treatment strategies by identifying critical nodes and edges that regulate disease-related processes. Targeting molecules with high centrality or control over network flow can produce greater therapeutic efficacy with reduced off-target effects. One key therapeutic application includes target identification. Analysis of network topology highlights essential molecules within pathogenic pathways that are amenable to pharmaceutical intervention. Another key therapeutic application includes drug repurposing. Network-based algorithms match known drugs to new indications by comparing network contexts of drug targets across diseases. Yet another application includes designing combined therapies. Network modeling predicts synergistic effects of drug combinations by evaluating complementary or converging pathway disruptions. The methodologydisclosed herein includes the development of precision medicine. Patient-specific molecular networks enable stratification based on individual molecular’ profiles, allowing for tailored therapeutic interventions.
[0170] Unlike traditional diagnostic and therapeutic strategies focused on single molecules or isolated pathways, molecular network-based approaches account for the systemic nature of biological processes. Network information incorporates redundancy, compensation, and crosstalk mechanisms, thereby improving the robustness and specificity of medical applications. By leveraging the structural and functional complexity of molecular networks, medical diagnostics and therapeutics gain predictive power, mechanistic insight, and translational relevance, leading to improved patient outcomes and more efficient healthcare interventions. Thus, improved molecular attribution, interaction validation, and dynamic profiling are required to enhance the accuracy and utility of molecular network models in complex biochemical systems.
[0171] The system and the method disclosed herein are exemplified by the following non-limiting example.EXAMPLESExample 1
[0172] The following example is intended only to illustrate the disclosure. Other procedures, methodologies, techniques, reagents and conditions may alternatively be used as appropriate.
[0173] Reference networks obtained from 800 LC-MS / MS standards and 1,000 GC-MS reference spectra were used to generate unpruned molecular networks. Because these compounds span various molecular families there is a level of uncertainty in the data. For example, there can be multiple nodes for a single compound due to detection through different methodologies.
[0174] The network was constructed using the MASST (Mass Spectrometry Search Tool) methodology, which enables large-scale molecular networking based on tandem mass spectrometry (MS / MS) data. The construction of the network utilized all publicly available datasets accessible through the GNPS (Global Natural Products Social Molecular Networking) platform and the MassIVE (Mass Spectrometry Interactive Virtual Environment) repository. MASST operates by comparing MS / MS spectra across multiple datasets to identify spectral matches, which were then used to define molecular nodes and establish edges based on spectral similarity scores. The integration of data from GNPS / MassIVE ensured broad chemical coverageand enables the identification of conserved or shared molecular features across diverse biological or environmental samples.
[0175] Cosine scores greater than 0.7 were used. A clustering algorithm (Louvain method) is used to determine the naturally present molecular communities within the data. Each community is pruned via the maximum spanning tree algorithm by removing low-weight edges; the full connectivity is retained by keeping at least one strongest connection for each node. The count of nodes connected to an annotated node and thus potentially accessible for annotation propagation, across the “global” network of LC-MS / MS data is 8,453,822 nodes.
[0176] Using the methodology disclosed herein 8,040,389 nodes (95.1%) are connected to neighbors (413,433 or 4.9% are singletons), and 5,934,112 (70.2%) nodes are connected to an annotated neighbor and thus are potentially amenable for annotation propagation.
[0177] Network connections between identical structures (Tanimoto score of 1) were used a to form network connections to determine molecular communities. After the formation of network communities, the GNPS public library was used for spectral matching to determine the identity of the molecular communities. This library contains spectra for various forms of parent molecules, including in-source fragmentation (ISFs), different adducts, and the like, not just “pure” spectra. FIG. 4 is a depiction of the generation of the El GC-MS spectra network for the data of kimchi volatilome dataset as described in the section above. The hatching corresponds to the molecular communities detected in the data. From left to right in the FIG. 4 is the unpruned network; the unpruned network with hatching corresponding to the molecular communities; and the MCN generated by pruning with the maximum weight spanning tree.
[0178] FIG. 5 is a depiction of a MCN of the reference compounds for El GC-MS spectra, generated for the subset of 1000 compounds parsed from GNPS library as described in the section above. FIG. 6 depicts an example of MCN connectivity: a close-up portion of a molecular community of the MCN shown in FIG. 5. The hatching corresponds to the MCN community; the edge thickness corresponds to the cosine similarity score. The edge thickness corresponds to the cosine similarity score. A connectivity of amine group-containing compounds is shown.
[0179] FIG. 7 depicts an example of MCN connectivity: a close-up portion of a molecular community of the MCN shown in FIG. 5. The hatching corresponds to the MCN community; the edge thickness corresponds to the cosine similarity score. The edge thickness corresponds to the cosine similarity score. A connectivity of esters is shown.
[0180] FIG. 8 depicts an example of MCN connectivity - a close-up portion of a molecular community of the MCN shown in FIG. 5. The hatchings correspond to the MCN community; the edge thickness corresponds to the cosine similarity score. The hatchings correspond to the cosine similarity score: gray, high values (at or above 0.7); black, low values (below 0.7). A connectivity of halogen-containing compounds is shown. This cosine threshold is suggested for high confidence connection between structurally similar nodes in El GC-MS data.
[0181] Program architectures that implement machine learning and embody artificial intelligence may include supervised learning architecture where models are trained on labeled input-output pairs to learn a mapping function. These include linear regression, decision trees, support vector machines, and neural networks. Unsupervised learning architectures do not rely on labeled data and are used for discovering hidden patterns or groupings in data; examples include k-means clustering, hierarchical clustering, and principal component analysis. Semi-supervised learning combines both labeled and unlabeled data to improve learning efficiency, often using models like self-training neural networks. Reinforcement learning architectures involve agents that learn by interacting with environments through trial and error, optimizing behaviors based on rewards — Q-leaming and deep Q-networks (DQN) are typical examples.
[0182] Among neural network architectures, feedforward neural networks serve as the foundation for many tasks, while convolutional neural networks (CNNs) are optimized for image recognition and spatial data. Recurrent neural networks (RNNs) and their valiants such as long short-term memory (LSTM) and gated recurrent units (GRUs) are well-suited for time-series or sequential data like language modeling or speech recognition. Transformer architectures, including attention mechanisms, are widely used in modern natural language processing applications and are the foundation for large language models. Autoencoders are used for dimensionality reduction and anomaly detection, and generative adversarial networks (GANs) are designed for generating new data samples similar to training data.
[0183] Embodiments of Al using these architectures can include rule-based expert systems, predictive analytics engines, recommendation systems, natural language processors, computer vision modules, and autonomous agents. These may be deployed in cloud-based platforms, embedded systems, mobile applications, or robotic devices, depending on the intended application and computational requirements. Each architecture supports different embodiments of Al, tailored to the complexity, data type, and performance goals of the system.
[0184] All statements herein reciting principles, aspects, and embodiments of the disclosure, as well as specific examples thereof, arc intended to encompass both structural and functional equivalents thereof. Additionally, it is intended that such equivalents include both currently known equivalents as well as equivalents developed in the future, i.e., any elements developed that perform the same function, regardless of structure.
[0185] All statements herein reciting principles, aspects, and embodiments of the disclosure, as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Additionally, it is intended that such equivalents include both currently known equivalents as well as equivalents developed in the future, i.e., any elements developed that perform the same function, regardless of structure.
[0186] Various other components may be included and called upon for providing for aspects of the teachings herein. For example, additional materials, combinations of materials and / or omission of materials may be used to provide for added embodiments that are within the scope of the teachings herein. Adequacy of any particular element for practice of the teachings herein is to be judged from the perspective of a designer, manufacturer, seller, user, system operator or other similarly interested party, and such limitations are to be perceived according to the standards of the interested party.
[0187] In the disclosure hereof any element expressed as a means for performing a specified function is intended to encompass any way of performing that function including, for example, a) a combination of circuit elements and associated hardware which perform that function or b) software in any form, including, therefore, firmware, microcode or the like as set forth herein, combined with appropriate circuitry for executing that software to perform the function. Applicants thus regard any means which can provide those functionalities as equivalent to those shown herein. No functional language used in claims appended herein is to be construed as invoking 35 U.S.C. § 112(f) interpretations as “means-plus-function” language unless specifically expressed as such by use of the words “means for” or “steps for” within the respective claim.
[0188] When introducing elements of the present invention or the embodiment(s) thereof, the articles “a,” “an,” and “the” are intended to mean that there are one or more of the elements. Similarly, the adjective “another,” when used to introduce an element, is intended to mean one or more elements. The terms “including” and “having” are intended to be inclusive such that theremay be additional elements other than the listed elements. The term “exemplary” is not intended to be construed as a superlative example but merely one of many possible examples.
[0189] Various other components may be included and called upon for providing for aspects of the teachings herein. For example, additional materials, combinations of materials and / or omission of materials may be used to provide for added embodiments that are within the scope of the teachings herein. Adequacy of any particular element for practice of the teachings herein is to be judged from the perspective of a designer, manufacturer, seller, user, system operator or other similarly interested parly, and such limitations are to be perceived according to the standards of the interested party.
[0190] While the invention has been described with reference to some embodiments, it will be understood by those skilled in the art that various changes may be made and equivalents may be substituted for elements thereof without departing from the scope of the invention. In addition, many modifications may be made to adapt a particular situation or material to the teachings of the invention without departing from essential scope thereof. Therefore, it is intended that the invention not be limited to the particular embodiments disclosed as the best mode contemplated for carrying out this invention, but that the invention will include all embodiments falling within the scope of the appended claims.
Claims
CLAIMSWhat is claimed is:
1. A system for detection of a molecular community in a molecular network, the system comprising: an information handling system that comprises: at least one processor; and a non-transitory storage medium storing instructions readable and executable by the processor to perform a method comprising: receiving molecular data from an analytical device and / or a data repository; generating an unpruned molecular network from the molecular data; treating the unpruned network with a community detection algorithm to form one or more molecular communities; where each molecular community comprises nodes and edges; where all network connections are included in the molecular community and the molecular community has a higher density of connections internally as compared with the density of connections outside the molecular community; parsing the molecular communities to prevent the formation of singletons; and identifying at least one molecule present in the molecular community by comparing the molecular community with reference spectra from a library of spectra or by using a molecular structure prediction tool; where the molecular structure prediction tool generates an annotation with an associated metric that indicates the accuracy of this annotation.
2. The system of Claim 1, wherein the at least one processor is in electrical or optical communication with the analytical device that is operative to perform electron ionization (El), mass spectrometry (MS), nuclear magnetic resonance (NMR), tandem mass spectrometry (MS / MS), liquid chromatography-mass spectrometry (LC-MS), hydrophilic interaction chromatography (HILIC), reverse-phase chromatography (RPC), gas chromatography-mass spectrometry (GC-MS), electron ionization in conjunction with gas chromatography -mass spectrometry (GC-MS), chemical ionization in conjunction with gas chromatography-mass spectrometry (GC-MS), electrospray ionization mass spectrometry (ESI-MS), matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF MS), quadrupoletime-of-flight mass spectrometry (QTOF-MS), fourier transform ion cyclotron resonance mass spectrometry (FT- ICR MS), inductively coupled plasma mass spectrometry (ICP-MS), atmospheric pressure chemical ionization mass spectrometry (APCI-MS), proton transfer reaction mass spectrometry (PTR-MS), thermal ionization mass spectrometry (TIMS), desorption electrospray ionization mass spectrometry (DESI-MS), small angle xray scattering (SAXS), wide angle xray scattering (WAXS), acoustic and mechanical vibration, ion mobility, or a combination thereof.
3. The system of claim 1, wherein the molecular data includes metabolomics data and wherein the molecular’ community is a metabolite community.
4. The system of claim 1, where the information handling system comprises a machine learning training and inference system that predicts annotations of molecular features based on a topology of a molecular community network.
5. The system of claim 4, where the machine learning training and inference system is operative to develop a predictive model that defines at least one of a similarity scoring model, the sigmoid function, a particular community detection algorithm and pruning parameters to deploy based on an evaluation of initial molecular’ data.
6. The system of claim 1, wherein the community detection algorithm is a modularity maximization algorithm, a hierarchical clustering algorithm, a statistical inference algorithm, or any combination thereof.
7. The system of claim 1, wherein the modularity maximization algorithm is a Louvain algorithm.
8. The system of Claim 1, where the modularity obtained from the algorithm is 0.4 to 0.9 for non-synthetic metabolomics data.
9. The system of claim 1 , wherein the molecular community includes at least one maximum weight spanning tree.
10. The system of claim 1, wherein the at least one molecule that is identified is used for identifying a therapeutic course of action, developing a new drug, highlighting pathogenic pathways that are amenable to pharmaceutical intervention, matching known drugs to new indications by comparing network contexts of drug targets across diseases, predicting synergistic effects of drug combinations by evaluating complementary or converging pathway disruptions and / or developing patients-specific molecular drugs allowing for tailored therapeutic interventions.
11. The system of claim 1, where the molecular data contains over 1000 original nodes, at least over 85% of which are connected in the molecular network.
12. A computer program product stored on non-transitory machine readable media, the computer program product comprising machine executable instructions configured for detecting a molecular community, comprising: generating an unpruned molecular network from molecular data received from an analytical device or a data repository; treating the unpruned network with a community detection algorithm to form one or more molecular communities; where each molecular community comprises nodes and edges; where all network connections are included in the molecular community and the molecular community has a higher density of connections internally as compared with the density of connections outside the community; parsing the molecular communities to prevent the formation of singletons; and identifying at least one molecule present in the molecular community by comparing the molecular community with reference spectra from a library of spectra or by using a molecular structure prediction tool; where the molecular structure prediction tool generates a likely annotation with a metric that gives the accuracy of this annotation.
13. The computer program product of claim 12, wherein the data repository provides reference data obtained from liquid chromatography-tandem mass spectrometry, from gaschromatography mass spectrometry, or a combination thereof.
14. The computer program product of claim 12, wherein the community detection algorithm is a modularity maximization algorithm, a hierarchical clustering algorithm, a statistical inference algorithm, or any combination thereof.
15. The computer program product of claim 14, wherein the modularity maximization algorithm is a Louvain algorithm.
16. The computer program product of claim 12, further comprising pruning the one or more molecular communities with a maximum weight spanning tree; where the maximum weight spanning tree is a spanning tree where a sum of weights assigned to the edges is maximized.
17. The computer program product of claim 16, where a weight of each edge is equal to its cosine similarity score.
18. The computer program product of claim 12, further comprising generating a predictive model that predicts annotations of molecular features based on a topology of a molecular community network.
19. The computer program product of claim 18, further comprising generating a predictive model that is operative to classify molecular’ features, predict molecular’ interactions or annotate unknown compounds based on spectral similarity of the molecular- data with molecular annotations obtained from the reference spectra.
Citation Information
Patent Citations
Methods and Systems for Cell State Quantification
US20130029879A1
Metabolomic Signatures for Predicting, Diagnosing, and Prognosing Various Diseases Including Cancer
US20200200754A1
Method and apparatus for identifying heterogeneous graph and property of molecular space structure and computer device
US20210043283A1
Filtering genetic networks to discover populations of interest
US20210257060A1
System and method for de novo drug discovery
US20220188652A1
Cited By
Method for monitoring macroalgae biodiversity of seaweed field based on environmental DNA technology
CN122428028A
Method for monitoring macroalgal biodiversity of seaweed field based on environmental DNA technology
CN122428028B