Systems and methods for co-discovering graphical structure and functional relationships within data

By selecting and pruning kernels for relationships between nodes and their ancestors, the system effectively discovers hypergraph relationships within complex datasets, addressing the limitations of existing methods in capturing higher-order data relationships.

WO2025117750A1PCT designated stage expired Publication Date: 2025-06-05CALIFORNIA INST OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/057763
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-27
Filing Date
2024-11-27
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing methods for discovering hypergraphs struggle to capture higher-order relationships in data, leading to incomplete identification of intricate relationships and dependencies in complex datasets.

Method used

The system processes input data by selecting a kernel for relationships between nodes and their candidate ancestors, and then pruning these ancestors to identify minimal ancestors, thereby discovering hypergraph relationships.

Benefits of technology

This approach enables the efficient discovery of hypergraph relationships, allowing for the generation of outputs that effectively capture complex patterns and structures in data, such as control signals for physical devices like vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024057763_05062025_PF_FP_ABST
    Figure US2024057763_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for hypergraph discovery in accordance with embodiments of the invention include a method of processing input data, comprising receiving input data at a data processing system, providing the input data to a discovered hypergraph, and generating an output using the discovered hypergraph, wherein the discovered hypergraph is characterized by a plurality of relationships between a plurality of nodes that each represent a variable of a plurality of variables, where the plurality of relationships were discovered by selection of a kernel for a relationship between at least one node of the plurality of nodes and a set of candidate ancestors, and pruning of the set of candidate ancestors to identify a set of minimal ancestors.
Need to check novelty before this filing date? Find Prior Art

Description

Systems and Methods for Co-discovering Graphical Structure and Functional Relationships Within Data CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The current application claims the benefit of and priority under 35 U.S.C. § 119(e) to U.S. Provisional Patent Application No. 63 / 603,035 entitled “Computational Hypergraph Discovery” filed November 27, 2023. The disclosure of U.S. Provisional Patent Application No. 63 / 603,035 is hereby incorporated by reference in its entirety for all purposes. STATEMENT OF FEDERAL SUPPORT

[0002] This invention was made with government support under Grant No. DE- SC0023163 awarded by DOE and under Grant No. FA9550-20-1-0358 awarded by AFOSR. The government has certain rights in the invention. FIELD OF THE INVENTION

[0003] The present invention generally relates to hypergraph discovery and, more specifically, discovering graphical structure and functional relationships within data. BACKGROUND

[0004] Previous approaches to discovering hypergraphs have typically involved analyzing relationships between nodes in a dataset to identify complex patterns and structures. Traditional methods for analyzing relationships in data have included techniques such as graph theory, network analysis, and clustering algorithms. These methods may not capture higher-order relationships that exist in the data. As a result, these approaches may not fully capture the intricate relationships and dependencies present in complex datasets, limiting the ability to uncover hidden patterns and structures.

[0005] Other approaches have focused on using machine learning techniques to discover relationships in data. These methods typically involve training models to learn relationships between inputs and outputs but may struggle to efficiently learn relationships between unrelated variables.SUMMARY OF THE INVENTION

[0006] Systems and methods for hypergraph discovery in accordance with embodiments of the invention are illustrated. One embodiment includes a method of processing input data, comprising receiving input data at a data processing system, providing the input data to a discovered hypergraph, and generating an output using the discovered hypergraph, wherein the discovered hypergraph is characterized by a plurality of relationships between a plurality of nodes that each represent a variable of a plurality of variables, where the plurality of relationships were discovered by selection of a kernel for a relationship between at least one node of the plurality of nodes and a set of candidate ancestors, and pruning of the set of candidate ancestors to identify a set of minimal ancestors.

[0007] In a further embodiment, the plurality of relationships were further discovered by selection of a different second kernel for a second relationship between a second node of the plurality of nodes and a corresponding set of candidate ancestors.

[0008] In still another embodiment, the kernel is at least one selected from the group consisting of a linear kernel, a quadratic kernel, and a nonlinear kernel.

[0009] In a still further embodiment, generating the output includes generating a set of control signals for controlling a physical device.

[0010] In yet another embodiment, the physical device is a vehicle.

[0011] In a yet further embodiment, receiving the set of input data includes normalizing the set of input data.

[0012] In another additional embodiment, normalizing the data includes performing affine transformations.

[0013] In a further additional embodiment, the selection of a kernel is based on a set of training data, the set of training data includes a plurality of samples, a first sample includes values for a first subset of the plurality of variables, and a second sample includes values for a different second subset of the plurality of variables.

[0014] In another embodiment again, selection of a kernel comprises for each of a plurality of kernels computing a signal-to-noise ratio (SNR) for the relationship between the selected node and the set of candidate ancestors, and when the SNR falls below agiven threshold, selecting the kernel for the relationship between the selected node and the set of candidate ancestors.

[0015] In a further embodiment again, the plurality of kernels includes at least one selected from the group consisting of a linear kernel, a quadratic kernel, and a nonlinear kernel.

[0016] In still yet another embodiment, computing the SNR comprises selecting a noise prior parameter, and performing a regression analysis with the selected noise prior parameter.

[0017] In a still yet further embodiment, performing the regression analysis includes predicting a value of the selected node based on the set of candidate ancestors.

[0018] In still another additional embodiment, the SNR is a ratio of a mean and a variance for predicted values of the selected node based on the set of candidate ancestors.

[0019] In a still further additional embodiment, the given threshold is a fixed value.

[0020] In still another embodiment again, at least one of the plurality of kernels is at least one selected from the group consisting of a Gaussian kernel and a Matérn kernel.

[0021] In a still further embodiment again, the set of candidate ancestors includes all of the plurality of nodes other than the selected node.

[0022] In yet another additional embodiment, pruning the set of candidate ancestors comprises identifying a least important ancestor of the set of candidate ancestors, computing a signal-to-noise (SNR) for the relationship between the selected node and the set of candidate ancestors without the least important ancestor, when the computed SNR exceeds a given threshold, select the set of candidate ancestors without the least important ancestor as the set of minimal ancestors, and when the computed SNR falls below a given threshold, select the set of candidate ancestors as the set of minimal ancestors.

[0023] In a yet further additional embodiment, pruning the set of candidate ancestors comprises for each of a plurality of subsets of the set of candidate ancestors, computing a noise-to-signal ratio (NSR) for the relationship between the selected node and the subset of candidate ancestors, identifying a particular subset of the plurality ofsubsets as the set of minimal ancestors based on the computed NSRs for the plurality of subsets.

[0024] In yet another embodiment again, computing the NSR comprises selecting a noise prior parameter, and performing a regression analysis with the selected parameter.

[0025] In a yet further embodiment again, performing the regression analysis includes predicting a value of the selected node based on the set of candidate ancestors.

[0026] In another additional embodiment again, the NSR is a ratio of a mean and a variance for predicted values of the selected node based on the set of candidate ancestors.

[0027] In a further additional embodiment again, the regression analysis is a Kernel Ridge Regression.

[0028] In still yet another additional embodiment, the noise prior parameter is selected based on characteristics of the kernel.

[0029] In a further embodiment, selecting the noise prior parameter includes maximizing the variance of an eigenvalue histogram when the kernel is a universal kernel.

[0030] In still another embodiment, identifying the particular subset includes identifying an inflection point in the NSRs as a function of a number of ancestors in the subset of candidate ancestors.

[0031] In a still further embodiment, selecting the kernel and pruning the set of candidate ancestors is based on a same threshold.

[0032] In yet another embodiment, pruning the set of candidate ancestors comprises identifying a set of redundant candidate ancestors of the set of candidate ancestors, and removing at least one candidate ancestor of the set of redundant candidate ancestors from the set of candidate ancestors, wherein redundant candidate ancestors make redundant contributions to the computed NSR.

[0033] In a yet further embodiment, identifying the set of redundant candidate ancestors is based on the determined relationships between the plurality of nodes.

[0034] In another additional embodiment, the plurality of relationships were further discovered by identifying clusters of nodes of the plurality of nodes based on relationships between each node and the set of candidate ancestors of the node, and identifyingrelationships between the clusters of nodes, wherein the plurality of relationships comprises the relationships between each node and the set of candidate ancestors of the node, and the relationships between the clusters of nodes.

[0035] One embodiment includes a data processing system, comprising at least one processor, a memory, wherein machine readable instructions stored in the memory configure the processor to perform a method comprising receiving input data at a data processing system, providing the input data to a discovered hypergraph, and generating an output using the discovered hypergraph, wherein the discovered hypergraph is characterized by a plurality of relationships between a plurality of nodes that each represent a variable of a plurality of variables, where the plurality of relationships were discovered by selection of a kernel for a relationship between at least one node of the plurality of nodes and a set of candidate ancestors, and pruning of the set of candidate ancestors to identify a set of minimal ancestors.

[0036] One embodiment includes a method for hypergraph discovery. The method includes steps for receiving a set of input data includes a plurality of samples, each sample includes values for at least a subset of a plurality of variables, determining relationships between a plurality of nodes based on the set of input data, each node of the plurality of node representing a variable of the plurality of variables, wherein determining relationships for each node of the plurality of nodes comprises selecting a kernel for a relationship between a selected node and a set of candidate ancestors, and pruning the set of candidate ancestors to identify a set of minimal ancestors, and generating a hypergraph of the input data based on the relationships between the plurality of nodes.

[0037] In a further additional embodiment, receiving the set of input data includes normalizing the set of input data.

[0038] In another embodiment again, normalizing the data includes performing affine transformations.

[0039] In a further embodiment again, the set of input data includes a plurality of samples, a first sample includes values for a first subset of the plurality of variables, and a second sample includes values for a different second subset of the plurality of variables.

[0040] In still yet another embodiment, selecting a kernel comprises for each of a plurality of kernels computing a signal-to-noise ratio (SNR) for the relationship between the selected node and the set of candidate ancestors, and when the SNR falls below a given threshold, selecting the kernel for the relationship between the selected node and the set of candidate ancestors.

[0041] In a still yet further embodiment, the plurality of kernels includes at least one selected from the group consisting of a linear kernel, a quadratic kernel, and a nonlinear kernel.

[0042] In still another additional embodiment, computing the SNR comprises selecting a noise prior parameter, and performing a regression analysis with the selected noise prior parameter.

[0043] In a still further additional embodiment, performing the regression analysis includes predicting a value of the selected node based on the set of candidate ancestors.

[0044] In still another embodiment again, the SNR is a ratio of a mean and a variance for predicted values of the selected node based on the set of candidate ancestors.

[0045] In a still further embodiment again, the given threshold is a fixed value.

[0046] In yet another additional embodiment, at least one of the plurality of kernels is at least one selected from the group consisting of a Gaussian kernel and a Matérn kernel.

[0047] In a yet further additional embodiment, the set of candidate ancestors includes all of the plurality of nodes other than the selected node.

[0048] In yet another embodiment again, pruning the set of candidate ancestors comprises identifying a least important ancestor of the set of candidate ancestors, computing a signal-to-noise (SNR) for the relationship between the selected node and the set of candidate ancestors without the least important ancestor, when the computed SNR exceeds a given threshold, select the set of candidate ancestors without the least important ancestor as the set of minimal ancestors, and when the computed SNR falls below a given threshold, select the set of candidate ancestors as the set of minimal ancestors.

[0049] In a yet further embodiment again, pruning the set of candidate ancestors comprises for each of a plurality of subsets of the set of candidate ancestors, computing a noise-to-signal ratio (NSR) for the relationship between the selected node and the subset of candidate ancestors, identifying a particular subset of the plurality of subsets as the set of minimal ancestors based on the computed NSRs for the plurality of subsets.

[0050] In another additional embodiment again, computing the NSR comprises selecting a noise prior parameter, and performing a regression analysis with the selected parameter.

[0051] In a further additional embodiment again, performing the regression analysis includes predicting a value of the selected node based on the set of candidate ancestors.

[0052] In still yet another additional embodiment, the NSR is a ratio of a mean and a variance for predicted values of the selected node based on the set of candidate ancestors.

[0053] In a further embodiment, the regression analysis is a Kernel Ridge Regression.

[0054] In still another embodiment, the noise prior parameter is selected based on characteristics of the kernel.

[0055] In a still further embodiment, selecting the noise prior parameter includes maximizing the variance of an eigenvalue histogram when the kernel is a universal kernel.

[0056] In yet another embodiment, identifying the particular subset includes identifying an inflection point in the NSRs as a function of a number of ancestors in the subset of candidate ancestors.

[0057] In a yet further embodiment, selecting the kernel and pruning the set of candidate ancestors is based on a same threshold.

[0058] In another additional embodiment, pruning the set of candidate ancestors comprises identifying a set of redundant candidate ancestors of the set of candidate ancestors, and removing at least one candidate ancestor of the set of redundant candidate ancestors from the set of candidate ancestors, wherein redundant candidate ancestors make redundant contributions to the computed NSR.

[0059] In a further additional embodiment, identifying the set of redundant candidate ancestors is based on the determined relationships between the plurality of nodes.

[0060] In another embodiment again, the method further includes steps for identifying clusters of nodes of the plurality of nodes based on relationships between each node and the set of minimal ancestors of the node, and identifying relationships between the clusters of nodes, wherein generating the hypergraph of the input data is further based on the relationships between the clusters of nodes.

[0061] One embodiment includes an apparatus comprising a set of one or more processors, a set of one or more peripherals, and a non-transitory machine readable medium containing program instructions that are executable by the set of processors to perform a method that includes steps for receiving a set of input data includes a plurality of samples, each sample includes values for at least a subset of a plurality of variables, determining relationships between a plurality of nodes based on the set of input data, each node of the plurality of node representing a variable of the plurality of variables, wherein determining relationships for each node of the plurality of nodes comprises selecting a kernel for a relationship between a selected node and a set of candidate ancestors, and pruning the set of candidate ancestors to identify a set of minimal ancestors, generating a hypergraph of the input data based on the relationships between the plurality of nodes, and controlling the apparatus based on the generated hypergraph.

[0062] In a further embodiment again, the set of input data includes data collected from at least one peripheral of the set of peripherals.

[0063] In still yet another embodiment, the set of peripherals includes at least one selected from the group consisting of a gyroscope, an accelerometer, an odometer, a LiDAR system, an environmental sensor, and a motion sensor.

[0064] One embodiment includes a non-transitory machine readable medium containing program instructions that are executable by a set of one or more processors to perform a method that includes steps for receiving a set of input data includes a plurality of samples, each sample includes values for at least a subset of a plurality of variables, determining relationships between a plurality of nodes based on the set of input data, each node of the plurality of node representing a variable of the plurality of variables,wherein determining relationships for each node of the plurality of nodes comprises selecting a kernel for a relationship between a selected node and a set of candidate ancestors, and pruning the set of candidate ancestors to identify a set of minimal ancestors, and generating a hypergraph of the input data based on the relationships between the plurality of nodes.

[0065] Additional embodiments and features are set forth in part in the description that follows, and in part will become apparent to those skilled in the art upon examination of the specification or may be learned by the practice of the invention. A further understanding of the nature and advantages of the present invention may be realized by reference to the remaining portions of the specification and the drawings, which forms a part of this disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] The description and claims will be more fully understood with reference to the following figures and data graphs, which are presented as exemplary embodiments of the invention and should not be construed as a complete recitation of the scope of the invention.

[0067] Fig.1 illustrates representations of Type 1, Type 2, and Type 3 problems.

[0068] Fig. 2 conceptually illustrates a process for hypergraph discovery in accordance with an embodiment of the invention.

[0069] Fig. 3 conceptually illustrates a process for kernel selection in accordance with an embodiment of the invention.

[0070] Fig. 4 conceptually illustrates a process for pruning ancestors based on a threshold in accordance with an embodiment of the invention.

[0071] Fig. 5 conceptually illustrates a process for pruning ancestors based on a difference between successive noise-to-signal ratios in accordance with an embodiment of the invention.

[0072] Fig. 6 illustrates examples of challenges that can be addressed with hypergraph discovery in accordance with various embodiments of the invention.

[0073] Fig. 7 illustrates a resultant computational graph of the recovery of a chemical reaction network.

[0074] Fig. 8 illustrates an example of a hypergraph discovery process in accordance with an embodiment of the invention.

[0075] Fig. 9 illustrates an example of pruning ancestors in accordance with an embodiment of the invention.

[0076] Fig. 10 illustrates an example of an analysis of NSR in accordance with an embodiment of the invention.

[0077] Fig. 11 illustrates an example of result data from an analysis of a FPUT system.

[0078] Fig. 12 illustrates another example of an analysis of NSR in accordance with an embodiment of the invention.

[0079] Fig. 13 illustrates an example of hypergraph discovery as a manifold discovery problem and hypergraph representation and the hypergraph representation of an affine manifold in accordance with an embodiment of the invention.

[0080] Fig. 14 illustrates an example of feature map generalization in accordance with an embodiment of the invention.

[0081] Fig.15 illustrates an example of histograms of eigenvalues ofఊ.

[0082] Fig.16 illustrates examples of the recovery of functional dependencies from data in accordance with an embodiment of the invention.

[0083] Fig.17 illustrates an example of hypergraph discovery with COVID-19 data in accordance with an embodiment of the invention.

[0084] Fig. 18 illustrates an example of the evolution of the observed signal-to- noise ratio when pruning ancestors in accordance with an embodiment of the invention.

[0085] Fig. 19 illustrates an example of hypergraph discovery in biological cellular signaling networks in accordance with an embodiment of the invention.

[0086] Fig. 20 illustrates an example of a hypergraph discovery system that performs hypergraph discovery in accordance with an embodiment of the invention.

[0087] Fig. 21 illustrates an example of a hypergraph discovery element that executes instructions to perform processes that perform hypergraph discovery in accordance with an embodiment of the invention.DETAILED DESCRIPTION

[0088] Many complex data analysis problems within and beyond the scientific domain involve discovering graphical structures and functional relationships within data. Nonlinear variance decomposition with Gaussian Processes in accordance with a number of embodiments of the invention simplifies and automates this process. Other methods, such as Artificial Neural Networks, lack this variance decomposition feature. Information- theoretic and causal inference methods suffer from super-exponential complexity with respect to the number of variables. Systems and methods in accordance with certain embodiments of the invention perform this task in polynomial complexity. This unlocks the potential for applications involving the identification of a network of hidden relationships between variables without a parameterized model at an unprecedented scale, scope, and complexity.

[0089] Most problems within and beyond the scientific domain can be framed into one of the following three levels of complexity of function approximation. Type 1: Approximate an unknown function given (possibly noisy) input / output data. Type 2: Consider a collection of variables and functions, some of which are unknown, indexed by the nodes and hyperedges of a hypergraph (a generalized graph where edges can connect more than two vertices). Given partial observations of the variables of the hypergraph (satisfying the functional dependencies imposed by its structure), approximate all the unobserved variables and unknown functions. Type 3: Expanding on Type 2, if the hypergraph structure itself is unknown, use partial observations of (subsets of) the variables of the hypergraph to discover its structure (the hyperedges and possibly the missing vertices) and approximate its unknown functions and / or unobserved variables. These hypergraphs offer a natural platform for organizing, communicating, and processing computational knowledge. While most Computational Science and Engineering and Scientific Machine Learning challenges can be as the data-driven discovery of unknown functions in a computational hypergraph whose structure is known (Type 2), many require the data-driven discovery of the structure (connectivity) of the hypergraph itself (Type 3). Despite their prevalence, these Type 3 challenges have been largely overlooked due to their inherent complexity.

[0090] Although Gaussian Process (GP) methods are sometimes perceived as well-founded but old technology limited to Type 1 curve fitting, their scope has recently been expanded to Type 2 problems. Systems and methods in accordance with numerous embodiments of the invention introduce an interpretable Gaussian Process (GP) framework for Type 3 problems that does not require randomization of the data, nor access to or control over its sampling, nor sparsity of the unknown functions in a known or learned basis. Its polynomial complexity, which contrasts sharply with the super- exponential complexity of causal inference methods, is enabled by the nonlinear analysis of variance capabilities of GPs used as a sensing mechanism. Hypergraph discovery in accordance with various embodiments of the invention allows for hypergraphs to be generated in a computationally efficient manner, where the hypergraphs enable the efficient computation of outputs based on a set of inputs. In a variety of embodiments, hypergraph discovery can enable such efficiencies by identifying ancestors using one or more kernels and pruning.

[0091] Systems and methods in accordance with many embodiments of the invention are based on a kernel generalization of (1) Row Echelon Form reduction from linear systems to nonlinear ones and / or (2) variance-based analysis. In many embodiments, variables can be linked via GPs, and those contributing to the highest data variance can unveil a hypergraph’s structure. The scope and efficiency of processes in accordance with a number of embodiments of the invention are illustrated below with applications to (algebraic) equation discovery, network discovery (gene pathways, chemical, and mechanical), and raw data analysis.

[0092] The scope of Type 1, 2, and 3 problems is immense. Numerical approximation, Supervised Learning, and Operator Learning can all be formulated as type 1 problems i.e., as approximating unknown functions given (possibly with noisy and infinite / high-dimensional) inputs / output data. The common GP-based solution to these problems is to replace the underlying unknown function by a GP and compute its MAP estimator given available data. Type 2 problems include solving and learning (possibly stochastic) ordinary or partial differential equations, Deep Learning, dimension reduction, reduced-ordered modeling, system identification, closure modeling, etc.

[0093] The scope of Type 3 problems extends well beyond Type 2 problems and includes applications involving model / equation / network discovery and reasoning with raw data. While most problems in Computational Sciences and Engineering (CSE) and Scientific Machine Learning (SciML) can be framed as Type 1 and Type 2 challenges, many problems in science can only be categorized as Type 3 problems, i.e., discovering the structure / connectivity of the graph itself from data prior to its completion. Despite their prevalence, these Type 3 challenges have been largely overlooked due to their inherent complexity.

[0094] Type 1, Type 2, and Type 3 problems can be formulated as completing or discovering hypergraphs where nodes represent variables and edges represent functional dependencies. Representations of Type 1, Type 2, and Type 3 problems are illustrated in Fig. 1. In this figure, the graph for Type 1 has only two variables and one unknown function. The graph for Type 2 has multiple variables and (some possibly unknown) functions, and the connectivity of the graph is known. The graph for Type 3 has an unknown connectivity (functional dependencies between variables may be unknown) and this is the focus of this work. Current methods for solving Type 1 and 2 problems include Deep Learning (DL) methods, which benefit from extensive hardware and software support but have limited guarantees.

[0095] Other potential avenues for addressing Type 3 problems include causal inference methods, probabilistic graphs, and sparse regression methods. However, it is important to note that their application to these problems necessitates additional assumptions. Causal inference models, for instance, typically assume randomized data and some level of access to the data generation process or its underlying distributions. Unlike causal inference models, systems and methods in accordance with certain embodiments of the invention seek to encode the functional dependencies between variables into the structure of the graph. Sparse regression methods, on the other hand, rely on the assumption that functional dependencies have a sparse representation within a known basis. Systems and methods in accordance with many embodiments of the invention do not rely on such assumptions. Furthermore, while the complexity of Bayesian causal inference methods may grow super-exponentially with the number of variables, the complexity of processes in accordance with several embodiments of the invention isthat of parallel computations of polynomial complexities bounded between (best case) and(worst case).

[0096] A detailed analysis of the computational demands of systems and methods in accordance with certain embodiments of the invention as a function of the number of variables, denoted as , and the number of samples, , pertaining to these variables follows. In the worst case, the proposed approach necessitates, for each of the variables: for , regressing a function mapping variables to the variable of interest and performing a mode decomposition to identify the variable with the minimal contribution to the signal. Since these two steps have the same cost, it follows that, in the worst case, the total computational complexity of the proposed method iswhich corresponds to the product of the number of double-looping operations,, and the cost of kernel regression from samples which, without acceleration, is(i.e., the cost of inverting a dense kernel matrix).

[0097] However, if kernel scalability techniques are utilized, such as when the kernel has a rank (for example, if the kernel is linear) or is approximated by a kernel of rank (e.g., via a random feature map), then this worst-case bound can be reduced toby reducing the complexity of each regression step fromtoNote that the statistical accuracy of the proposed approach requires thatif the dependence of the unknown functions on their inputs is not sparse. Moreover, in the absence of kernel scalability techniques, the worst-case memory footprint of the method isdue to the necessity of handling dense kernel matrices. However, once the functional ancestors of each variable are determined, these matrices can be discarded. Consequently, only one such matrix needs to be retained in memory at any given time. 1. Processes for Hypergraph Discovery

[0098] A process for hypergraph discovery in accordance with an embodiment of the invention is conceptually illustrated in Fig. 2. Process 200 receives (205) input data. Input data in accordance with many embodiments of the invention includes sample values for multiple different variables. In many embodiments, input data includes nodes for a hypergraph, where each node represents a different variable. Nodes of the input data in accordance with certain embodiments of the invention may be disconnected. Input datain accordance with numerous embodiments of the invention include samples, where different samples may include values for different subsets (or all) of the variables. In a number of embodiments, input data is encoded into a matrix of values for each (or for subsets) of the variables (or nodes). Processes in accordance with a number of embodiments of the invention normalize the data via various methods, including (but not limited to) affine transformations.

[0099] Process 200 determining (210) a relationship between a selected node and a set of candidate ancestors. In a number of embodiments, candidate ancestors for each node may include all the remaining nodes from the input data. In a variety of embodiments, determining a relationship includes selecting a kernel for the relationship. Kernels in accordance with numerous embodiments of the invention can include (but are not limited to) linear, quadratic, and / or nonlinear kernels. Processes in accordance with several embodiments of the invention can select a kernel in an iterative manner, determining whether a kernel can define a function between the candidate ancestors and the selected node with a value that meets (or falls below / above) a threshold. In several embodiments, processes may evaluate multiple kernels and select a best kernel (e.g., best fit). In certain embodiments, the value includes (but is not limited to) a signal-to-noise ratio (SNR), a noise-to-signal ratio (NSR), etc. Thresholds in accordance with many embodiments of the invention can include (but are not limited to) a fixed value (e.g., 0.5) or a computed value. Selecting a kernel in accordance with many embodiments of the invention includes approximating the value of the selected node as a sum of noise and a function of the values of other variables in a given set of ancestors. Kernel selection and determining an SNR to decide whether a node has ancestors in accordance with several embodiments of the invention is described in greater detail below (e.g., with reference to Fig.3).

[0100] Process 200 prunes (215) the set of candidate ancestors to identify a set of minimal ancestors. In certain embodiments, pruning (or removing) candidate ancestors includes calculating an SNR (or NSR) for the selected node based on the selected kernel and the input data. Processes in accordance with certain embodiments of the invention can perform the pruning in an iterative manner, removing variables that are least contributing to the signal until the value (e.g., SNR, NSR, etc.) meets (or fallsbelow / above) a threshold. In several embodiments, thresholds can include (but are not limited to) a fixed value (e.g., 0.5) or a computed value (e.g., based on the number of candidate ancestors). Processes in accordance with various embodiments of the invention may use a first threshold for selecting a kernel and a second different threshold for pruning the set of candidate ancestors. In some embodiments, pruning ancestors includes computing an NSR for multiple subsets and selecting the set of minimal ancestors based on the computed NSRs. Selecting the set of minimal ancestors in accordance with a number of embodiments of the invention is based on an inflection point in the NSRs as a function of the number of ancestors in the subset. In several embodiments, pruning ancestors includes identifying redundant candidate ancestors and removing redundant candidate ancestors from the set of candidate ancestors. Redundant candidate ancestors in accordance with a number of embodiments of the invention may be ancestors that make the same or otherwise redundant contributions to the SNR (or NSR). In several embodiments, identifying redundant candidate ancestors is based on the determined relationships. Pruning ancestors in accordance with many embodiments of the invention is described in greater detail below (e.g., with reference to Figs.4-5).

[0101] Process 200 determines (220) whether there are more nodes for which ancestors need to be identified. When the process determines (220) there are more nodes, the process returns to step 210 and selects a kernel for the next selected node. When the process determines (220) there are no more nodes, the process generates (225) the hypergraph based on the selected kernels and relationships between nodes and their minimal ancestors.

[0102] In certain embodiments, similar processes may be used at multiple hierarchies (or levels). For example, processes in accordance with a number of embodiments of the invention may identify clusters and identify ancestors with processes similar to those described throughout this description to identify relationships between the identified clusters. In various embodiments,

[0103] While specific processes for hypergraph discovery are described above, any of a variety of processes can be utilized for hypergraph discovery as appropriate to the requirements of specific applications. In certain embodiments, steps may be executed or performed in any order or sequence not limited to the order and sequence shown anddescribed. In a number of embodiments, some of the above steps may be executed or performed substantially simultaneously where appropriate or in parallel to reduce latency and processing times. In some embodiments, one or more of the above steps may be omitted.

[0104] Kernel selection in accordance with a number of embodiments of the invention can be used to determine relationships between nodes of a hypergraph. A process for kernel selection in accordance with an embodiment of the invention is conceptually illustrated in Fig.3. Process 300 identifies (305) a candidate kernel. Kernels in accordance with numerous embodiments of the invention can include (but are not limited to) linear, quadratic, and / or nonlinear kernels.

[0105] Process 300 computes (310) an SNR for a node and a set of candidate ancestors based on the identified kernel. Computing an SNR in accordance with several embodiments of the invention includes selecting a noise prior parameter and performing a regression analysis. In a variety of embodiments, the noise prior parameter is selected based on characteristics of the kernel (e.g., whether the kernel is a universal kernel or whether it is derived from a finite-dimensional feature map). Performing a regression analysis can include predicting a value of the node based on the set of candidate ancestors.

[0106] Process 300 determines (315) whether the computed SNR meets a threshold. The threshold in accordance with some embodiments of the invention can be a fixed value (e.g., 0.5). When the process determines (315) that the threshold has been met, the process selects (320) the identified kernel. In various embodiments, the identified kernel can be used to prune the set of candidate ancestors. When the process determines (315) that the threshold has not been met, the process determines (325) whether there are more kernels to analyze. When the process determines (325) that there are more kernels, the process returns to step 305 to identify the next kernel. When the process determines (325) that there are no more kernels, the process removes (330) all candidate ancestors for the node.

[0107] While specific processes for kernel selection are described above, any of a variety of processes can be utilized for kernel selection as appropriate to the requirements of specific applications. In certain embodiments, steps may be executed or performed inany order or sequence not limited to the order and sequence shown and described. In a number of embodiments, some of the above steps may be executed or performed substantially simultaneously where appropriate or in parallel to reduce latency and processing times. In some embodiments, one or more of the above steps may be omitted.

[0108] Pruning ancestors in accordance with many embodiments of the invention removes ancestors to increase the signal-to-noise ratio (SNR) of the functional relationship between each node and its ancestors. A process for pruning ancestors based on a threshold in accordance with an embodiment of the invention is conceptually illustrated in Fig.4. In several embodiments, process 400 may be performed once a kernel has been selected and candidate ancestors for a given node have been identified. Process 400 identifies (405) a least important ancestor from the candidate ancestors. In various embodiments, the least important ancestor is the ancestor that contributes the least to the signal.

[0109] Process 400 computes (410) the SNR for a node and the candidate ancestors without the identified least important ancestor. Computing an SNR in accordance with several embodiments of the invention includes selecting a noise prior parameter and performing a regression analysis. In a variety of embodiments, the noise prior parameter is selected based on characteristics of the kernel (e.g., whether the kernel is a universal kernel or whether it is derived from a finite-dimensional feature map). Performing a regression analysis can include predicting a value of the node based on the set of candidate ancestors (without the identified least important ancestor). Computing SNRs is described in greater detail below.

[0110] Process 400 determines (415) whether to prune the identified ancestor. In a number of embodiments, processes determine whether to prune an ancestor based on whether the computed SNR falls below (or meets) a threshold. Processes in accordance with many embodiments of the invention seek to identify the minimal set of ancestors that maintains an SNR above the threshold.

[0111] When the process determines (415) to prune the identified ancestor, the process returns to step 405 to identify the next least important ancestor and repeat the process. When the process determines (415) that the identified ancestor should not bepruned (e.g., when the SNR drops below a threshold), the process identifies (420) the remaining candidate ancestors as the ancestors for the given node.

[0112] In certain embodiments, ancestors may be pruned by identifying an optimal set of ancestors. A process for pruning ancestors based on a difference between successive noise-to-signal ratios in accordance with an embodiment of the invention is conceptually illustrated in Fig. 5. Process 500 computes (505) a noise-to-signal ratio (NSR) for a node and its candidate ancestors. Computing an NSR in accordance with a number of embodiments of the invention includes selecting a noise prior parameter and performing a regression analysis with the selected parameter. In a variety of embodiments, the noise prior parameter is selected based on characteristics of the kernel (e.g., whether the kernel is a universal kernel or whether it is derived from a finite- dimensional feature map). Performing a regression analysis can include predicting a value of the node based on the set of candidate ancestors. Computing NSRs is described in greater detail below.

[0113] Process 500 identifies (510) a subset of the candidate ancestors. In certain embodiments, the subset of the candidate ancestors is identified by removing one or more of the least important candidate ancestors.

[0114] Process 500 computes (515) the NSR of the subset of candidate ancestors. Process 500 determines (520) whether there are more subsets to process. When the process determines (520) there are more subsets to process, the process returns to step 505 to identify the next least important ancestor and repeat the process. When the process determines (515) that the identified ancestor should not be pruned (e.g., when the SNR drops below a threshold), the process identifies (520) the remaining candidate ancestors as the ancestors for the given node. In a variety of embodiments, identifying the particular subset includes identifying an inflection point in the NSRs as a function of a number of ancestors in the subset of candidate ancestors.

[0115] While specific processes for pruning ancestors are described above, any of a variety of processes can be utilized for pruning ancestors as appropriate to the requirements of specific applications. In certain embodiments, steps may be executed or performed in any order or sequence not limited to the order and sequence shown and described. In a number of embodiments, some of the above steps may be executed orperformed substantially simultaneously where appropriate or in parallel to reduce latency and processing times. In some embodiments, one or more of the above steps may be omitted. 2. Examples of Hypergraph Discovery (Type 3) Problems

[0116] While most problems in Computer Science and Engineering can be framed and resolved as computational graph or hypergraph completion tasks, some scientific challenges necessitate first uncovering the hypergraph’s structure before completing it. Systems and methods in accordance with certain embodiments of the invention can be used to tackle challenges in various areas. Examples of these challenges that illustrate the scope and efficiency of the proposed approach are described below. Examples of challenges that can be addressed with hypergraph discovery in accordance with various embodiments of the invention are illustrated in Fig.6.

[0117] One such problem is the recovery of chemical reaction networks from concentration snapshots. As an example consider the hydrogenation of ethyleneinto ethane A proposed mechanism for this chemical reaction, represented byisThe problem is that of recovering the underlying chemical reaction network from snapshots of concentrations ଶ , , ଶ ସ and ଶ ହ and their time derivativesThe proposed approach leads to a perfect recovery of thecomputational graph and a correct identification of quadratic functional dependencies between variables. A resultant computational graph of the recovery of a chemical reaction network is illustrated in Fig.7.

[0118] Another problem is data analysis. As an example, consider the analysis of COVID-19 data from GOOGLE. Data from a single country, France, was used to ensure consistency in the data and to avoid considering cross-border variations that are not directly reflected in the data. Thirty one variables were selected that describe the state of the country during the pandemic, spanning over 500 data points, with each data point corresponding to a single day. These variables were categorized as the following datasets: Epidemiology dataset: Includes quantities such as new infections, cumulative deaths, etc. Hospital dataset: Provides information on the number of admitted patients, patients in intensive care, etc. Vaccine dataset: Indicates the number of vaccinated individuals, etc. Policy dataset: Consists of indicators related to government responses, such as school closures or lockdown measures, etc. The problem is then to analyze this data and identify possible hidden functional relations between these variables.

[0119] Another problem is equation discovery. As an example, letandbe six variables such that therandom variables, andsatisfy Then, given observations of (samplesfrom) those variables, consider the problem of discovering hidden functional relationships or dependencies between these variables. (see Fig.6(a)).

[0120] An open problem in computational biology is to identify pathways and interactions between genes from their expression levels. One example stems from phosphorylated proteins and phospholipids in the human immune system T-cells. This dataset consists of proteins and samples of their expressions. Fig. 6(b) presents the (reference) causal dependencies found using experiments and structure learning algorithms for Bayesian networks, i.e., the directed acyclic graph from the cell- signaling data.

[0121] Another example (see Fig.6(c)) is the modeling of land surface interactions in weather prediction, and the problem is to discover possibly hidden functional dependencies between state variables for a finite number of snapshots of those variables.

[0122] Another example (see Fig. 6(d)) is the recovery and analysis of economic networks, and the problem is to discover functional dependencies between the economic markers of different agents or companies, which is significant to systemic risk analysis.

[0123] Another example (see Fig. 6(e)) concerns a graph analysis of the stock market, and the problem is to discover functional dependencies between the price returns of different stocks.

[0124] Another example (see Fig. 6(f)) is social network analysis / discovery, and the problem is to discover functional dependencies between quantitative markers associated with each individual (a vector of numbers per individual) in situations where the connectivity of the network may be hidden.

[0125] Another example is in autonomous vehicles, particularly in space. When vehicles are launched out into space, it can be difficult to train and account for the various situations that a vehicle may face. In addition, due to the distances, it can be difficult and slow to send additional information or to make adjustments to the vehicle. In several embodiments, vehicles may utilize hypergraph discovery to learn and adjust when current models are no longer accurate. By identifying new relationships between variables based on its new experiences, an autonomous vehicle may update its own programming and better adjust to unexpected environments. 3. Overview of the proposed approach for Type 3 problems

[0126] In this section, an overview of the proposed approach for Type 3 problems is presented. An example of a hypergraph discovery process in accordance with an embodiment of the invention is illustrated in Fig.8. For ease of presentation, consider the simple setting of Fig. 8(a), given samples on the variables. After measurements / collection, these variables are normalized to have zero mean and unit variance. The objective is to uncover the underlying dependencies between them.a. A signal-to-noise ratio to decide whether or not a node has ancestors.

[0127] Hypergraph discovery in accordance with certain embodiments of the invention identifies ancestors for each node in the graph. For a specific node, say, as depicted in Fig.8(b). Determining whether^has ancestors is akin to asking if can beexpressed as a function of In other words, can a function (living in a pre-specified space of functions that could be of controlled regularity) be found such that:To answer this question, regress with a centered whose covariancefunction is an additive kernel of the formwhere is a smoothing kernel,is the white noise covarianceoperator. This is equivalent to assuming the GP to be the sum of two independent GPs, i.e.,is a smoothing / signal GP andis a noise GP. Writingೞfor the Reproducing Kernel Hilber Space (RKHS) induced by the kernel this is also equivalent to approximating with a minimizer ofwhere is the Euclidean norm onis the input data on obtained as an - vector (or an N x M matrix) whose entries (or rows)^are the samples on, is the output data on obtained as an -vector whose entries are obtained from the samples on^, and is a -vector whose entries are the evaluations^. At the minimumquantifies the data variance explained by the signalandquantifies the data variance explained by the noiseThis allows us to define the signal-to-noise ratiothen, as illustrated in Fig. 8(c), has no ancestors, i.e., cannot be^ ^approximated as a function of. Conversely,, then deduce that^has ancestors, i.e.,^can be approximated as a function ofଶ ^.

[0128] Although many of the examples described herein refer to a signal-to-noise (or noise-to-signal) ratio, one skilled in the art will recognize that these ratios maybe used interchangeably and / or other measures of functions can be used in a variety of applications without departing from this invention. b. Selecting the signal kernel^.

[0129] In some embodiments, selecting the signal kernel is performed by selecting the kernel ^ to be linearquadratic(or fully nonlinear (not linear and not quadratic) to identifyas linear, quadratic, or nonlinear. Hypergraph discovery processes in accordance with many embodiments of the invention may iterate through kernels in a specified order (e.g., linear, then quadratic, then nonlinear), in order to identify the simplest kernel that meets a threshold. Nonlinears may also be referred to as interpolatory as they may be selected as having an infinite-dimensional feature map. They may also be referred to as universal kernels. Kernels may also be referred to as interpolatory when the number of data points is smaller than the dimension of the feature map.

[0130] In the case of a nonlinear kernel, the following kernel may be employed:where is a universal kernel, such as (but not limited to) a Gaussian or a Matérn kernel, with all parameters set to , and ^ assigned the default value . In some embodiments,^may be selected as the first kernel that surpasses a signal-to-noise ratio (e.g., 0.5). Ifno kernel reaches this threshold, it may be determined that^lacks ancestors. c. Pruning ancestors based on signal-to-noise ratio.

[0131] Once it is established that^has ancestors (e.g., via kernel selection), its set of ancestors may be pruned iteratively in accordance with a variety of embodimentsof the invention. An example of pruning ancestors in accordance with an embodiment of the invention is illustrated in Fig. 9. In numerous embodiments, nodes with the least contribution to the signal-to-noise ratio can be removed, stopping before that ratio drops below a threshold (e.g., ). To describe this, assume that^is as in (Eq.1). Then^is an additive kernel that can be decomposed into two parts: ^^ ଶdoes not depend onଶandଶ ^ ^depends onଶ. This decomposition allows the expression of as the sum of two components:does not depend onଶ,ଶdepends onଶand Furthermore,and quantifies the contribution ofଶto the signal data variance.

[0132] Following the procedure illustrated in Fig. 9, if, for example,ସis found to have the least contribution to the signal data variance, the signal-to-noise ratio can be recomputed withoutସin the set of ancestors for^. If that ratio is below ,ସis not removed from the list of ancestors, andଶ ଷ ସ ହ ^is the final set of ancestors of^. If this ratio remains above ,ସis removed. This iterative process continues, and stops before the signal-to-noise ratio drops below to identify the final list of ancestors of^.

[0133] In a variety of embodiments, hypergraph discovery does not use a fixed threshold on the signal-to-noise ratio to prune ancestors, but rather employs an inflection point in the noise-to-signal ratio as a function of the number of ancestors.To put it simply, after ordering the ancestors in decreasing contribution to the signal, the final number of ancestors is determined as the maximizer of4. Hypergraph Discovery

[0134] In this section, an accessible overview of key components of hypergraph discovery in accordance with numerous embodiments of the invention is discussed. Processes in accordance with a variety of embodiments of the invention focus on determining the edges within a hypergraph. To achieve this, processes in accordance with certain embodiments of the invention consider each node individually, find its ancestors, and establish edges from these ancestors to the node in question. While processes are described throughout this specification are often described for a single node, one skilled in the art will recognize that processes can be applied iteratively to all nodes within the graph. Processes in accordance with a number of embodiments of the invention include:

[0135] 1. Initialization: Assume that all other nodes are potential ancestors of the current node.

[0136] 2. Selecting a Kernel: Choose a kernel function, such as linear, quadratic, or fully nonlinear kernels (refer to Example 1). The kernel selection process is analogous to the subsequent pruning steps, involving the determination of a parameter , regression analysis, and evaluation based on signal-to-noise ratios. In various embodiments, if the signal-to-noise ratio is insufficient for all possible kernels, the process terminates, indicating that the node has no ancestors.

[0137] 3. Pruning Process: While there are potential ancestors left to consider:

[0138] (a) Identify the Least Important Ancestor: Ancestors are ranked based on their contribution to the signal.

[0139] (b) Noise prior: Determine the value of .

[0140] (c) Regression Analysis: Predict the node’s value using the current set of ancestors, excluding the least active one (i.e., the one contributing the least to the signal). Processes in accordance with several embodiments of the invention employ Kernel Ridge Regression with the selected kernel function and parameter .

[0141] (d) Evaluate Removal: Compute the regression signal-to-noise ratio. If the signal-to-noise ratio falls below a certain threshold, terminate the process and return thecurrent set of ancestors. If the signal-to-noise ratio is sufficient, remove the least active ancestor and continue the pruning process.

[0142] Throughout this description, processes in accordance with several embodiments of the invention may be further refined and generalized, as discussed in the following subsections.

[0143] Complexity Reduction with Kernel PCA Variant. To improve efficiency, processes in accordance with various embodiments of the invention introduce a variant of Kernel PCA . This modification significantly reduces the computational complexity of hypergraph discovery, making it primarily dependent on the number of principal nonlinear components in the underlying kernel matrices rather than the number of data points.

[0144] Generalizing Descendants and Ancestors with Kernel Mode Decomposition. In several embodiments, the concept of descendants and ancestors can be extended to cover more complex functional dependencies between variables, including implicit ones, as discussed in greater detail below. This generalization can be achieved through a Kernel-based adaptation of Row Echelon Form Reduction, initially designed for affine systems, and leveraging the principles of Kernel Mode Decomposition.

[0145] Parameter Selection. The choice of the parameter can be a critical aspect of hypergraph discovery. Processes in accordance with a number of embodiments of the invention provide a structured approach for selecting based on the characteristics of the kernel matrix^. Specifically, when^is derived from a finite-dimensional feature map , the regression residual can be employed to determine as follows:as discussed below. Alternatively, when^is a universal kernel, can be selected by maximizing the variance of the eigenvalue histogram of asdiscussed in greater detail below.

[0146] Ancestor pruning. Hypergraph discovery in accordance with certain embodiments of the invention does not use a fixed threshold on the signal-to-noise ratio to prune ancestors, but rather employs an inflection point in the noise-to-signal ratioas a function of the number of ancestors. To put it simply, after ordering the-27-ancestors in decreasing contribution to the signal, the final number of ancestors is determined as the maximizer of

[0147] An example of hypergraph discovery processes in accordance with numerous embodiments of the invention is illustrated with an application to the recoveryof a mechanical network. The Fermi-Pasta-Ulam-Tsingou (FPUT) system is a prototypicalchaotic dynamical system. It is composed of a set of masses indexed by with equilibrium position with . Each mass is tethered to its two adjacent masses by a nonlinear spring, and the displacement of the mass^adheres to the equation:whereଶis a velocity depending on and the strength of the springs, , and . Fixed boundary conditions can be used by adding two more masses, withି^ ெ.

[0148] The selection of this model is driven by several aspects: its chaotic behavior, established and sparse equations, and the presence of linear / nonlinear equations contingent on the choice of . The system can be simulated with two different choices of : a nonlinear case whereଶ, and a linear case where .

[0149] A total of 1000 snapshots were taken from multiple trajectories and the observed variables are the positions, velocities, and accelerations of all the underlying masses. In the graph discovery phase, every other node is initially deemed a potential ancestor for a specified node of interest. The node with the least signal contribution was then iteratively removed. The step resulting in the largest surge in the noise-to-signal ratio is inferred as one eliminating a crucial ancestor, thereby pinpointing the final ancestor set.

[0150] An example of noise-to-signal ratioas a function of the number of proposed ancestors and noise-to-signal ratio incrementsas a function of the number of ancestors in accordance with an embodiment of the invention are illustrated in Fig. 10. -test quantiles are also plotted in chart 1005. In the absence of signal, the noise-to-signal ratio should fall within the shaded area with probability . Fig.10 includes chart 1005, which shows a plot of the noise-to-signal ratioas a function of the number of proposed ancestors for the variable and^with Z-test quantiles (in the absence of signal, the noise-to-signal ratio should fall within the shaded area with probability 0.9). Removing a node essential to the equation of interest causes the noise-to-signal ratio to markedly jump from approximately 25% to99%. Chart 1010 shows a plot of the noise-to-signal ratio increments^as a function of the number of ancestors for the variable ^. Note thatthe increase in the noise-to-signal ratio is significantly higher compared to previous removals when an essential node was removed. Therefore, while solely relying on a fixed threshold to decide when to cease the removals might prove challenging, evaluating the increments in noise-to-signal ratios can offer a clear guideline for efficiently and reliably pruning ancestors.

[0151] An example of result data from an analysis of a FPUT system is illustrated in Fig.11. The recovered full graph, depicted in Fig.11(a), is remarkably accurate despite the nonlinear nature of the model and the fact that the prior only encodes that the nonlinearity is smooth. Therefore, hypergraph discovery in accordance with some embodiments of the invention does not require a dictionary or extensive knowledge of the structure of the unknown functions. Notably, velocity variables are accurately identified as non-essential and omitted from the ancestors of position and acceleration variables. Fig. 11(b), which omits velocity variables for clarity, further elucidates the accurate recovery of dependencies. The dependencies are the simplest and clearest possible. They match exactly those of the original equations except for the boundary particles for which valid equivalent equations are recovered.

[0152] While performing these experiments, a set threshold of 60% on the noise- to-signal ratio could be used to obtain the desired equations. However, this choice could be seen as arbitrary or guided by knowledge of the true equations. However, the choice to retain or discard an ancestor becomes clear when examining the noise-to-signal ratio’s evolution against the number of ancestors. In the graph discovery phase, every other node is initially deemed a potential ancestor for a specified node of interest. The node with the least signal contribution is then iteratively removed. The step resulting in thelargest surge in the noise-to-signal ratio is inferred as one eliminating a crucial ancestor, thereby pinpointing the final ancestor set. In the context of the FPUT problem, removing a node essential to the equation of interest causes the noise-to-signal ratio to markedly jump from approximately 25% to 99%. Additionally, the increase in the noise-to-signal ratio during the removal process can be analyzed. Two representative trajectories illustrating this are presented in Fig. 10 and 12. Notably, in the final iteration where an essential node was removed, the increase in the noise-to-signal ratio is significantly higher compared to previous removals. This observation indicates that the process halted at the appropriate juncture, as the removal of this last node was not sanctioned. As shown in Fig.12, while solely relying on a fixed threshold to decide when to cease the removals might prove challenging, evaluating the increments in noise-to-signal ratios offers a clear guideline for efficiently and reliably pruning ancestors.

[0153] Type 3 problems are challenging and can even be intractable if notformalized and approached properly. First, the problem suffers from the curse ofcombinatorial complexity in the sense that the number of hypergraphs associated with nodes blows up rapidly with . A lower bound on the number of hypergraphs associated with 3 nodes is the A003180 sequence, which answers the following question: given unlabeled vertices, how many different hypergraphs in total can be realized on them bycounting the equivalent hypergraphs only once? For , this lower bound is^ଷ.

[0154] Secondly, it is important to note that, even with an infinite amount of data, the exact structure of the hypergraph might not be identifiable. To illustrate this point, consider a problem with samples from a computational graph with variables and . The task is to determine the direction of functional dependency between and . Doesthe dependency go from to or from to ? In some cases, it may be clear when canonly be expressed as a function of or when can solely be written as a function of . However, in some situations, it becomes challenging to decide as it may be possible to write both as a function of and as a function of . Further complicating matters is the possibility of implicit dependencies between variables. There may also be instanceswhere neither can be derived as a function of , nor can be represented as a function of .

[0155] The field of causal inference and probabilistic modeling seeks to address distinct graph learning problems by introducing auxiliary assumptions. Within noise and probabilistic models, it is generally assumed that the data is randomized.

[0156] Causal inference methods broadly consist of two approaches: constraint and score-based methods. While constraint-based approaches are asymptotically consistent, they only learn the graph up to an equivalence class . Instead, score-based methods resolve ambiguities in the graph’s edges by evaluating the likelihood of the observed data for each graphical model. For instance, they may assign a higher evidence to over if the conditional distribution exhibits less complexity than . The complexity of searching over all possible graphs, however, grows super- exponentially with the number of variables. Thus, it is often necessary to use approximate, but more tractable, search-based methods or alternative criteria based on sensitivity analysis . For example, the preference could lean towards rather than if demonstrates less sensitivity to errors or perturbations in . In contrast, processes in accordance with several embodiments of the invention avoid the growth in complexity by performing a guided pruning process that assesses the contribution of each node to the signal. Hypergraph discovery in accordance with numerous embodiments of the invention is not limited to learning acyclic graph structures as it can identify feedback loops between variables.

[0157] Alternatively, methods for learning probabilistic undirected graphical models, also known as Markov networks, identify the graph structure by assuming the data is randomly drawn from some probability distribution. In this case, edges in the graph (or lack thereof) encode conditional dependencies between the nodes. A common approach learns the graph structure by modeling the data as being drawn from a multivariate Gaussian distribution with a sparse inverse covariance matrix, whose zero entries indicate pairwise conditional independencies . Recently, this approach has been extended using models for non-Gaussian distributions, as well as kernel-based conditional independence tests. In a variety of embodiments, functional dependencies are learned rather than causality or probabilistic dependence. It is not assumed that thedata is randomized nor are strong assumptions imposed, such as additive noise models, in the data-generating process.

[0158] Compare the hypergraph discovery framework to structure learning for Bayesian networks and structural equation models (SEM). Letௗbe a random variable with probability density function that follows the autoregressive factorizationgiven a prescribed variable ordering. Structure learning forBayesian networks aims to find the ancestors of variable^, often referred to as the set of parents in the sense that^. Thus,the variable dependence of the conditional density^is identified by finding the parent set so that^is conditionally independent of all remaining preceding variables given its parents, i.e.,^Finding ancestors that satisfy this condition requires performing conditional independence tests, which are computationally expensive for general distributions . Alternatively, SEMs assume that each variable^is drawn as a function of its ancestors with additive noise, i.e,^for some function and noise . For Gaussian noiseeach marginal conditional distribution in a Bayesian network is given by. Thus, finding the parents for such a model by maximum likelihood estimation corresponds to finding the parents that minimize the expected mean-squared error While thegraph structure identified in Bayesian networks is influenced by the specific sequence in which variables are arranged (a concept exploited in numerical linear algebra where Schur complementation is equivalent to conditioning GPs and a carefully ordering leads to the accuracy of the Vecchia approximationthe graph recovered by hypergraph discovery in accordance with various embodiments of the invention remains unaffected by any predetermined ordering of those variables.

[0159] In a number of embodiments, a formulation of the problem is used that remains well-posed even when the data is not randomized, i.e., the problem can be formulated as the following manifold learning / discovery problem.

[0160] Problem 1 Let be a Reproducing Kernel Hilbert Space (RKHS) of functions mapping^to . Let be a closed linear subspace of be a subsetof^such that if and only if for all .

[0161] Given the (possibly noisy and nonrandom) observation of elements,approximate . To understand why problem 1 serves as the appropriateformulation for hypergraph discovery, consider a manifold. Suppose this manifold can be represented by a set of equations, expressed as a collection of functionssatisfyingTo keep the problem tractable, assume a certain levelof regularity for these functions, necessitating they belong to a RKHS , ensuring the applicability of kernel methods for our framework. Given that any linear combination of the^will also be evaluated to zero on , the relevant functions are those within the span of the^, forming a closed linear subspace of denoted as .

[0162] An example of (a) hyper formulation as a manifold discovery problem and hypergraph representation (b) the hypergraph representation of an affine manifold is equivalent to its Row Echelon Form Reduction in accordance with an embodiment of the invention is illustrated in Fig. 13. The manifold can be subsequently represented by agraph or hypergraph (see Fig. 13(a)), whose ambiguity can be resolved through a deliberate decision to classify some variables as free and others as dependent. This selection could be arbitrary, informed by expert knowledge, and / or derived from probabilistic models or sensitivity analysis.

[0163] Systems and methods in accordance with some embodiments of the invention include a kernel generalization of variance-based sensitivity analysis guiding the discovery of the structure of the hypergraph. In a variety of embodiments, GP variance decomposition of the data leads to signal-to-noise and a Z-score that can be employed to determine whether a given variable can be approximated as a nonlinear function of a given number of other variables.

[0164] To describe the proposed solution to Problem 1, start with a simple example. In this example, is a space of affine functions of the form

[0165] As a particular instantiation (see Fig. 13(b)), assume to be the manifold of ଷ ( ) defined by the affine equationswhich is equivalent to selectingandin the problem formulation 1.

[0166] Then, irrespective of how the manifold is recovered from data, the hypergraph representation of that manifold is equivalent to the row echelon form reduction of the affine system, and this representation and this reduction require a possibly arbitrary choice of free and dependent variables. So, for instance, for the system (Eq. 2), ifଷis declared to be the free variables and^andଶto be the dependent variables, then the manifold can be represented via the equationswhich have the hypergraph representation depicted in Fig.13(b).

[0167] Now, in the regime where the number of data points is larger than the number of variables, the manifold can simply be approximated via a variant of PCA in accordance with many embodiments of the invention. Takefor ୀ uivalent toே∗. Since , can be identified exactlyThen the obtained manifold iswhereis the zero-eigenspacecan be written for the eigenvalues of ே (in decreasing order), and ^ ௗା^ for thecorresponding eigenvectors Hypergraph discovery in accordance with anumber of embodiments of the invention extends to the noisy case (when the data points are perturbations of elements of the manifold) by simply replacing the zero-eigenspace of the covariance matrix by the linear span of the eigenvectors associated with eigenvalues that are smaller than some threshold , i.e., by approximating with (Eq.4) where is such that . In this affine setting, (Eq. 4) allows the estimation of directly without RKHS norm minimization / regularizationbecause linear regression does not require regularization in the sufficiently large data regime. Furthermore, in some embodiments, the process of pruning ancestors can be replaced by that of identifying sparse elements

[0168] This simple approach can be generalized in accordance with numerous embodiments of the invention by generalizing the underlying feature map used to define the space of functions (writing^for the dimension of the range of )An example of feature map generalization in accordance with an embodiment of the invention is illustrated in Fig.14.

[0169] For instance, using the feature mapthen becomes a space of quadratic polynomials onௗ, i.e.,and, in the large data regime ( ^), identifying quadratic dependencies betweenvariables becomes equivalent to (1) adding nodes to the hypergraph corresponding to secondary variables obtained from primary variables^through known functions (for (Eq. 5), these secondary variables are the quadratic monomials^ ^, see Fig. 14(a)), and (2) identifying affine dependencies between the variables of the augmented hypergraph. The problem can, therefore, be reduced to the affine case. Indeed, as in the affine case, the manifold can then be approximated in the regime where the number of data points is larger than the dimension^of the feature map by (Eq. 4), where^ ேare the eigenvectors of( q ) whose eigenvalues are zero (noiseless case) or smaller than some threshold(noisy case).

[0170] Furthermore, the hypergraph representation of the manifold is equivalent to a feature map generalization of Row Echelon Form Reduction to nonlinear systems of equations. For instance, choosingଷas the dependent variable and^ ଶas the free variables,can be represented as in Fig. 14(b) where the round node represents the concatenated variable^ ଶand the dashed arrow represents a quadratic function. In several embodiments, the generalization also enables the representation of implicit equations by selecting secondary variables as free variables.

[0171] This feature-map extension of the affine case can evidently be generalized to arbitrary degree polynomials and to other basis functions. However, as the dimension^of the range of the feature map increases beyond the number of data points, theproblem becomes underdetermined: the data only provides partial information about the manifold, i.e., it is not sufficient to uniquely determine the manifold. Furthermore, if the dimension of the feature map is infinite, then the problem is always in that low data regime, and there is the additional difficulty that we cannot directly compute with that feature map. On the other hand, if^is finite (i.e., if the dictionary of basis functions is finite), then some elements of (some constraints defining the manifold) may not be representable or well approximated as equations of the formTo address these conflicting requirements, processes in accordance with numerous embodiments of the invention may be kernelized and / or regularized (as done in interpolation).

[0172] The kernel associated with the feature map. To describe this kernelization,

[0173] Observe that is a Hilbert space endowed with the inner productfor the kernel defined by and observe that is the RKHS defined by the kernel(which is assumed to containin Problem 1). Observe in particular that forsatisfies the reproducing property

[0174] Complexity Reduction with Kernel PCA Variant. In some embodiments, the feature-map PCA variant (characterizing the subspace ofsuch that ) can be kernelized as a variant of kernel PCA . To describe this, write for the matrix with entriesWrite for the nonzero eigenvalues ofndexed in decreasing order and write⋅,^for the corresponding unit-normalized eigenvectors, i.e., , ,

[0175] Write for the vector with entrieswriteand

[0176] Writefor the vector with entries.

[0177] Proposition 6.1 The subspace of functionssuch thatis equal to the subspacesuch thatFurthermore,with feature map representationwith we have the identity (where

[0178] f gே( q 3) indexed in decreasing order. Writefor the corresponding eigenvectors, i.e.,

[0179] Observing that

[0180] (Eq.9) implies that forusing the reproducing property (Eq.6) of in the last identity. Write

[0181] Using (Eq. 9) withimplies thatis an eigenvector of the matrix with eigenvalue implies thatTherefore, the,are unit-normalized. Summarizing, this analysis (closely related to the one found in kernel PCA ) shows that the nonzero eigenvalues of coincide withThe identity (Eq.12) then implies (Eq.7).

[0182] Remark As in PCA the dimension / complexity of the problem can be further reduced by truncating is identified as the smallest index such thatis some small threshold.

[0183] Kernel Mode Decomposition. When the feature mapis infinite- dimensional, the data only provides partial information about the constraints defining the manifold in the sense that or equivalently is a necessary, but not sufficient, condition for the zero level set of to be a valid constraint for the manifold (for to be such that for all ). This leads to the following problems: (1) How to regularize? (2) How to identify free and dependent variables? (3) How to identify valid constraints for the manifold? In some embodiments, in order to address such problems,hypergraph discovery may be based on the Kernel Mode Decomposition (KMD) framework.

[0184] A quick reminder on KMD in the setting of the following mode decomposition problem. So, in this problem, an unknown functionmaps some input spaceto the real line This function can be written as a sum of other unknown functions (ormodes), i.e.,

[0185] Assume each modeற^ to be an unknown element of somedefined by some kernel^. Then consider the problem in which given the data(withே ே), the modes composing the target functionare to be approximate.

[0186] Theorem 6.3 Using the relative error in the product normas a loss, the minimax optimal recovery ofି^where is the additive kernel

[0187] The GP interpretation of this optimal recovery result is as follows. Let^be independent centered GPs with kernels. Writefor the additive GP(Eq. 13) can be recovered by replacing the modesby independent centeredGs with kernels^and approximating the mode by conditioningon the available datais the additive GP obtained by summing the independent GPs, i.e.,

[0188] Furthermore can also be identified as the minimizer of^

[0189] The variational formulation (Eq. 14) can be interpreted as a generalization of Tikhonov regularization which can be recovered by selecting, to be a smoothing kernel (such as a Matérn kernel) and to be a white noisekernel.

[0190] Now, this abstract KMD approach is associated with a quantification of how much each mode contributes to the overall data or how much each individual GP^explains the data. More precisely, the activation of the mode or GP^can be quantified as whereThese activationsyୀthey can be thought of as a generalization of Sobol sensitivity indices to the nonlinear setting in the sense that they are associated with the following variance representation / decomposition (writing for the RKHS inner product induced by ):

[0191] Now, looking back to the original manifold approximation problem 1 in the kernelized setting. Given the data , we cannot regress an elementdirectly since the minimizer ofthe null function. To identify the functions, they need to be decomposed into modes that can be interpreted as a generalization of the notion of free and dependent variables. To describe this, assume that the kernel can be decomposed as the additive kernel

[0192] Then^implies that for all function can be decomposed.

[0193] Example 1 As a running example, take to be the following additive kernelthat is the sum of a linear kernel, a quadratic kernel and a fully nonlinear kernel. Take^to be the part of the linear kernel that depends only on^, i.e.,¬ᇱ^ ^ ^ ^Take^to be the part of the kernel that does not depend on^, i.e.,And take௭to be the remaining portion,

[0194] Therefore the following questions are equivalent: Given a functionin the RKHSdefined by the kernel ^ is there a function^in the the RKHSdefined by the kernel^ such thatGiven a functionೌis there a function in the RKHSdefined by the kernel such that and such that itsmode is zero?

[0195] Then, the natural answer to these questions is to identify the modes of the constraint^ andto be the minimizer of the following variational problem

[0196] This is equivalent to introducing the additive GPwhose modes are the independent GPs,௭ ௭, (use the label “n” in reference to “noise”), and then recovering^as

[0197] Takingfor running example 1, the questions are equivalent to asking whether there exists a function that does not depend on^(since^doesnot depend on^) such that^ ^ ଶTherefore, the mode can be thought of as a dependent mode (use the label “a” inreference to “ancestors”), the mode^as a free mode (use the label “s” in reference to “signal”), the mode௭as a zero mode.

[0198] While the numerical illustrations have primarily focused on the scenario where^takes the form of , and the aim is to express^as a function of othervariables, the generality of hypergraph discovery processes in accordance with a number of embodiments of the invention is motivated by their potential to recover implicit equations. For example, consider the implicit equationwhich can be retrieved by setting the mode of interest to be to depend onlyon the variable

[0199] Signal-to-noise ratio. Now, the following question: since the mode^(the minimizer of (Eq. 18)) always exists and is always unique, how is it know that it leads to a valid constraint? To answer that question, compute the activation of the GPs used to regress the data. Write for the activation of the signal Gfor the activation of the noise and then these allow a signal-to-noise ratio to bedefined as

[0200] Note that this corresponds to activation ratio of the noise GP defined in (Eq. 15). This ratio can then be used to test the validity of the constraint in the sense that ifas a prototypical example), then the data is mostly explained by the signal GP and the constraint is valid. , then thedata is mostly explained by the noise GP and the constraint is not valid.

[0201] Iterating by removing the least active modes from the signal. If the constraint is valid, then the activation of the modes composing the signal can be computed. To describe this, assume that the kernel can be decomposed as the additive kernel,which results in,, which results in the fact tha^can bedecomposed aswith, ೞ, The activation of the mode can then be quantified aswhich combined, leads to

[0202] For running example 1,^(Eq.17) can be decomposed as the sum of an affine kernel, a quadratic kernel, and a fully nonlinear kernel,.

[0203] As another example for the running example,^can be taken to be the sum of the portion of the kernel that does not depend on^andଶand the remaining portion, i^,ଶ ^ ^,^.

[0204] Then, these sub-modes can be ordered from most active to least active and create a new kernel^by removing the least active modes from the signal and adding them to the mode that is set to be zero. To describe this, letbe an ordering of the modes by their activation, i.e.,.

[0205] Writing for the additive kernel obtained from the leastactive modes (with as the value used for our numerical implementations), update the kernels ^ and ௭ by assigning the least active modes from௭, ,(zero the least active modes).

[0206] Finally, the process in accordance with a variety of embodiments of the invention can be iterative. This iteration can be thought of as identifying the structure of the hypergraph by placing too many hyperedges and removing them according to the activation of the underlying GPs.

[0207] For the running example 1, in an attempt to identify the ancestors of the variable if the sub-mode associated with the variableଶis found to be least active, then try to remove ଶ from the list of ancestors and try to identify as a function ofଷ toThis is equivalent to selectingand / to assess whether there exists a function^that does not depend on

[0208] Alternative determination of the list of ancestors. Systems and methods in accordance with certain embodiments of the invention determine the list of ancestors of a given node using a fixed threshold (e.g., ) to prune nodes. In some embodiments, an alternative approach that mimics the strategy employed in Principal Component Analysis (PCA) for deciding which modes should be kept and which ones should be removed is utilized. The PCA approach in accordance with several embodiments of the invention is to order the modes in decreasing order of eigenvalues / variance and (1) either keep the smallest number modes holding / explaining a given fraction (e.g., ) of the variance in the data, and / or (2) use an inflection point / sharp drop in the decay of the eigenvalues to select which modes should be kept. Hypergraph discovery processes in accordance with numerous embodiments of the invention employ an alternative determination of the least active mode by iteratively removing the mode that leads to the smallest increase in noise-to-signal ratio. In a variety of embodiments, the mode is removed such that,

[0209] For running example 1, in attempting to find the ancestors of the variable^, this is equivalent to removing the variables or node whose removal leads to thesmallest loss in signal-to-noise ratio (or increase in noise-to-signal ratio) by selecting

[0210] Next, this process can be iterated, and (a) the noise-to-signal ratio, and (b) the increase in noise-to-signal ratio can be plotted as a function of the number of ancestors ordered according to this iteration. Fig. 10 illustrates this process and shows that the removal of an essential node leads to a sharp spike in increase in the noise-to- signal ratio (the noise-to-signal ratio jumps from approximately 50-60% to 99%). Theidentification of this inflection point can be used as a method for effectively and reliably pruning ancestors.

[0211] Hypergraph discovery in accordance with many embodiments of the invention may be described as follows: (1) For each variable of the hypergraph, try to approximate its values as a sum of noise and a function of the values of other variables in a given set of ancestors. (2) If the signal-to-ratio exceeds a given threshold, then iteratively remove variables that are least contributing to the signal until the signal-to- noise ratio drops a given threshold.

[0212] Hypergraph discovery methods in accordance with a number of embodiments of the invention receive data (e.g., encoded into a matrix ) and a set of nodes as an input and produces, for each node , its set of minimal ancestors^and the simplest possible function^such thatIt employs the default threshold (e.g., ) on the signal-to-noise ratios for its operations. In several embodiments, processes normalize the data (e.g., via an affine transformation) so that the samples^are of mean zero and variance . Given a node with index (for ease of presentation), processes in accordance with a variety of embodiments of the invention select a signal kernel of the form, (where may be selected to be a vanilla RBF kernel such asGaussian or Matérn), with ଷ for the linear kernel,ଶଷfor the quadratic kerneland ଶ ଷ for the fully nonlinear (interpolative)kernel. In a variety of embodiments, processes compute the signal-to-noise ratio withnd with the selected kernel. In many embodiments, the value of is selectedautomatically by maximizing the variance of the histogram of eigenvalues of (with theselected kernel andIn a number of embodiments, the value of is re-computed whenever a node is removed from the list of ancestors, and isnonlinear. Processes in accordance with many embodiments of the invention iteratively identify the ancestor node contributing the least to the signal and remove that node from the set of ancestors of the node if the removal of that node does not send the signal-to- noise ratio below the default threshold (e.g., ).

[0213] In several embodiments, rather than using a default static threshold (e.g., 0.5), processes compute the noise-to-signal ratio (NSR), represented as .This ratio can be calculated as a function of the number of ancestors, which may be ordered based on their decreasing contribution to the signal. Further descriptions of computing the NSR in accordance with a number of embodiments of the invention are described throughout this specification. In various embodiments, the final number of ancestors can be determined by finding the value that maximizes the difference between successive noise-to-signal ratios,

[0214] In various embodiments, the signal-to-noise ratio (SNR) depends on the prior on the level of noise. The signal-to-noise ratio (Eq. 19) depends on the value of , which is the variance prior on the level of noise. The goal of this subsection is to answer the following two questions: (1) How to select (2) How to obtain a confidence level forthe presence of a signal? Or equivalently for a hyperedge of the hypergraph? To answer these questions, the signal-to-noise ratio is analyzed in the following regression problem in which the unknown function can be approximated based on noisy observationsof its values at collocation points, and the entries ^ of arei.i.d ). Assumingଶto be unknown and writing for a candidate for its value, recall that the GP solution to this problem is approximateby interpolating the data with the sum of two independent GPs, i.e.,where s the GP prior for the signal s the GP priorfor the noise in the measurements. can also be identified as a minimizer ofthe activation of the signal GP can be quantified as^, the activation of the noise GP can be quantified as ^ Define the noise-to-signal ratiowhich admits the following representer formula,Observe that when applied to the setting where mode^always exists and is always unique, this signal-to-noise ratio is calculated with^and^.

[0215] The following proposition follows from (Eq.24).

[0216] Proposition 8.1 It holds true that , and if has fullrank,

[0217] Therefore, if the signalற ଶand the level of noiseare both unknown, how should be selected to decide whether the data is mostly signal or noise?

[0218] In certain embodiments, selecting depends on whether the feature-map associated with the base kernel is finite-dimensional or not. In a number of embodiments, if the feature-map associated with the base kernel is finite-dimensional, then can be estimated from the data itself when the number of data-points is sufficiently large (at least larger than the dimension of the feature-space ). A prototypical example (when trying to identify the ancestors of the variable^) is^=( q In the general setting, assume thatwhere the range of is finite- dimensional. Assume thatறbelongs to the RKHS defined by , i.e., assume that it is of the form or some in the feature-space. Then (Eq. 21) reduces to

[0219] When the feature map is finite-dimensional, select

[0220] In certain embodiments, if the feature-map associated with the base kernel is infinite-dimensional (or has more dimensions than data points) then it can interpolate the data exactly and the previous strategy cannot be employed since the minimum of (Eq. 27) is zero. A prototypical example (when trying to identify the ancestors of the variableIn this situation, the level of noise may not be estimated, but a prior may be selected such that the resulting noise-to-signal ratio can effectively differentiate noise from signal. To describe this, observe that the noise-to- signal ratio (Eq.24) admits the representer formulaఊ↓ ఊ↑

[0221] Writefor the eigenpairs of^ areordered in decreasing order. Then the eigenpairs ofwhere

[0222] Note that the^are contained in and also ordered in decreasing order.

[0223] Writing^for the orthogonal projection of onto^,

[0224] It follows that if the histogram of the eigenvalues ofఊis concentrated near or near , then the noise-to-signal ratio is non-informative since the prior dominates it. To avoid this phenomenon, processes in accordance with certain embodiments of the invention select so that the eigenvalues ofఊare well spread out in the sense that the histogram of its eigenvalues has maximum or near-maximum variance. An example of histograms of eigenvalues ofఊare illustrated in Fig.15. In this example, histogram 1505 is more spread out and is a good choice for , while histogram 1510 is not spread out and a poor choice for . If the eigenvalues have an algebraic decay, then this is equivalent to taking to be the geometric mean of those eigenvalues.

[0225] In various embodiments, an optimizer can be used to obtain by maximizing the sample variance of If this optimization fails, processes inaccordance with some embodiments of the invention can default to the median of the eigenvalues. This can ensure a balanced, well-spread spectrum for , with half of the eigenvalues^being lower and half being higher than the median.

[0226] The purpose of this section is to present a rationale for the choices for in accordance with many embodiments of the invention. We present an asymptotic analysis of the signal-to-noise ratio in the setting of a simple linear regression problem. According to (Eq. 28), must scale linearly in ; this scaling is necessary to achieve a ratio that represents the signal-to-noise per sample. Without it (if remains bounded as a function of ), this scaling of the signal-to-noise would converge towards asTo see how, consider a simple example in which we seek to linearly regress the variable as a function of the variable , both taken to be scalar (in which case). Assume that the samples are of the form, the^are i.i.d. random variables, and the^^ satisfyே^. Then, the signal-to-noise ratio isand is a minimizer of

[0227] If is bounded independently from , then)converges towards zero as,which is undesirable as it does not represent a signal-to-noise ratio per samplewhich does not converge to, which is alsoundesirable. If is taken awhich converges towards which has,therefore, the desired properties.

[0228] When the kernel can interpolate the data exactly, (Eq. 27) should not be used to estimate the level of noise . For a finite-dimensional feature map , with data can be decomposed into a signal part ^ and noise part ^, s.t.While ^ belongs to the linear span of eigenvectors ofassociated withnon-zero eigenvalues,^also activates the eigenvectors associated with with the null space of and the projection of onto that null-space is what allows to be derived. Since in the interpolatory case, all eigenvalues are strictly positive, we need to choose which eigenvalues are associated with noise differently, as is described in the previous section. With a fixed which contributes in (Eq. 33) to yield a low noise-to-signal ratio. Similarly, i this eigenvalue yields a high noise- to-signal ratio. Thus, the choice of assigns a noise level to each eigenvalue. While in the finite-dimensional feature map setting, this assignment is binary, here soft thresholding can be performed using to indicate the level of noise of eacheigenvalue. This interpretation sheds light on the selection of in equation (Eq.28). Let represent the feature map associated withAssuming the empirical mean of iszero, the matrix corresponds to an unnormalized kernel covariance matrixConsequently, its eigenvalues correspond to times the variances of theacross various eigenspaces. After conducting Ordinary Least Squares regressionin the feature space, if the noise variance is estimated asଶ, then any eigenspace of the normalized covariance matrix whose eigenvalue is lower thanଶcannot be recovered due to the noise. Given this, the soft thresholding cutoff can be set to beଶfor the unnormalized covariance matrix

[0229] If the noise is only comprised of noise, then an interval of confidence can be obtained on the noise-to-signal ratio. To describe this consider the problem of testing the null hypothesis(there is no signal) against the alternative hypothesis(there is a signal). Under the null hypothesis ^, the distribution of the noise-to-signal ratio (Eq.29) is known and it follows that of the random variable

[0230] Therefore, the quantiles of can be used as an interval of confidence on the noise-to-signal ratio ifis true. More precisely, selectingsuch thatwithas a prototypical example, the noise-to-signal ratio (Eq. 29) would be expected, under, to be larger than with probability The estimation ofrequires Monte-Carlo sampling.

[0231] In some embodiments, an alternative approach (in the large data regime) to using the quantileis to use the Z-scoreafter estimating and via Monte-Carlo sampling. In particular, if ^ is true thenఈshould occur with probability with

[0232] Remark 8.2 Although the quantile^or the Z-score can be employed to produce an interval of confidence on the noise-to-signal ratio under^, they should not be used as thresholds for removing nodes from the list of ancestors. Indeed, observing a noise-to-signal ratio (Eq.29) below the threshold^does not imply that all the signal has been captured by the kernel; it only implies that some signal has been captured by the kernel . To illustrate this point, consider the setting where one tries to approximate the variable^as a function of the variable . If^is not a function ofଶ, but ofଶandଷ,as in^ ଶ ଷ, then applying the proposed approach with encoding the values of^, encoding the values ofଶ, and the kernel depending onଶcould lead to a noise-to-signal ratio below^due to the presence of a signal inଶ. Therefore, although the variableଷis missing in the kernel , a possibly low noise-to-signal ratio may still be observed due to the presence of some signal in the data. Summarizing, if the data onlycontains noise thenshould occur with probability. If the event^()ା (୬) is observed in the setting of / ( q ) where the ancestors ofare identified, then it can only be deduced that^contain some signal but perhapsnot all of it (processes in accordance with many embodiments of the invention may use this a criterion for pruningଶ).

[0233] Examples of the recovery of functional dependencies from data are illustrated in Fig.16. These examples illustrate the application of the proposed approach to the recovery of functional dependencies from data satisfying hidden algebraic equations. In all these examples, there are or variables and samples from those variables. For , the variables are . For , the variables are The samples from the variablesସarerandom variables, and the samples from arefunctionally dependent on the other variables.

[0234] In the example of Fig.16(a), and the samples from^andଶsatisfy the equationsFig. 16(a) shows the hypergraph recovered by selecting a kernel and pruning based on a difference between successive noise-to-signal ratios. The recovery is accurate and illustrates the one-to-one mappings between^, as well as betweenଶandଶ.

[0235] In Fig.16(b), and the samples from^ ଶandଷsatisfy the equationsFig. 16(b) shows the hypergraph recovered by hypergraph discovery processes as described in this description. Hypergraph discovery processes in accordance with a variety of embodiments of the invention select the quadratic kernel and Fig.16(b) shows the recovered graph (which is exact). The direct correspondence between^and^is correctly identified, as well as betweenଷandଷ. Yet, the efficiency of hypergraph discovery processes may be compromised when duplicate variables are present. For instance, given the quadratic kernel, even thoughଶ can trace back its origin to either^andଶor^andଶ, the hypergraph discovery process recognizes as itsancestors. This underscores the significance of eliminating redundant variables in accordance with a variety of embodiments of the invention when aiming to derive the sparsest and the most understandable graph.

[0236] In the example of Fig.16(c), and the samples fromଶsatisfythe equationsFig.16(c) shows the hypergraph recovered by the proposed approach using hypergraph discovery processes as described in this description. The process selects the nonlinear kernel and Fig. 16(c) shows the recovered graph (which is exact). The underlying functional dependencies are inherently nonlinear and were not explicitly represented in our feature maps. By utilizing the fully nonlinear kernel (Eq.16), the hypergraph discovery process identified the correct functional dependencies.

[0237] In the example of Fig. satisfy the equationsFig. 16(d) shows the hypergraph recovered with a hypergraph discovery process in accordance with several embodiments of the invention. Given the kernel selections (linear, quadratic, and fully nonlinear), one might anticipate an approximate recovery of the equations via the fully nonlinear kernel, especially as the equation appears to be cubic. However, the hypergraph discovery process astutely identifies an alternative combination, revealing hidden quadratic dependencies between variables and permitting an exact recovery with the quadratic kernel. Notably, upon expanding , itbecomes evident that the cubic terms can be portrayed using linear and quadratic combinations of other nodes. Such an automatic simplification in equations is undeniably profound. Such automatic discovery of simplifying equations is very powerful. Yet, it is worth mentioning that to achieve this exact quadratic recovery; more ancestors must be considered than originally denoted in (Eq.39). This emphasizes the notion that functional dependencies are profoundly influenced by the chosen function form, and hypergraph discovery processes in accordance with a variety of embodiments of the invention discern one of the potential dependencies.

[0238] Referring back to the example of recovery of chemical reaction networks, the proposed mechanism for the hydrogenation of ethyleneg modeled by the following system of differential equationsThe primary variables are the concentrationsand their time derivatives The computational hypergraph encodes thefunctional dependencies (Eq. 40) associated with the chemical reactions. The hyperedges of the hypergraph are assumed to be unknown and the primary variables are assumed to be known. Given samples from the graph of the formthe objective is to recover the structure of the hypergraph given by (Eq.40), representing the functions by hyperedges. A dataset of the form (Eq. 41) was created by integrating 50 trajectories of (Eq. 40) for different initial conditions, and each equispaced 50 times from to . It is imposed that the information that the derivative variables are functions of the non-derivative variables to avoid ambiguity in the recovery, as (Eq.40) is not the unique representation of the functional relation between nodes in the graph. In this example, processes in accordance with many embodiments of the invention were performed with weights for linear, quadratic, and nonlinear, respectively. The proposed approach leads to a perfect recovery of the computation graph illustrated in Fig. 7. The approach obtains a perfect recovery of the computational graph and a correct identification of the relations being quadratic.

[0239] Hypergraph discovery in accordance with a number of embodiments of the invention was applied as a statistical analysis tool to scrutinize COVID-19 open data. In this example, categorical data are treated as scalar values, with all variables scaled to achieve a mean of 0 and a variance of 1. Three distinct kernel types were implemented: linear, quadratic, and Gaussian, with a length scale of 1 for the latter. A weight ratio of is assigned between kernels, signifying that the quadratic kernel is weighted tentimes less than the linear kernel. Lastly, the noise parameter, , is determined using processes as described throughout this description.

[0240] An example of hypergraph discovery with COVID-19 data in accordance with an embodiment of the invention is illustrated in Fig. 17. Initially, a complete graph is constructed using all variables. An example of a complete, full recovered graph in accordance with an embodiment of the invention is illustrated in Fig. 17(a). The construction in this example was done using only linear and quadratic kernels. Processes in accordance with a variety of embodiments of the invention can use only linear and quadratic kernels to regularise the problem by choosing a feature map whose dimension is not too high.

[0241] A notably clustered section within the graph of Fig. 17(a) is observed upon examination. Fig.17(b) offers an in-depth view of this cluster, revealing its representation of the government’s pandemic response. Specifically, Fig. 17(b) shows the node 1705 corresponding to the variable “schools closing” revealing that the government either implemented multiple restrictive measures simultaneously or lifted them in unison. A few representative nodes (e.g., “schools closing”) for the cluster are selected, ones with few ancestors and serving as an ancestor to numerous nodes. In this example, although mask mandates also has few ancestors and serves as an ancestor to numerous nodes, the SNR level for the relationship is low and was on the verge of being identified as noise). Representative nodes in accordance with numerous embodiments of the invention are selected based on various factors, such as (but not limited to) a number of ancestors for a given node, a number of nodes to which a given node is an ancestor, a SNR level for a kernel to a given node, etc. The vaccination policy and stay-at-home requirements are deemed an apt choice in this context, especially given their impact on the population.

[0242] In some embodiments, hypergraph discovery processes may identify redundant information by identifying loops in the hypergraph. Identifying redundant information in accordance with numerous embodiments of the invention may be performed as part of the pruning process (e.g., prior to pruning for a given node, after all nodes have been pruned, etc.). Examples of loops (or clusters) in the example of Fig.17 are illustrated in Figs.17(c-d). New vaccinations form a cluster (Fig. 17(d)), displaying a linear relationship between nodes, signaling redundant information. The hospitalizationcluster (Fig. 17(c)) reveals a nonlinear relationship between nodes. When nodes are related in this way, processes in accordance with a number of embodiments of the invention can eliminate (or consolidate) the redundant nodes. In numerous embodiments, eliminating redundant nodes can be performed prior to or as a part of ancestor pruning. Eliminating redundant nodes is vital for two reasons: firstly, it improves the graph’s readability, especially with 31 variables; secondly, it avoids hindering graph discovery. Eliminating redundant nodes leads to the sparse graph shown in Fig. 17(e), which is interpretable and amenable to (both quantitative and qualitative) analysis. In an extreme case, treating two identical variables as distinct would result in one variable’s ancestor simply being its duplicate, yielding an uninformative graph.

[0243] An example of the evolution of the observed signal-to-noise ratio when pruning ancestors in accordance with an embodiment of the invention is illustrated in Fig. 18. In this case, it is very clear where to stop pruning, as the jumping noise ratio is significant. Chart 1805 shows the noise-to-signal ratio (and chart 1810 its increments) as a function of the number of ancestors of the “cumulative number of hospitalized patients” variable. Even for this real dataset, the proposed approach gives a clear signal for stopping the pruning process.

[0244] Subsequently, the hypergraph discovery process is rerun, with reduced variables due to eliminating redundancy, ushering us into a predominantly noisy regime. With fewer variables available, the Gaussian kernel may also be used. Two indicators are employed to navigate our discovery process: the signal-to-noise ratio and the Z-test. The former quantifies the degree to which the regression is influenced by noise, while the latter signals the existence of any signal. Results of in the graph presented in Fig.17(e). Here, some ancestor relationships are logical when interpreting using our experience of the COVID-19 pandemic. For instance, the number of new hospitalized and new deceased appear closely linked. Indeed, most recorded deaths are from people in a hospital. Additionally, it is observed that many ancestors are needed to explain the number of people tested. This number depends not only on the state of the epidemic but also on the government’s and the population’s reaction. Furthermore, many tests are conducted during holidays such as Christmas, regardless of the current pandemic status. This widespread testing complicates the ability to predict testing based solely on presentvariables. However, it is vital to acknowledge that not all relationships in the graph necessarily suggest causal connections. Our method determines graph edges based on functional relationships, not causal ones. Ultimately, this graph demonstrates sparsity, facilitating a superior understanding of the data’s underlying structure. Furthermore, the method allows us to identify and eliminate redundant information, enhancing clarity regarding the most pivotal trends observed.

[0245] Lastly, hypergraph discovery in accordance with various embodiments of the invention was applied to discover a hierarchy of functional dependencies in biological cellular signaling networks. In this experiment, single-cell data was used, consisting of the phosphoproteins and phospholipids levels in the human immune system T- cells that were measured using flow cytometry. This dataset was studied from a probabilistic modeling perspective in previous works. Existing methods have been used to learn a directed acyclic graph to encode causal dependencies or to learn an undirected graph of conditional independencies between the molecule levels by assuming the underlying data follows a multivariate Gaussian distribution. The latter analysis encodes acyclic dependencies but does not identify directions. In some embodiments, hypergraph discovery identifies functional dependencies without imposing strong distributional assumptions on the data.

[0246] To identify the ancestors of each node, processes in accordance with certain embodiments of the invention can learn the dependencies using only linear and quadratic kernels. An example of hypergraph discovery in biological cellular signaling networks in accordance with an embodiment of the invention is illustrated in Fig.19. The first graph 1905 of Fig. 19 identifies the resulting graph learned given a subset of samples chosen uniformly at random from the dataset consisting of 11 proteins and 7446 samples of their expressions. The graph identified by a hypergraph process consists of four disconnected clusters where the molecule levels in each cluster are closely related by linear or quadratic dependencies (all connections are linear except for the connection between Akt and PKA, which is quadratic). These edges match a subset of the edges found in the gold standard model. With perfect dependencies that have no noise, one can define constraints that reduce the total number of variables in the system. For this noisydataset, these dependencies can be treated as forming groups of similar variables and introduce a hierarchical approach to learn the connections between groups.

[0247] Hypergraph discovery processes in accordance with several embodiments of the invention can identify relationships (or connections) between nodes and / or clusters at multiple levels. In the second graph 1910 of Fig.19, solid arrows indicate strong intra- cluster connections identified in the first level, and dashed lines indicate weaker connections between nodes and clusters identified in the second level. The width and grayscale intensities of each edge correspond to its signal-to-noise ratio.

[0248] In a variety of embodiments, processes run a hypergraph discovery process after grouping the molecules into clusters to learn the connections between groups of variables within each cluster with nonlinear kernels and obtain the graph 1910 in Fig.19. For each node in the graph, the ancestors of each node were identified by constraining the dependence to be a subset of the clusters. In other words, when identifying the ancestors of a given node in cluster , hypergraph discovery processes in accordance with several embodiments of the invention are only permitted to (1) use ancestors that donot belong to cluster , and (2) include all or none of the variables in each cluster ( incluster is listed as an ancestor if and only if all other nodes in cluster are also listed as ancestors). In this example, the ancestors were identified using a Gaussian (fully nonlinear) kernel and the number of by ancestors were selected manually based on the inflection point in the noise-to-signal ratio.

[0249] In graph 1910, each edge is weighted based on its signal-to-noise ratio. There is a stronger dependence of the Jnk, PKC, and P38 cluster on the PIP3, Plcg, and PIP2 cluster, which closely matches the gold standard model. As compared to approaches based on acylic DAGs, however, graph 1910 also contains feedback loops between the various molecule levels.

[0250] While the Bayesian network analysis rely on the control of the sampling of the underlying variables (the simultaneous measurement of multiple phosphorylated protein and phospholipid components in thousands of individual primary human immune system cells, and perturbing these cells with molecular interventions), the reconstruction obtained by hypergraph discovery did not use this information and recovered functional dependencies rather than causal dependencies. Interestingly, the information recoveredappears to complement and enhance other findings (e.g., the linear and noiseless dependencies between variables in the JNK cluster is not something that could easily be inferred from the graph produced in other processes).

[0251] In numerical experiments, hypergraph discovery in accordance with a number of embodiments of the invention maintains robustness with respect to the selection of the signal kernel, i.e., the^in the kernel (Eq.1). Furthermore, the noise prior variance identified by hypergraph discovery is very close to the regularization term that minimizes cross-validation mean square loss. More precisely, the identification of functional dependencies between variables relies on regressing input-output datafor the signal function with a kernel and a nugget termg for a subset of the input-output data,ି^ ^ for the corresponding kernel ridge regressor, and ^ for the minimizer of errorthe value of selected by our algorithm is very close to ^. This suggeststhat one could also use cross-validation (applied to recover the signal function between a given node and a group of nodes in computing a noise-to-signal ratio) to select in accordance with several embodiments of the invention.

[0252] In several embodiments, forming clusters from highly interdependent variables helps to obtain a sparser graph. Additionally, the precision of the pruning process can be enhanced in accordance with various embodiments of the invention by avoiding the division of node activation within the cluster among its separate constituents. From the algebraic equation recovery examples above, we can only hope to recover equivalent equations, and the type of equations recovered depends on the kernel employed for the signal.

[0253] Systems and methods in accordance with several embodiments of the invention provide a comprehensive Gaussian Process framework for solving Type 3 (hypergraph discovery) problems, which is interpretable and amenable to analysis. The breadth and complexity of Type 3 problems significantly surpass those encountered in Type 2 (hypergraph completion), and the initial numerical examples presented serve as a motivation for the scope of Type 3 problems and the broader applications made possible by this approach. Hypergraph discovery in accordance with certain embodiments of theinvention is designed to be fully autonomous, yet it offers the flexibility for manual adjustments to refine the graph’s structure recovery. It aims to incorporate a distinct kind of information into the graph’s structure, namely, the functional dependencies among variables rather than their causal relationships.

[0254] Additionally, hypergraph discovery in accordance with certain embodiments of the invention eliminates the need for a predetermined ordering of variables, a common requirement in acyclic probabilistic models where determining an optimal order is an NP- hard problem usually tackled using heuristic approaches. Furthermore, processes in accordance with certain embodiments of the invention can actually be utilized to generate such an ordering by quantifying the strength of the connections it recovers. In various embodiments, the Uncertainty Quantification properties of the underlying Gaussian Processes can be employed to quantify uncertainties in the structure of the recovered graph. While the uncertainty quantification (UQ) properties of Gaussian Processes (GPs) lack a direct equivalent in Deep Learning, hypergraph discovery in accordance with some embodiments of the invention encompasses richly hyperparametrized kernels to facilitate its extension to Deep Learning type connections between variables. This can be achievable by interpreting neural networks as hyperparameterized feature maps. A. Systems for Hypergraph Discovery 5. Hypergraph Discovery System

[0255] An example of a hypergraph discovery system that performs hypergraph discovery in accordance with an embodiment of the invention is illustrated in Fig. 20. Network 2000 includes a communications network 2060. The communications network 2060 is a network such as the Internet that allows devices connected to the network 2060 to communicate with other connected devices. Server systems 2010, 2040, and 2070 are connected to the network 2060. Each of the server systems 2010, 2040, and 2070 is a group of one or more servers communicatively connected to one another via internal networks that execute processes that provide cloud services to users over the network 2060. One skilled in the art will recognize that a hypergraph discovery system may exclude certain components and / or include other components that are omitted for brevity without departing from this invention.

[0256] For purposes of this discussion, cloud services are one or more applications that are executed by one or more server systems to provide data and / or executable applications to devices over a network. The server systems 2010, 2040, and 2070 are shown each having three servers in the internal network. However, the server systems 2010, 2040 and 2070 may include any number of servers and any additional number of server systems may be connected to the network 2060 to provide cloud services. In accordance with various embodiments of this invention, a hypergraph discovery system that uses systems and methods that perform hypergraph discovery in accordance with an embodiment of the invention may be provided by a process being executed on a single server system and / or a group of server systems communicating over network 2060.

[0257] Users may use personal devices 2080 and 2020 that connect to the network 2060 to perform processes that perform hypergraph discovery in accordance with various embodiments of the invention. In the shown embodiment, the personal devices 2080 are shown as desktop computers that are connected via a conventional “wired” connection to the network 2060. However, the personal device 2080 may be a desktop computer, a laptop computer, a smart television, an entertainment gaming console, or any other device that connects to the network 2060 via a “wired” connection. The mobile device 2020 connects to network 2060 using a wireless connection. A wireless connection is a connection that uses Radio Frequency (RF) signals, Infrared signals, or any other form of wireless signaling to connect to the network 2060. In the example of this figure, the mobile device 2020 is a mobile telephone. However, mobile device 2020 may be a mobile phone, Personal Digital Assistant (PDA), a tablet, a smartphone, or any other type of device that connects to network 2060 via wireless connection without departing from this invention.

[0258] As can readily be appreciated the specific computing system used to perform hypergraph discovery is largely dependent upon the requirements of a given application and should not be considered as limited to any specific computing system(s) implementation.6. Hypergraph Discovery Element

[0259] An example of a hypergraph discovery element that executes instructions to perform processes that perform hypergraph discovery in accordance with an embodiment of the invention is illustrated in Fig. 21. Hypergraph discovery elements in accordance with many embodiments of the invention can include (but are not limited to) one or more of mobile devices, cameras, and / or computers. Hypergraph discovery element 2100 includes processor 2105, peripherals 2110, network interface 2115, and memory 2120. One skilled in the art will recognize that a hypergraph discovery element may exclude certain components and / or include other components that are omitted for brevity without departing from this invention.

[0260] The processor 2105 can include (but is not limited to) a processor, microprocessor, controller, or a combination of processors, microprocessor, and / or controllers that performs instructions stored in the memory 2120 to manipulate datastored in the memory. Processor instructions can configure the processor 2105 to performprocesses in accordance with certain embodiments of the invention. In various embodiments, processor instructions can be stored on the memory 2120 (or non- transitory machine readable medium).

[0261] Peripherals 2110 can include any of a variety of components for capturing data, such as (but not limited to) cameras, microphones, displays, and / or sensors (e.g., accelerometers, gyroscopes, radar, LiDAR, environmental sensors, odometers, etc.). In a variety of embodiments, peripherals can be used to gather inputs and / or provide outputs. Hypergraph discovery element 2100 can utilize network interface 2115 to transmit and receive data over a network based upon the instructions performed by processor 2105. Peripherals and / or network interfaces in accordance with many embodiments of the invention can be used to gather inputs that can be used to perform hypergraph discovery.

[0262] Memory 2120 includes a hypergraph discovery application 2125, model data 2130, and training data 2135. Hypergraph discovery applications in accordance with several embodiments of the invention can be used to perform hypergraph discovery. In numerous embodiments, hypergraph discovery applications may support otherapplications with discovered hypergraphs. For example, in certain embodiments, hypergraph discovery applications provide support for control applications in autonomous apparatuses by providing dynamic updates to hypergraphs used to control the apparatus based on changing conditions.

[0263] Model data in accordance with numerous embodiments of the invention can store various parameters and / or weights for various models that can be used for various processes as described in this specification. Model data in accordance with many embodiments of the invention can be updated through training on variable data captured on a hypergraph discovery element or can be trained remotely and updated at a hypergraph discovery element. In some embodiments, model data can include hypergraphs with nodes for various system variables (e.g., sensor data and outputs) and the updated connections and functions between them.

[0264] In a variety of embodiments, variable data can include data collected by a training element across a number of different variables. As described throughout this description, variable data can include data across a variety of different types, such as (but not limited to) video data, audio data, stock data, chemical data, sensor data, environmental data, social network data, system control data, etc.

[0265] Although a specific example of a hypergraph discovery element 2100 is illustrated in this figure, any of a variety of hypergraph discovery elements can be utilized to perform processes for hypergraph discovery similar to those described herein as appropriate to the requirements of specific applications in accordance with embodiments of the invention.

[0266] Although specific methods of hypergraph discovery are discussed above, many different methods of hypergraph discovery can be implemented in accordance with many different embodiments of the invention. It is therefore to be understood that the present invention may be practiced in ways other than specifically described, without departing from the scope and spirit of the present invention. Thus, embodiments of the present invention should be considered in all respects as illustrative and not restrictive. Accordingly, the scope of the invention should be determined not by the embodiments illustrated, but by the appended claims and their equivalents.

Claims

WHAT IS CLAIMED IS:

1. A method of processing input data, comprising: receiving input data at a data processing system; providing the input data to a discovered hypergraph; and generating an output using the discovered hypergraph; wherein the discovered hypergraph is characterized by a plurality of relationships between a plurality of nodes that each represent a variable of a plurality of variables, where the plurality of relationships were discovered by: selection of a kernel for a relationship between at least one node of the plurality of nodes and a set of candidate ancestors; and pruning of the set of candidate ancestors to identify a set of minimal ancestors.

2. The method of claim 1, wherein the plurality of relationships were further discovered by selection of a different second kernel for a second relationship between a second node of the plurality of nodes and a corresponding set of candidate ancestors.

3. The method of claim 1 or 2, wherein the kernel is at least one selected from the group consisting of a linear kernel, a quadratic kernel, and a nonlinear kernel.

4. The method of any of claims 1-3, wherein generating the output comprises generating a set of control signals for controlling a physical device.

5. The method of claim 4, wherein the physical device is a vehicle.

6. The method of any of claims 1-5, wherein receiving the set of input data comprises normalizing the set of input data.

7. The method of any of claims 1-6, wherein normalizing the data comprises performing affine transformations.

8. The method of any of claims 1-7, wherein: the selection of a kernel is based on a set of training data; the set of training data comprises a plurality of samples; a first sample comprises values for a first subset of the plurality of variables; and a second sample comprises values for a different second subset of the plurality of variables.

9. The method of any of claims 1-8, wherein selection of a kernel comprises: for each of a plurality of kernels: computing a signal-to-noise ratio (SNR) for the relationship between the selected node and the set of candidate ancestors; and when the SNR falls below a given threshold, selecting the kernel for the relationship between the selected node and the set of candidate ancestors.

10. The method of claim 9, wherein the plurality of kernels comprises at least one selected from the group consisting of a linear kernel, a quadratic kernel, and a nonlinear kernel.

11. The method of claim 9 or 10, wherein computing the SNR comprises: selecting a noise prior parameter; and performing a regression analysis with the selected noise prior parameter.

12. The method of claim 11, wherein performing the regression analysis comprises predicting a value of the selected node based on the set of candidate ancestors.

13. The method of claim 12, wherein the SNR is a ratio of a mean and a variance for predicted values of the selected node based on the set of candidate ancestors.

14. The method of any of claims 9-13, wherein the given threshold is a fixed value.

15. The method of any of claims 9-14, wherein at least one of the plurality of kernels is at least one selected from the group consisting of a Gaussian kernel and a Matérn kernel.

16. The method of claim 1-15, wherein the set of candidate ancestors comprises all of the plurality of nodes other than the selected node.

17. The method of claim 1-16, wherein pruning the set of candidate ancestors comprises: identifying a least important ancestor of the set of candidate ancestors; computing a signal-to-noise (SNR) for the relationship between the selected node and the set of candidate ancestors without the least important ancestor; when the computed SNR exceeds a given threshold, select the set of candidate ancestors without the least important ancestor as the set of minimal ancestors; and when the computed SNR falls below a given threshold, select the set of candidate ancestors as the set of minimal ancestors.

18. The method of claim 1-17, wherein pruning the set of candidate ancestors comprises: for each of a plurality of subsets of the set of candidate ancestors, computing a noise-to-signal ratio (NSR) for the relationship between the selected node and the subset of candidate ancestors; and identifying a particular subset of the plurality of subsets as the set of minimal ancestors based on the computed NSRs for the plurality of subsets.

19. The method of claim 18, wherein computing the NSR comprises: selecting a noise prior parameter; and performing a regression analysis with the selected parameter.

20. The method of claim 19, wherein performing the regression analysis comprises predicting a value of the selected node based on the set of candidate ancestors.

21. The method of claim 20, wherein the NSR is a ratio of a mean and a variance for predicted values of the selected node based on the set of candidate ancestors.

22. The method of any of claims 19-21, wherein the regression analysis is a Kernel Ridge Regression.

23. The method of any of claims 19-22, wherein the noise prior parameter is selected based on characteristics of the kernel.

24. The method of claim 23, wherein selecting the noise prior parameter comprises maximizing the variance of an eigenvalue histogram when the kernel is a universal kernel.

25. The method of claim 18-24, wherein identifying the particular subset comprises identifying an inflection point in the NSRs as a function of a number of ancestors in the subset of candidate ancestors.

26. The method of claim 1-25, wherein selecting the kernel and pruning the set of candidate ancestors is based on a same threshold.

27. The method of claim 18-26, wherein pruning the set of candidate ancestors comprises: identifying a set of redundant candidate ancestors of the set of candidate ancestors; and removing at least one candidate ancestor of the set of redundant candidate ancestors from the set of candidate ancestors, wherein redundant candidate ancestors make redundant contributions to the computed NSR.

28. The method of claim 27, wherein identifying the set of redundant candidate ancestors is based on the determined relationships between the plurality of nodes.

29. The method of claim 1-28, wherein the plurality of relationships were further discovered by: identifying clusters of nodes of the plurality of nodes based on relationships between each node and the set of candidate ancestors of the node; and identifying relationships between the clusters of nodes, wherein the plurality of relationships comprises: the relationships between each node and the set of candidate ancestors of the node; and the relationships between the clusters of nodes.

30. A data processing system, comprising: at least one processor; a memory; and wherein machine readable instructions stored in the memory configure the processor to perform the method of any of claims 1-29.

31. A method for hypergraph discovery, the method comprising: receiving a set of input data comprising a plurality of samples, each sample comprising values for at least a subset of a plurality of variables; determining relationships between a plurality of nodes based on the set of input data, each node of the plurality of node representing a variable of the plurality of variables, wherein determining relationships for each node of the plurality of nodes comprises: selecting a kernel for a relationship between a selected node and a set of candidate ancestors; and pruning the set of candidate ancestors to identify a set of minimal ancestors; and generating a hypergraph of the input data based on the relationships between the plurality of nodes.

32. The method of claim 31, wherein receiving the set of input data comprises normalizing the set of input data.

33. The method of any of claims 31-32, wherein normalizing the data comprises performing affine transformations.

34. The method of any of claims 31-33, wherein: the set of input data comprises a plurality of samples; a first sample comprises values for a first subset of the plurality of variables; and a second sample comprises values for a different second subset of the plurality of variables.

35. The method of any of claims 31-34, wherein selecting a kernel comprises: for each of a plurality of kernels: computing a signal-to-noise ratio (SNR) for the relationship between the selected node and the set of candidate ancestors; and when the SNR falls below a given threshold, selecting the kernel for the relationship between the selected node and the set of candidate ancestors.

36. The method of claim 35, wherein the plurality of kernels comprises at least one selected from the group consisting of a linear kernel, a quadratic kernel, and a nonlinear kernel.

37. The method of any of claims 35-36, wherein computing the SNR comprises: selecting a noise prior parameter; and performing a regression analysis with the selected noise prior parameter.

38. The method of claim 37, wherein performing the regression analysis comprises predicting a value of the selected node based on the set of candidate ancestors.

39. The method of claim 38, wherein the SNR is a ratio of a mean and a variance for predicted values of the selected node based on the set of candidate ancestors.

40. The method of any of claims 35-39, wherein the given threshold is a fixed value.

41. The method of any of claims 35-40, wherein at least one of the plurality of kernels is at least one selected from the group consisting of a Gaussian kernel and a Matérn kernel.

42. The method of claim 31-41, wherein the set of candidate ancestors comprises all of the plurality of nodes other than the selected node.

43. The method of claim 31-42, wherein pruning the set of candidate ancestors comprises: identifying a least important ancestor of the set of candidate ancestors; computing a signal-to-noise (SNR) for the relationship between the selected node and the set of candidate ancestors without the least important ancestor; when the computed SNR exceeds a given threshold, select the set of candidate ancestors without the least important ancestor as the set of minimal ancestors; and when the computed SNR falls below a given threshold, select the set of candidate ancestors as the set of minimal ancestors.

44. The method of claim 31-43, wherein pruning the set of candidate ancestors comprises: for each of a plurality of subsets of the set of candidate ancestors, computing a noise-to-signal ratio (NSR) for the relationship between the selected node and the subset of candidate ancestors; and identifying a particular subset of the plurality of subsets as the set of minimal ancestors based on the computed NSRs for the plurality of subsets.

45. The method of claim 44, wherein computing the NSR comprises: selecting a noise prior parameter; and performing a regression analysis with the selected parameter.

46. The method of claim 45, wherein performing the regression analysis comprises predicting a value of the selected node based on the set of candidate ancestors.

47. The method of claim 46, wherein the NSR is a ratio of a mean and a variance for predicted values of the selected node based on the set of candidate ancestors.

48. The method of claim 45-47, wherein the regression analysis is a Kernel Ridge Regression.

49. The method of claim 45-48, wherein the noise prior parameter is selected based on characteristics of the kernel.

50. The method of claim 49, wherein selecting the noise prior parameter comprises maximizing the variance of an eigenvalue histogram when the kernel is a universal kernel.

51. The method of claim 44-50, wherein identifying the particular subset comprises identifying an inflection point in the NSRs as a function of a number of ancestors in the subset of candidate ancestors.

52. The method of claim 31-51, wherein selecting the kernel and pruning the set of candidate ancestors is based on a same threshold.

53. The method of claim 44-53, wherein pruning the set of candidate ancestors comprises: identifying a set of redundant candidate ancestors of the set of candidate ancestors; and removing at least one candidate ancestor of the set of redundant candidate ancestors from the set of candidate ancestors, wherein redundant candidate ancestors make redundant contributions to the computed NSR.

54. The method of claim 53, wherein identifying the set of redundant candidate ancestors is based on the determined relationships between the plurality of nodes.

55. The method of claim 31-54 further comprising: identifying clusters of nodes of the plurality of nodes based on relationships between each node and the set of minimal ancestors of the node; and identifying relationships between the clusters of nodes, wherein generating the hypergraph of the input data is further based on the relationships between the clusters of nodes.

56. An apparatus comprising: a set of one or more processors; a set of one or more peripherals; and a non-transitory machine readable medium containing program instructions that are executable by the set of processors to perform the methods of any of claims 31-55, wherein the method further comprises controlling the apparatus based on the generated hypergraph.

57. The apparatus of claim 56, wherein the set of input data comprises data collected from at least one peripheral of the set of peripherals.

58. The apparatus of claim 57, wherein the set of peripherals comprises at least one selected from the group consisting of a gyroscope, an accelerometer, an odometer, a LiDAR system, an environmental sensor, and a motion sensor.

59. A non-transitory machine readable medium containing program instructions that are executable by a set of one or more processors to perform the method of any of claims 31-55.

Citation Information

Patent Citations

  • Probability hypergraph-driven geoscience knowledge graph reasoning optimization system and method

    CN114547325A