SYSTEM AND METHOD FOR MONITORING RELATED METRICS - Patent application
Patent Information
- Application Number
- JP2024553717
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-03-09
- Filing Date
- 2023-03-07
- Publication Date
- 2025-12-09
AI Technical Summary
Current approaches to monitoring business-related metrics, such as KPIs, are limited as they do not provide a comprehensive understanding of the relationships between metrics, the data used to generate them, and the performance of entities that generate the underlying data.
A system and method that create a feature graph with nodes representing concepts, datasets, metadata, models, metrics, and edges representing statistically significant relationships between them, allowing users to identify and track relevant metrics, set alert rules, and receive recommendations for additional metrics to monitor.
Enables businesses to more accurately and comprehensively monitor KPIs and assess data quality, providing users with a deeper understanding of metric relationships and data-driven insights for informed decision-making.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Detailed Description of the Invention
[0001] [CROSS REFERENCE TO RELATED APPLICATIONS] This application claims the benefit of U.S. Provisional Application No. 63 / 318,170, filed March 9, 2022, and entitled “System and Methods for Monitoring Related Metrics,” the entire contents of which are incorporated by reference.
[0002] Please note that references to a "system" in the context of an architecture, or a system architecture or platform herein, refer to the architecture, platform, and processes for performing statistical searching and other forms of data organization described in U.S. patent application Ser. No. 16 / 421,249, entitled "Systems and Methods for Organizing and Finding Data," filed May 23, 2019 (now U.S. Patent No. 11,354,587, issued June 7, 2022), which claims priority to U.S. Provisional Patent Application No. 62 / 799,981, entitled "Systems and Methods for Organizing and Finding Data," filed February 1, 2019, the entire contents of which are incorporated herein by reference. [background] Data-driven organizations track key performance indicators (called KPIs) and other metrics to measure the status of the organization and to support strategic decision making. KPIs and metrics are increasingly becoming part of news reports (e.g., the level and percentage change of the Dow Jones Industrial Average, the S&P 500 Index, stock prices of underlying companies, or the level and change of weekly new unemployment claims). Current approaches to monitoring such metrics rely on dashboards, data catalogs, and KPI trackers to provide users with information about specific KPIs.
[0003] Although useful, traditional approaches have limitations and disadvantages. For one, traditional approaches provide information about KPIs in relative isolation from other factors. Furthermore, traditional approaches do not track and monitor key metrics in the context of the modeling and statistical correlation work performed by modern data science and analytics teams. This limits the ability of users to understand the significance of changes in KPIs and how such changes may relate to or affect other metrics. This prevents users from gaining a more complete and accurate understanding of the relationships between various metrics, the data used to generate the metrics, and the performance of the company (or other entity) that generated the underlying data.
[0004] Developing tools that automate the process of evaluating statistical relationships within and between data sets and generating metrics and decisions based on the data sets requires dedicated resources that may not be readily available or affordable for many businesses. Embodiments of the systems and methods described herein are directed to solving these and related problems, both individually and collectively. [overview] As used herein, the terms "invention", "the invention", "this invention", "the present invention", "the disclosure" or "the disclosure" are intended to broadly refer to all subject matter and claims disclosed in this document, drawing or figure. Descriptions containing these terms do not limit the meaning or scope of the disclosed subject matter or claims. The embodiments encompassed by this disclosure are defined by the claims, not by this summary. This summary is a broad overview of various aspects of the disclosure and introduces some of the concepts that are further described in the Detailed Description section below. This summary is not intended to identify key, essential or required features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to the entire specification, in appropriate portions, and to any and all figures or drawings, and claims.
[0005] Embodiments of the disclosure relate to systems and methods for improving a business or other entity's ability to monitor business-related metrics (such as, for example, KPIs) and assess the quality of the underlying data used to generate the metrics. In some embodiments, the disclosed systems and methods may comprise elements, components, functions, operations, or processes configured and operative to provide one or more of the following:
[0006] Creating a feature graph comprising a set of nodes and edges, where: ○ A node may represent, by way of non-limiting examples, one or more of the following: a concept, a topic, a dataset, metadata, a model, a metric, a variable, a measurable quantity, an object, a property, a feature, or a factor.
[0007] ■In some embodiments, nodes may be created in response to, by way of non-limiting examples, discovering or gaining access to datasets, metadata, models, generating output from a trained model, generating metadata about a dataset, or developing an ontology or other form of hierarchical relationships.
[0008] o An edge represents, by way of non-limiting example, a relationship between a first node and a second node, such as a statistically significant relationship, a dependency, or a hierarchical relationship. ■In some embodiments, an edge may be created connecting a first node and a second node and representing a statistically valid relationship between the two nodes as determined by statistical analysis, machine learning models, or research.
[0009] A label associated with an edge may indicate an aspect of the relationship between the two nodes connected by the edge, such as, by way of non-limiting example, metadata on which the relationship between the two nodes is based, or a dataset supporting a statistically significant relationship between the two nodes.
[0010] - Providing a user with user interface display screens, tools, features, and selectable elements to enable the user to perform one or more of the following functions:
[0011] Identifying metrics of interest (e.g., KPIs) to monitor or track. ■The metric of interest may be generated by a trained model, formula, equation, or rule set and may be based on, generated from, or derived from underlying data that is a function of time.
[0012] Defining rules that describe when an alert should be generated regarding the behavior of a specified metric. ■Such rules may be based on, by way of non-limiting example, absolute values, changes to values, percentage changes, percentage changes over time, or above or below a threshold.
[0013] o Defining how the results of applying the rules are identified or shown on a user interface display. ■ This may depend, for example, on user preferences and / or the value or type of changes to the metric.
[0014] ○ Allowing a user to select the metric for which an alert was generated and, in response, providing information regarding, by way of non-limiting example, the change in the metric's value over time, the satisfied or enabled rules that resulted in the alert, the metric's relationship to other metrics (if relevant), and available information about the dataset, machine learning model, rules, formulas, or other factors used to generate the metric.
[0015] ●Generating recommendations for the user regarding different metrics or sets of metrics that may be worth monitoring, data sets that may be useful for scrutiny, metadata that may be relevant to the identified metrics, or other aspects of the underlying data or metrics that may be of potential interest to the user.
[0016] Here, the recommendations may result (at least in part) from output generated by trained machine learning models, statistical analysis, research, or other forms of collection or evaluation of data.
[0017] In one embodiment, the disclosure relates to a system for improving a business or other entity's ability to monitor business-related metrics (such as, for example, KPIs) and assess the quality (and thus accuracy and reliability) of the underlying data. The system may include a set of computer-executable instructions stored in (or on) one or more non-transitory computer-readable media and an electronic processor or coprocessor. The instructions, when executed by the processor or coprocessor, cause the processor or coprocessor (or an apparatus or device of which they are a part) to perform a set of operations that implement an embodiment of one or more of the disclosed methods.
[0018] In one embodiment, the disclosure relates to one or more non-transitory computer-readable media that include a set of computer-executable instructions that, when executed by an electronic processor or coprocessor, cause the processor or coprocessor (or an apparatus or device of which it is a part) to perform a set of operations that implement an embodiment of one or more of the disclosed methods.
[0019] In some embodiments, the systems and methods described herein may provide services through a SaaS or multi-tenant platform. The platform provides access to multiple entities, each of which has a separate account and associated data storage. Each account may correspond to, for example, a user, a set of users, an entity that provides a data set for evaluation and use in generating business-related metrics, or an organization. Each account may have access to one or more services, which are instantiated in the accounts and which implement one or more of the methods or functions described herein.
[0020] Other objects and advantages of the described system and method will become apparent to those skilled in the art upon review of the detailed description and the included drawings. The same reference characters and descriptions throughout the drawings indicate similar, but not necessarily identical, elements. While the exemplary embodiments described herein are susceptible to various modifications and alternative forms, specific embodiments have been shown by way of example in the drawings and are described in detail herein. However, it is not intended that the exemplary or specific embodiments described herein be limited to the forms described. Instead, the present disclosure covers all modifications, equivalents, and alternatives falling within the scope of the appended claims. [Brief description of the drawings]
[0021] DETAILED DESCRIPTION OF THE DRAWINGS An embodiment of the present disclosure will be described with reference to the following drawings. [Figure 1(a)] A block diagram illustrating a set of elements, components, functions, processes, or operations that may be part of a platform architecture 100 in which one embodiment of the disclosed system and method for metric monitoring may be implemented. [Figure 1(b)] 1 is a flowchart or flow diagram illustrating a process, method, function, or act for constructing a feature graph 150 using an implementation of one embodiment of the systems and methods disclosed herein. [Figure 1(c)] FIG. 1 is a flowchart or flow diagram illustrating a process, method, function, or operation for an example use case in which a feature graph is traversed to identify potentially relevant datasets and that may be implemented in one embodiment of the systems and methods disclosed herein. [Figure 1(d)] FIG. 1 illustrates an example of a portion of a feature graph data structure that may be used to organize and access data and information, and that may be created using implementation of an embodiment of the systems and methods disclosed herein. [Figure 2(a)]2 is a block diagram illustrating a set of elements, components, functions, processes, or operations that may be part of a platform architecture in which an embodiment of the disclosed system and method for metric monitoring may be implemented. Specifically, FIG. 2(a) illustrates how changes in characteristics from a data set stored within a cloud database service may be monitored using an implementation of the disclosed metric monitoring capabilities. [Figure 2(b)] FIG. 2(b) is a flow chart or diagram illustrating a set of elements, components, functions, processes, or operations that may be performed as part of a platform architecture in which an embodiment of the disclosed system and method for metrics monitoring may be implemented. In particular, FIG. 2(b) depicts certain steps in FIG. 2(a) with a greater focus on the different user interactions and software elements that contribute to how the metrics monitoring functionality is implemented and made available to the user. [Figure 2(c)] An example of a user interface display showing the most recent value, the percent change for that value, and identification of the subpopulation with the largest change (which can be calculated when a metric is created as a collection of values in a table where there are multiple subpopulations / dimensions in the data). [Figure 2(d)] 1 is an example of a user interface display illustrating a metric monitoring panel on a page for the metric Weekly Active Users. On the left side of the platform characteristics graph, metric monitoring is turned on for other metrics, and edges between nodes in the graph contain metadata that describe statistical relationships between the metrics. [Figure 2(e)] is an example of a user interface display illustrating a platform catalog view of metric monitoring, where metric monitoring is turned on for eight metrics on this page. [Figure 2(f)]1 is an example of a user interface display illustrating one or more notifications for a metric monitoring feature. [Figure 2(g)] is an example of a user interface display illustrating a simplified rule setting dialog: A condition that would be applied to this metric would be when the absolute value of the percent change is strictly greater than 4.5. [Figure 2(h)] A diagram illustrating elements, components, or processes that may be present in or executed by one or more computing devices, servers, platforms, or systems configured to implement methods, processes, functions, or operations in accordance with some embodiments. [Diagram 3] FIG. 1 illustrates an architecture for a multi-tenant or SaaS platform that may be used in implementing an embodiment of the systems and methods described herein. [Figure 4] FIG. 1 illustrates an architecture for a multi-tenant or SaaS platform that may be used in implementing an embodiment of the systems and methods described herein. [Diagram 5] FIG. 1 illustrates an architecture for a multi-tenant or SaaS platform that may be used in implementing one embodiment of the systems and methods described herein. Note that the same numbers are used throughout the disclosure and figures to reference similar components and features. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0022] [Detailed Description] The subject matter of the embodiments of the present disclosure has been described herein with specificity to meet legal requirements, but this description is not intended to limit the scope of the claims. The claimed subject matter may be implemented in other ways, may include different elements or steps, and may be used in conjunction with other existing or later developed technologies. This description should not be construed as implying any required order or arrangement between the various steps or elements, unless expressly noted as requiring a particular order of steps or arrangement of elements.
[0023] Embodiments of the disclosure are now more fully described with reference to the accompanying drawings, which form a part of the disclosure, and which show, by way of illustration, exemplary embodiments in which the disclosure may be practiced. The disclosure may, however, be embodied in different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy legal requirements and will convey the scope of the disclosure to those skilled in the art.
[0024] In particular, the present disclosure may be implemented in whole or in part as a system, as one or more methods, or as one or more devices. The disclosed embodiments may take the form of hardware-implemented embodiments, software-implemented embodiments, or embodiments combining software and hardware aspects. For example, in some embodiments, one or more of the operations, functions, processes, or methods described herein may be implemented by one or more suitable processing elements (such as, but not limited to, a processor, microprocessor, CPU, GPU, TPU, or controller) that are part of a client device, a server, a network element, a remote platform (such as, for example, a SaaS platform), an "in the cloud" service, or other form of computing or data processing system, device, or platform.
[0025] One or more processing elements may be programmed with a set of executable instructions (e.g., software instructions), where the instructions may be stored on (or within) one or more suitable non-transitory computer-readable data storage media or elements. In some embodiments, the set of instructions may be communicated to a user through a transfer of the instructions or through an application that executes the set of instructions (e.g., over a network, such as the Internet). In some embodiments, the set of instructions or the application may be available to an end user through access to a SaaS platform or a service offered through such a platform.
[0026] In some embodiments, one or more of the operations, functions, processes, or methods described herein may be implemented in specialized forms of hardware, such as programmable gate arrays or application specific integrated circuits (ASICs). It should be noted that an embodiment of the disclosure may be implemented as an application, a subroutine that is part of a larger application, a "plug-in," an extension to the functionality of a data processing system or platform, or in any other suitable form. Therefore, the following detailed description is not to be taken in a limiting sense.
[0027] As mentioned, in some embodiments, the systems and methods described herein may provide services through a SaaS or multi-tenant platform. The platform provides access to multiple entities, each of which has a separate account and associated data storage. Each account may correspond to, for example, a user, a set of users, an entity, or an organization. Each account may have access to one or more services, which are instantiated in the accounts and which implement one or more of the methods or functions described herein.
[0028] Embodiments of the disclosure relate to systems and methods for improving a business or other entity's ability to monitor business-related metrics (such as, for example, KPIs) and assess the quality of the underlying data used to generate the metrics.
[0029] As a general principle, it is desirable for the data used to make decisions to be relevant (or, in some cases, "sufficiently" relevant) to the task being performed or decision being made. Making reliable data-driven decisions or predictions requires not just data on the desired outcome of the decision or the target of the prediction, but also data on the variables that are statistically associated with that outcome or target (ideally all, but at least the variables that are most strongly statistically associated). Unfortunately, using traditional approaches, it is difficult to discover which variables have been demonstrated to be statistically associated with the outcome or target, and to access data on those variables in order to better assess the reliability of the decision made based on those variables.
[0030] In many situations, discovering and accessing data is made more efficient by representing the data in a particular format or structure. The format or structure may include labels for one or more columns, rows, or fields in a data record. Traditional approaches to identifying and discovering data of interest are typically based on terms that semantically match labels in (or referencing or for) a dataset. While this method is useful for discovering and accessing data about topics (e.g., targets or outcomes) that may be relevant, it does not address the problem of discovering and accessing data about variables that induce, affect, predict, or are otherwise statistically associated with the topic of interest.
[0031] An embodiment of the system and method disclosed herein may include constructing or creating a graph database. In the context of this disclosure, a graph is a set of objects presented together if they have some type of close or relevant relationship. An example is two data representing nodes and connected by a path. A node may be connected to many nodes, and many nodes may be connected to a particular node. A path or line connecting one or more first nodes and a second node is called an "edge." An edge may be associated with one or more values, and such values may represent, as non-limiting examples, characteristics of the connected nodes or metrics or measures of relationships between one or more nodes (e.g., statistical parameters, etc.). The graph format may make it easier to identify certain types of relationships, such as more central or less meaningful ones of a set of variables or relationships. Graphs typically occur in two main types: "undirected", where the relationships they represent are symmetric, and "directed", where the relationships are not symmetric (in the case of directed graphs, arrows may be used instead of lines to indicate the manner of the relationships between the nodes).
[0032] In some embodiments, information and data are represented in the form of a data structure referred to herein as a "feature graph." A feature graph is a graph or diagram that includes nodes and edges, where edges serve to "connect" a node to one or more other nodes. Nodes in a feature graph may represent, by way of example, variables (i.e., measurable quantities), objects, properties, characteristics, or factors. Edges in a feature graph may represent measures of statistical association between a node and one or more other nodes.
[0033] The associations may be expressed in numerical and / or statistical terms and may range from, by way of example, observed (or possibly unsupported) relationships, to measured correlations, to causal relationships. The information and data used to construct the feature graph may be obtained from, by way of non-limiting example, one or more of scientific papers, experiments, results of machine learning models, human or machine observations, or unsupported evidence of associations between two variables.
[0034] As one example, a feature graph may be constructed by accessing a set of sources that contain information about statistical associations between a topic of study and one or more variables considered in the study. The information contained in the sources is used to construct a data structure or data representation that includes nodes and edges connecting the nodes. The edges may be associated with information about statistical relationships between two nodes. One or more nodes may have a dataset associated therewith that is accessible using a link or other form of address or access element. An embodiment may include functionality that allows a user to describe a data structure and perform searches through the data structure to identify datasets that may be relevant for training a machine learning model that is used in making a particular decision or classification.
[0035] Thus, embodiments may generate a data structure that includes nodes, edges, and links to datasets. The nodes and edges represent concepts, topics of interest, or topics of previous research. The edges represent information about statistical relationships between the nodes. The links (or other forms of address or access elements) provide access to datasets that establish (or support, substantiate, etc.) statistical relationships between one or more variables that were part of the study, or between variables and concepts or topics.
[0036] One of the responsibilities for data science and data engineering teams is managing "data quality," which refers to the suitability and applicability of collected or acquired data for use in data analysis and machine learning (ML) modeling. Assessing data quality can include not only gathering information or facts about the data, such as the source, the date of collection, and information about the collection process, but also examining different statistical properties of the data. These statistical properties can be used to identify data sets that are "better" (i.e., more accurate or more reliable) candidates for use in training a model or evaluating the performance of a business or other entity.
[0037] Conventional tools exist that provide users with detailed information about the data itself, and tools that automate the process for validating data quality. However, assessing the statistical properties of a data set typically involves writing custom computer code to either query a database or otherwise access the data, and then applying rules or heuristics (using additional custom code) to determine whether the accessed data (or a subset contained within the data) is within the boundaries of the rules or heuristics. This places a burden on many entities and requires the allocation of resources to which the entities may not have access or may not be affordable.
[0038] Data quality can also affect the evaluation of machine learning models. Machine learning (ML) involves the study of algorithms and statistical models that computer systems use to perform specific tasks without explicit instructions, but rather rely on identifying patterns and applying inference processes. Machine learning algorithms build mathematical "models" based on sample data (known as "training data") and information about what the data represents (called labels or annotations) to make predictions, classifications, or decisions without being explicitly programmed to perform the task.
[0039] Machine learning algorithms are used in a wide variety of applications, including email filtering and computer vision, where it is difficult or not feasible to develop traditional algorithms to perform the task effectively. Due to the importance of the ML models used for the tasks, researchers and developers of machine learning based applications spend time and resources to build the most "accurate" predictive models for their use cases. The evaluation of the performance of a model and the importance of each feature in the model is typically represented by a particular metric used to characterize the model and its performance. These metrics may include, for example, model accuracy, confusion matrix, precision (P), recall (R), specificity, F1 score, precision-recall curve, ROC (Receiver Operating Characteristic) curve, or PR vs. ROC curve. Each metric may provide a slightly different way of evaluating a model or a particular aspect of the model's performance.
[0040] A key element of decision making in modern "data-driven" businesses is the identification of KPIs ("key performance indicators" or "key metrics"). Many corporate leadership teams are focused on maintaining the growth of, or otherwise using, KPIs as the primary "signals" or indicators of their company's health or performance. The importance of a KPI to business decisions and the quality of the data used to generate that KPI are related, because the utility of a KPI and the legitimacy of using that KPI as an indicator of a company's or team's performance depend on the applicability of the KPI and statistical (or other) measures of the accuracy and / or reliability of the underlying data used to calculate the KPI. Companies may invest in analysts and engineers to build "dashboards" and other analytical tools that highlight the levels and changes in their company's KPIs and inform decision makers about those changes.
[0041] Due to the criticality of the data used in determining KPIs and / or training a model and its potential impact on model performance, the characteristics of the dataset may be an important factor in selecting training data and interpreting results from a trained model. This may be particularly important in a business setting where data generated by the business is used as training data or as input to a trained model to generate metrics of interest to the enterprise. For example, a trained model may be used to generate KPIs that represent aspects of the business's operations, such as, by way of non-limiting examples, revenue growth, profit margins, marketing costs, or sales conversion rates.
[0042] In some embodiments, the described user interface (UI) and user experience (UX) may be implemented as part of an underlying data analytics platform, such as the systems platform referenced herein and described in U.S. Patent Application Serial No. 16 / 421,249, entitled "Systems and Methods for Organizing and Finding Data" (now issued U.S. Patent No. 11,354,587). The disclosed platform may discover, store, and in some cases generate statistical relationships between data, concepts, variables, or other features. These relationships may be generated from machine learning models or programmatically run correlations.
[0043] The disclosed metric monitoring functionality leverages the system's data organization and analytics platform to provide a way to show KPI levels and changes, similar to how traditional approaches such as dashboards, data catalogs, and KPI trackers may do. However, instead of this functionality running in isolation, metadata about the "status" of a metric (e.g., its level and change over time) may be displayed along with its relationship to other metrics being measured or otherwise monitored. The metric monitoring functionality shows each metric's level and changes in the context of that level along with changes in other metrics. However, in contrast to traditional approaches, this context is not based purely on concurrency (which may lead to spurious associations between metrics and inaccurate causal inferences), but rather on statistical relationships driven by the platform's underlying cataloging of machine learning models and correlation-based associations.
[0044] Although the metric monitoring capabilities are designed to be part of the disclosed platform, those skilled in the art (e.g., software engineers who understand graph databases and HTTP requests) should find the disclosure to be both enabling and possible for the implementation of the metric monitoring capabilities in a programming language of their choice. Because the purpose of metric monitoring is to track changes in important KPIs / metrics, it presumes that there is a source of data that is updated in an event-driven or otherwise automated manner (as is often the case with data sets stored in cloud database services). The frequency with which these data are updated is not critical. Thus, while metric monitoring may be of value to users in the financial services sector, where data is presumed to be updated almost continuously, it may also be used by individuals performing scientific research and working with government data (often published by government agencies) that may be updated quarterly, annually, or even decadally.
[0045] 1(a) is a block diagram illustrating a set of elements, components, functions, processes, or operations that may be part of a platform architecture 100 in which an embodiment of the disclosed system and method for metric monitoring may be implemented. A brief description of an example architecture is provided below.
[0046] architecture In some embodiments, elements or components of the architecture illustrated in FIG. 1(a) may be distinguished based on their functionality and / or based on how access to said elements or components is provided. Functionally, the system architecture 100 distinguishes between:
[0047] ○ Information / Data Access and Retrieval (illustrated as Applications 112, Add / Edit 118, and Open Science 103) - These are sources of information and descriptions of experiments, studies, machine learning models, or observations that provide the data, variables, topics, concepts, and statistical information that serve as the basis for generating feature graphs or similar data structures.
[0048] o Database (illustrated as System DB 108) - An electronic data storage medium or entity, utilizing an appropriate data structure or schema and data retrieval protocol / methodology.
[0049] Applications (illustrated as Applications 112 and Websites 116) - These execute in response to instructions or commands received from public users (Public 102), customers 104, and / or administrators 106. Applications may perform one or more processes, operations, or functions, including but not limited to:
[0050] ■ searching the system DB 108 or feature graph 110 to retrieve variables, data sets, and other information that are relevant to the user query; ■ Identifying a particular node or relationship in the feature graph; ■ writing data to the system DB 108 so that the data may be accessed by the public 102 or by others other than the customer or business 104 that owns or controls access to the data (note that in this context, the customer 104 serves as an element of the information or data retrieval architecture or source); ■ Generating a feature graph from a given dataset; Characterizing a particular feature graph according to one or more metrics or measures of complexity, relative degree of statistical significance, or other aspect or characteristic; and / or ■Generating and accessing recommendations for datasets used to train machine learning models.
[0051] ●In terms of access to and capabilities of the system 100, the system's architecture distinguishes between elements or components accessible to the public 102, elements or components accessible to a defined customer, business, organization or set of businesses or organizations (e.g. an industry consortium or a "data collaboration" in the social sector) 104, and elements or components accessible to administrators of the system 106.
[0052] Information / data about or demonstrating statistical associations between topics, concepts, factors, or variables may be retrieved (i.e., accessed and obtained) from multiple sources, which may include (but are not limited to or required to include) journal articles, technical and scientific publications and databases, digital "notebooks" for research and data science, experimentation platforms (e.g., for A / B testing), data science and machine learning platforms, and / or public websites (elements / websites 116) where users can enter observed (or unsupported) statistical relationships between observed variables and topics, concepts, or targets.
[0053] o Information and Data Retrieval Components of the architecture may, for example, scan or "read" published or otherwise accessible scientific journal articles (e.g., by using optical character recognition, or OCR) using natural language processing (NLP), natural language understanding (NLU), and / or computer vision to process images (as suggested by input / source processing element 120), identify words and / or images that indicate that a statistical association has been measured (e.g., by recognizing the term "increasing" or another relevant term or description), and in response, retrieve information and data about the association and about the dataset that measures the association (e.g., provides support for the association) (as suggested by the element labeled "Open Science" 103 in the figure and by step or stage 202 of FIG. 1(a)).
[0054] o Other components of the information and data retrieval architecture (not shown) may provide a way for users to enter code into their digital "notebooks" (e.g., Jupyter Notebook) and retrieve metadata output of machine learning experiments (e.g., measures of "feature importance" of features used in a given model) and information about the datasets used in the experiment.
[0055] o Note that in some embodiments, retrieval of information and data generally occurs periodically or continuously, providing the system 100 with new information to store, structure, and thereby release to users.
[0056] ●In some embodiments, the type of algorithm and model (e.g., logistic regression), model parameters, numerical values (e.g., 0.725), units (e.g., log loss), statistical properties (e.g., p-value = 0.03), feature importance, feature rank, model performance (e.g., AUC score), and other statistical values related to the association are identified and stored after extraction.
[0057] Given that researchers and data scientists may use different words or terms to describe the same or closely similar concepts, variable names (e.g., "aerobic exercise") may be stored as they are retrieved and then semantically grounded (i.e., linked or associated) to a public domain ontology (e.g., Wikidata) to facilitate clustering of variables (and associated statistical associations) based on common, or typically synonymous or closely related, terms and concepts.
[0058] For example, a variable labeled "log_house_sale_price" by a given user may be semantically linked by the system (and further validated by the user) to the topic "real estate prices" in Wikidata, which has a unique ID.
[0059] As disclosed herein, a central database ("System DB" 108 in the figures) stores the retrieved information and data and its associated data structures (i.e., nodes, edges, values). An instance or projection of the central database containing all or a subset of the information and data stored in the System DB is made available to a particular customer, business, or organization 104 (or group thereof) for their use, typically in the form of a "Feature Graph" 110.
[0060] Because access to a specialized feature graph may be restricted to certain individuals associated with a given business or organization, the specialized feature graph may be used to represent information and data about variables and statistical associations that may be considered private or proprietary to a given business or organization 104 (such as, by way of non-limiting example, employment data, financial data, product development data, business metrics, or R&D data).
[0061] Each customer or user is provided with their own instance of the System DB in the form of a feature graph. The feature graph typically reads data from the System DB concurrently (and in most cases frequently) to ensure that users of the feature graph have access to the most current information, data, and knowledge stored in the System DB.
[0062] ● Applications 112 may be developed ("built") on top of the feature graph 110 to perform a desired function, process, or operation. An application may read data from the feature graph, write data to the feature graph, or perform both functions. One example of an application is a recommender system (referred to herein as a "data recommender") for a dataset. Customers 104 using the feature graph 110 can "write" information and data to the system DB 108 using appropriate applications 112, which may be useful if the customer wants to share certain information and data with a broader group of users outside their organization or with the public.
[0063] o The application 112 may be integrated with a data platform and / or a machine learning (ML) platform 114 of the customer 104. An example of a data platform is Google Cloud Storage. The ML (or data science) platform may include software such as Jupyter Notebook.
[0064] ■Such data platform integrations may, for example, allow users to access features in a customer's data store or other data repository (e.g., those recommended by a data recommender application). As another example, data science / ML platform integrations may allow users to query feature graphs, for example, from within a notebook.
[0065] o It is noted that in addition to or instead of integrating with the customer's data platform and / or machine learning (ML) platform, access to the application may be provided to the customer by the administrator using a suitable service platform architecture, such as Software-as-a-Service (SaaS) or similar multi-tenant architecture. Further description of the main elements or features of such an architecture are provided herein with reference to Figures 3-5.
[0066] ● In some embodiments, a web-based application may be accessible to the public 102. On the website (represented by www.xyz.com 116), users may be enabled to read from and write to the system DB 108 (as suggested in the figure by the add / edit functionality 118) in a manner similar to that experienced with websites such as Wikipedia.
[0067] ● In some embodiments, data stored in the system DB 108 and published to the public at www.xyz.com 116 may be made available to the public in a manner similar to that experienced with websites such as Wikipedia.
[0068] Once the information and data are accessed and processed for storage in a database (which may include both raw and processed data and information, as well as data and information stored in the form of a data model), a feature graph may be constructed that includes a specified set of variables, topics, targets, or factors. A feature graph for a particular user may include all data and information in the platform database 108, or a subset thereof. For example, a feature graph (110 in FIG. 1(a)) for a particular customer 104 may be constructed based on selecting data and information from the system DB 108 that meets conditions such as a given domain (e.g., public health) being applicable to a domain of interest (e.g., media) of the customer. In developing, generating, or constructing a feature graph for a particular customer or user, the data in the database 108 may be filtered to improve performance by removing data that may not be relevant to the problem, concept, or topic being investigated.
[0069] In some embodiments or uses, the data used to generate the feature graph may be proprietary to an organization or user, for example, the data used to build the feature graph may be obtained from an experiment, a set of customers or users, or a particular database of protected data, as non-limiting examples.
[0070] Fig. 1(b) is a flowchart or flow diagram illustrating processes, methods, functions, or acts for constructing 150 a feature graph using an implementation of an embodiment of the systems and methods disclosed herein. Fig. 1(c) is a flowchart or flow diagram illustrating processes, methods, functions, or acts that may be implemented in an embodiment of the systems and methods disclosed herein for an example use case in which the feature graph is traversed to identify potentially relevant data sets and / or to perform other functions of interest (e.g., those resulting from the execution of a particular application, such as that suggested by element 112 in Fig. 1(a)).
[0071] As shown in the diagram (specifically FIG. 1(b)), the feature graph is constructed or created by identifying and accessing a set of sources containing information and data regarding statistical associations between variables or factors used in a study (as suggested by step or stage 152). This type of information may be retrieved periodically or continuously to provide information about the variables, statistical associations, and the data used to support the associations (as suggested by 154). As disclosed herein, this information and data is processed to identify the variables used or described in the sources and the statistical associations between one or more of the variables and one or more other variables.
[0072] Continuing with FIG. 1(b), sources of data and information are accessed at 152. The accessed data and information is processed to identify variables and statistical associations found in one or more sources 154. As noted, such processing may include image processing (e.g., OCR, etc.), natural language processing (NLP), natural language understanding (NLU), or other forms of analysis to aid in understanding the contents of journal articles, research notebooks, lab logs, or other records of a study or investigation.
[0073] Further processing may include linking certain of the variables (as suggested by step or stage 156) to an ontology (e.g., the International Classification of Diseases) or other set of data that provides semantic equivalents or similar terms to the terms used for the variables. This helps to extend the variable names used in a particular study to a larger set of substantially equivalent or similar entities or concepts that may have been used in other studies. Once identified, the variables (which, as noted, may be known by different names or labels) and statistical associations are stored (158) in a database, e.g., system DB 108 of FIG. 1(a).
[0074] The accessed information and results of processing the data are then structured or represented according to a particular data model (as suggested by step or stage 160). This model will be described in more detail herein, but generally includes the elements used to construct a feature graph (i.e., nodes representing topics or variables, edges representing statistical associations, and measures including metrics or evaluations of the statistical associations). The data model is then stored (162) in a database and can be accessed to build or create a feature graph for a particular user or set of users.
[0075] As noted, the process or operations described with reference to FIG. 1(b) allows for the construction of a graph containing nodes and edges linking certain ones of the nodes (an example of which is illustrated in FIG. 1(d)). The nodes represent topics, targets, or variables of study or observation, and the edges represent statistical associations between a node and one or more other nodes. Each statistical association may be associated with one or more of a numerical value, a type of model or algorithm, and a statistical property that describes the strength, confidence, or reliability of the statistical association between the nodes (i.e., variables, factors, or topics) connected by the edge. It is noted that the numerical value, type of model or algorithm, and statistical property associated with an edge may indicate, by way of non-limiting examples, a correlation, a predictive relationship, a cause-effect relationship, or a weakly supported observation.
[0076] 1(c) is a flow chart or diagram illustrating a process, method, function or operation 190 that may be used to build a feature graph for a user, according to one embodiment of the disclosed systems and methods. In one embodiment, this may include the following steps or stages (some of which overlap with those described with reference to FIG. 1(b)):
[0077] ● Identifying and accessing source data and information (as suggested by step or stage 191). o In one embodiment, this may refer to data and information available to the public from journals, research periodicals, or other publications that describe studies or investigations.
[0078] o In one embodiment, this may represent proprietary data and information such as experimental results generated by the organization, research topics of interest to the organization, or data collected by the organization from customers or clients.
[0079] ● Processing the accessed data and information (as suggested by step or stage 192). In one embodiment, this may include identification and extraction of information regarding one or more of the topics of a study or investigation, the variables or parameters considered in the study or investigation, and data or datasets used to establish statistical associations between one or more variables and / or between variables and topics, along with measures of such statistical associations in the form of metrics, relationships, or similar quantities.
[0080] In one embodiment, this process may be performed automatically or semi-automatically by using trained models that utilize language model or language embedding techniques to identify interesting or relevant data and information.
[0081] ● Storing the processed data and information in a database (as suggested by step or stage 193). In one embodiment, the database may include one or more partitions that separate data obtained from an organization, set of sources, or set of populations into separate data sets to be used to generate the feature graph.
[0082] ■This can be a useful approach when the data set is obtained from a proprietary study, a specific population, or otherwise complies with regulations or constraints (e.g., privacy or security regulations, etc.).
[0083] o In some embodiments, the processed data and information may be stored according to a particular data schema that includes specific labels or fields. ● Receiving user input indicating topics of interest (as suggested by step or stage 194) and, in response thereto, generating a feature graph.
[0084] In one embodiment, user input may specify sources, dates, thresholds, or other forms of constraints that are used as a filtering mechanism for the data and information used to generate the feature graph.
[0085] ● Traversing the feature graph (as suggested by step or stage 195) to evaluate the data, information, and metadata used to generate the feature graph. ○ This may involve filtering the data and information represented by the feature graph according to rules, constraints, thresholds, or other conditions prior to the evaluation process.
[0086] o This may involve evaluating data, information and metadata in a process flow determined by a particular application or set of controls or instructions. ■In one embodiment, this may include, by way of non-limiting example, aggregating statistical data and / or metadata, identifying statistically relevant or significant relationships, or generating specified metrics or indices of relationships or variable values.
[0087] ■In one embodiment, this may involve evaluating the aggregated data using rule sets or conditions to identify potentially important variables or relationships, or to alert a user so that certain conditions can be addressed.
[0088] ■In one embodiment, this may involve performing some type of network analysis on the nodes in the layer to identify network characteristics. ● Presenting the results of the graph traversal and evaluation to the user (as suggested by step or stage 196).
[0089] o In one embodiment, this may involve separating the topics, variables, and data used to generate the feature graph into distinct layers of nodes and connecting edges between the nodes and layers.
[0090] In one embodiment, this may include showing the user the relationship between two nodes that have a certain characteristic (e.g., strength, recentness, above a threshold, or more trustworthy, etc.).
[0091] o In one embodiment, this may involve presenting the user with a list or table specifying concepts or topics that affect or are affected by the input concept or topic, with metadata about the properties of this relationship.
[0092] o In one embodiment, this may involve associating a set of variables or topics with a metric and showing the value and / or changes of that metric to the user. In one embodiment, this may include representing the relationship between two variables, two topics, or between a variable and a topic using one or more metrics or indicators (e.g., flags, alerts, or colors) related to the statistical relationship between the entities.
[0093] 1(d) is a diagram illustrating an example of a portion of a feature graph data structure 198 that may be used to organize and access data and information, and that may be created using implementation of an embodiment of the systems and methods disclosed herein. A description of the elements or components of the feature graph 198 and the associated data model implemented is provided below.
[0094] Feature Graph As noted, a feature graph (in the context of this disclosure, the term "feature graph" is used because the embodiments assemble graphs from entities connected through statistical relationships between variables (measures of interest), referred to herein as features, instead of semantic co-occurrences (as in traditional "knowledge graphs")) is a way of structuring, representing, and storing statistical relationships between topics and their associated variables, factors, or categories. The core elements or components (i.e., "building blocks") of a feature graph are variables (identified as V1, V2, etc. in FIG. 1(d)) and statistical associations (identified as connecting lines or edges between the variables). Variables may be linked or associated with "concepts" (one example of which is identified in the figure as C1), which are typically semantic concepts or topics that are not directly measurable or measurable in a useful way by themselves (e.g., the variable "number of robberies" may be linked to the concept "crime"). A variable is an empirical object or factor that can be measured. In statistics, association is defined as "a statistical relationship, whether causal or not, between two random variables." Statistical associations arise from one or more steps or stages of what is often called the scientific method and may be characterized as, for example, weak, strong, observed, measured, correlative, causal, or predictive.
[0095] o As an example and with reference to Fig. 1(d), a statistical search of an input variable V1 retrieves: (i) variables (e.g., V6, V2) that are statistically associated with V1 (in some embodiments, a variable may be retrieved only if the statistical association value exceeds a defined threshold), (ii) variables (e.g., V5, V3, V4) that are statistically associated with the variable (in some embodiments, a variable may be retrieved only if the statistical association value exceeds a defined threshold), (iii) variables (e.g., V7) that are semantically related by a common concept (e.g., C1) to one or more variables (e.g., V2) that are statistically associated with the input variable V1, and (iv) a data set (e.g., D6, D2, D5, D3, D4, D7, D8) that measures the association of variables associated with the variable (e.g., V8) that is statistically associated with the variable or that demonstrates the statistical association of the retrieved variables.
[0096] ■Note that, in contrast to the disclosed embodiment, a semantic search of an input variable V1 retrieves (1) the variable V1 and (2) a dataset (e.g., D1) that measures that variable.
[0097] The feature graph is populated with information and data about statistical associations extracted from (for example) journal articles, scientific and technical databases, research and data science digital "notebooks," experiment logs, data science and machine learning platforms, public websites where users can enter observed or perceived statistical relationships, proprietary business information, and / or other possible sources.
[0098] As noted, by using techniques from Natural Language Processing (NLP), Natural Language Understanding (NLU), and / or Image Processing (OCR, video / image processing, and recognition), the information and data retrieval architecture components (an example of which is shown in Figure 1(a)) can scan or "read" published scientific journal articles, identify words or images that indicate that a statistical association (e.g., "increase") has been measured, and retrieve information and data about that association and about data sets that measure or confirm that association.
[0099] o Information and Data Extraction Another component of the architecture provides data scientists and researchers with a way to enter code into their digital "notebooks" (e.g., Jupyter Notebook) and extract metadata output of machine learning experiments (e.g., "feature importance" measurements for features used in the model) and information about the datasets used in the experiments. Note that information and data extraction occurs periodically, and in some cases continuously, providing the system with new information to store, structure, and expose to users.
[0100] In one embodiment, datasets are associated with variables in the feature graph using links to the URIs or other forms of access or addresses of the associated datasets / buckets / pipelines.
[0101] ○ This allows users of the feature graph to retrieve datasets based on the previously proven or determined predictive power of that data for a specified target or topic (rather than datasets that are semantically related to a specified target or topic and potentially less relevant or irrelevant for the topic based on semantic co-occurrence between sources, as in traditional knowledge graphs).
[0102] For example, using one embodiment of the systems and methods disclosed herein, if a data scientist searches for "vandalism" as a target topic or goal of research, rather than a dataset measuring instances of vandalism, they would retrieve datasets for topics that have been shown to predict that target or topic, such as "household income," "lightness," and "traffic density" (and evidence of their statistical association to the target).
[0103] ● The numerical value (e.g., 0.725) and statistical properties (e.g., p-value=0.03) of the association may be retrieved and stored in the system DB 108 and made available as part of the constructed feature graph. As noted, given that researchers and data scientists may use different words to describe the same or similar concepts or topics, the variable names (e.g., "cardio") may be retrieved and stored and semantically grounded in public domain ontologies (e.g., Wikidata), dictionaries, thesauri, or similar sources) to facilitate clustering of the variables (and associated statistical associations) based on common or similar concepts (e.g., synonymous terms or terms that are understood to be interchangeable in an industry).
[0104] ●In one sense, system 100 uses mathematical, language-based, and visual methods to represent the epistemological and fundamental characteristics of available data and information, such as (as non-limiting examples) the quality, rigor, veracity, reproducibility, and completeness of the information and / or data supporting a given statistical association.
[0105] For example, a given statistical association may be associated with a particular score, label, and / or icon in the user interface based on its scientific quality (overall and / or with respect to a particular parameter, such as "peer reviewed") to present to the user information that the user can use to make a decision on whether to further investigate the association. In some embodiments, statistical associations retrieved by searching the feature graph may be filtered based on their "scientific quality" score. In certain embodiments, the calculation of the quality score may combine data stored within the feature graph (e.g., the statistical significance of a given association or the extent to which the association has been documented) with data stored outside the feature graph (e.g., the number of citations of the journal article from which the association was retrieved or the h-index of the article's authors).
[0106] For example, statistical associations with features that contain high and significant "feature importance" scores measured in models that have high area under the curve (AUC) scores, have partial dependence plots (PDPs), and are documented for reproducibility may be considered "strong" (and, as expected, more reliable) statistical associations in the feature graph and may be given a distinguishing color or icon in the graphical user interface.
[0107] o Note that in addition to retrieving variables and statistical associations for a topic or concept, an embodiment may also retrieve other variables used in an experiment or study to contextualize the statistical associations for the user. This may be useful if the user wants to know (for example) whether a particular variable was controlled in the experiment or what other variables (or features) were included in the model.
[0108] Data Model A major object in the feature graph (or system DB) will typically contain one or more of the following, along with an indication of information that may help define the object:
[0109] • Variables (or characteristics)--What are we measuring and in what population? ● Concept--What is the topic, hypothesis, idea, or theory you are studying? Neighborhood--What is the subject (which is typically broader than a concept) that you are measuring? Statistical Association--What is the mathematical basis and value of the relationship? • Model (or experiment)--What is the source of this measurement? Datasets--What datasets were used to suggest or measure relationships (e.g., model training data) or to measure variables? These objects are related, as illustrated in the example feature graph in FIG. 1(d).
[0110] • Variables are linked to other variables through statistical associations. • Statistical associations arise from models and are supported by data sets. • Variables are linked to concepts, and concepts are linked to neighborhoods (or parts of them).
[0111] As noted, with reference to FIG. 1(d), one use of the feature graph is to allow a user to search the feature graph for one or more data sets that contain variables that have been demonstrated to be statistically associated with a target topic, variable, or concept of study. An example use is as follows.
[0112] ●A user inputs a target variable (as suggested by process 170 in Figure 1(b)) and wants to retrieve a dataset that can be used to train a model to predict the target variable, i.e., a dataset linked to variables that are statistically associated with the target variable.
[0113] 1(d), a statistical search input V1 (in this case a variable) causes an algorithm (e.g., a breadth-first search (BFS)) to traverse the feature graph (as suggested by step or stage 174 of FIG. 1(b)) and return (as suggested by step or stage 176 of FIG. 1(b)):
[0114] ■ Variables statistically associated with V1 (e.g., V6, V2), • In some embodiments, variables may be removed only if their statistical association value exceeds a defined threshold.
[0115] ■ Variables that are statistically associated with the variable in question (e.g., V5, V3, V4), • In some embodiments, variables may be removed only if their statistical association value exceeds a defined threshold.
[0116] ■ One or more variables (e.g., V2) that are statistically related to the input variable V1, and a variable (e.g., V7) that is semantically related to the input variable V1 through a common concept (e.g., C1), and ■ A variable that is statistically associated with the variable in question (e.g., V8), and ■ Data sets that measure or establish the statistical significance of the extracted variables (e.g., D6, D2, D5, D3, D4, D7, D8).
[0117] After traversing the feature graph to retrieve potentially relevant datasets, the datasets may be “filtered,” ranked, or otherwise ordered based on application or use case (as suggested by step or stage 178 of FIG. 1(b)).
[0118] o Data sets retrieved through the described traversal process may subsequently be filtered based on criteria entered by users with their searches and / or based on criteria entered by administrators of the instance of the software. Exemplary search dataset filters may include one or more of the following:
[0119] ■ Populations and Keys: Are the variables of interest measured in the populations and keys (e.g., unique identifiers for users, species, cities, or companies, for example) that are of interest to the user? This impacts the user's ability to link the data into training sets for use with machine learning algorithms.
[0120] ■ Compliance: Does this data set meet applicable regulatory considerations (e.g., GDPR compliance or HIPAA regulations)? ■ Interpretability / Explainability: Are the variables interpretable or understandable by humans? ■ Immediately available: Is this variable immediately available to users of the model? In one embodiment, a user may input a concept (represented by C1 at 198 in FIG. 1(d)) such as "crime," "property," or "high blood pressure." In response, the systems and methods disclosed herein may use a combination of semantic and / or statistical search techniques to identify one or more of the following:
[0121] A concept (C2) that is semantically related to C1 (note that this step can be optional); ● Variables that are semantically related to C1 and / or C2 (V X ), Variable V X A variable that is statistically associated with each of the One or more measures of the identified statistical associations, as well as Variable V X and / or the variable VX A data set that establishes or supports the statistical association of a variable that is statistically associated with each of the following:
[0122] Figure 2(a) is a block diagram illustrating a set of elements, components, functions, processes, or operations that may be part of a platform architecture in which an embodiment of the disclosed system and method for metrics monitoring may be implemented. Figure 2(b) is a flow chart or diagram illustrating a set of elements, components, functions, processes, or operations that may be performed as part of a platform architecture in which an embodiment of the disclosed system and method for metrics monitoring may be implemented. In particular, Figure 2(b) depicts certain of the steps in Figure 2(a) with a greater focus on the different user interactions and software elements that contribute to how the metrics monitoring functionality is implemented and made available to the user.
[0123] 2(a) illustrates how changes in features from a dataset stored in a cloud database service (or “data warehouse” 204) can be monitored using an implementation of the disclosed metric monitoring capabilities. The blocks representing elements, functions, or operations in the left column (represented by element 202) (e.g., dataset metadata 206) are examples of how features and metrics are represented on a system platform (along with measured statistical relationships between features), while the blocks representing elements, functions, or operations on the right (represented by element 203) illustrate user interactions, user inputs, and software calculations or other executed code that the platform may use to process and store metadata about the dataset and its features.
[0124] In some embodiments, the steps, stages, functions, operations, or process flows illustrated in FIG. 2(a) may include a process step in which the platform's data warehouse retrieval integration computes relevant metadata and sends the metadata (typically via an HTTP request) to the platform's backend API. The backend service stores the metadata in the platform's graph database (e.g., element 108 of FIG. 1(a)), which contains data supporting the functionality of the feature graph. The feature graph is what a user sees and interacts with using the platform's front-end user interface and the platform's generated user interface.
[0125] A user can interact with the platform's front-end user interface to identify features of interest, and when the feature has the desired form (i.e., the feature has a numerical value associated with a timestamp), the user can define a metric for monitoring, connect the metric to the feature, and activate the metric monitoring function. Metric monitoring provides the user with a visual display (on the feature graph) based on the value or value changes in the metric (and in the platform's underlying data), and may generate alerts and notifications via email or within the platform application itself.
[0126] As noted, a metric monitoring feature or capability would show changes in metrics in the context of one another. For example, as suggested in FIG. 2(a), a user of the platform would be able to see changes in metric 1 (208) along with changes in metric 2 (210) along with a description of the statistical relationship (as suggested by data 209 and 211, respectively) measured between the metrics. The platform's context for showing changes in both metrics not only displays the current levels and changes in the metrics, but may also use outputs from machine learning models and other statistical relationships between underlying features connected to the metrics to generate and display data and information to the user.
[0127] Figure 2(b) illustrates certain of the steps in Figure 2(a), focusing more on the user interactions and software elements that contribute to how the metrics monitoring functionality is implemented and made available to the user. Each step, stage, element, function, or operation in this diagram corresponds to a software component (or software service) of the disclosed platform that contributes to enabling the metric monitoring capability to be used by the user. In the example illustrated in Figure 2(b), the components shown are (in order from top to bottom in the diagram):
[0128] ● As suggested by step, stage, action, process or function 250, users can add data sets for tracking on the platform through integration with a database service (data warehouse).
[0129] ● As suggested by step, stage, action, process, or function 252, the platform's retrieval service computes the relevant dataset and feature metadata and submits an HTTP request to the platform's backend API.
[0130] ●As suggested by step, stage, action, process, or function 254, the platform's backend API processes the data payload included in the request and prepares a dataset and / or feature metadata for storage.
[0131] ●As suggested by step, stage, action, process, or function 256, the platform's backend services store the dataset and / or feature metadata, as well as statistical relationships, in a graph database.
[0132] ●As suggested by step, stage, action, process, or function 258, the platform's backend services connect the new metadata from the retrieval process to existing metadata in the graph database, so that datasets and features are connected to existing objects, if applicable (note that this is an optional step and depends on the existing graph database content).
[0133] ● Platform metadata is made available on the platform front end, allowing users to see connections between objects (in one example, datasets and features) that are part of the feature graph. Users can also make connections between features and metrics that they are using to track their KPIs or key metrics, as suggested by step, stage, action, process, or function 260.
[0134] ●When the features are in the correct form (e.g., the data has associated time indexes, as suggested by element 264), as suggested by step, stage, operation, process, or function 262, the platform may show the features and metrics along with their latest values and recent changes and prompt the user to turn on metric monitoring.
[0135] o The platform or system may also prompt the user to turn on metric monitoring and suggest important objects to monitor if their features and metrics have a significant relationship with the currently monitored metric.
[0136] ● As implied by the steps, stages, actions, processes, or functions 266, users can set rules for metric monitoring that control the visual display / differentiation presented for the monitored metrics and generate alerts and notifications on the platform via email. These rules are written to the platform backend and stored in the feature graph.
[0137] The conditions set by the user are then evaluated to generate displayed visual differentiation, alerts, and / or notifications, as suggested by steps, stages, actions, processes, or functions 268. The platform backend also tracks the status of metric monitoring, as described above, to uncover significant or important relationships between metrics, and to make recommendations.
[0138] ●As suggested by the steps, stages, actions, processes or functions 270 and the control loops connecting the steps, stages, actions, processes or functions 254, these steps or processes are performed iteratively such that new information or data extracted results in changes to the data that the user is interested in monitoring.
[0139] In some embodiments, the disclosed platform includes software as part of its architecture to automatically retrieve data from remote databases, process the data, and write computed metadata (including metadata about statistical relationships between features in a dataset) to platform data storage. The architecture is based on microservices designed to run in a scheduled and / or event-driven manner. However, this form of implementation may not be required if updated data is "pulled" from the source and written to a storage location where metric monitoring software and functionality can access it. As noted, for purposes of implementing metric monitoring functionality, it is desirable for the data to be retrieved in a manner that correlates values of interest in the data with specific time intervals or other forms of indexing.
[0140] For example, an associative array in JavaScript can be used to associate a data value with a particular timestamp object, i.e., {"2010-01-01 00:00:00Z":10.4, "2010-01-02 00:00:00Z":11.2}, where the "key" of the associative array represents a timestamp in the "UTC" time standard, and the number following the key represents the data value associated with that timestamp. This is one non-limiting example of a data structure that can hold a numerical value and associate the number with a particular timestamp.
[0141] Embodiments may include specific techniques for interpolating and aggregating data across different time intervals and specifying data values to be associated with a time interval. The metric monitoring functionality disclosed herein will assist the user regardless of the method used to "make decisions" about the time interval or index associated with each value. However, since users will typically rely on data to understand how the metrics of interest are changing over time, the methodology for doing this should be transparent to the user.
[0142] Where electronic storage of data is performed with a timestamp associated with the value of the data, in one embodiment software implementing metric monitoring functionality may include the following data organization operations or processes:
[0143] When timestamps are sorted in "descending" time order, the "current" or "latest" value is the value associated with the first timestamp. The "previous" value is the value associated with the next-to-last timestamp in the "descending" time order (see elements 209 and 211 in Figure 2(a)).
[0144] When only one value is present, the "previous" value is given the value "not available", "N / A" or "not a number" and the percent change is shown as "not available" (or "N / A" or "not a number"). When neither of these two values is a number, both values are given the value "not available", or "N / A" or "not a number", as in the percent change.
[0145] ● Otherwise, the percent change is calculated as the current value minus the previous value divided by the previous value. If the previous value is zero, the platform may represent the percent change as "Inf" instead of "infinity".
[0146] On the platform, values are stored in a graph database and available to a backend API via HTTP requests. Although the percent change can be calculated for the user using "frontend" techniques, in some embodiments, metric monitors write the percent change values to metric objects in the graph database. This is desirable and recommended because users may want to query the backend API to get information about the process or status of the metric monitor.
[0147] Another aspect of the implementation of metric monitoring capabilities is the configuration and evaluation of "rules" for monitoring (as suggested by functions, operations, or processes 212 and 213 in FIG. 2(a)). In one embodiment, part of the platform architecture includes parameterization of comparison / alert rules, where monitoring rules are represented by "triples" of "field", "operator", and "value".
[0148] o "Field" refers to a field of a metric monitoring object stored in the graph database. This field could be "Latest Value", "Percent Change", or other metadata that can be used by the metric monitoring capability to allow users to monitor KPIs or metrics. This field is designed to be flexible; latest value and percent change are commonly tracked values, but a user might want to track "All-Time High" or "52-Week Low" as an example, in the case of two commonly tracked financial metrics.
[0149] o The "Value" field is a value that the user can specify (and may have a default value) and that serves as the basis for comparison in the rule. Because metric monitoring is numerical in nature, it is expected that users will specify this "value" in numerical terms.
[0150] o The "Operator" field represents how a mathematical comparison will be made between the value of the monitored metric's "Field" and a user-specified "Value" (which, as mentioned, may be suggested to the user by the metric monitoring feature). For example, the operator may be specified as "absolute value is greater than", meaning that the absolute value taken from the value referenced in "Field" will be compared with the "Value" entered to see if the absolute value is greater than that "Value".
[0151] ■ The definition of an "operator" is preferably flexible enough to encompass monitoring rules that may involve calculations or "aggregations" of values stored in "fields." Implementation of this capability may include an enumeration of operators, where a predefined software function (if the programming language being utilized allows) implements each operator.
[0152] ● The metric monitoring capability includes visual elements that allow users to quickly see the levels and changes of their monitored metrics. In one implementation of metric monitoring, metrics that require attention or are in the "alert" phase are depicted either with a non-default color selected by the user, with a specified format (such as italics or bold, as examples), or with an icon (for users who prefer not to distinguish user interface elements with color or format). The color or format selection is saved as part of the monitoring rule.
[0153] The metric monitoring capability may include a user interface that allows the user to specify the desired monitoring rules. In one embodiment, this is a language-based "drop-down menu" feature that allows the user to choose from a set of available "fields", "operators", and then set a "value" to specify the rule. These defined triples (based on user input) are stored in the graph database as properties associated with the metrics of interest.
[0154] One implementation of metric monitoring may also allow a user to see what the results of the monitor will be as they specify or define the rules. For example, if a monitoring rule is to set a visual element to green when the latest value is greater than 0, then the latest value field on the monitored data will be set to green if the latest value of the metric is in fact greater than 0. If a monitoring rule is to set a visual element to blue when the percent change is less than 10%, then the percent change value on the monitored data will be blue if this condition is met. This will change back to a default color or appearance if the user subsequently changes the value in the rule to a comparison value where this condition no longer holds.
[0155] The difference between the metric monitoring capabilities disclosed herein and other cataloging, dashboard, or analytics tools is that users can see their monitoring information in its full context, along with the results of modeling or other sources of data showing statistical relationships. This is a property of the disclosed platform, and the implementation details for showing relationships involving monitored metrics are related to how the disclosed platform was designed and implemented.
[0156] o In this regard, the disclosed platform is built on a graph database, whereby each metric object being monitored has a potentially rich network of connections or "edges" with other objects. The visual element of metric monitoring is particularly useful to the user when there are many relationships in the graph and many are being monitored. In this case, the user can see the different connections and understand how and why their chosen metric has the indicated "pattern" of statistical variation.
[0157] In one embodiment, implementing a metric monitoring function involves not only specifying a data structure to which monitoring rules can be applied, but also having a storage technique that allows metrics of interest to be correlated across different metadata.
[0158] ●In some embodiments, implementations of the metric monitoring functionality, in addition to the features or capabilities described, may also include the ability for users to discover or be informed of optimal (or more optimal) rules, thereby learning more about the systems and relationships represented by their data.
[0159] ● Note that in the absence of predefined business rules or published goals for (for example) KPIs / metrics, a user may not be aware of how to best define rules for metric monitoring. In one embodiment, assistance with this may be provided by a recommendation function that operates to suggest values / metrics for monitoring based on metadata collected about the features and metrics in question.
[0160] o As one non-limiting example, a critical value may be suggested when the value for a feature or metric rarely rises above or falls below a certain numerical boundary, where a user would expect to be alerted or notified only a percentage of the time. Alternatively, the feature and metric in question may be similar to another feature or metric, and the recommended rule may be to monitor both metrics in the same way.
[0161] ■The disclosed platform, graph database (System DB) and backend infrastructure gives users the ability to view data and metadata from multiple sources as a system. This design allows developers and users to quickly query for features, variables and relationships (nodes and edges in the graph) that have similar statistical properties and / or similar characteristics within their metadata.
[0162] ■ This information is specific to the disclosed platform and can be used to discover plausible candidates for metric monitoring even in the absence of user-defined metric monitoring rules or other predefined business rules. For example, a "built-in" recommendation feature can suggest monitoring rules based on many of these statistical characteristics or properties.
[0163] ■ Recommendation feature implementations can include queries and code that identify actual KPIs, such as measures of active users (which often predict sales and revenue). In some embodiments, these metrics can be based on one or more of the following: (1) statistical characteristics (e.g., highly predictive of other features or strongly correlated with other measures important to the business); (2) metadata, including feature or variable names, i.e., presence as a feature in multiple data sets, or tracked over relatively longer time intervals; or (3) usage measures, such as how many times users visit a variable or feature page relative to others.
[0164] ■The recommendation feature can suggest "smart" monitoring rules based on the statistical properties or metadata of a metric. Training data on how to implement these rules can also be sourced from a public version of the platform, where users can set up metric monitoring rules on data from various sources, and the effectiveness of the rules (how often the rules are triggered and how users respond to those alerts) can drive iterative improvements to the performance of the recommendation rules.
[0165] In one embodiment, a "building block" for recommendation functionality is to index similarities in statistical properties, rather than just measuring metadata similarity across different features and metrics. In contrast, in a typical data warehouse, generating statistical relationships between features for every feature is often a difficult and computationally expensive activity.
[0166] ● Such recommendation functionality may be implemented using suggested, rule-based similarity expressions or relationships. As a non-limiting example, the first suggested rule could be to set the same rule for every semantically similar metric. One way to implement this would be to index the value of the metric's name (and possibly other metadata about that metric) in the search service, and when a user sets monitoring rules for different metrics, have a similarity score calculated for each of the other monitored metrics, and then whatever default rules exist, the rule associated with the most similar metric is suggested.
[0167] o Another possible implementation feature is to suggest monitoring for metrics that are not part of the dataset retrieval / update process. ■As one non-limiting example, model performance metrics, if updated periodically, may look similar to the timestamp-indexed value arrays used for the metric monitoring functionality (which may be represented by timestamp-indexed value arrays, as noted). These may be stored as metadata associated with the model object and are available to users of the disclosed platform. A user interface for the platform may present these time-indexed model performance metrics as an additional feature that can be connected to and monitored by other metrics.
[0168] ■When a model performance metric has a timestamp associated with it, a separate software service or function may operate to look for other arrays of data with the same timestamp index (this may result from the use of methods to interpolate or extrapolate between instances of time, if necessary) and calculate time series analysis values to develop robust relationships between features indexed by time.
[0169] The disclosed metric monitoring functionality is intended to provide users with the full statistical context and relationships of their monitored KPIs or other metrics. To do so, the platform front-end draws a feature graph that is built using the platform's architecture and the metadata it collects and identifies. Visual cues from the metric monitoring functionality, combined with visual cues from the feature graph, help users develop a deeper and more complete understanding of how the data in the graph is related.
[0170] User interface (UI) displays associated with a metric monitoring capability are generated from data stored on the platform backend. When a metric monitoring capability or feature is activated, the platform frontend applies the defined monitoring rule(s) to the most recent value of the metric and any relevant previous values, which may result in changes to the view presented to the user by the platform.
[0171] In one embodiment, before rendering the visual representation of the metric nodes in the feature graph, either part of the platform or for a specific metric page generated by the platform, front-end JavaScript code is used to process the defined rules, which are typically stored on the metric object itself. As mentioned, a rule may be expressed as a collection of:
[0172] Value (i.e., a critical value or threshold against which the value of the metric will be compared); A field (the source of the metric value to be compared as part of the rule, e.g., the level of the most recent value, or the percentage change between the most recent value and the previous value), and • The operator (how the relevant field should be compared to the rule's value, e.g. "greater than or equal to" or "strictly less than").
[0173] Rules can be selected or defined in one or more places in the platform architecture where metadata about metrics can be edited. In one embodiment, this includes the metric page, the metric "card" (where the metric is referenced as part of another object, such as in a model or dataset), and in the matching console where users can match metrics to features. In one embodiment, rule setting can consist of three steps:
[0174] ● Setting a "rule", which means choosing a threshold or condition for when a metric level or change determines that the user should be alerted. Specify how any rule "violations" or alerts should be displayed visually (e.g., either through color, formatting, or icon material); and How alerts should be communicated to users (e.g. users may be able to choose how they are notified, such as via email or on-platform notifications, and how frequently these alerts should be communicated).
[0175] Once a rule is defined, the rule definition may be displayed on the metrics page. In one embodiment, the metric monitoring function may be performed regardless of whether rules are set. If rules are not set, the presentation of the metric will not trigger an alert (either via a notification or visually on the platform), but each time the metric is displayed (e.g., in a platform graph, on a metrics page, and / or in a catalog where the metric is tracked), the latest value, the previous value, and the percent change between the two values may be displayed.
[0176] Metric values are generated by the platform front end using a graph query that finds appropriate values of the features used to measure the selected metric. When only one feature with time-specific (indexed) data is connected / associated with a metric, that feature is used for the metric monitor value. When multiple features with time-specific data are connected to a metric, the first feature connected to the metric is, by default, the feature used for the metric monitor value (although the user may change this default to another feature). In one embodiment, the features that provide values for the metric monitor may be displayed at the top of the metrics page with a link to the feature, allowing the user to review each of the features used to generate the metric monitor data.
[0177] The disclosed platform and data model captures information about datasets and models to help users manage, discover, and use statistical relationships generated from correlations and associations produced by machine learning models. The platform data model is built using a graph architecture that designates features, datasets, models, and other objects as nodes, and the platform stores edges between those objects and objects created by the platform that encode information about the relationships.
[0178] The platform tracks (and may calculate) relationship strengths based on statistical properties of the datasets and models. In one embodiment, the platform may be periodically updated with scientific standards for how to assess relationship strengths, including standard measures of statistical significance (e.g., calculated confidence intervals and various forms of statistical hypothesis testing), statistical "rules of thumb" (e.g., traditionally accepted levels of effect size, as defined by Cohen (1962)), and other sources of specific domain knowledge coded into the platform's backend and machine learning pipeline.
[0179] The platform's processing of discovered and learned statistical relationships, sourced from correlations and machine learning models calculated by the platform, results in feature graphs that form the basis of the metric monitoring capabilities and functions. The disclosed metric monitoring capabilities and functions provide users with periodically updated metric values from different data sources and can inform users of important or significant changes in metric levels or metric growth rates. Thus, the feature graphs can be used to inform users about changes in KPIs / metrics that can or should be expected. Correlations and machine learning models added to the platform, including data from the current time interval, can be incorporated into the statistical relationship measurements. This has the effect of enabling the platform to constantly "learn" and improve the knowledge and data that users can access and utilize in making decisions.
[0180] As disclosed, the data used to generate the user interface displays for the platform is stored in a graph database. The graph database includes feature nodes that can be connected to nodes that summarize statistical information about each of the features, and edges between the features and "association" nodes that aggregate and summarize the statistical relationships between the features. Feature nodes may also have edges to metric nodes, where users (and the platform) store metadata about the metrics and tracking or supporting information about the metrics.
[0181] In some embodiments, the disclosed systems and methods provide users with the ability to monitor business-related metrics (such as, for example, KPIs) and more efficiently assess the quality of the underlying data used to generate those metrics. This ability is expected to enable users to make more informed decisions regarding the operation of their business. In some embodiments, this may include implementation of one or more of the following features or capabilities:
[0182] Creating a feature graph comprising a set of nodes and edges, where: ○ A node represents one or more of the following (as non-limiting examples): a concept, a topic, a dataset, metadata, a model, a metric, a variable, a measurable quantity, an object, a property, a feature, or a factor.
[0183] ■In some embodiments, nodes may be created in response to (by way of non-limiting examples) discovering (or gaining access to) a dataset, metadata, or model, generating output from a trained model, generating metadata about a dataset, or developing an ontology or other form of hierarchical relationships.
[0184] o An edge represents (by way of non-limiting example) a relationship between a first node and a second node, for example a statistically significant relationship, a dependency, or a hierarchical relationship. ■In some embodiments, an edge may be created connecting a first node and a second node to represent a statistically valid relationship between the two nodes as determined by a machine learning model or other form of evaluation.
[0185] ○ A label associated with an edge may indicate an aspect of the relationship between the two nodes connected by the edge, such as (as non-limiting examples) the metadata on which the relationship between the two nodes is based, or a dataset that supports a statistically significant relationship between the two nodes.
[0186] - Providing a user with user interface displays, tools, features, and selectable elements to enable the user to perform one or more of the following functions or actions:
[0187] Identifying metrics of interest (e.g., KPIs) to monitor or track. ■Here, the metric of interest may be generated by (as non-limiting examples) a trained model, formula, equation, or rule set, and may further be based on, generated from, or derived from underlying data that is a function of time (i.e., indexed by time).
[0188] Defining rules that describe when alerts or notifications should be generated regarding the behavior of a specified metric. ■This may be based on (as non-limiting examples) an absolute value, a change to a value, a percentage change, a percentage change over time, or a threshold.
[0189] o Defining how the results of applying the rules should be identified or indicated on a user interface display, such as (as non-limiting examples) by color, icon, or formatting.
[0190] ○ Allowing a user to select the metric for which an alert was generated and, in response, being provided with information regarding one or more of the metric's value change over time, the satisfied or enabled rules that resulted in the alert or notification, the metric's relationship to other metrics (if relevant), and available information about (by way of non-limiting examples) the datasets, machine learning models, rules, or other factors used to generate the metric.
[0191] ●Generating recommendations for the user regarding one or more of different metrics or sets of metrics that may be worth monitoring, data sets that may be useful for scrutiny, metadata that may be relevant to the identified metrics, or other aspects of the underlying data or metrics that may be of potential interest to the user.
[0192] ○ Here, the recommendations may arise (at least in part) from output generated by trained machine learning models, statistical analysis, research, comparison with other metrics or datasets, or other forms of evaluation.
[0193] The disclosed metric monitoring capabilities and functionality improve the process of KPI (or other metric) monitoring and data quality analysis in an integrated manner. The metric monitoring capabilities provide data quality monitoring that measures statistical characteristics of a data set, such as, but not limited to, the percentage of missing values in the data or changes in summary statistics (e.g., minimum, maximum, or average), and allow users to visualize and understand changes in the data in a contextual environment.
[0194] In some embodiments, users may receive alerts or notifications indicating changes in data, where these changes are compared across data sets from different sources and displayed along with relevant metadata about the data sources and / or monitored metrics. In contrast to traditional dashboards that display KPIs in a detached manner, the disclosed systems and methods also display monitored metrics in a graphical format or representation as part of (or in conjunction with) a feature graph. This allows important statistical relationships between metrics to be recognized and allows users to identify important metric "co-movements." This capability provides users with an efficient and effective way to assess current levels and / or growth rates of metrics and to forecast future levels and growth rates of related metrics.
[0195] As described, an embodiment of the disclosed system and method for monitoring metrics and evaluating statistical associations of underlying data sets can be used in conjunction with the referenced platform operated by the assignee. This platform can be used to illuminate to users the underlying relationships that govern tasks, teams, enterprises, and communities. In one sense, the task of a data team is to generate understanding through the collection and analysis of data. The disclosed platform can be used to aggregate that information and display the environment and context of the resulting knowledge to users. Similarly, a team may measure KPIs or other metrics to gauge the relative health of a particular part of their team, enterprise, or community. The disclosed metric monitoring functionality provides the team with a better and more complete understanding of the health of the team (or enterprise, or community) as reflected or indicated by a set of metrics.
[0196] In one embodiment, the "system" platform or platforms referenced herein and described in U.S. Patent Application Serial No. 16 / 421,249, entitled "Systems and Methods for Organizing and Finding Data" (now issued U.S. Patent No. 11,354,587) include an "extraction" tool that performs automated extraction of metadata and statistical characteristics from a data set (as part of the software integration with the database service). This automated extraction capability allows the platform to store time-indexed statistical metadata. In one embodiment, when a time-indexed feature (such as a variable or parameter, for example) exists, the user can indicate through the user interface that this is a metric that they would like to monitor. If a metric is monitored, the user can be shown the current "level" of the data used to measure or determine the value of that metric, in addition to the previous value and (in some embodiments) the percentage change between the previous value and the current value.
[0197] In one embodiment, the metric monitoring function does not rely on the automatic extraction function. Instead, the user can be provided with the same tools to "monitor" metrics as features exist with time indexes. This can include metrics that are not actually stored in the database, such as values of performance metrics of a machine learning model, or values of different important features in the model. These values can also be set for monitoring by the user.
[0198] As disclosed, a user may specify "rules" for metric monitoring based (for example) on either the level (value of the metric) and / or the percentage change between the current and previous values of the metric. When a user is prompted to specify a rule, the metric monitoring capability may also (or alternatively) recommend rules based on similarly monitored metrics, where similarity may be determined by one or more of (as non-limiting examples) statistical characteristics of the metric, a semantic analysis of the metric's name, or metric monitoring rules previously specified by the user.
[0199] Such a "recommendation" might include a prompt to the user in the form "The recommended threshold for mean change is 2.2% (which occurs in 5% of observations)." The form of user-defined or platform-suggested rules depends on the structure and values of the data, but commonly includes rules based on (by way of example):
[0200] the value of the data (e.g., the data is positive, at least zero, negative, greater than / greater than a particular value, or less than / equal to a particular value); "Absolute" change in the value of the data (e.g., the numerical change is strictly zero, or the numerical change is less than / below a particular value, or the numerical change is less than / below a particular value in absolute value), or ● The percent change of the data from its previous value (eg, the percent change is zero or the percent change is greater than a particular value).
[0201] In one embodiment, a user may specify multiple rules and can specify whether to be notified / alerted when a particular rule is "violated" or when all rules are "violated," where a rule is "violated" when the condition specified by the rule exists or is satisfied. That is, if a user sets a rule for a metric that should be monitored when its value is negative, the rule is said to be "violated" whenever the value of the metric becomes negative, i.e., the condition set within the rule is satisfied.
[0202] Based on one or more rules, the platform may display whether a value (if the rule is based on a value) or a change in a value (if the rule is based on the most recent change in a value) is in "violation" of a configured rule. Such a "violation" represents an "alert" or notification generation condition, and in response, the platform may change the display of the value (or value change) in a manner specified by the user. As mentioned, the user may be provided with options as to how the display changes, for example, by setting a color for the alert condition and / or choosing an icon to be shown along with the value or value change.
[0203] In one embodiment, the default changes to the display of the metric are to show the value (or the change in value, depending on the rule applied) in red when the rule is in an alert state (when the rule is "violated"), and in green when the rule is not in an alert state. When there are no rules applied, the monitor may display a default color, which may be black. These settings can be changed by the user, along with the accessibility parameters that the user sets on the platform's display.
[0204] In some embodiments, the metric monitoring feature can provide a user with monitoring of objects that the user is not yet familiar with. As a non-limiting example, a team may set up the metric monitoring feature with a focus on KPIs and specific rules. Because the platform captures metadata and relationships between metrics, a different metric (or set of metrics) or performance metric from a machine learning model added to the platform may be a "good" predictor or leading indicator of the metric being monitored. In this situation, the platform's metric monitoring feature may suggest that this metric should be monitored and provide recommendations for more comprehensive and improved monitoring based on the machine-learned relationships in the metadata added to the platform.
[0205] This capability builds on functionality built into the disclosed platform. As part of constructing the feature graph via data fetching (e.g., a metadata fetching service that periodically queries a cloud database service), the platform has a software process that automatically calculates statistical relationships between different features and measures the relative strength of the relationships according to a calibration process. As part of the calibration process, closely related metrics can be identified via queries, and when a newly added metric is closely related to a currently monitored metric, this information can be stored within the graph itself. The platform can then prompt the user, with appropriate role-based access with certain suggestions, to open the monitoring model and apply monitoring rules to the newly added metric. Over time, the calibration process will continue to identify new metrics in the same manner and can also identify existing metrics that are highly related to the set of metrics already being monitored.
[0206] As one non-limiting example use case, consider the following scenario. A "business" user might be using the platform to track a set of 16 core KPIs / metrics that the enterprise leadership team has defined and identified as important to the enterprise's operations and business strategy. Using the platform's integration with database and data warehouse services, statistical metadata about the data sets and characteristics can be updated so that the 16 core metrics can be connected to a regularly updating source of data. Members of the enterprise's data team can set up appropriate metric monitoring rules to track and alert users when tracked metrics reach critical levels or growth rates.
[0207] The determined correlations, or the output of machine learning models calculated using the data connected to these metrics, are viewable and navigable on a platform-generated feature graph, so that a "map" of the company's core metrics becomes viewable, navigable, and shareable. Business users may periodically access the platform to review the core metric levels and / or see how the data team's work is creating additional statistical relationships between the company's core metrics (or improving existing statistical relationships between the company's core metrics).
[0208] Metric monitoring capabilities allow users to track important metrics they use to measure the operational status of the enterprise, and platform feature graphs allow users to find connections and / or relationships between metrics. For example, a user may select a UI element connecting two metrics to discover a colleague's model that explored how one metric can be used to "predict" another, because knowing these relationships can provide a more accurate and reliable understanding of operational status. For example, metadata from models and correlations can quantify the predictive relationship between the average waiting time for an order and the likelihood that a customer will reorder from the enterprise, thereby improving the enterprise's decision-making in several areas (e.g., marketing, fulfillment, or inventory management).
[0209] A user of a public version of the platform (available, for example, through www.system.com) may encounter a metric monitoring feature through browsing a portion of the platform feature graph that the user is interested in. For example, the public version of the platform may have a metric defined as "Global Nitrogen Dioxide Emissions." This metric may be connected to a feature that is part of a dataset published by NASA that measures global atmospheric emission levels, and a user may have used that feature as the basis for monitoring the metric of global nitrogen dioxide emissions.
[0210] The public platform UI will then show global nitrogen dioxide emissions as a metric, and users can visit the metric page to get information about changes in levels or growth reported from metadata extracted from NASA's published datasets. When connections to other metrics are made, created, or discovered by the platform (whether through specific machine learning modeling or based on statistical correlations calculated between features in the dataset and other features tracked over time on the platform), the connections will be displayed in the graph. This will allow users to see if other metrics are related to nitrogen dioxide emissions. Users will be able to use the user interface to see the levels and recent changes for the related metrics, and can use links provided in the platform feature graphs to access statistical and / or scientific evidence for the relationships displayed in the graphs (and, if desired, to observe the extent to which the relationships are getting stronger or weaker over time).
[0211] In some embodiments, this information can be made available to other applications via HTTP API requests (such as via gRPC, REST, and / or GraphQL requests, etc.) For example, a call to the metrics endpoint will return platform metadata about metrics, and a call to the metrics / associations endpoint will return metadata about which metrics are related to a given metric (as well as details about statistical relationships, e.g., evidence demonstrating the relationship and the type of model or correlation that contributes to the relationship).
[0212] In one embodiment, for metrics that are relevant to the metric monitoring function, the metadata made available may include one or more of the following: Name, description, ● Creation time, ●Update time, ● Creator, ● Updater, The measured feature, Metric monitoring status, Metric monitoring rules, and ● Associations that include the metric.
[0213] Other (or less) metadata may be provided if the platform is configured to do so. As another example use case, the data generating views or displays provided by the platform can be used by data journalists covering financial markets. In this use case, the data journalists can query for metrics that have levels or recent changes above a predefined threshold and then use queries to find related metrics. The information included in response to these queries will provide statistical context for why the metric of interest is at a particular level (or has had a particular magnitude of change) and will provide statistical evidence for why other historically related metrics can be expected to move in a particular direction. For example, a data journalist might see that the price of silver traded in a particular commodity market has experienced a significant drop. Modeling or correlations calculated using the price of silver will then inform the journalist as to what other market factors have been recently (or historically) associated with the change in the price of silver and what further changes in the market may entail.
[0214] A further description of the platform implementation and capabilities follows. ● The platform stores data about features that have values associated with specific times, e.g., weekly / monthly sales or revenue, annual values for GDP of different countries, or daily stock closing prices for different listed equities. When this type of data is added to the platform, it can be stored with a series of index values that correspond to the specific time recorded for each value (i.e., stored as a timestamp), and the value itself. When these values are numerical, it is possible to track their levels and changes, because the platform knows how to order the data chronologically and can calculate growth rates between specific values.
[0215] ● The platform's data model distinguishes between "features" (which are collections of data or sets of measurements) and "metrics" (which are user-defined objects of interest that a user wants to measure and track). For example, a user interested in measuring sales at a business might define "total monthly sales" as the metric of interest, and the values of that metric are features (or transformations of features) generated from electronic data records stored by the business.
[0216] ● The platform architecture and functionality includes a technique for connecting metrics with features in a feature graph. The platform allows a user to specify that a particular feature (or features) provides the value used to determine a given metric, which allows other users to understand that the metric is being measured or evaluated using the connected features. The platform architecture then allows connections to be made between metrics and features using relationships inferred from machine learning models and / or from statistical relationships calculated directly from the data (e.g., correlations between measures).
[0217] The disclosed metric monitoring features use these aspects of the platform to provide metric monitoring capabilities and contextual information to users. The monitoring capabilities are based on retrieving data from various sources and arranging the data along a commonly stored timestamp-based index. This index is available on features from a dataset on the platform, and when a user connects / associates such a feature with a numerical value with a metric, a visual interface for that metric will show (in some embodiments) the latest and previous values and the percentage change between the values.
[0218] Metric monitoring provides contextual information about metrics because the platform established relationships between metrics when models and datasets were added to the platform. Additionally, the common timestamp index allows the platform to automatically compute time series analysis to generate statistically robust relationships between tracked metrics along the time dimension.
[0219] The metric monitoring capability can be utilized on data collected from different types of sources, including data generated from the platform itself. As an example, for models added to the platform that users update periodically (e.g., via manual model updates, automatically scheduled model updates using online machine learning tools or services, or periodic updates from deployed machine learning model services such as AWS Sagemaker), model performance metrics can be collected according to periodic time intervals. This type of data can also be added to the metrics for monitoring, and statistical relationships can be established (through correlation analysis or explicit modeling) between tracked model performance metrics and other measured metrics on the platform. This allows users of the platform to use metric monitoring to manage their model performance and metrics in the context of their other collected data (since these metrics are often KPIs or key metrics for data science teams).
[0220] In one embodiment, when metric monitoring is available for a feature in a dataset or other data with a time-based index, a visual interface change or display (showing the recent level and percent change of the data) may be used to inform the user that this is data that can be tracked or monitored. The visual interface may also allow the user to set specific rules so that the user can monitor these changes with a greater degree of visual distinction and receive alerts and notifications for value changes for the metric. The user can configure the metric monitoring functionality by setting these rules, which are defined not only in terms of comparing the most recent level of the metric or changes between recent values using a set of predefined comparison operators, but also in terms of options for how to visually indicate when a metric "violates" or "satisfies" the condition expressed by the rule (and how to notify the user that a "violation" has occurred). Once the rules are set, a visual indicator on the feature graph is set to reflect the chosen color or format (or marked with an icon for users with color vision concerns), which distinguishes the monitored metrics from metrics that can be monitored but do not have rules set for them (which remain with the default color or format).
[0221] In one embodiment, and either as part of or separate from metric monitoring, the platform may generate visualizations that show how the underlying feature graphs have changed over time, or changes that have occurred between different sets of sources. This can be useful to identify whether previously identified statistical relationships have been substantiated by subsequent work, or whether what was thought to be a plausible relationship should now be interpreted differently.
[0222] This capability complements metric monitoring by highlighting relationship values that have changed over a user-specified time interval. Users can use metric monitoring to quickly identify important metrics and how their values have changed over time, and can use this type of capability (e.g., as presented in the form of a visualization) to identify whether the value of a key metric has changed because the value of a (statistically) closely related metric has changed, or whether the underlying statistical relationship is stronger or weaker than previously thought. This capability can be automatically available to platform users, replacing the exploratory modeling that a data analyst or scientist might perform in response to changes in a key metric.
[0223] In one example embodiment of the rule configuration process, default rules are pre-populated for the user depending on which field (e.g., current value, previous value, percent change) on the metric is used to set the monitoring rule. Default rules can be configured for different teams using the platform because each business or team account will typically have a separate work area for data and models. This allows configuration settings, including metric monitoring rules, to be stored separately for each separate account of a business or team. For business and team accounts, monitoring rules are typically set using a rule of thumb level (e.g., a standard rule for a metric may be to alert red when the percent change in value is 5% or more in absolute value). When an account already has metric monitoring set for a different metric, the platform can recommend that future alerts be set according to the settings already existing for a metric that is semantically similar (i.e., has the same or sufficiently similar name, description, or type). For example, a team may have set up a metric monitoring rule that displays a "yellow" alert when the value of "Product X in Stock" falls below 100, and a suggested rule for "Product Y in Stock" or "Product X in Production" for that user or team might be to set the same rule as was set for "Product X in Stock."
[0224] Rules may also be suggested when metrics are statistically similar. For example, if "Production of Product X" is known to be statistically related to "Inventory of Product X" due to machine learning models or other determined statistical associations, the rule suggested for "Production of Product X" can be the same as for the related metric, or can be configured to suggest a rule that will occur with a similar likelihood as the alert set for "Inventory of Product X". Metric monitoring can be used to discover or "learn" and apply monitoring rules; this ability provides an advantage over traditional solutions that require rules to be set in isolation, without considering context, for different metrics in the same system.
[0225] As noted, current solutions for monitoring metrics or managing metadata for machine learning models focus on datasets and models in isolation. In contrast, the disclosed platform architecture and its focus on connecting metadata from datasets, models, and other data-oriented operations in one place and in a feature graph means that the metric monitoring functionality is not limited to a particular type of metadata. Furthermore, although metric monitoring has been described with reference to actual feature levels or percentage changes in a dataset, the monitoring functionality can be applied to other metadata collected on the platform that is associated with a corresponding time element.
[0226] Traditional solutions to metadata management or data cataloging may track the number of observations in a particular data set and provide alerts or notifications when this number changes, but existing solutions do not collect and store statistical relationships between different ones of the tracked metadata. For example, a team might use metric monitoring to track daily model performance for a model deployed "in production," actively monitoring five KPI metrics (after setting appropriate rules). The platform's feature graphs will show the movement of these five metrics with contextual highlighting (or other indications) based on the metric's value (or change) compared to thresholds set in the metric monitoring rules.
[0227] Conventional approaches to metric monitoring do not provide a monitoring framework flexible enough to combine metric behavior from disparate sources, such as model performance data generated from deployed machine learning models, with metrics tracked from different data sources. The disclosed platform is designed as a knowledge management tool for the entire data stack, and metric monitoring on the platform is a context-driven monitoring and alerting tool to understand the behavior of important metrics when the sources of those metrics are distributed.
[0228] As described, in some embodiments, the platform may perform its own automated machine learning modeling on the metadata available to the platform. Because metadata for metrics on the platform can be indexed to the same time span, the platform can "know" or "learn" statistical relationships between daily model performance (stored in the feature graph) and other metrics on the platform that are retrieved from a database service (or added by a user) and that have a time index.
[0229] This capability can enable discovery of new and meaningful metrics that the team is not currently monitoring and / or suggest more effective rules for monitoring metrics that highlight key inflection points for model success (e.g., via tracked model performance metrics) or metric levels / changes that predict known critical values for other metrics. This can be done unobtrusively through recommendations presented in the rules setting panel (e.g., by suggesting "better" rules and explaining to the user what the platform is "learning" through its automated machine learning).
[0230] As an example of this capability and its benefit to a user, the platform can be used to take metric monitoring data (including a time-indexed indication of whether a metric is in an "alert" status) and run a classification model that "predicts" whether a given metric is in an alert status using previous values ("lagged" values) of other metrics. The results of this model can be used to identify a "better" threshold for the metric being monitored (this is the case when a particular level or change in a metric is a good predictor of a different metric being in an "notify" or "alert" status), or to identify whether the level / change in a model performance metric is a predictor of the alert status of other metrics (which suggests that the user may want to set up metric monitoring for that model performance metric).
[0231] In some embodiments, the number of statistical comparisons the platform automatically performs may be limited to avoid highlighting spurious correlations and for computational efficiency reasons. Because the platform metadata contains knowledge about the metrics being monitored and the metrics that have high usage on the platform (whether in models or in users' browsing behavior), the automated rule generation and recommendation functionality can focus on metrics and objects that have relatively high interest and relatively high statistical significance on the platform.
[0232] As mentioned, after constructing a feature graph for a particular user or set of users, the graph can be traversed to identify variables of interest to the topic or goal of a study, model, or investigation, and, if desired, to retrieve data sets that confirm or validate the relevance of the variables or measure the variables of interest. It is noted that the process by which the feature graph is traversed can be controlled in one of two ways: (a) explicit tuning of search parameters by the user, or (b) algorithm-based tuning of parameters for variable / data retrieval.
[0233] Returning to FIGS. 2(a) and 2(b), as noted, FIG. 2(a) illustrates how changes in features from a dataset stored within a cloud database service (or “data warehouse” 204) may be monitored using an implementation of the disclosed metric monitoring capabilities. In the example display shown in the figure, dataset metadata 206 is illustrated for two statistically related features, shown as Feature 1 and Feature 2. A first metric (Metric 1 208) is defined and its most recent value is displayed (209). Rules governing the display of alerts or notifications are displayed (212), and the resulting information for Metric 1 is shown in display section 214. Similarly, a second metric (Metric 2 210) is defined and its most recent value is displayed (211), rules governing the display of alerts or notifications are displayed (213), and the resulting information for Metric 2 is shown in display section 215.
[0234] Continuing with the description of the back-end processing on the platform that supports the generation of the displays shown in element or section 202 as shown in element or section 203, data warehouse integration process 220 operates to "fetch" datasets and features from data warehouse 204 and computes or accesses relevant metadata. This fetch process sends an http request to the platform's back-end API that contains the dataset and feature metadata. The metadata includes statistical relationships between features (as suggested by process 222).
[0235] The platform backend writes the dataset, features, and relationship metadata to the platform graph database (as suggested by process 224). The user can view the dataset, features, and relationships on the available website. Once the features have time indexes associated with values (as an example, the example of Feature1 and Feature2 shown in 206) and the user has associated Feature1 and Feature2 with Metric1 (208) and Metric2 (210), the user can then activate or select a metric monitoring function (as suggested by process 226).
[0236] A user can activate or select the metric monitoring feature and then define monitoring rules that specify visual alerts (among other aspects) and set up email / application notifications (as suggested by process 228). In response, metrics available on the platform front end reflect the statistical relationships between features. Users can view monitored metrics with detailed metadata and full statistical context (e.g., levels, percent changes, feature history, alerts, and relationships), as suggested by process 230.
[0237] 2(c) through 2(g) are example user interface displays that may be generated by a platform or system configured to discover or determine and represent statistically meaningful relationships between specified metrics, datasets, and machine learning models in accordance with embodiments of the disclosed platform and system.
[0238] Figure 2(c) is an example of a user interface display illustrating the most recent value (314,779), the percent change to that value (-4%), and identification of the subpopulation with the largest change (which can be calculated when a metric is defined as an aggregation of values in a table where there are multiple subpopulations / dimensions in the data).
[0239] FIG. 2(d) is an example of a user interface display illustrating a metric monitoring panel on a page for a defined metric: weekly active users. The data source for weekly average users (wau) is connected and has a time index, so monitoring is available. The user can set / define rules for monitoring by selecting the +Monitor button, and then specify the color of the monitor and the frequency of email alerts. On the platform characteristics graph on the left side of the figure, metric monitoring is turned on for other metrics, and the edges between nodes in the graph contain metadata that describe the statistical relationships between the metrics. Knowing which metrics are in alert status and understanding the relationships between metrics allows the user to understand the statistical drivers of KPIs / key metrics within the context of their own dataset.
[0240] FIG. 2(e) is an example of a user interface display illustrating a platform catalog view of metric monitoring, where metric monitoring has been turned on for eight metrics on the displayed page. While other solutions for data monitoring may have similar views in some respects (or, in the case of dashboard tools, other chart views), the advantage of the metric monitoring approach can be seen in the collection of evidence for a given metric at the bottom of each "card" or section. Each metric has been used in different models (some of which are predicted outcomes for the models), and metadata about not only each metric, but also relationships between any metrics that have been included in the same machine learning model, or other statistical relationships established by the user or by automated machine learning, is visible by clicking on any of the cards.
[0241] FIG. 2(f) is an example of a user interface display illustrating one or more notifications for metric monitoring functionality. Not only the most recent and up-to-date values (along with percent change) are displayed, but also values for related metrics. These relationships are created from metadata taken from machine learning models added to the platform, from relationships added directly by users, and from automated machine learning applied to user-added feature metadata, retrieved from database services, or generated from periodic updates from tracked models deployed during creation.
[0242] Figure 2(g) is an example of a user interface display illustrating a simplified rule setting dialog. The condition to be applied to this metric is when the absolute value of the percent change is strictly greater than 4.5. In this example, there is one default color difference, and the color indication is red because the percent change (73.10%) is strictly greater than 4.5%.
[0243] 2(h) is a diagram illustrating elements, components, or processes that may be present in or executed by one or more computing devices, servers, platforms, or systems 280 configured to implement a method, process, function, or operation according to some embodiments. In some embodiments, the disclosed systems and methods may be implemented in the form of one or more apparatuses (e.g., a server or client device that is part of a system or platform) that include a processing element and a set of executable instructions. The executable instructions may be part of a software application (or applications) and may be located within a software architecture.
[0244] In general, an embodiment of the disclosure may be implemented using a set of software instructions designed to be executed by an appropriately programmed processing element (such as, by way of non-limiting example, a GPU, TPU, CPU, microprocessor, processor, controller, or computing device). In a complex application or system, such instructions are typically arranged into "modules," each of which typically performs a particular task, process, function, or operation. The entire set of modules may have their operations controlled or coordinated by an operating system (OS) or other form of organizational platform.
[0245] The modules and / or sub-modules may include any suitable set of computer-executable code or instructions, such as computer-executable code corresponding to a programming language. For example, programming language source code may be compiled into computer-executable code. Alternatively, or in addition, the programming language may be an interpreted programming language, such as a scripting language.
[0246] 2(h), system 280 may represent one or more of a server, a client device, a platform, or other form of computing or data processing device. Each of modules 282 includes a set of executable instructions that, when executed by a suitable electronic processor (such as, for example, that shown in the figure by “physical processor 298”), causes system (or server, or device) 280 to operate to perform a particular process, operation, function, or method.
[0247] Modules 282 may include one or more sets of instructions for performing the methods or functions described with reference to the figures and the functional and operational disclosure provided in the specification. These modules may include the modules shown, but may also include more or fewer modules than those shown. Furthermore, the modules and the sets of computer-executable instructions included therein may be executed (in whole or in part) by the same processor or by two or more processors. When executed by two or more processors, the coprocessors may be included in different devices, e.g., a processor in a client device and a processor in a server.
[0248] The modules 282 are stored within memory 281, which typically includes an operating system module 284 that contains (among other functions) instructions used to access and control the execution of instructions encapsulated within other modules. The modules 282 within memory 281 are accessed for the purpose of transferring data and executing instructions through the use of a "bus" or communication lines 290, which also allows a processor 298 to communicate with the modules for the purposes of accessing and executing the instructions. The bus or communication lines 290 also allow the processor 298 to interact with other elements of the system 280, such as input or output devices 292, communication elements 294 for exchanging data and information with devices external to the system 280, and additional memory devices 296.
[0249] Each module or sub-module may correspond to a particular function, method, process, or operation implemented by execution (in whole or in part) of instructions in that module or sub-module. Each module or sub-module may include a set of computer-executable instructions that, when executed by a programmed processor or co-processor, causes the processor or co-processor (or one or more devices or one or more servers in which they are included) to perform a particular function, method, process, or operation. As noted, the device in which the processor or co-processor is included may be a client device, or a remote server or platform, or both. Thus, a module may include instructions executed (in whole or in part) by a client device, or a server or platform, or both. Such functions, methods, processes, or operations may include those used to implement one or more aspects of the disclosed systems and methods, for example to:
[0250] Creating a feature graph (as suggested by module 284) comprising a set of nodes and edges, where: ○ A node may represent, by way of non-limiting examples, one or more of the following: a concept, a topic, a dataset, metadata, a model, a metric, a variable, a measurable quantity, an object, a property, a feature, or a factor.
[0251] o An edge represents a relationship between a first node and a second node, such as, by way of non-limiting example, a statistically significant relationship, a dependency, or a hierarchical relationship. A label associated with an edge may indicate some aspect of the relationship between the two nodes connected by the edge, such as, by way of non-limiting example, the metadata on which the relationship between the two nodes is based, or a dataset that supports a statistically significant relationship between the two nodes.
[0252] ● Providing user interface displays, tools, features, and selectable elements to the user (as suggested by module 286) to enable the user to perform one or more of the following functions:
[0253] Identifying metrics of interest (e.g., KPIs) to monitor or track. Defining rules that describe when alerts should be generated regarding the behavior of identified metrics.
[0254] o Defining how the results of applying the rules should be identified or shown on a user interface display. ○ Allowing a user to select the metric for which an alert was generated and, in response, providing information such as, by way of non-limiting example, the change in the metric's value over time, the satisfied or enabled rules that resulted in the alert, the metric's relationship to other metrics (if relevant), and available information about the dataset, machine learning model, rules, or other factors used to generate the metric.
[0255] ●Generating recommendations for the user regarding different metrics or sets of metrics that may be worth monitoring (as suggested by module 288), data sets that may be useful for scrutiny, metadata that may be relevant to the identified metrics, or other aspects of the underlying data or metrics that may be of potential interest to the user.
[0256] Here, the recommendations may arise (at least in part) from output generated by trained machine learning models, statistical analysis, research, or other forms of evaluation. In some embodiments, the functionality and services provided by the systems and methods disclosed herein may be made available to multiple users by accessing accounts maintained by a server or service platform. Such a server or service platform may be referred to as a form of Software-as-a-Service (SaaS). Figure 3 illustrates a SaaS system in which an embodiment may be implemented. Figure 4 illustrates elements or components of an example operating environment in which an embodiment may be implemented. Figure 5 illustrates additional details of elements or components of the multi-tenant distributed computing services platform of Figure 4 in which an embodiment may be implemented.
[0257] In some embodiments, the systems or services disclosed or described herein may be implemented as microservices, processes, workflows, or functions that are executed in response to a user's submission of a response. The microservices, processes, workflows, or functions may be executed by a server, data processing element, platform, or system. In some embodiments, data analytics and other services may be provided by a services platform located "in the cloud." In such embodiments, the platform may be accessible through APIs and SDKs. Functions, processes, and capabilities may be provided as microservices within the platform. Interfaces to the microservices may be defined by REST and GraphQL endpoints. An administration console may enable a user or administrator to securely access underlying request and response data, manage accounts and access, and in some cases, change processing workflows or configurations.
[0258] It should be noted that, although Figures 3-5 depict a multi-tenant or SaaS architecture that may be used to deliver business-related or other applications and services to multiple accounts / multiple users, such architectures may also be used to deliver other types of data processing services and provide access to other applications. In some embodiments, a platform or system of the type depicted in Figures 3-5 may be operated by a third-party provider to provide a particular set of business-related applications, while in other embodiments, the platform may be operated by a provider and different businesses may provide applications or services for users through the platform.
[0259] 3 is a diagram illustrating a system 300 in which an embodiment may be implemented or may relay access to an embodiment of the disclosed or described service. By taking advantage of a business service system (e.g., a multi-tenant data processing platform, etc.) hosted by an application service provider (ASP), users of the services described herein may include individuals, businesses, stores, organizations, etc. Users may access the services using any suitable client, including, but not limited to, a desktop computer, a laptop computer, a tablet computer, a scanner, a smartphone, etc. Users interconnect with the service platform across the Internet 308 or another suitable communications network or combination of networks. Examples of suitable client devices include a desktop computer 303, a smartphone 304, a tablet computer, or a laptop computer 305.
[0260] The platform 310 may be hosted by a third party and may include a set of services to assist users in accessing the data processing and metric monitoring services 312 and web interface server 314, as described herein, connected as shown in Figure 3. It should be appreciated that either or both of the services 312 and web interface server 314, although depicted as separate units in Figure 3, may be implemented on one or more different hardware systems and components. The services 312 may include one or more functions or operations to enable users to access the feature graphs and perform the metric monitoring functions disclosed herein.
[0261] By way of example, in some embodiments, the set of features, operations, or services made available through platform 310 may include the following: ● Account management services 318. Examples are given below.
[0262] o The process or service that authenticates a user (in conjunction with the submission of proof of the user's identity using a client device). o A process or service that creates containers or instantiations of services or applications that will be made available to users.
[0263] A feature graph generation service 320. An example of this is shown below. A process or service that generates or accesses the disclosed feature graph, which comprises a set of nodes and a set of edges connecting certain ones of the nodes.
[0264] - User interface display and tool generation services 322, examples of which are given below. A process or service that generates one or more user interface displays, and user interface tools and elements that enable a user to:
[0265] ■Identifying metrics of interest (e.g., KPIs) for monitoring or tracking. ■ Defining rules that describe when alerts should be generated regarding the behavior of identified metrics.
[0266] ■ Defining how the results of applying the rules are identified or shown on the user interface display. ○ Allowing a user to select the metric for which an alert was generated and, in response, providing information such as, by way of non-limiting example, the change in the metric's value over time, the satisfied or enabled rules that resulted in the alert, the metric's relationship to other metrics (if relevant), and available information about the dataset, machine learning model, rules, or other factors used to generate the metric.
[0267] - A recommendation generation service 324, an example of which is shown below. A process or service that generates recommendations for a user regarding different metrics or sets of metrics that may be worth monitoring, data sets that may be useful for vetting, metadata that may be relevant to identified metrics, or other aspects of the underlying data or metrics that may be of potential interest to the user.
[0268] - Management Services 326. Examples are shown below. Processes or services that enable a service provider and / or platform to manage and configure the processes and services provided to users, such as by changing, by way of non-limiting examples, how a user's data is modeled, how metrics are calculated, or how the resulting metrics and recommendations are presented to a particular user.
[0269] It should be noted that, in addition to the operations or functions enumerated, an application module or sub-module may include computer-executable instructions that, when executed by a programmed processor, cause the system or device to perform functions related to the operation of the service platform, such functions may include, but are not limited to, those related to user registration, user account management, data security between accounts, allocation of data processing and / or data storage capacity, providing access to data sources other than the system DB (e.g., ontologies or reference materials, etc.).
[0270] The platform or system shown in FIG. 3 may be hosted on a distributed computing system consisting of at least one, but possibly multiple, "servers." A server is a physical computer dedicated to providing data storage and an execution environment for one or more software applications or services intended to serve the needs of users of other computers in data communication with the server over a public network, such as the Internet. The server and the services it provides may be referred to as "hosts" and remote computers, and the software applications running on the remote computers being serviced may be referred to as "clients." Depending on the computing services it provides, the server may be referred to as, for example, a database server, a data storage server, a file server, a mail server, a print server, or a web server. In most cases, a web server is a combination of hardware and software that helps deliver content, typically by hosting websites, to client web browsers that access the web server over the Internet.
[0271] 4 is a diagram illustrating elements or components of an example operating environment 400 in which an embodiment may be implemented. As shown, a variety of clients 402 incorporating and / or being embedded in a variety of computing devices may communicate with a multi-tenant service platform 408 through one or more networks 414. For example, a client may incorporate and / or be embedded in a client application (i.e., software) implemented at least in part by one or more of the computing devices. Examples of suitable computing devices include personal computers, server computers 404, desktop computers 406, laptop computers 407, notebook computers, tablet computers or personal digital assistants (PDAs) 410, smartphones 412, cell phones, and consumer electronic devices incorporating one or more computing device components, such as one or more electronic processors, microprocessors, central processing units (CPUs), or controllers. Examples of suitable networks 414 include networks that utilize wired and / or wireless communication technologies and networks (eg, the Internet) that operate according to any suitable network and / or communication protocol.
[0272] The distributed computing service / platform 408 (which may also be referred to as a multi-tenant data processing platform) may include multiple processing tiers, including a user interface tier 416, an application server tier 420, and a data storage tier 424. The user interface tier 416 may maintain multiple user interfaces 417, including graphical user interfaces and / or web-based interfaces. The user interfaces may include not only a default user interface through which the service provides access to applications and data for users or “tenants” of the service (depicted in the figures as “Service UI”), but also one or more user interfaces specialized / customized according to user-specific requirements (e.g., represented in the figures by “Tenant A UI”, …, “Tenant Z UI”, which may be accessed via one or more APIs).
[0273] The default user interface may include user interface components that enable a tenant to manage the tenant's access to and use of features and capabilities provided by the service platform. This may include accessing tenant data, initiating instantiations of particular applications, causing the execution of particular data processing operations, etc. Each application server or processing tier 422 illustrated in the figure may be implemented with a set of computers and / or components, including computer servers and processors, and may perform various functions, methods, processes, or operations, as determined by the execution of a software application or set of instructions. The data storage tier 424 may include one or more data stores, including a service data store 425 and one or more tenant data stores 426. The data stores may be implemented with any suitable data storage technology, including a structured query language (SQL)-based relational database management system (RDBMS).
[0274] The service platform 408 may be multi-tenant and operated by an entity to provide a set of business-related or other data processing applications, data storage, and functionality to multiple tenants. For example, the applications and functionality may include providing web-based access to functionality used by a business to provide services to end users, thereby allowing a user with a browser and an Internet or intranet connection to view, enter, process, or modify certain types of information. Such functionality or applications are typically implemented by one or more modules of software code / instructions maintained on and executed by one or more servers 422 that are part of the platform's application server tier 420. As noted with respect to FIG. 3, the platform system shown in FIG. 4 may be hosted on a distributed computing system comprised of at least one, but typically multiple, "servers."
[0275] As mentioned, rather than building and maintaining such a platform or system itself, a business may utilize a system provided by a third party. The third party may implement such a business system / platform in the context of a multi-tenant platform, where individual instantiations of the business's data processing workflows are provided to users, with each business representing a tenant of the platform. One advantage to such a multi-tenant platform is the ability of each tenant to customize their own instantiation of the data processing workflow to the tenant's particular business needs or operational methods. Each tenant may be a business or entity that uses the multi-tenant platform to provide business services and functionality to multiple users.
[0276] FIG. 5 is a diagram illustrating additional details of elements or components of the multi-tenant distributed computing services platform of FIG. 4, in which an embodiment may be implemented. The software architecture shown in FIG. 5 represents an example of an architecture that may be used to implement an embodiment of the invention. In general, an embodiment of the invention may be implemented using a set of software instructions designed to be executed by appropriately programmed processing elements (e.g., CPUs, GPUs, microprocessors, processors, controllers, or computing devices, etc.). In complex systems, such instructions are typically arranged into "modules," each of which performs a particular task, process, function, or operation. The entire set of modules may have their operations controlled or coordinated by an operating system (OS) or other form of organizational platform.
[0277] As noted, FIG. 5 is a diagram illustrating additional details of elements or components 500 of a multi-tenant distributed computing services platform in which an embodiment may be implemented. The example architecture includes a user interface layer or tier 502 having one or more user interfaces 503. Examples of such user interfaces include graphical user interfaces and application programming interfaces (APIs). Each user interface may include one or more interface elements 504. For example, a user may interact with the interface elements to access functionality and / or data provided by the application and / or data storage layers of the example architecture. Examples of graphical user interface elements include buttons, menus, checkboxes, drop-down lists, scroll bars, sliders, spinners, text boxes, icons, labels, progress bars, status bars, toolbars, windows, hyperlinks, and dialog boxes. The application programming interfaces may be local or remote and may include interface elements such as various controls, parameterized procedure calls, programmatic objects, messaging protocols, etc.
[0278] The application layer 510 may include one or more application modules 511, each of which has one or more sub-modules 512. Each application module 511 or sub-module 512 may correspond to a function, method, process, or operation implemented by the module or sub-module (e.g., functions or processes related to data processing and providing services to users of the platform). Such functions, methods, processes, or operations may include those used to implement one or more aspects of the disclosed systems and methods, such as for one or more of the processes, functions, or operations disclosed or described herein.
[0279] An application module and / or sub-modules may include any suitable set of computer-executable code or instructions (e.g., as may be executed by a suitably programmed processor, microprocessor, GPU, TPU, or CPU), such as computer-executable code corresponding to a programming language. For example, programming language source code may be compiled into computer-executable code. Alternatively, or in addition, the programming language may be an interpreted programming language, such as a scripting language, etc. Each application server (e.g., as represented by element 422 in FIG. 4) may include a respective application module. Alternatively, different application servers may include different sets of application modules. Such sets may be disjoint or may have common elements with each other.
[0280] The data storage layer 520 may include one or more data objects 522, each having one or more data object components 521, such as attributes and / or behaviors. For example, a data object may correspond to a table in a relational database, and a data object component may correspond to a column or field of such a table. Alternatively, or in addition, a data object may correspond to a data record having fields and associated services. Alternatively, or in addition, a data object may correspond to persistent instances of programmatic data objects, such as structures and classes. Each data store in the data storage layer may include each data object. Alternatively, different data stores may include different sets of data objects. Such sets may be disjoint or common to each other.
[0281] It should be noted that the example computing environments depicted in Figures 3-5 are not intended to be limiting examples. Additional environments in which an embodiment of the disclosure may be implemented in whole or in part include devices (including mobile devices), software applications, systems, appliances, networks, SaaS platforms, infrastructure-as-a-service (IaaS) platforms, or other configurable components that may be used by multiple users for data entry, data processing, application execution, or data review.
[0282] The present disclosure includes the following provisions and embodiments. 1. A method for monitoring one or more metrics, comprising: constructing or accessing a feature graph, the feature graph including a set of nodes and a set of edges, each edge in the set of edges connecting a node in the set of nodes to one or more other nodes, each node representing a variable known to be statistically associated with a topic, and each edge representing a statistical association between a node and the topic or between a first node and a second node; A user interface display and a user interface tool are generated to allow a user to: Identifying metrics for monitoring; defining rules describing when an alert should be generated regarding the behavior of the identified metrics; defining how the results of applying the rules are to be shown on the user interface display; enabling the user to select a metric for which an alert was generated, and in response providing information regarding one or more of the change in value of the metric over time, the rule that resulted in the alert, the relationship of the metric to other metrics, and information regarding the dataset, machine learning model, rules, or factors used to generate the metric; and enabling one or more of A method comprising:
[0283] 2. The method of clause 1, further comprising generating recommendations for the user relating to one or more of different metrics or sets of metrics to monitor, data sets that may be useful for scrutiny, metadata that may be associated with metrics, or aspects of the underlying data or metrics.
[0284] 3. The method of claim 1, wherein constructing the feature graph further comprises: accessing one or more sources, each source containing information regarding a statistical association between a topic discussed in said source and one or more variables considered in discussing said topic; processing the accessed information from each source to identify the one or more variables considered and, for each variable, to identify information regarding the statistical association between the variable and the topic; storing the results of processing the one or more accessed sources in a database, the stored results including, for each source, a reference to each of the one or more variables, a reference to the topic, and information regarding the statistical association between each variable and the topic; A method comprising:
[0285] 4. The method of clause 3, further comprising storing an element providing access to a dataset, the dataset including data used to establish the statistical association between each variable and the topic or data representing measures of one or more of the variables.
[0286] 5. The method according to clause 4, further comprising: traversing the feature graph to identify one or more data sets associated with one or more variables that are statistically associated with a topic of interest to a user or that are statistically associated with a topic that is semantically related to the topic of interest; filtering and ranking the identified one or more data sets; presenting the results of filtering and ranking the identified one or more data sets to the user; and A method comprising:
[0287] 6. The method of clause 3, wherein the one or more sources includes at least one source containing proprietary data. 7. The method of clause 6, wherein said proprietary data is derived from business, research or experimentation.
[0288] 8. The method of clause 1, wherein the recommendations are generated by one or more of a trained model or statistical analysis. 9. A system comprising: one or more electronic processors configured to execute a set of computer-executable instructions; One or more non-transitory computer-readable media containing the set of computer-executable instructions that, when executed, cause the one or more electronic processors or an apparatus or device including the processors to: constructing or accessing a feature graph, the feature graph including a set of nodes and a set of edges, each edge in the set of edges connecting a node to one or more other nodes in the set of nodes, each node representing a variable known to be statistically associated with a topic, and each edge representing a statistical association between a node and the topic or between a first node and a second node; A user interface display and a user interface tool are generated to allow a user to: Identifying metrics for monitoring; defining rules describing when an alert should be generated regarding the behavior of the identified metrics; defining how the results of applying the rules are to be shown on the user interface display; enabling the user to select a metric for which an alert was generated, and in response providing information regarding one or more of the change in value of the metric over time, the rule that resulted in the alert, the relationship of the metric to other metrics, and information regarding the dataset, machine learning model, rules, or factors used to generate the metric; and enabling one or more of one or more non-transitory computer-readable media for causing A system comprising:
[0289] 10. The system of clause 9, wherein the instructions cause the one or more electronic processors, or an apparatus or device including the processor, to generate recommendations for the user relating to one or more of different metrics or sets of metrics to monitor, data sets that may be useful for scrutiny, metadata that may be associated with metrics, or aspects of the underlying data or metrics.
[0290] 11. The system of claim 9, wherein constructing the feature graph further comprises: accessing one or more sources, each source containing information regarding a statistical association between a topic discussed in said source and one or more variables considered in discussing said topic; processing the accessed information from each source to identify the one or more variables considered and, for each variable, to identify information regarding the statistical association between the variable and the topic; storing the results of processing the one or more accessed sources in a database, the stored results including, for each source, a reference to each of the one or more variables, a reference to the topic, and information regarding the statistical association between each variable and the topic; A system comprising:
[0291] 12. The system of clause 11, further comprising storing an element providing access to a dataset, the dataset including data used to establish the statistical association between each variable and the topic, or data representing measures of one or more of the variables.
[0292] 13. A system according to clause 12, wherein the instructions are configured to cause the one or more electronic processors or an apparatus or device including the processor to: traversing the feature graph to identify one or more data sets associated with one or more variables that are statistically associated with a topic of interest to a user or that are statistically associated with a topic that is semantically related to the topic of interest; filtering and ranking the identified one or more data sets; presenting the results of filtering and ranking the identified one or more data sets to the user; and A system that allows the user to:
[0293] 14. The system of clause 11, wherein the one or more sources include at least one source containing proprietary data, and further wherein the proprietary data is derived from business, research, or experimentation.
[0294] 15. One or more non-transitory computer-readable media containing a set of computer-executable instructions that, when executed by one or more programmed electronic processors, cause the processors or an apparatus or device including the processors to: constructing or accessing a feature graph, the feature graph including a set of nodes and a set of edges, each edge in the set of edges connecting a node to one or more other nodes in the set of nodes, each node representing a variable known to be statistically associated with a topic, and each edge representing a statistical association between a node and the topic or between a first node and a second node; A user interface display and a user interface tool are generated to allow a user to: Identifying metrics for monitoring; defining rules describing when an alert should be generated regarding the behavior of the identified metrics; defining how the results of applying the rules are to be shown on the user interface display; enabling the user to select a metric for which an alert was generated, and in response providing information regarding one or more of the following: the change in value of the metric over time, the rule that resulted in the alert, the relationship of the metric to other metrics, and information regarding the dataset, machine learning model, rules, or factors used to generate the metric; and enabling one or more of A non-transitory computer-readable medium for causing
[0295] 16. The non-transitory computer readable medium of clause 15, wherein the instructions cause the one or more electronic processors, or an apparatus or device including the processor, to generate recommendations for the user regarding one or more of different metrics or sets of metrics to monitor, data sets that may be useful for scrutiny, metadata that may be associated with metrics, or aspects of the underlying data or metrics.
[0296] 17. The non-transitory computer-readable medium of clause 15, wherein constructing the feature graph further comprises: accessing one or more sources, each source containing information regarding a statistical association between a topic discussed in said source and one or more variables considered in discussing said topic; processing the accessed information from each source to identify the one or more variables considered and, for each variable, to identify information regarding the statistical association between the variable and the topic; storing the results of processing the one or more accessed sources in a database, the stored results including, for each source, a reference to each of the one or more variables, a reference to the topic, and information regarding the statistical association between each variable and the topic; 1. A non-transitory computer-readable medium comprising:
[0297] 18. The non-transitory computer readable medium of clause 17, further comprising storing an element providing access to a dataset, the dataset including data used to establish the statistical association between each variable and the topic, or data representing measures of one or more of the variables.
[0298] 19. The non-transitory computer-readable medium of clause 18, wherein the instructions are configured to cause the one or more electronic processors or an apparatus or device including the processor to: traversing the feature graph to identify one or more data sets associated with one or more variables that are statistically associated with a topic of interest to a user or that are statistically associated with a topic that is semantically related to the topic of interest; filtering and ranking the identified one or more data sets; presenting the results of filtering and ranking the identified one or more data sets to the user; and A non-transitory computer-readable medium for causing
[0299] 20. The non-transitory computer-readable medium of clause 17, wherein the one or more sources include at least one source containing proprietary data, and further wherein the proprietary data is derived from business, research, or experimentation.
[0300] The disclosed systems and methods can be implemented in the form of control logic using computer software in a modular or integrated manner. Those skilled in the art will know and appreciate, based on the disclosure and teachings provided herein, other ways and / or methods of implementing the present invention using hardware and combinations of hardware and software.
[0301] Machine learning (ML) is increasingly being used to enable data analysis and decision support in multiple industries. To benefit from the use of machine learning, a machine learning algorithm is applied to a set of training data and labels to generate a "model" that represents what the application of the algorithm has "learned" from the training data. Each element (or instance or example in the form of one or more parameters, variables, characteristics, or "features") of the set of training data is associated with a label or annotation that defines how the element should be classified by the trained model. A machine learning model in the form of a neural network is a set of layers of connected neurons that operate to make decisions (e.g., classification) on samples of input data. Once the model is trained (i.e., the weights connecting the neurons have converged and become stable or within an acceptable amount of variation), it will operate on new elements of input data to generate the correct label or classification as output.
[0302] In some embodiments, certain of the methods, models, or functions described herein may be implemented in the form of a trained neural network, where the network is implemented by execution of a set of computer-executable instructions or a representation of a data structure. The instructions may be stored in (or on) a non-transitory computer-readable medium and executed by a programmed processor or processing element. The set of instructions may be communicated to a user through a transfer of instructions or through an application that executes the set of instructions (e.g., over a network, e.g., the Internet, etc.). The set of instructions or application may be utilized by an end user through access to a SaaS platform or a service provided through such a platform. A trained neural network, a trained machine learning model, or any other form of decision-making or classification process may be used to implement one or more of the methods, functions, processes, or operations described herein. It should be noted that a neural network or deep learning model may be characterized in the form of a data structure in which data representing a set of layers containing nodes are stored, and connections are created (or formed) between the nodes in different layers that operate on inputs and provide decisions or values as outputs.
[0303] In general, a neural network can be viewed as a system of interconnected artificial "neurons" or nodes that exchange messages between each other. The connections have numerical weights that are "tuned" during the training process so that a properly trained network will respond correctly when presented with an image or pattern to be recognized (for example). In this characterization, the network consists of multiple layers of feature-detecting "neurons", each layer having neurons that respond to different combinations of inputs from previous layers. The network is trained by using a "labeled" data set of inputs with a wide variety of representative input patterns that are linked to their own intended output responses. Training uses a generic method to iteratively determine weights for intermediate and final feature neurons. In terms of the computational model, each neuron calculates a dot product of the input and weight, adds a bias, and applies a nonlinear trigger or activation function (e.g., using a sigmoid response function).
[0304] Any of the software components, processes, or functions described in this application may be implemented as software code to be executed by a processor using any suitable computer language, such as Python, Java, JavaScript, C, C++, or Perl, using conventional or object-oriented techniques. The software code may be stored as a sequence of instructions or commands in (or on) a non-transitory computer-readable medium, such as a random access memory (RAM), a read-only memory (ROM), a magnetic medium such as a hard drive, or an optical medium such as a CD-ROM. In this context, a non-transitory computer-readable medium is nearly any medium suitable for storing data or sets of instructions, except for transitory waveforms. Any of such computer-readable media may reside on or within a single computing device or may be present on or within different computing devices within a system or network.
[0305] According to one example implementation, the term processing element or processor, as used herein, may be a central processing unit (CPU) or may be conceptualized as a CPU (e.g., a virtual machine, etc.). In this example implementation, the CPU or the device in which the CPU is embedded may be coupled to, connected to, and / or in communication with one or more peripheral devices, such as a display, etc. In another example implementation, the processing element or processor may be embedded within a mobile computing device, such as a smartphone or tablet computer, etc.
[0306] The non-transitory computer-readable storage medium referred to herein may include a number of physical drive units, such as Redundant Array of Independent Disks (RAID), flash memory, USB flash drive, external hard disk drive, thumb drive, pen drive, key drive, High-Density Digital Versatile Disc (HD-DVD) optical disk drive, internal hard disk drive, Blu-Ray optical disk drive, or Holographic Digital Data Storage (HDDS) optical disk drive, Synchronous Dynamic Random Access Memory (SDRAM), or similar devices or other forms of memory based on similar technology, etc. Such computer-readable storage media enable a processing element or processor to access computer-executable process steps and application programs, etc., stored on removable and non-removable memory media to offload data from or upload data to a device. As stated, with respect to the embodiments described herein, the non-transitory computer-readable medium may include nearly any structure, technology, or method, except for a transitory waveform or similar medium.
[0307] Certain implementations of the disclosed technology are described herein with reference to block diagrams of systems and / or flowcharts or flow diagrams of functions, operations, processes, or methods. It will be understood that one or more blocks of the block diagrams, or one or more stages or steps of the flowcharts or flow diagrams, and combinations of blocks of the block diagrams and stages or steps of the flowcharts or flow diagrams, respectively, can be implemented by computer-executable program instructions. It should be noted that in some embodiments, one or more of the blocks, stages, or steps may not necessarily be performed in the order presented, or may not necessarily be performed at all.
[0308] These computer executable program instructions may be loaded into a general purpose computer, special purpose computer, processor, or other programmable data processing apparatus to produce a particular instance of a machine, whereby the instructions executed by the computer, processor, or other programmable data processing apparatus produce means for implementing one or more of the functions, operations, processes, or methods described herein. These computer program instructions may also be stored in a computer readable memory capable of causing the computer or other programmable data processing apparatus to function in a particular manner, whereby the instructions stored in the computer readable memory produce an article of manufacture including instruction means for implementing one or more of the functions, operations, processes, or methods described herein.
[0309] Although certain implementations of the disclosed technology have been described in connection with what are presently believed to be the most practical and various implementations, it is to be understood that the disclosed technology is not to be limited to the disclosed implementations. Instead, the disclosed implementations are intended to cover various modifications and equivalent arrangements that fall within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.
[0310] This written description uses examples to disclose certain implementations of the disclosed technology and enables any person skilled in the art to practice certain implementations of the disclosed technology, including making and using any device or system, and performing any incorporated methods. The patentable scope of certain implementations of the disclosed technology is defined in the claims, and may include other examples that occur to those skilled in the art. Such other examples are intended to be within the scope of the claims if they have structural and / or functional elements that do not differ from the language of the claims, or if they include structural and / or functional elements that have insubstantial differences from the language of the claims.
[0311] All references cited in this specification, including publications, patent applications, and patents, are hereby incorporated by reference to the same extent as if each reference was individually and specifically indicated to be incorporated by reference and / or was set forth in full herein.
[0312] Use of the terms "a," "an," and "the" and similar referents herein and in the following claims should be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The terms "having," "comprising," and "including" and similar referents herein and in the following claims should be construed as open-ended (e.g., meaning "including, but not limited to"), unless otherwise noted. Recitation of ranges of values herein is merely intended to serve as a shorthand method of individually referring to each separate value falling within the range, inclusive, unless otherwise indicated herein, and each separate value is incorporated herein as if it were individually recited. All methods described herein can be performed in any suitable order, unless otherwise indicated herein or clearly contradicted by context. The use of any and all examples or exemplary language (e.g., "such as") provided herein is intended merely to better illuminate embodiments of the invention and does not impose limitations on the scope of the invention unless specifically claimed. No language in the specification should be construed as indicating any non-claimed element as essential to each embodiment of the invention.
[0313] As used in this application (i.e., the claims, figures, and specification), the term "or" is used to inclusively refer to items in alternative and combination. Different arrangements of the components depicted in the drawings or described herein, as well as components and steps not shown or described, are possible. Similarly, some features and subcombinations are useful and may be used without reference to other features and subcombinations. The embodiments have been described for purposes of illustration and not limitation, and alternative embodiments will become apparent to the reader of this specification. Thus, the disclosed embodiments are not limited to those described or depicted in the drawings, and various embodiments and modifications may be made without departing from the scope of the following claims.
Claims
1. 1. A method for monitoring one or more metrics, comprising: constructing or accessing a feature graph, the feature graph comprising a set of nodes and a set of edges, each edge in the set of edges connecting a node to one or more other nodes in the set of nodes, and further wherein each node represents a variable known to be statistically associated with a topic, and each edge represents a statistical association between a node and the topic or between a first node and a second node; Generate user interface displays and user interface tools to allow users to: Identifying metrics for monitoring; defining rules that describe when alerts should be generated regarding the behavior of the identified metrics; defining how the results of applying the rules are to be shown on the user interface display; enabling the user to select a metric for which an alert was generated, and in response providing information regarding one or more of the change in value of the metric over time, the rule that resulted in the alert, the relationship of the metric to other metrics, and information regarding the dataset, machine learning model, rule, or factor used to generate the metric; and enabling one or more of A method comprising:
2. 10. The method of claim 1 further comprising: generating recommendations for the user regarding one or more of different metrics or sets of metrics to monitor, data sets that may be useful for vetting, metadata that may be relevant to the metrics, or aspects of the underlying data or metrics.
3. 10. The method of claim 1, Constructing the feature graph further comprises: accessing one or more sources, each source containing information regarding a statistical association between a topic discussed in said source and one or more variables considered in discussing said topic; processing the accessed information from each source to identify the one or more variables considered and, for each variable, to identify information regarding the statistical association between the variable and the topic; storing results of processing the accessed one or more sources in a database, the stored results including, for each source, a reference to each of the one or more variables, a reference to the topic, and information regarding the statistical association between each variable and the topic; A method comprising:
4. 4. The method of claim 3, further comprising:
1. A method comprising: storing an element that allows access to a dataset, the dataset including data used to establish the statistical association between each variable and the topic or data representing measures of one or more of the variables.
5. 5. The method of claim 4, further comprising: traversing the feature graph to identify one or more datasets associated with one or more variables that are statistically associated with a topic of interest to a user or that are statistically associated with a topic that is semantically related to the topic of interest; filtering and ranking the identified one or more data sets; presenting the results of filtering and ranking the identified one or more datasets to the user; and A method comprising:
6. 4. The method of claim 3, The method, wherein the one or more sources include at least one source containing proprietary data.
7. 7. The method of claim 6, The method, wherein the proprietary data is derived from business, research, or experimentation.
8. 3. The method of claim 2, The method, wherein the recommendations are generated by one or more of a trained model or statistical analysis.
9. 1. A system comprising: one or more electronic processors configured to execute a set of computer-executable instructions; one or more non-transitory computer-readable media containing the set of computer-executable instructions, which, when executed, cause the one or more electronic processors or an apparatus or device including the electronic processor(s): constructing or accessing a feature graph, the feature graph comprising a set of nodes and a set of edges, each edge in the set of edges connecting a node to one or more other nodes in the set of nodes, and further each node representing a variable known to be statistically associated with a topic, and each edge representing a statistical association between a node and the topic or between a first node and a second node; Generate user interface displays and user interface tools to allow users to: Identifying metrics for monitoring; defining rules that describe when alerts should be generated regarding the behavior of the identified metrics; defining how the results of applying the rules are to be shown on the user interface display; enabling the user to select a metric for which an alert was generated, and in response providing information regarding one or more of the change in value of the metric over time, the rule that resulted in the alert, the relationship of the metric to other metrics, and information regarding the dataset, machine learning model, rule, or factor used to generate the metric; and one or more non-transitory computer-readable media that cause the A system comprising:
10. 10. The system of claim 9, The instructions cause the one or more electronic processors or an apparatus or device including the electronic processor to generate recommendations for the user regarding one or more of different metrics or sets of metrics to monitor, data sets that may be useful for review, metadata that may be associated with metrics, or aspects of the underlying data or metrics.
11. 10. The system of claim 9, Constructing the feature graph further comprises: accessing one or more sources, each source containing information regarding a statistical association between a topic discussed in said source and one or more variables considered in discussing said topic; processing the accessed information from each source to identify the one or more variables considered and, for each variable, to identify information regarding the statistical association between the variable and the topic; storing results of processing the accessed one or more sources in a database, the stored results including, for each source, a reference to each of the one or more variables, a reference to the topic, and information regarding the statistical association between each variable and the topic; A system comprising:
12. 12. The system of claim 11, further comprising: Storing an element that allows access to a dataset, the dataset including data used to establish the statistical association between each variable and the topic or data representing measures of one or more of the variables.
13. 13. The system of claim 12, The instructions may be configured to cause the one or more electronic processors or an apparatus or device including the electronic processor to: traversing the feature graph to identify one or more datasets associated with one or more variables that are statistically associated with a topic of interest to a user or that are statistically associated with a topic that is semantically related to the topic of interest; filtering and ranking the identified one or more datasets; presenting the results of filtering and ranking the identified one or more datasets to the user; and A system that allows the following to be performed.
14. 12. The system of claim 11, The system, wherein the one or more sources include at least one source containing proprietary data, and further wherein the proprietary data is derived from business, research, or experimentation.
15. One or more non-transitory computer-readable media containing a set of computer-executable instructions that, when executed by one or more programmed electronic processors, cause the electronic processors, or an apparatus or device including the electronic processors, to: constructing or accessing a feature graph, the feature graph comprising a set of nodes and a set of edges, each edge in the set of edges connecting a node to one or more other nodes in the set of nodes, and further each node representing a variable known to be statistically associated with a topic, and each edge representing a statistical association between a node and the topic or between a first node and a second node; Generate user interface displays and user interface tools to allow users to: Identifying metrics for monitoring; defining rules that describe when alerts should be generated regarding the behavior of the identified metrics; defining how the results of applying the rules are to be shown on the user interface display; enabling the user to select a metric for which an alert was generated, and in response providing information regarding one or more of the change in value of the metric over time, the rule that resulted in the alert, the relationship of the metric to other metrics, and information regarding the dataset, machine learning model, rule, or factor used to generate the metric; and enabling one or more of A non-transitory computer-readable medium for causing
16. 16. The non-transitory computer-readable medium of claim 15, The instructions are on a non-transitory computer-readable medium that cause the one or more electronic processors or an apparatus or device including the electronic processor to generate recommendations for the user regarding one or more of different metrics or sets of metrics to monitor, data sets that may be useful for scrutiny, metadata that may be associated with metrics, or aspects of the underlying data or metrics.
17. 16. The non-transitory computer-readable medium of claim 15, Constructing the feature graph further comprises: accessing one or more sources, each source containing information regarding a statistical association between a topic discussed in said source and one or more variables considered in discussing said topic; processing the accessed information from each source to identify the one or more variables considered and, for each variable, to identify information regarding the statistical association between the variable and the topic; storing in a database results of processing the accessed one or more sources, the stored results including, for each source, a reference to each of the one or more variables, a reference to the topic, and information regarding the statistical association between each variable and the topic; 1. A non-transitory computer-readable medium comprising:
18. 20. The non-transitory computer-readable medium of claim 17, further comprising:
1. A non-transitory computer-readable medium comprising: storing an element that allows access to a dataset, the dataset including data used to demonstrate the statistical association between each variable and the topic, or data representing measures of one or more of the variables.
19. 20. The non-transitory computer-readable medium of claim 18, The instructions may be written to the one or more electronic processors or to an apparatus or device that includes the electronic processor(s): traversing the feature graph to identify one or more datasets associated with one or more variables that are statistically associated with a topic of interest to a user or that are statistically associated with a topic that is semantically related to the topic of interest; filtering and ranking the identified one or more datasets; presenting the results of filtering and ranking the identified one or more datasets to the user; and A non-transitory computer-readable medium for causing
20. 20. The non-transitory computer-readable medium of claim 17, The one or more sources include at least one source containing proprietary data, and further, the proprietary data is derived from business, research, or experimentation.