Interpret Machine Learning Results Using Feature Analysis

By grouping the input parameters of the machine learning model into the feature group and analyzing the contribution value of the feature group, the problem of user distrust of the machine learning model results is solved, providing easier understanding information, and improving the efficiency of the model usage.

CN112990250BActive Publication Date: 2025-07-01BUSINESS OBJECTS SOFTWARE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202011460448.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-12
Filing Date
2020-12-11
Publication Date
2025-07-01
Estimated Expiration
2040-12-11

AI Technical Summary

Technical Problem

The results of machine learning models are difficult to interpret and trust in prior art, especially when the contribution of the input parameters (features) of the model to the results is unclear.

Method used

By grouping input parameters of machine learning models into feature groups, the contribution values ​​of feature groups are analyzed to provide more understandable and manipulated information. The formation of feature groups can be based on dependencies between features, context contribution values, and relationships in the data model.

Benefits of technology

Increases user trust in machine learning model results, provides a higher level of information, enables users to better understand the contribution of features to the results, and take measures to adjust the model to improve performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112990250B_ABST
    Figure CN112990250B_ABST
Patent Text Reader

Abstract

Techniques and solutions for analyzing the results of a machine learning model are described. Results of a data set including a first plurality of features are obtained. A plurality of feature groups are defined. At least one feature group includes a second plurality of features among the first plurality of features. The second plurality of features is less than all of the first plurality of features. The feature groups may be defined based on determining dependencies between features among the first plurality of features, including using context contribution values. A group context contribution value of a feature group may be determined by aggregating context contribution values of constituent features of the feature group.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to interpreting machine learning models, including results provided by machine learning models and operations of machine learning models. Particular embodiments relate to analyzing features used as inputs to a machine learning model to identify relationships between the features, including grouping features into feature groups in an embodiment. Background Art

[0002] Machine learning is increasingly being used to make or assist in making various decisions, or otherwise analyze data. Machine learning techniques can be used to analyze data faster or more accurately, which could be done by humans. In some cases, it is impractical to manually analyze a data set. Thus, machine learning has facilitated the rise of "big data" by providing a path for the practical application of such data.

[0003] However, machine learning can be difficult to understand, even for experts in the field. The situation becomes more complex when machine learning is applied to a particular application in a particular domain. That is, a computer scientist may understand the algorithms used in machine learning techniques, but may not understand the subject domain well enough to ensure that the model is accurately trained or that the results provided by machine learning are correctly evaluated. Conversely, a domain expert may be proficient in a given subject domain, but may not understand how machine learning algorithms work.

[0004] Accordingly, if users do not understand how a machine learning model works, they may have no confidence in the results provided by machine learning. If users have no confidence in the results of machine learning, they may be much less likely to use machine learning and the aforementioned advantages that can be obtained. Thus, there is room for improvement. Summary of the Invention

[0005] This Summary of the Invention is provided to introduce in a simplified form a few concepts that will be further described in the Detailed Description below. This Summary of the Invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0006] Techniques and solutions for analyzing the results of a machine learning model are described. Results are obtained for a data set including a first plurality of features. A plurality of feature groups are defined. At least one feature group includes a second plurality of features of the first plurality of features. The second plurality of features is less than all of the first plurality of features. The feature groups can be defined based on determining dependencies between features of the first plurality of features (including using context contribution values). A group context contribution value can be determined for a feature group by aggregating the context contribution values of the constituent features of the feature group.

[0007] A method of forming feature groups is provided. A training data set is received. The training data set includes values of a first plurality of features. The training data set is used to train a machine learning algorithm to provide the machine learning algorithm. The trained machine learning algorithm is used to process an analysis data set to provide a result. A plurality of feature groups are formed. At least one of the feature groups includes a second plurality of features of the first plurality of features. The second plurality of features is a proper subset of the first plurality of features.

[0008] According to another embodiment, a method of forming feature groups using dependencies between features in a data set is provided. A training data set is received. The training data set includes values of a first plurality of features. The training data set is used to train a machine learning algorithm to provide the trained machine learning algorithm. The trained machine learning algorithm is used to process an analysis data set to provide a result. Context contribution values are determined for a second plurality of features of the first plurality of features. Dependencies between features in the second plurality of features are determined. A plurality of feature groups are formed at least in part based on the determined dependencies. At least one of the plurality of feature groups includes a third plurality of features of the first plurality of features. The third plurality of features is a proper subset of the first plurality of features.

[0009] According to another aspect, a method for determining feature group contribution values is provided. A first plurality of features used in a machine learning algorithm are determined. A plurality of feature groups are formed, such as using analysis of machine learning results, semantic analysis, statistical analysis, data lineage, or combinations thereof. At least one of the plurality of feature groups includes a second plurality of features of the first plurality of features. The second plurality of features is a proper subset of the first plurality of features. The machine learning algorithm is used to determine the result of the analysis data set. For at least a portion of the feature groups, the contribution values of the features in each feature group to the result are aggregated to provide a feature group contribution value.

[0010] The present disclosure also includes a computing system and a tangible non-transitory computer-readable storage medium configured to perform the above methods or including instructions for performing the above methods. As described herein, various other features and advantages can be incorporated into the technology as desired. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 is a schematic diagram showing how values (for training a model or for classification) used as inputs to a machine learning model are associated with features.

[0012] Figure 2 is a schematic diagram showing how values (for training a model or for classification) used as inputs to a machine learning model are associated with features and how different features contribute to the result to different degrees.

[0013] Figure 3 It is a diagram showing how multiple star schemas are related data models.

[0014] Figure 4 It is a schematic diagram of a database schema showing the relationships between at least some of the database tables in the schema.

[0015] Figure 5 It is a schematic diagram showing the relationships between table elements that can be included in a data dictionary or otherwise used to define database tables.

[0016] Figure 6 It is a schematic diagram showing the components of a data dictionary and the components of a database layer.

[0017] Figure 7 It presents an example data access operation that provides query results by accessing and processing data from multiple data sources, including operations that join the results from multiple tables.

[0018] Figure 8 It is a matrix showing the dependency relationship information between features used as inputs to a machine learning model.

[0019] Figure 9 It is a plot showing the relationships between features used as inputs to a machine learning model.

[0020] Figure 10 It is a schematic diagram showing how at least some of the features used as inputs to a machine learning model are assigned to feature groups.

[0021] Figure 11 It is an example user interface screen presenting feature groups and their contributions to the results provided by a machine learning model.

[0022] Figure 12 It is a diagram schematically showing how a data set is processed to train and use a machine learning model, and how the features used as inputs in these processes are analyzed and used to form feature groups.

[0023] Figure 13A It is a flowchart of an example method for forming feature groups.

[0024] Figure 13B It is a flowchart of an example method for forming feature groups by at least partially analyzing the dependencies between features used as inputs to a machine learning model.

[0025] Figure 13C It is a flowchart of an example method for forming feature groups and calculating their contributions to the results provided by a machine learning model.

[0026] Figure 14 FIG. 1 is a schematic diagram of an example computing system in which some of the described embodiments may be implemented.

[0027] Figure 15 FIG. 2 is an example cloud computing environment that may be used in conjunction with the techniques described herein. DETAILED DESCRIPTION

[0028] Example 1 - Overview

[0029] Machine learning is increasingly being used to make or assist in making various decisions or otherwise analyze data. Machine learning techniques can be used to analyze data faster or more accurately than can be done by humans. In some cases, it is impractical for a human to analyze a data set. Thus, machine learning has facilitated the rise of "big data" by providing a path for the practical use of such data.

[0030] However, even for experts in the field, machine learning can be difficult to understand. The situation is more complex when machine learning is applied to a particular application in a particular domain. That is, a computer scientist may understand the algorithms used in machine learning techniques, but may not understand the subject domain well enough to ensure that the model is accurately trained or that the results provided by the machine learning are correctly evaluated. Conversely, a domain expert may be proficient in a given subject domain, but may not understand how machine learning algorithms work.

[0031] Accordingly, if users do not understand how a machine learning model works, they may not have confidence in the results provided by the machine learning. If users do not have confidence in the results of machine learning, they may be much less likely to use machine learning and the aforementioned advantages that can be obtained.

[0032] As an example, a machine learning model typically may use dozens, hundreds, or thousands of input parameters, which also may be referred to as features or variables. It may be difficult for a user to understand how a given variable affects or contributes to the results, such as a prediction, provided by the machine learning model. At least in some cases, the contribution of a particular variable to a particular type of result (e.g., the results typically provided by a machine learning model), or to a particular result of a particular set of input features, can be quantified. However, once a given machine learning model uses variables other than a few variables, it may be difficult for a user to understand the contribution of an individual variable to the model. If a user does not understand how a variable contributes, they may not trust the model or the results.

[0033] In addition, even if a user fully trusts the model and its results, if the user does not understand what contributes to the results, actionable information may not be provided to the user, thereby reducing the use of the machine learning model. Take such a machine learning model as an example, where the model provides a successful prediction for a given set of conditions. For a specific set of conditions, assume that the machine learning model provides a prediction of a 75% chance of achieving a successful result. To make a final decision, the user may find it helpful to understand which factors tend to indicate success or which factors tend to indicate failure. This may be because, for a specific environment, humans may weight a factor more or less than the machine learning model, or humans may be able to take measures to mitigate adverse variables. For future behavior, the user may want to know the steps that can be taken to increase the success rate. If the machine learning model takes into account, say, 1000 variables, it may be difficult for an individual to understand how a single variable or combination of variables contributes to the result and how different values of the variables may affect future results. Therefore, there is room for improvement.

[0034] The present disclosure facilitates the design of machine learning models and the analysis of machine learning results by grouping at least some input parameters (which may be referred to as features) of a machine learning model into one or more feature groups. Examples of machine learning techniques that can use the disclosed techniques include, but are not limited to, logistic regression, Bayesian algorithms (e.g., Naive Bayes), k-nearest neighbors, decision trees, random forests, gradient boosting frameworks, support vector machines, and various types of neural networks.

[0035] The contribution to the machine learning result can be determined for at least some of the features (such as the features in the feature group) used in a specific machine learning technique (e.g., for a specific application). For a given feature group, the individual contributions of the features in the feature group can be summed or otherwise combined to provide a contribution or significance value of the feature group to the machine learning result. The significance value of a feature group can be the sum or other aggregation of the contributions (or significance values) of the features in the feature group. In a specific example, the significance value of a feature group is calculated as the average of the significance values of the features in the feature group.

[0036] As an example, consider a machine learning model that uses features A - Z as input. Suppose the feature group is defined to include features A - D. If it is determined that feature A contributes 4% to the result, feature B contributes 5% to the result, feature C contributes 2% to the result, and feature D contributes 3% to the result, then the overall contribution of this feature group is 14%. If features A - Z are divided into four groups, it is much easier to compare the contributions of each group to the machine learning result than to compare 26 individual unorganized features.

[0037] Thus, the disclosed techniques can provide a higher level of information about machine learning models that is more easily understandable and operable by humans. If a user wishes to obtain more information about how a given feature contributes to its feature group or the overall result / model, the user can drill down into the feature group to view the contributions of its constituent features. Having the property of being organized by feature groups and first viewing the overall contribution of a feature group to a prediction can then help the user understand how individual features contribute to the prediction and how they can be adjusted in the future to change the prediction.

[0038] In some cases, a given feature is included in a single feature group. In other cases, a given feature can be included in multiple feature groups, although in such cases, the contributions of all features / feature groups of the machine learning model may exceed 100%. In some scenarios where a feature is included in multiple groups, the feature is only "active" for a single group at a time. For example, if a feature is assigned to groups A and B, if the feature is active in group A, it is inactive in group B. Alternatively, during a specific scenario, either group A or group B can be set to be active or inactive. In another scenario, when a feature is included in multiple feature groups, different techniques can be used - such as manually forming groups, using mutual information, using schema information, etc. - to determine the multiple groups. While in some cases, all the groups used in a specific analysis are of the same type (e.g., determined using mutual information), in other cases, the groups used in a specific analysis can include groups of different types (e.g., the analysis can have manually defined groups, groups determined using schema information, and groups determined using mutual information). Or, multiple considerations can be used to define a single group - such as using a combination of schema information and mutual information to define the group.

[0039] Similarly, in some cases, all features of a machine learning model are included in at least one feature group, while in other cases, one or more features do not need to be included in a feature group.

[0040] Feature groups can be defined in a variety of ways. In some cases, a feature group can be defined based on the structure of one or more data sources (sometimes referred to as data lineage). For example, at least some of the inputs to a machine learning model can be associated with data stored in a star schema. Data for a particular table (such as a particular dimension table) can be putatively assigned to a feature group. Similarly, tables with relationships between attributes (such as foreign key relationships or associations) can indicate that these tables or at least the related attributes should be included in the feature group or considered for inclusion in the feature group.

[0041] Data access patterns can also be used to suggest feature groups. For example, data for a machine learning model can be obtained by joining two or more tables in a relational database. The tables included in the join or other data access operation can be included as separate feature groups.

[0042] A user can manually assign features to a feature group, including changing features in a suggested feature group, splitting a feature group into subgroups, or combining two or more groups into a larger feature group. For example, a feature group can initially be selected based on data relationships (e.g., in common or related tables) or data access considerations (e.g., joins). Then, the user can determine that features should be added to or removed from these groups, determine that two groups should be combined, divide a single group into two or more groups, etc.

[0043] In other cases, a feature group can be determined or modified by evaluating relationships between features. For example, if a first feature is determined to be related to a second feature, the first and second features can be indicated as belonging to a common feature group. A feature group suggested by other means (such as based on data relationships or manual selection) can be modified based on evaluating relationships between features. In the case of a feature group based on a dimension table, it can be determined that one or more attributes of the dimension table are not significantly related. Such attributes can be removed from the group. Similarly, it can be determined that features of another dimension table are significantly related, and such a dimension can be added to the feature group.

[0044] A feature group can be modified or filtered according to other criteria. In particular, if the predictive power of a feature is below a threshold, that feature can be omitted from the feature group, otherwise they may belong to the feature group. Alternatively, features can be filtered in this way before being analyzed for membership in a feature group, including presenting the user with a selection of features that meet the threshold for possible inclusion in the feature group.

[0045] In addition to being used to help analyze the results provided by a machine learning model, feature groups can also be used to develop or train a machine learning model. For example, a machine learning model can be tailored for a specific use by selecting or emphasizing (e.g., weighting) feature groups of interest. Similarly, a machine learning model can be made more accurate or efficient by eliminating the consideration of features with low predictive power / features that do not belong to relevant feature groups.

[0046] In some embodiments, feature groups are analyzed for causality cycles. That is, considering feature group A and feature group B, features in group A are allowed to affect (or cause) features in group B. However, it may not be equally allowed for features in group B to affect features in group A. If a causality cycle is observed, the feature groups can be reconfigured to remove such a cycle. However, in other cases, feature groups can be defined without considering causality cycles.

[0047] In some cases, it may be useful to include or exclude groups with causal dependencies when determining what features / feature groups to use to train or retrain a machine learning model. For example, it may be useful to isolate the effect of a particular feature by excluding from the analysis (or training of the machine learning model) a feature group that has a causal dependency on features in the selected feature group.

[0048] As described above, while some aspects of the present disclosure define feature groups based on the analysis of machine learning results, other aspects can be used to define feature groups without using such analysis (such as defining feature groups based on semantic analysis, data lineage, or non-machine learning statistical methods). In yet another aspect, one or more machine learning-based feature group definition methods can be used in combination with one or more non-machine learning-based feature group definition methods.

[0049] Example 2 - Example of Training with Features and Using a Machine Learning Model

[0050] Figure 1 Schematically depicts how multiple features 110 are used as inputs to a machine learning model 120 to provide a result 130. Typically, the types of features 110 used as inputs to provide the result 130 are those used to train a machine learning algorithm to provide the machine learning model 120. Training and classification can use discrete input instances of the features 110, where each input instance has values for at least some of the features. Typically, the features 110 and their respective values are provided in a way that uses the particular features in a specific manner. For example, each feature 110 can be mapped to a variable used in the machine learning model.

[0051] The result 130 can be a qualitative or quantitative value, such as a numerical value indicating the likelihood that a certain condition holds or a numerical value indicating the relative strength of the result (e.g., a high number indicates a stronger / more valuable result). For a qualitative result, the result 130 can be, for example, a label applied based on the input features 110 of a particular input instance.

[0052] Note that for any of these results, typically the result 130 itself does not provide information on how the result was determined. Specifically, the result 130 does not indicate how much any given feature 110 or set of features contributed to the result. However, in many cases, one or more features 110 will contribute positively to the result, and one or more features can argue against the result 130 and, conversely, can contribute to another result not chosen by the machine learning model 120.

[0053] Thus, for many machine learning applications, the user may not know how a given result 130 relates to the input features specifically used by the machine learning model. As described in Example 1, if the user is unsure what features 110 contributed to the result 130, or unsure how they contributed or to what extent, they may have less confidence in the result. Additionally, the user may not know how to change any given feature 110 to try and obtain a different result 130.

[0054] In at least some cases, it is possible to determine how feature 110 contributes to the results of a machine learning model (for an individual classification result that is an average or other statistical measure over multiple input instances of the machine learning model 120). In particular, "Consistent Individualized Feature Attribution for Tree Ensembles" by Lundberg et al. (available at https: / / arxiv.org / abs / 1802.03888 and incorporated herein by reference) describes how to compute SHAP (Shapley additive explanation) values for the attributes used in a machine learning model, thus allowing determination of the relative contribution of feature 110. However, other context interpretable metrics (which may also be referred to as context contribution values) can be used, such as those computed using LIME (local interpretable model-agnostic explanation) techniques (described in "‘Why Should I Trust You?’Explaining the Predictions of Any Classifier" by Ribeiro et al. (available at https: / / arxiv.org / pdf / 1602.04938.pdf and incorporated herein by reference)). Generally, a context contribution value is a value that considers the contribution of a feature to a machine learning result in the context of other features used in generating the result, rather than simply considering, for example, the impact of an individual feature on the result in isolation.

[0055] Context SHAP values can be computed using the following formula as defined and used by Lundberg et al. as described by Lundberg et al.:

[0056]

[0057] The single-variable (or overall) SHAP contribution (the impact of the feature on the result, without considering the relationship of the feature in context to other features used in the model), φ1, can be computed as:

[0058]

[0059] where:

[0060]

[0061] and

[0062]

[0063] The above values can be converted to a probability scale using the following formula:

[0064]

[0065] where s is the sigmoid function:

[0066]

[0067] Figure 2 is generally similar to Figure 1 , but shows how to calculate contribution values 140 (such as those calculated using the SHAP method) for feature 110. As explained in Example 1, a large number of features 110 are used with many machine learning models. In particular, if the contribution values 140 of each (or most or many) of the features 110 are relatively small, it may be difficult for a user to understand how any feature contributes to the results provided by the machine learning model (including a particular result 130 for a particular set of values of feature 110).

[0068] Similarly, it may be difficult for a user to understand how different combinations of features 110 together affect the results of machine learning model 120.

[0069] Example 3 – Example relationships between data models and their components

[0070] As explained in Example 1, the disclosed techniques relate to grouping at least some of the features used by a machine learning model (e.g., features 110 used with the machine learning model 120 in Figure 1 and Figure 2 ). The grouping can be based on or at least partially based on relationships between the features. For example, at least some of the features 110 can be associated with a data model (such as a database used in a relational database system or other data storage). In a particular example, data for OLAP analysis can be stored in conjunction with an OLAP cube definition, where the cube definition can be defined relative to data stored in multiple tables (such as tables in a star schema). Both the cube definition and the star schema can be used as data models from which relationships between features can be extracted and used to form groups of features for the disclosed techniques.

[0071] Figure 3 Schematically depicts a data model 300 including two star schemas 310, 320. Star schema 310 includes a central fact table 314 and three-dimensional tables 318. Star schema 320 includes a central fact table 324 and four dimension tables 328.

[0072] To obtain data from multiple star schemas, a dimension table common to two fact tables is used to bridge the two schemas. In some cases, this bridging can occur if one dimension table is a subset of another (e.g., one table includes all the attributes of the other plus one or more additional attributes). In other cases, bridging can occur as long as there is at least one attribute shared or consistent between the two star schemas.

[0073] For example, in Figure 3 , dimension table 318a is the same as dimension table 328a (except for record IDs or other ways of identifying tuples that do not convey substantial information). Alternatively, rather than having duplicate tables, dimension tables 318a and 328a can be the same table but represented as members of multiple star schemas. Each attribute in dimension tables 318a, 328a can be used as a path between the facts in fact table 314 and the facts in fact table 324. However, each of these paths is different because different attributes are linked together. Which attributes are used to link dimension tables 318a and 328a can be important. For example, the operations implementing the path (e.g., as specified by an SQL statement) can be different. Additionally, some of these paths can use indexed attributes while others do not, which can affect the execution speed of a particular path.

[0074] In Figure 3 's example scenario, another way to obtain facts from fact tables 314 and 324 is by using attribute 340 of dimension table 318b and attribute 344 of dimension table 328b.

[0075] The various information in data model 300 can be used to determine which features (attributes of the tables in star schemas 310, 320) can be placed into feature groups. In some cases, the data model 300 as a whole can suggest that an attribute in data model 300 should be placed into a common feature group, or can be a factor considered when determining whether to place such an attribute into a common feature group or one of multiple feature groups. For example, if the attributes used in training or using a machine learning model come from multiple data sources, it may make sense to place the attributes from data model 300 into a common feature group. Or, when assigning attributes from multiple data sources to feature groups, being part of data model 300 can be a factor weighing for or against including a given feature in a given feature group.

[0076] Membership in the sub-elements of the data model 300 (e.g., whether an attribute / feature is part of star schema 310 or 320, or is in a separate table 314, 318, 324, 328 of the star schema) can be handled in a similar manner. Thus, a feature group can be proposed for one or more of star schema 310 or 320, or for one or more of tables 314, 318, 324, 328. Alternatively, the membership being such star schema 310, 320 or tables 314, 318, 324, 328 can be a factor in determining whether a given attribute / feature of the data model 300 should be included in a given feature group, even if the feature group does not correspond to a data model or a unit (or element, such as a table, view, attribute, OLAP cube definition) of the data model.

[0077] The relationships between individual attributes in the data model 300 can also be used to determine the feature groups to be formed, or to evaluate the membership of features in a feature group. For example, the attributes of table 350 of star schema 320 can be considered for inclusion in a feature group of the star schema, or in another feature group where membership in star schema 320 is a positive factor. Tables 328a and 324 of star schema 320 can be evaluated in a similar manner to table 350. However, if tables 328a and 324 are related by relationship 354 (e.g., can be having one or more common attributes, including foreign key relationships or associations), their membership in another feature group can be considered.

[0078] When two or more tables are related, a feature group or feature group membership can be proposed based on the related tables, related attributes, or a combination thereof. For example, relationship 354 can be used to propose that all attributes of tables 324, 328 should be part of a feature group, or to evaluate possible feature group membership. Alternatively, only the attributes 356 directly linked by relationship 354 can be evaluated in this manner. Or, one of the attributes 356 and its associated tables 324, 328 can be evaluated in this manner, but only the linked attributes of the other table are considered for inclusion in a given feature group.

[0079] Example 4 – Example relationships between tables in a data model

[0080] Figure 4 Additional details are provided on how the attributes of different tables are related, and how these relationships are used to define feature groups or evaluate potential membership in a feature group. Figure 4Table 404 represents a car, table 408 represents a license holder (e.g., a driver with a driver's license), table 412 provides an accident history, and table 416 represents a license number (e.g., associated with a license plate).

[0081] Each of tables 404, 408, 412, and 416 has multiple attributes 420 (although in some cases, a table may have only one attribute). For a particular table 404, 408, 412, 416, one or more of the attributes 420 can be used as a primary key - uniquely identifying a specific record in a tuple and designated as the main way to access tuples in the table. For example, in table 404, the Car_Serial_No attribute 420a is used as the primary key. In table 416, the combination of attributes 420b and 420c is used together as the primary key.

[0082] A table can reference a record associated with the primary key of another table by using a foreign key. For example, the license number table 416 has an attribute 420d of the car serial number in table 416, which is a foreign key and is associated with the corresponding attribute 420a of table 404. The use of foreign keys has multiple purposes. A foreign key can link specific tuples in different tables. For example, the foreign key value 8888 of attribute 420d will be associated with a specific tuple in table 404 that has the value of attribute 420a. A foreign key can also be used as a constraint, where a record with (or changed to have) a foreign key value that does not exist as a primary key value in the referenced table cannot be created. A foreign key can also be used to maintain database consistency, where a change to a primary key value can be propagated to the table whose attribute is the foreign key.

[0083] A table can have other attributes, or combinations of attributes, that can be used to uniquely identify a tuple, but they are not the primary key. For example, table 416 has an alternate key formed by attribute 420c and attribute 420d. Thus, the unique tuple in table 416 can be accessed either by using the primary key (e.g., as a foreign key in another table) or through an association with the alternate key.

[0084] In Figure 4In the scenario described above, it can be seen that there are multiple paths between the tables. For example, consider the operation of collecting data from Table 416 and Table 408. One path is to move from Table 416 to Table 412 using the foreign key 420e. Then, Table 408 can be reached through the foreign key relationship between the attribute 420l of Table 412 and the primary key 420m of Table 408. Alternatively, since Table 416 has the attribute 420d which serves as the foreign key of the primary key 420a of Table 404, and the attribute 420 is also associated with the replacement key of the attribute 420g of Table 408, Table 408 can be reached from Table 416 through Table 404.

[0085] In the above scenario, both paths have the same length, but are linked to different attributes of Table 412. Figure 4 The scenario described above is relatively simple. Thus, it can be seen that as the number of tables in the data model increases, the number of possible paths may increase significantly. Additionally, even between two tables, there may be multiple different paths. For example, Table 408 can access the tuples of Table 416 through the foreign key attributes 420h, 420i of Table 408, access the primary key attributes 420b, 420c of Table 416, or use the association to the replacement key of Table 416 provided by the attribute 420j of the reference attribute 420k of Table 416. Although the final paths from Table 408 to Table 416 are different, the paths differ in that different attributes 420 are connected.

[0086] If Tables 404, 408, 412, and 416 are represented graphically, each table can be a node. The paths between Tables 404, 408, 412, and 416 can be one-way or two-way edges. However, different paths between the tables form different edges. Again, using the paths between Table 408 and Table 416 as an example, the path through the foreign key attributes 420h, 420i and the path through the association attribute 420j are different edges.

[0087] In a similar manner to that described in Example 4, Tables 404, 408, 412, and 416 can be used to suggest and populate feature groups. Similarly, foreign keys, associations, or other relationships between tables (e.g., using defined views, triggers, using common SQL statements, including their individual attributes) can be used to suggest or populate feature groups. In addition to using the relationships that exist between tables to suggest or populate feature groups, the number, type, or direction of the relationships between tables can also be considered, such as placing more emphasis on foreign key relationships when determining the membership of a feature group.

[0088] Example 5 – Example Relationships between Elements of a Database Schema

[0089] In some cases, data model information can be stored in a data dictionary or similar repository, such as an information schema. The information schema can store information that defines an overall data model or schema, tables in the schema, attributes in the tables, and relationships between the tables and their attributes. However, data model information can include additional types of information as Figure 5 shown.

[0090] Figure 5 FIG. 6 is a schematic diagram showing the elements of a database schema 500 and how they are interrelated. These interrelationships can be used to define feature groups or to evaluate membership in a feature group. At least in some cases, the database schema 500 can be maintained outside of the database layer of a database system. That is, for example, the database schema 500 can be independent of the underlying database that includes a schema for the underlying database. Typically, the database schema 500 is mapped to a schema (e.g., Figure 4 schema 400) of the database layer such that records or portions thereof (e.g., specific values of specific fields) can be retrieved via the database schema 500.

[0091] The database schema 500 can include one or more packages 510. The packages 510 can represent organizational components for classifying or categorizing other elements of the schema 500. For example, the package 500 can be replicated or deployed to various database systems. The packages 510 can also be used to enforce security restrictions, such as by limiting access to specific schema elements by specific users or specific applications. The packages 510 can be used to define feature groups. Or, if attributes are members of a given package 510, this may more or less indicate that they should be included in another feature group.

[0092] The packages 510 can be associated with one or more domains (i.e., specific types of semantic identifiers or semantic information) 514. In turn, the domains 514 can be associated with one or more packages 510. For example, domain 1 (514a) is only associated with package 510a, while domain 2 (514b) is associated with both package 510a and package 510b. At least in some cases, the domains 514 can specify which packages 510 can use the domain. For example, a domain 514 associated with materials used in a manufacturing process can be used by a process control application but not by a human resources application.

[0093] At least in some embodiments, although multiple packages 510 can access a domain 514 (and database objects incorporated into the domain), the domain (and optionally other database objects, such as tables 518, data elements 522, and fields 526, which will be described in more detail below) is primarily assigned to one package. Assigning the domain 514 and other database objects to a single package can help create logical (or semantic) relationships between database objects. InFigure 5 In it, the allocation from domain 514 to package 510 is shown as a solid line, while the access permission is shown as a dashed line. Thus, domain 514a is allocated to package 510a, and domain 514b is allocated to package 510b. Package 510a can access domain 514b, but package 510b cannot access domain 514a.

[0094] Note that at least some database objects (such as table 518) can include database objects associated with multiple packages. For example, table 518 (Table 1) can be allocated to package A and have fields allocated to packages A, B, and C. The use of the fields in Table 1 allocated to packages A, B, and C creates a semantic relationship among packages A, B, and C, where if a field is associated with a specific domain 514, this semantic relationship can be further explained (i.e., the domain can provide further semantic context for database objects associated with objects of another package rather than being allocated to a common package).

[0095] As will be explained in more detail, domain 514 can represent the most granular unit from which database table 518 or other schema elements or objects can be built. For example, domain 514 can be associated with at least a data type. Each domain 514 is associated with a unique name or identifier and typically with a description that provides the semantic meaning of the domain, such as a human-readable text description (or an identifier that can be associated with a human-readable text description). For example, one domain 514 can be an integer value representing a phone number, while another domain can be an integer value representing a part number, and another integer domain can represent a social security number. Thus, domain 514 can maintain a common and consistent use (e.g., semantic meaning) across schema 500. That is, for example, whenever a domain representing a social security number is used, the corresponding field can be recognized as having that meaning, even if the field or data element has different identifiers or other characteristics for different tables.

[0096] Since domains 514 can be used to help provide a common and consistent semantic meaning, they can be used to define feature groups. Alternatively, domains 514 can be used to decide whether an attribute associated with a domain should be part of a given feature group, where the feature group is not defined entirely based on domains.

[0097] Schema 500 may include one or more data elements 522. Each data element 522 is typically associated with a single domain 514. However, multiple data elements 522 may be associated with a particular domain 514. Although not shown, multiple elements of table 518 may be associated with the same data element 522, or may be associated with different data elements having the same domain 514. Data elements 522 can be used in particular to allow customization of domain 514 for a particular table 518. Thus, data elements 522 can provide additional semantic information for the elements of table 518.

[0098] Table 518 includes one or more fields 526, at least a portion of which map to data elements 522. Fields 526 can map to the schema of the database layer, or table 518 can be mapped to the database layer in another form. In any case, in some embodiments, fields 526 are mapped to the database layer in some way. Alternatively, the database schema can include semantic information equivalent to the elements of schema 500 that include domains 514.

[0099] In some embodiments, one or more of fields 526 do not map to domain 514. For example, fields 526 can be associated with primitive data components (e.g., primitive data types such as integers, strings, booleans, character arrays, etc., where primitive data components do not include semantic information). Or, the database system can include one or more tables 518 that do not include any fields 526 associated with domain 514. However, the disclosed techniques include schema 500 (which can be separate from or incorporated into the database schema), which includes multiple tables 518 having at least one field 526 that is directly or through data element 522 associated with domain 514.

[0100] Because data elements 522 can indicate common or related attributes, they can be used to define groups of features, or to evaluate membership in a group of features, such as those described for domain 514 and package 510.

[0101] Example 6 – Example data dictionary

[0102] Schema information (such as that associated with Figure 5The information associated with schema 500 can be stored in a repository (such as a data dictionary). As previously mentioned, at least in some cases, the data dictionary is independent of but mapped to an underlying relational database. This independence can allow the same database schema 500 to be mapped to different underlying databases (e.g., databases using software from different vendors, or different software versions or products from the same vendor). The data dictionary can be persisted (such as maintained in stored tables) and can be maintained in memory in whole or in part. The in-memory version of the data dictionary can be referred to as the dictionary buffer.

[0103] Figure 6 A database environment 600 with a data dictionary 604 is shown, which can access a database layer 608 (such as through mapping). The database layer 608 can include a schema 612 (e.g., INFORMATION_SCHEMA in PostgreSQL) and data 616, such as data associated with a table 618. The schema 612 includes various technical data items / components 622 that can be associated with a field 620, such as a field name 622a (which may or may not correspond to a human-readable description of the purpose of the field, or otherwise explicitly describe the semantics of the value of the field), a field data type 622b (e.g., integer, variable character, string, boolean), a length 622c (e.g., the size of the numbers allowed for the value in the field, the length of a string, etc.), the number of decimal places 622d (optionally, for a suitable data type, such as a floating point number with a length of 6, specifying whether the value represents XX.XXXX or XXX.XXX), a position 622e (e.g., the position in the table where the field should be displayed, such as the first display field, the second display field, etc.), optionally, a default value 622f (e.g., "NULL", "0", or some other value), a null flag 622g indicating whether the field allows null values, a primary key flag 622h indicating whether the field is the primary key of the table or is used for the primary key of the table, and a foreign key element 622i that can indicate whether the field 620 is associated with the primary key of another table, and optionally, the identifier of the table / field referenced by the foreign key element. A particular schema 612 can include more, fewer, or different technical data items 622 than Figure 6 shown.

[0104] All or part of the technical data items 622 can be used to define a feature group or evaluate the membership of a feature in a feature group. In particular, the foreign key element 622i can be used to identify other tables (and their specific attributes) that may be related to a given field, where the other table or field can be considered for membership in a feature group.

[0105] Table 618 is associated with one or more values 626. The values 626 are typically associated with fields 620 defined using one or more of the technical data elements 622. That is, each row 628 typically represents a unique tuple or record, and each column 630 is typically associated with the definition of a particular field 620. Table 618 is typically defined as a collection of fields 620 and is given a unique identifier.

[0106] The data dictionary 604 includes one or more packages 634, one or more domains 638, one or more data elements 642, and one or more tables 646, which can at least generally correspond to Figure 5 the similarly captioned components 510, 514, 522, 518 of Figure 5 As explained in the discussion of

[0107] each package 634 includes one or more (typically, multiple) of the domains 638. Each domain 638 is defined by a plurality of domain elements 640. The domain elements 640 can include one or more names 640a. The names 640a are used to uniquely identify a particular domain 638 in some cases. A domain 638 includes at least one unique name 640a and can include one or more names that may or may not be unique. The names (which can be unique or not) can include versions of the name or description of the domain name 638 of different lengths or levels of detail. For example, the name 640a can include text that can be used as a label for the domain 638 and can include short, medium, and long versions, as well as text that can be designated as a heading. Alternatively, the name 640a can include a primary name or identifier and a short description or field label that provides human - understandable semantics for the domain 638.

[0108] The domain element 640 may also include at least information similar to the information that may be included in the schema 612. For example, the domain element 640 may include a data type 640b, a length 640c, and a number of decimal places 640d associated with the relevant data type, which may correspond to the technical data elements 622b, 622c, 622d, respectively. The domain element 640 may include conversion information 640e. The conversion information 640e may be used to convert (or mutually convert) the values input to the domain 638 (optionally, including the values modified by the data element 642). For example, the conversion information 640 may specify that a number in the form XXXXXX should be converted to XXX-XX-XXXX, or that the number should have a decimal point or comma separating groups of digits (e.g., formatting 1234567 as 1,234,567.00). In some cases, the field conversion information for multiple domains 638 may be stored in a repository, such as a field catalog.

[0109] The domain element 640 may include one or more value restrictions 640f. The value restrictions 640f may specify, for example, whether negative values are allowed or not, or a specific range or threshold of values acceptable to the domain 638. In some cases, when an attempt is made to use a value that does not conform to the value restriction 640 for the domain 638, an error message or a similar indication may be provided. The domain element 640g may specify one or more packages 634 that are allowed to use the domain 638.

[0110] The domain element 640h may specify metadata that records the creation or modification events associated with the domain element 638. For example, the domain element 640h may record the identity of the user or application that last modified the domain element 640h, as well as the time when the modification occurred. In some cases, the domain element 640h stores a larger history (including the complete history) of the creation and modification of the domain 638.

[0111] The domain element 640i may specify the original language associated with the domain 638 that includes the name 640a. For example, the domain element 640i may be useful when it is necessary to determine whether the name 640a should be converted to another language, or how such a conversion should be accomplished.

[0112] The data element 642 can include a data element field 644, where at least some of the data element fields 644 can be at least substantially similar to the domain element 640. For example, the data element field 644a can correspond to at least a part of the name domain element 640a, such as being (or including) a unique identifier of a specific data element 642. The field label information described for the name domain element 640a is shown as being divided into a short description label 644b, a medium description label 644c, a long description label 644d, and a header description 644e. As described for the name domain element 640a, the labels and headers 644b - 644e can be maintained in one language or multiple languages.

[0113] The data element field 644f can specify the domain 638 to be used with the data element 642, thereby incorporating the characteristics of the domain element 640 into the data element. The data element field 644g can represent the default value of the data element 642, and can be at least similar to the default value 622f of the schema 612. The created / modified data element field 644h can be at least substantially similar to the domain element 640h.

[0114] The table 646 can include one or more table elements 648. At least a part of the table element 648 can be at least similar to the domain element 640, such as the table element 648a being at least substantially similar to the domain element 640a or the data element field 644a. The description table element 648b can be similar to the description and header labels described in connection with the domain element 640a, or the label and header data element fields 644b - 644e. The table 646 can be associated with a type using the table element 648c. Example table types include transparent tables, clustered tables, and pooled tables, such as those used in the database products provided by SAP SE in Walldorf, Germany TM provided.

[0115] The table 646 can include one or more field table elements 648d. The field table element 648d can define a specific field of a specific database table. Each field table element 648d can include an identifier 650a for the specific data element 642 for that field. The identifiers 650b - 650d can specify whether the field is the primary key or part of the primary key of the table (identifier 650b), or whether it has a relationship with one or more fields of another database table, such as being a foreign key (identifier 650c) or an association (identifier 650d).

[0116] The created / modified table element 648e can be at least substantially similar to the domain element 640h.

[0117] The package 634, the field 638, the data element 642, and their specific components (e.g., components 640, 644) can be used to define the properties of a feature group or to evaluate the membership in a feature group, such as those described in Example 5. Similarly, the table 646 and its elements (specifically, the type 648c, the primary key 650b, the foreign key 650c, and the association 650d) can be used to define the properties of a feature group or to evaluate the membership in a feature group. For example, features having a common value for the original language domain element 640i can be proposed to form a feature group or may more or less likely be included in another feature group.

[0118] Example 7 – Example relationships between database objects based on data access operations

[0119] Data requests can also be used to identify feature groups or to evaluate the membership of features in a feature group. As an example, data requests (such as those specified in a query language statement) can be used to identify feature groups and their constituent features.

[0120] Figure 7 An example logical query plan 700 for a query involving multiple query operations (including several join operations) is shown. In some cases, the overall query plan 700 can identify possible feature groups or be used to evaluate the membership in a feature group. For example, in addition to the data associated with the query plan 700, the features of a machine learning model can come from one or more sources. Thus, in some cases, the features from the query plan 700 (including in terms of their contribution / capability to predicting the results of a machine learning model) can be relevant.

[0121] Additional feature groups or membership evaluation criteria can be suggested by one or more operations in the query plan 700. For example, a join operation can indicate possible feature groups or membership evaluation criteria, where at least some of the features of the data sources associated with that join are included or considered to be included in such a feature group.

[0122] The query plan 700 includes a join 710 that joins the results of the joins 714, 716. The join 714 itself joins the results of a join 724 from tables 720 and tables 728, 730. Similarly, the join 716 includes the results of a join 738 of tables 734 and tables 742, 744.

[0123] At each level of the query plan 700, a join operation can suggest a feature group or criteria that can be used to evaluate membership in a feature group. For example, join 714 can be referenced to define a feature group or membership criteria, which can then include features associated with tables 720, 728, 730. Similarly, join 716 can be referenced to define a feature group or membership criteria that includes features associated with tables 734, 742, 744.

[0124] Moving down in the query plan 700, a feature group or membership criteria can be defined based on join 724 (tables 728, 730) or join 738 (tables 742, 744). Individual data sources, tables 720, 728, 730, 734, 742, 744 can also be considered as possible feature groups or membership criteria.

[0125] Each of the joins 710, 714, 716, 724, 738 includes one or more join conditions. A join condition can be a relationship between a feature or intermediate result of one table (e.g., the result of another join) and a feature or intermediate result of another table. A join can also include a filter condition (e.g., a predicate) and other operations defined for features of one or more of the data sources being joined. One or more features included in a join condition can be considered as defining a feature group or at least partially used to evaluate membership in a feature group.

[0126] Similarly, the query plan 700 can include operations other than the joins 710, 714, 716, 724, 738. These operations can include a predicate 750 (e.g., a filter condition), a sort operation 752 (e.g., ascending sort), and a projection operation 754 (e.g., selecting specific fields of the results returned by an earlier operation). These operations 750, 752, 754 that include specific features used in the operations (e.g., fields used in predicate evaluation or for sorting) can be used to define a feature group or as membership criteria. As an example, tables 720, 728, 730, 734, 742, 744 can all have features (e.g., attributes / fields / columns) used in projection 754, and the feature / projection can be defined as a feature group or used to evaluate membership in a feature group.

[0127] Example 8 – Example relationships between features

[0128] In some embodiments, a feature group can be determined by evaluating the relationships between features. These relationships can be determined by various techniques, including using various statistical techniques. One technique involves determining the mutual information of feature pairs, which identifies the dependencies of features on each other. However, other types of relationship information can be used to identify related features, as can various clustering techniques.

[0129] Figure 8 A plot 800 (e.g., a matrix) showing the mutual information of ten features is shown. Each square 810 represents the mutual information or correlation or dependency of a pair of different features. For example, square 810a reflects the dependency between feature 3 and feature 4. The squares 810 can be associated with discrete numerical values indicating any dependency between variables, or these values (including providing a heat map of the dependency) can be binned.

[0130] As shown, the plot 800 shows squares 810 with different fill patterns, where the fill pattern indicates the strength of the dependency between feature pairs. For example, a greater dependency can be represented by a darker fill value. Thus, square 810a can indicate a strong correlation or dependency, square 810b can indicate little or no dependency between features, and squares 810c, 810d, 810e can indicate a medium level of dependency.

[0131] Dependency information can be used to define feature groups and to determine membership in a feature group. For example, features that have a dependency on other features within at least a given threshold can be considered part of a common feature group. Referring to plot 800, it can be seen that feature 10 has varying degrees of dependency on features 1, 3, 4, 6, 7. Thus, features 1, 3, 4, 6, 7, and 10 can be defined as a feature group. Or, if the threshold is set such that feature 4 does not meet the mutual relationship threshold, then feature 4 can be excluded. In other embodiments, features that have at least a threshold dependency on features 3, 4, 5, 6, 7 can be added to the feature group associated with feature 10.

[0132] Various criteria can be defined for suggesting feature groups, including the minimum or maximum number of feature groups, or the minimum or maximum number of features within a feature group. Similarly, a threshold can be set for features that are considered likely to be included in a feature group (e.g., where features that do not meet the threshold of any other feature can be omitted from plot 800). A threshold can also be set for the feature dependencies that qualify for membership in a feature group (that is, if the dependency of feature 1 on feature 2 meets the threshold, then feature 1 or feature 2 can be included in a feature group that is not defined based on the dependency of feature 1 or feature 2).

[0133] In some cases, the feature groups identified using relevance / mutual dependency can be manually adjusted before or after recognition. For example, before determining the feature groups, the user can specify that certain features must be considered when determining the feature groups, or exclude certain features from the determination of the feature groups. Alternatively, a feature group including one or more features can be defined, and mutual information can be used to populate the group. Information (such as drawing 800) can be presented to the user to evaluate whether any manual selection is correct. For example, the dependency value may indicate that the user's manual assignment of the feature group is incorrect. The user can manually adjust the feature group / features used to determine the feature group after being presented with drawing 800 or otherwise having mutual information available (such as adding a feature that the user believes should be in the feature group to the feature group even though the dependency information does not indicate that the feature should be in the feature group).

[0134] Various methods for determining relevance can be used, such as mutual information. Generally, mutual information can be defined as where X and Y are random variables with joint distribution P(X,Y) and marginal distributions P X 、P Y . The mutual information can include various types of mutual information, such as metric-based mutual information, conditional mutual information, multivariate mutual information, directed information, normalized mutual information, weighted mutual information, adjusted mutual information, absolute mutual information, and linear correlation. The mutual information can include calculating Pearson's correlation, including using Pearson's chi-squared test, or using G-test statistics.

[0135] When used to evaluate a first feature relative to a specified (target) second feature, supervised correlation can be used: scorr(X,Y) = corr(ψ X ,ψ Y ), where scorr is Pearson's correlation and (binary classification).

[0136] In some examples, a modified X 2 test can be used to calculate the dependency between two features:

[0137]

[0138] where:

[0139]

[0140] O xy is the observed count for observing X = x and Y = y, while E xy is the expected count when X and Y are independent.

[0141] Note that this test produces signed values, where a positive value indicates that the observed count is higher than expected, and a negative value indicates that the observed count is lower than expected.

[0142] Similarly, dependent features can be considered to be included in a feature group or used to define a feature group. The dependencies between features can also be used to otherwise interpret the results provided to a machine learning model, either for individual features or as part of an analysis that groups at least some features into feature groups.

[0143] In yet another embodiment, the mutual relationship between features (which can be related to the variability in the SHAP values of the features) can be calculated as:

[0144]

[0145] where φ ii is the main SHAP contribution of feature i (excluding the mutual relationship), and φ ij + φ ji is the contribution of the mutual relationship between variables i and j and the strength of the mutual relationship between features can be calculated as:

[0146]

[0147] Example 9 - An example display for showing the relationship between features

[0148] Mutual information or other types of dependency or correlation information determined using the techniques described in Example 8 (or used without being presented to the user in a visual form, such as simply providing the feature groups resulting from the analysis) can be presented to the user in different formats. For example, Figure 9 plot 900 shows relationship 910 between features 914, where features 914 can be features whose relationship strength meets a threshold.

[0149] Relationship 910 can be encoded with information indicating the relative strength of the relationship. As shown, relationship 910 is shown with different line weights and styles, where various combinations of styles / weights can be associated with different strengths (e.g., ranges or bins of strength). For example, for a given line weight, highly dashed lines can indicate a weaker relationship, while increasing line weights can indicate a stronger relationship / dependency. In other cases, relationship 910 can be shown in different colors to indicate the strength of the relationship.

[0150] Using the information in Plot 900 (and / or Plot 800), the user can adjust the feature groups, such as by adding or removing feature groups. That is, Plot 900 can represent an overall analysis of the features used in a machine learning model or a subset of these features. It is possible that all features 914 should be included in the feature group. Alternatively, the user may wish to, for example, change a threshold such that features 914 with a weaker relationship 910 are omitted from an updated version of Plot 900. Or, the user may wish to manually add features to a feature group that are not shown in Plot 900 (or, not shown as being linked by relationship 910), or may wish to manually remove features from a feature group.

[0151] In particular, it may be helpful for the user to evaluate the mutual information results to confirm that the results given the meaning of the different features make sense. This helps ensure that desired relationships are not overlooked and that features correlated by spurious relationships are not included in the feature group.

[0152] Example 10 - An example of assigning features to feature groups

[0153] Figure 10 is a schematic diagram showing how at least a portion of feature 1010 (e.g., Figure 1 and 2 feature 110) can be assigned to feature group 1014 (including based on one or more of the techniques described in Examples 1 - 9). It can be seen that feature group 1014 can include a different number of features. Although not shown, feature group 1014 may include a single feature.

[0154] Typically, each feature 1010 is included in a single feature group 1014. However, in some cases, a given feature 1010 can be included in multiple feature groups 1014. For example, feature 1010a (feature_1) is shown as a member of feature groups 1014a and 1014b. Although feature groups 1014 can have the same number of features 1010, typically feature groups 1014 are allowed to have different numbers of features, where user input, statistical methods, data relationships, other information, or a combination thereof (including as described in Examples 1 - 9) is used to determine the identity / number of features in the group.

[0155] In some cases, all features 1010 are assigned to feature groups 1014. However, one or more of the feature groups 1014 can simply be designated as "leftover" features that are not specifically assigned to another feature group. Or, as Figure 10 shown, some of features 1010, feature 1010b do not need to be assigned to feature group 1014.

[0156] SHAP values that can aggregate features (such as for feature groups), including:

[0157]

[0158] In some embodiments, the relative importance of a feature group can be defined as:

[0159]

[0160] Example 11 – Example display screen showing feature group information

[0161] Figure 11 Illustrates how feature groups can be used to provide information on how such feature groups contribute to a machine learning model (including their contribution to a specific result provided by the machine learning model for a specific set of feature input values).

[0162] Figure 11 Presents an example user interface screen 1100 that provides result 1108 in a predictive form. Result 1108 can be an indication of how likely the result (e.g., success) is for a specific set of input values.

[0163] Panel 1112 of screen 1100 lists multiple feature groups 1120. For each feature group 1120, the contribution 1124 of the feature group to result 1108 is shown. If the contribution is from a feature not assigned to a feature group or from a feature group not shown in panel 1112, the contribution 1124 can be normalized or otherwise calculated such that the sum of all contributions 1124 is 100%, or at least the sum of the contributions is less than 100%. A visual indicator 1132 (such as a bar) can be shown to help visually convey the contribution 1124 of feature group 1120.

[0164] The user can choose to expand or collapse feature group 1120 to view the contributions 1130 from individual features 1134 within the feature group. Typically, the sum of the contributions 1130 within feature group 1120 will equal the contribution 1124 of the feature group. However, other metrics can be used. For example, certain feature groups 1120 or features 1134 can be weighted more or less than other feature groups or features. Or, a measure of the importance of feature group 1120 can be presented that takes into account multiple features within the feature group. For example, the contribution or importance of feature group 1120 can be calculated as the average (mean value) of the features 1134 in the group. In this case, the contribution of feature group 1120 is weighted by considering the number of features in the group. However, in other cases, the contribution to feature group 1120 is not weighted because if the group has a higher absolute contribution to result 1108, the group may have higher importance even if it also has more features 1134 than other feature groups.

[0165] As described in Example 1, a user may adjust a machine learning model based on Feature Group 1124. Retraining the machine learning model based on the selected Feature 1134 / Feature Group 1120 may improve the performance of the model for at least some scenarios, where the improved performance may be one or both of improved accuracy or improved performance (e.g., speed or efficiency such as processing fewer features / data, using less memory or processor resources).

[0166] As Figure 11 shown, a user may select box 1140 for Feature Group 1120 or an individual Feature 1134. By selecting icon 1144, the user may use the selected Feature 1134 / Feature Group 1120 to retrain the model.

[0167] Example 12 – Example of Construction and Use of Feature Groups

[0168] Figure 12 is a schematic diagram showing how a feature group may be determined from a data set 1210 and optionally used to train (e.g., retrain) a machine learning model. The data set 1210 is obtained from one or more data sources (not shown) (including as described in Example 111). The data set 1210 includes a plurality of features 1214, at least a portion of which are used to train a machine learning model or provide classification results using a trained classifier.

[0169] The data set 1210 may be partitioned into multiple parts, including a part used as training data 1218 and a part used as classification data 1220 (or more generally, as an analysis data set). The training data 1218 may be processed by a machine learning algorithm 1224 to provide a trained classifier 1228. The classification data 1220 may be processed by the trained classifier 1228 to provide results 1232. The results 1232 and a portion of the data set 1210 may be processed to provide feature contributions 1236. The feature contributions 1236 may include context contribution values such as context or overall SHAP or LIME values, overall contribution values (including as defined in Example 13), or combinations thereof.

[0170] Feature groups can be extracted at 1240. Extracting feature groups at 1240 can include analyzing feature contributions to determine relevant / irrelevant features. Extracting feature groups at 1240 can include extracting feature groups or determining membership in a feature group based on other considerations (such as relationships between features in one or more data sources used to form dataset 1210, data access requests when obtaining data for the dataset, or any predefined feature groups that may have been provided). Extracting feature groups at 1240 can also include applying one or more clustering techniques to the features, including based on data associated with the features in dataset 1210, results 1232, feature contributions determined at 1236, or combinations of these factors.

[0171] Feature groups can be viewed or adjusted at 1244. Viewing and adjusting feature groups at 1244 can include a user manually viewing or adjusting the feature groups, which can include adding or removing feature groups, or adding or removing features from a feature group. In other cases, viewing and adjusting of feature groups can be performed automatically (such as by using rules to analyze relationships between features in a feature group and rejecting features that do not meet a threshold for membership in the feature group, or removing feature groups that do not meet criteria for forming a feature group (e.g., features in a feature group do not meet a minimum prediction contribution threshold set for the feature group, do not meet a threshold number of features in the feature group, features in the feature group are not sufficiently related / dependent, other factors, or combinations of these or other factors)).

[0172] Feature groups determined as a result of extraction at 1240 and inspection / adjustment at 1244 can optionally be used to retrain the trained classifier 1228. Or the feature groups can be used to train a machine learning algorithm 1224 with the same dataset 1210 or a new dataset, as described for classifying classification data 1220 using the trained classifier 1228, after which feature contributions 1236 can be extracted, feature groups can be extracted at 1240, and inspection / adjustment can be performed again at 1244. Then, the process can continue as needed.

[0173] Although shown as including a data set 1210 divided into training data 1218 and classification data 1220, in some embodiments, the training data and classification data need not be from a training data set. At least in some cases, whether a common data set is required can depend on the particular technique used for the machine learning algorithm 1224. Additionally, one or more of the feature contributions determined at 1236, the feature group extraction at 1240, and the feature group inspection / adjustment at 1244 can be performed using the results 1232 from multiple data sets 1210 (which may have been previously used as classification data or may be the classification data portion of a data set). For example, aggregated SHAP values can be determined from the results 1232 of multiple sets of classification data 1220. Other techniques such as cross-validation can be used to help determine whether two sets of results 1232 are suitable for use in steps 1236, 1240, 1244. In some cases, at least a portion of the training data 1218 can be used as classification data 1220.

[0174] Additionally, it should be understood that not all of the disclosed techniques require the use of a machine learning model 1224 to identify feature groups. As explained in Examples 1 and 3 - 7, feature groups can be determined manually or based on data sources associated with the features or relationships between features / data sources. Similarly, trained ML models can be used to identify feature groups using techniques such as mutual information or supervised correlation to cluster features. However, if desired, feature groups determined using such other techniques can be used with the machine learning algorithm 1224 / trained classifier 1228, such as to determine how to train the ML algorithm 1224 or to interpret the results 1232 provided by the trained classifier.

[0175] Although many examples discuss the use of feature groups with classification machine learning tasks / algorithms, it should be further understood that the disclosed techniques (including the identification and use of feature groups) can be used with other types of machine learning algorithms, including those for multi-class classification or regression.

[0176] Example 13 – Example Calculations and Comparisons of Context and Overall Feature Contributions

[0177] In some cases, information about the importance of a feature for a machine learning model (or a specific result provided using such a model) can be accessed by comparing the contextual importance of the feature (such as determined by SHAP or LIME calculations) and the overall importance of the feature based on a univariate prediction model (e.g., the strength of the association between the feature and the result). That is, for example, the difference between the contextual importance and the overall importance can indicate relationships that cannot be revealed by typical overall importance analysis of single-features or even by using contextual analysis.

[0178] In a specific example, the overall importance of feature X (e.g., using univariate SHAP values) can be calculated as:

[0179]

[0180] where x i is the value of feature X for observation i, and N is the number of observations in the test set.

[0181] In a specific example, using the SHAP technique, the contextual importance of feature X (using contextual SHAP values) can be calculated as:

[0182]

[0183] where φ X,i is the contextual SHAP value of feature X and observation i.

[0184] The overall importance and the contextual importance of one or more features can be compared (including using bar charts or scatter plots). Features with a large difference between the overall importance and the contextual importance can be marked for review. In some cases, a threshold difference can be set, and features that meet the threshold can be presented to the user. Additionally, a low correlation (e.g., <<1) between the contextual importance and the overall importance of a feature can indicate that the relationship of the feature with the result is contrary to what is expected from the overall (univariate) SHAP values.

[0185] The statistics of the features identified using this technique can be further analyzed, such as analyzing the statistics of different values of the features. The difference in the number of data points in the dataset with a specific value, and the association of that value with a specific result, can provide information that can be used to adjust the machine learning model or change behavior to influence the result, where the result can be an actual (simulated world) result or a result obtained using the machine learning model.

[0186] The patterns of the values of features can also be compared using global and contextual metrics. For example, the change between the global metric and the contextual metric may occur only in one or a specified subset of the values of the feature. Alternatively, the differences may be observed more consistently (e.g., for all values or for a larger subset of the values). The consistency of the differences can be calculated as the Pearson correlation between the contextual SHAP values of the feature and its global association with the target (outcome) (univariate SHAP values). A value close to 1 indicates high consistency, while a value close to 0 indicates low or no consistency. Negative values may indicate the presence of an anomalous relationship, such as Simpson’s Paradox.

[0187] In a specific example, the global importance and the contextual importance of a feature can be presented on a plot that shows the consistency between the global importance and the contextual importance over a range of values. For example, the plot can have importance values presented on the Y-axis and the values of a specific feature on the X-axis, where the contextual importance values and the global importance values are plotted (including showing the variation in the consistency of the contextual values and the global values over various observations in the dataset (e.g., for multiple input instances)).

[0188] Example 14 – Example of Determination and Use of Causal Relationship Information

[0189] As described in Example 1, causal relationship information can be determined and used to analyze a machine learning model and potentially to modify the model. For example, an initial set of results provided by the model can be analyzed, including determining features that influence other features. At least some features that are dependent on other features can be excluded from retaining the model. In this way, the impact of a specific feature (or group of features) on the outcome can be isolated by reducing or eliminating the influence of the dependent features. Similar analysis and training can be performed using groups of features.

[0190] Various methods (including the methods described in Examples 1 - 13) can be used to determine the relationships between features. For a given feature and its dependent features, the feature can be further classified, including manually. This further classification can include determining whether the feature is an actionable feature or a non-actionable feature. An actionable feature can be a feature that can potentially be influenced or changed (such as by someone using the results of an ML model or during a scenario modeled using an ML model). As an example, it can be recognized that features such as gender or country of birth have an impact on features such as occupation or education. However, while occupation or education can be actionable (e.g., can help someone have a different occupation or have a different educational status), features such as gender or country of birth are non-actionable.

[0191] Features can be classified into feature groups using characteristics such as whether they are actionable or non-actionable features. For example, for a given feature group category, additional sub-categories can be formed for actionable and non-actionable features. In such cases, one or more of the sub-categories can be used as the feature group without using the original parent group.

[0192] When analyzing a particular problem, a user can choose to use relevant features / feature groups to train a machine learning model or choose a suitably trained ML model. For example, if interested in a particular actionable feature, the user may wish to use data for that actionable feature and relevant features / feature groups to train the model. In some cases, the relevant features / feature groups can be those that do not depend on the feature of interest (or the feature group of which the feature of interest is a member). In other cases, the relevant features / feature groups can be or can include features that depend on the feature of interest (or the group of which the feature of interest is a member). Similarly, in some cases, the relevant features / feature groups can be actionable features, while in other cases, the relevant features / feature groups are non-actionable features. Various combinations of actionable / non-actionable and dependent / non-dependent features / feature groups can be used as needed.

[0193] Using these techniques can provide various advantages, including helping the user understand the relative contributions of actionable / non-actionable features / feature groups. At least in some cases, actionable features can be a relatively small subset of the features used in an ML model. In any case, being able to focus on actionable features can help the user better understand how different outcomes are achieved for a given scenario modeled by the ML model.

[0194] In addition, using causal relationship information, by excluding intermediate results of the feature of interest, the user can exclude dependent features from the analysis (e.g., by excluding them from model training) to better understand the overall impact of features such as actionable features on the outcome.

[0195] Example 15 – Example method for training and using a classifier

[0196] Figure 13A is a flowchart of an example method 1300 for forming feature groups. At 1310, a training data set is received. The training data set includes values for a first plurality of features. At 1314, the training data set is used to train a machine learning algorithm to provide a trained machine learning algorithm. At 1318, the trained machine learning algorithm is used to process an analysis data set to provide a result. At 1322, a plurality of feature groups are formed. At least one of the feature groups includes a second plurality of features from the first plurality of features. The second plurality of features is a proper subset of the first plurality of features.

[0197] Figure 13B It is a flowchart of an example method 1340 for forming feature groups using dependencies between features in a dataset. At 1344, a training dataset is received. The training dataset includes values of a first plurality of features. At 1348, the training dataset is used to train a machine learning algorithm to provide a trained machine learning algorithm. At 1352, the trained machine learning algorithm is used to process an analysis dataset to provide a result. At 1356, context contribution values are determined for a second plurality of features among the first plurality of features. At 1360, dependencies between features among the second plurality of features are determined. At 1364, a plurality of feature groups are formed at least in part based on the determined dependencies. At least one feature group among the plurality of feature groups includes a third plurality of features among the first plurality of features. The third plurality of features is a proper subset of the first plurality of features.

[0198] Figure 13C It is a flowchart of an example method 1370 for determining feature group contribution values. At 1374, a first plurality of features used in a machine learning algorithm are determined. At 1378, a plurality of feature groups are formed, such as using analysis of machine learning results, semantic analysis, statistical analysis, data lineage, or a combination thereof. At least one feature group among the plurality of feature groups includes a second plurality of features among the first plurality of features. The second plurality of features is a proper subset of the first plurality of features. At 1382, the machine learning algorithm is used to determine a result for an analysis dataset. At 1386, for at least a portion of the feature groups, contribution values of the features of the corresponding feature groups to the result are aggregated to provide feature group contribution values.

[0199] Example 16 - Computing System

[0200] Figure 14 Depicts a general example of a suitable computing system 1400 in which the described innovations may be implemented. Computing system 1400 is not intended to impose any limitation on the scope of use or functionality of the present disclosure, as the innovations may be implemented in different general-purpose or special-purpose computing systems.

[0201] Reference Figure 14 , computing system 1400 includes one or more processing units 1410, 1415 and memories 1420, 1425. At Figure 14In it, the basic configuration 1430 is included within the dashed lines. The processing units 1410, 1415 execute computer-executable instructions, such as those for implementing the techniques described in Examples 1-15. The processing unit can be a general-purpose central processing unit (CPU), a processor in an application-specific integrated circuit (ASIC), or any other type of processor. In a multiprocessing system, multiple processing units execute computer-executable instructions to improve processing power. For example, Figure 14 A central processing unit 1410 and a graphics processing unit or coprocessing unit 1415 are shown. The tangible memories 1420, 1425 can be volatile memories (e.g., registers, caches, RAM) accessible by the processing units 1410, 1415, non-volatile memories (e.g., ROM, EEPROM, flash memory, etc.), or some combination of both. The memories 1420, 1425 store software 1480 that implements one or more innovations described herein in the form of computer-executable instructions suitable for execution by the processing units 1410, 1415.

[0202] The computing system 1400 can have additional features. For example, the computing system 1400 includes a storage 1440, one or more input devices 1450, one or more output devices 1460, and one or more communication connections 1470. An interconnection mechanism (not shown) (such as a bus, a controller, or a network) interconnects the components of the computing system 1400. Typically, an operating system software (not shown) provides an operating environment for other software executing in the computing system 1400 and coordinates the activities of the components of the computing system 1400.

[0203] The tangible storage 1440 can be removable or non-removable and includes magnetic disks, tapes, or cartridges, CD-ROMs, DVDs, or any other medium that can be used to store information in a non-transitory manner and can be accessed within the computing system 1400. The storage 1440 stores instructions for the software 1480 to implement one or more innovations described herein.

[0204] The input devices 1450 can be touch input devices, such as a keyboard, a mouse, a pen, or a trackball, voice input devices, scanning devices, or other devices that provide input to the computing system 1400. The output devices 1460 can be a display, a printer, a speaker, a CD burner, or another device that provides output from the computing system 1400.

[0205] One or more communication connections 1470 enable communication with another computing entity via a communication medium. The communication medium conveys information such as computer-executable instructions, audio or video input or output, or other data in a modulated data signal. A modulated data signal is a signal having one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, the communication medium may use electrical, optical, RF, or other carriers.

[0206] These innovations may be described in the general context of computer-executable instructions, such as those included in program modules, executing in a computing system on a target real or virtual processor. Generally, program modules or components include routines, programs, libraries, objects, classes, components, data structures, etc. that perform particular tasks or implement particular abstract data types. In various embodiments, the features of program modules may be combined or separated as desired among program modules. The computer-executable instructions for program modules may be executed in a local or distributed computing system.

[0207] The terms "system" and "device" are used interchangeably herein. Unless the context clearly indicates otherwise, neither term implies any limitation as to the type of computing system or computing device. Generally, a computing system or computing device may be local or distributed and may include any combination of dedicated hardware and / or general-purpose hardware with software implementing the functions described herein.

[0208] In various examples described herein, a module (e.g., a component or an engine) may be "encoded" to perform certain operations or provide certain functions, indicating that the computer-executable instructions for the module may be executed to perform these operations such that the operations are performed or otherwise the features are provided. Although the functions described with respect to software components, modules, or engines may be performed as discrete software units (e.g., programs, functions, class methods), it need not be implemented as discrete units. That is, the function may be incorporated into a larger or more general program, such as one or more lines of code in a larger or general program.

[0209] For presentation, the detailed description uses terms such as "determine" and "use" to describe computer operations in a computing system. These terms are high-level abstractions of operations performed by a computer and should not be confused with acts performed by a human. The actual computer operations corresponding to these terms vary depending on the implementation.

[0210] Example 17 – Cloud Computing Environment

[0211] Figure 15FIG. 1500 depicts an example cloud computing environment in which the described techniques may be implemented. The cloud computing environment 1500 includes cloud computing services 1510. The cloud computing services 1510 may include various types of cloud computing resources such as computer servers, data repositories, network resources, and the like. The cloud computing services 1510 may be centralized (e.g., provided by an enterprise or organization's data center) or distributed (e.g., provided by various computing resources located at different locations such as different data centers and / or located in different cities or countries).

[0212] The cloud computing services 1510 are utilized by various types of computing devices (e.g., client computing devices) such as computing devices 1520, 1522, and 1524. For example, the computing devices (e.g., 1520, 1522, and 1524) may be computers (e.g., desktop or laptop computers), mobile devices (e.g., tablets or smartphones), or other types of computing devices. For example, the computing devices (e.g., 1520, 1522, and 1524) may utilize the cloud computing services 1510 to perform computing operations (e.g., data processing, data storage, etc.).

[0213] Example 18 – Embodiments

[0214] Although, for the sake of presentation, some operations of the disclosed methods are described in a particular order, it should be understood that such description includes rearrangement unless the particular language set forth requires a particular order. For example, operations described in sequence may in some cases be rearranged or performed concurrently. In addition, for simplicity, the figures may not show the various ways in which the disclosed methods may be used in conjunction with other methods.

[0215] Any disclosed method may be implemented as computer-executable instructions or a computer program product stored on one or more computer-readable storage media (such as tangible, non-transitory computer-readable storage media) and executed on a computing device (e.g., any available computing device including a smart phone or other mobile device including computing hardware). A tangible computer-readable storage medium is any available tangible medium that can be accessed in a computing environment (e.g., one or more optical media disks such as DVDs or CDs, volatile memory components such as DRAM or SRAM, or non-volatile memory components such as flash memory or a hard disk drive). By way of example, and with reference Figure 14 FIG. 14, computer-readable storage media includes memories 1420 and 1425 and storage 1440. The term computer-readable storage medium does not include signals and carriers. In addition, the term computer-readable storage medium does not include a communication connection (e.g., 1470).

[0216] Any computer-executable instructions for implementing the disclosed technology and any data created and used during the implementation of the disclosed embodiments can be stored on one or more computer-readable storage media. The computer-executable instructions can be, for example, part of a dedicated software application or a software application accessed or downloaded via a web browser or other software application (such as a remote computing application). Such software can be executed, for example, on a single local computer (e.g., any suitable commercially available computer) or in a network environment using one or more networked computers (e.g., via the Internet, wide area network, local area network, client-server network (such as a cloud computing network) or other such network).

[0217] For clarity, only certain selected aspects of the software-based implementation are described. Other details known in the art are omitted. For example, it should be understood that the disclosed technology is not limited to any particular computer language or program. For example, the disclosed technology can be implemented by software written in C, C++, C#, Java, Perl, JavaScript, Python, Ruby, ABAP, SQL, XCode, GO, Adobe Flash or any other suitable programming language, or in some examples, by a markup language such as html or XML or a combination of a suitable programming language and a markup language. Similarly, the disclosed technology is not limited to any particular computer or hardware type. Certain details of suitable computers and hardware are well known and need not be elaborated in this disclosure.

[0218] In addition, any software-based embodiment (including, for example, computer-executable instructions for causing a computer to perform any of the disclosed methods) can be uploaded, downloaded, or remotely accessed via a suitable communication method. Such suitable communication methods include, for example, the Internet, World Wide Web, intranet, software application, cable (including fiber optic cable), magnetic communication, electromagnetic communication (including radio frequency, microwave, and infrared communication), electronic communication, or other such communication methods.

[0219] The disclosed methods, apparatuses, and systems should not be construed as being limited in any way. Instead, this disclosure is directed, both individually and in various combinations and sub-combinations with each other, to all novel and non-obvious features and aspects of the various disclosed embodiments. The disclosed methods, apparatuses, and systems are not limited to any particular aspect or feature or combination thereof, and the disclosed embodiments do not require the presence of any one or more particular advantages or the solving of problems.

[0220] The techniques in any example can be combined with the techniques described in any one or more other examples. Given the many possible embodiments in which the principles of the disclosed techniques can be applied, it should be recognized that the illustrated embodiments are examples of the disclosed techniques and should not be regarded as limiting the scope of the disclosed techniques. Instead, the scope of the disclosed techniques includes what is covered by the scope and spirit of the following claims.

Claims

1. A computing system, comprising: a memory; one or more processing units, coupled to the memory; and one or more computer-readable storage media storing instructions that, when loaded into the memory, cause the one or more processing units to perform operations for: receiving a training data set, the training data set including values of a first plurality of features; using the training data set to train a machine learning algorithm to provide a trained machine learning model; using the trained machine learning model to process an analysis data set to provide a result; forming a plurality of feature groups, at least one of the feature groups including a second plurality of features of the first plurality of features, forming the plurality of feature groups including grouping at least some of the first plurality of features according to a data model of a database used in a relational database system, the second plurality of features being a proper subset of the first plurality of features, wherein forming the plurality of feature groups further includes determining a contextual contribution of at least a portion of the first plurality of features to the result and detecting a causal loop within the feature groups; and responsive to determining the contextual contribution and detecting the causal loop, determining which features and feature groups to use when retraining the machine learning model.

2. The computing system according to claim 1, wherein, The contextual contribution is calculated as a SHAP value.

3. The computing system according to claim 1, wherein, The contextual contribution is calculated as a LIME value.

4. The computing system according to claim 1, wherein the operations further include: for at least a portion of the first plurality of features, determining an overall contribution of a corresponding feature; and for a third plurality of features selected from the first plurality of features, comparing the overall contribution of a given feature in the third plurality of features with the contextual contribution of the given feature.

5. The computing system according to claim 4, wherein the operations further include: comparing the overall contributions of features of a plurality of input instances of the data set to determine a consistency value of a corresponding feature in the third plurality of features.

6. The computing system according to claim 1, wherein the operations further include: aggregating the contributions of features associated with the plurality of feature groups to provide an aggregated contribution value of a feature group in the plurality of feature groups.

7. The computing system according to claim 6, wherein the operations further include: calculating an importance value of at least one feature group in the plurality of feature groups as an average of the contribution values of the features belonging to the at least one feature group.

8. The computing system according to claim 7, wherein the operations further include: performing a presentation for display on a user interface screen that displays at least a portion of the plurality of feature groups and the importance values of the corresponding feature groups of the at least a portion of the plurality of feature groups.

9. The computing system according to claim 6, wherein the operations further include: performing a presentation for display on a user interface screen that displays at least a portion of the plurality of feature groups and the features that are members of the corresponding feature groups.

10. The computing system according to claim 9, wherein the operations further include: Receive user input to add at least one of the first plurality of features to a feature group among the plurality of feature groups, or to remove at least one of the first plurality of features from a feature group among the first plurality of feature groups.

11. The computing system according to claim 1, wherein the operation further comprises: Adjusting the trained classifier at least in part based on one or more of the plurality of feature groups.

12. The computing system according to claim 1, wherein, Forming the plurality of feature groups further comprises: Analyzing a data model associated with a third plurality of features among the first plurality of features, the third plurality of features being a subset of the first plurality of features; Determining a plurality of data model elements from the data model; Defining at least one feature group at least in part based on the data model elements among the plurality of data model elements; Determining a fourth plurality of features selected from the first plurality of features as members of the data model elements; and Assigning at least a portion of the fourth plurality of features to the at least one feature group.

13. The computing system according to claim 1, wherein, Forming the plurality of feature groups further comprises: Analyzing a plurality of data access operations for obtaining at least a portion of the data in the data set; Determining one or more data sources accessed by the plurality of data access operations; Defining at least one feature group at least in part based on the data access operations among the plurality of data access operations; Determining a fourth plurality of features selected from the first plurality of features as members of the one or more data sources; and Assigning at least a portion of the fourth plurality of features to the at least one feature group.

14. The computing system according to claim 1, wherein, Forming the plurality of feature groups further comprises: Determining dependency information for a plurality of feature pairs among the first plurality of features; and Forming at least one feature group among the plurality of feature groups at least in part based on determining features among the first plurality of features that depend on the first feature using the dependency information of the first feature among the first plurality of features.

15. The computing system according to claim 14, wherein, The dependency information includes a signed chi-square test.

16. The computing system according to claim 14, wherein, Determining the dependency information includes determining feature pairs that satisfy a threshold dependency level.

17. One or more computer-readable storage media storing computer-executable instructions for causing a computing system to perform a process, the process comprising: Receiving a training data set, the training data set including values of a first plurality of features; Training a machine learning algorithm using the training data set to provide a trained machine learning algorithm; Processing an analysis data set using the trained machine learning algorithm to provide a result; Determining context contribution values for a second plurality of features among the first plurality of features; Determining dependencies between features among the second plurality of features; Forming a plurality of feature groups, at least in part based on the determined dependencies, the forming including grouping at least some of the first plurality of features according to a data model of a database used in a relational database system, at least one of the plurality of feature groups including a third plurality of features of the first plurality of features, the third plurality of features being a proper subset of the first plurality of features, wherein forming the plurality of feature groups further includes determining a contextual contribution of at least a portion of the first plurality of features to the result and detecting causal loops within the feature groups; and Determining which features and feature groups to use when retraining a machine learning model in response to determining the contextual contribution and detecting the causal loops.

18. A method implemented in a computing system including a memory and one or more processors, the method including: Determining a first plurality of features used in a machine learning algorithm; Using the machine learning algorithm to determine a result of analyzing a data set; Forming a plurality of feature groups, the forming including grouping at least some of the first plurality of features according to a data model of a database used in a relational database system, at least one of the plurality of feature groups including a second plurality of features of the first plurality of features, the second plurality of features being a proper subset of the first plurality of features, wherein forming the plurality of feature groups further includes determining a contextual contribution of at least a portion of the first plurality of features to the result and detecting causal loops within the feature groups; And Determining which features and feature groups to use when retraining a machine learning model in response to determining the contextual contribution and detecting the causal loops.