Automatic feature subset selection using feature ranking and scalable automatic search

By ranking and non-sequential search of features, combining multiple feature scoring functions and parallel processors, the problems of low feature selection efficiency and overfitting in the prior art are solved, and the effect of accelerating training and improving model performance is achieved.

CN112840361BActive Publication Date: 2025-05-09ORACLE INT CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980067625.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-20
Filing Date
2019-10-01
Publication Date
2025-05-09
Estimated Expiration
2039-10-26

AI Technical Summary

Technical Problem

The prior art has problems of inefficiency, overfitting trends and difficulty in effectively managing feature subsets when dealing with feature selection in machine learning models.

Method used

By ranking features individually and combining feature subsets based on ranking order, a non-sequential search method is adopted to reduce the number of evaluations, and the feature selection process is accelerated by multiple feature scoring functions and parallel processors.

Benefits of technology

The acceleration of the training and inference process of machine learning models is implemented, prevents overfitting, and improves insight into complex data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112840361B_ABST
    Figure CN112840361B_ABST
Patent Text Reader

Abstract

The present invention relates to dimensionality reduction for machine learning (ML) models. The technology herein is to individually rank features and combine features based on their rankings to achieve an optimal combination of features, which can accelerate training and / or reasoning, prevent overfitting, and / or provide insights into a somewhat mysterious data set. In an embodiment, for each feature of a training data set, a computer calculates a relevance score based on a relevance scoring function and statistical information of the values ​​of the feature that appear in the training data set. A ranking based on the relevance score of the feature is calculated for each feature. Based on the ranking of the features, a sequence of different subsets of the features is generated. For each different subset in the sequence of different feature subsets, a fitness score is generated based on training a machine learning (ML) model configured for the different subset.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to dimensionality reduction for machine learning (ML) models. The techniques herein are to individually rank features and combine features based on their rankings to achieve an optimal combination of features that can accelerate training and / or inference, prevent overfitting, and / or provide insights into a somewhat mysterious data set. Technical Field

[0003] The present invention relates to dimensionality reduction for machine learning (ML) models. The technology herein is to rank features individually and combine features based on their rankings to achieve an optimal combination of features that can accelerate training and / or inference, prevent overfitting, and / or provide insights into certain mysterious data sets. Background Art

[0004] The use of machine learning (ML) such as deep learning (DL) is rapidly spreading across industrial and commercial sectors and is becoming a ubiquitous tool in some enterprises. ML model training can take up a lot of time and / or computer space. ML models can be configured to process all features present in a dataset. However, processing each feature can consume a lot of computer resources. When ML models do not process some features, efficiency can be improved. However, not all features are logically and / or operationally equal. In fact, performance improvements may require judicious selection of features to process and features to ignore.

[0005] Feature selection can have a significant impact on model performance (e.g., accuracy, f1 score, etc.). Selecting the best subset of features for a model is difficult and, especially for complex datasets, can be very time-consuming. Because there are 2^n possible subset combinations (n ​​= number of features), it is not feasible to consider all possible subsets. For example, a dataset may have 100,000 features.

[0006] Below are various previous approaches, all of which have limited capabilities in managing the number / combinations of possible feature subsets for a given dataset. Filtering methods that are based on a collective analysis of feature values ​​and may ignore the ML model itself, such as when the filtering method or threshold is not appropriate for the given dataset or ML model in question, may tend to remove important features. Using embedded methods that leave feature selection to the learning ability of the ML model, thus handling many features, still leads to a tendency for long model training times or overfitting, thus reducing the two main motivations for feature selection. A serious challenge for wrapper approaches that treat ML models as black boxes is to define the feature subsets to be evaluated, and such black boxes must be empirically explored for their sensitivity to features. The main problem with most of the above methods is their sequential nature. Since feature subsets are evaluated sequentially, including model training and testing, horizontally scaling parallelism is impractical. In addition, the ideal number of features (i.e., the target subset size) is usually unknown in advance. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] In the attached picture:

[0008] Figure 1 is a block diagram depicting an example computer that individually ranks features and combines features based on their rankings to achieve an optimal combination of features that can accelerate training and / or inference and / or prevent overfitting of a machine learning (ML) model;

[0009] Figure 2 is a flow chart of an example computer process for individually ranking features and combining features based on their rankings to achieve an optimal combination of features that can accelerate training and / or inference and / or prevent overfitting of a machine learning (ML) model;

[0010] Figure 3 is a block diagram depicting an example computer utilizing multiple feature scoring functions to improve accuracy and utilizing multiple computational processors (e.g., CPU cores) for acceleration;

[0011] Figure 4 is a block diagram illustrating a computer system upon which embodiments of the present invention may be implemented;

[0012] Figure 5 is a block diagram illustrating a basic software system that may be used to control the operation of a computing system. DETAILED DESCRIPTION

[0013] In the following description, for the purpose of explanation, many specific details are set forth in order to provide a thorough understanding of the present invention. However, it will be apparent that the present invention can be practiced without these specific details. In other cases, well-known structures and devices are shown in block diagram form to avoid unnecessarily obscuring the present invention.

[0014] General Overview

[0015] The approach in this paper combines a feature ranking criterion with two key improvements: a) using feature ranking to define the order of features, and b) applying a non-sequential search to reduce the number of evaluations, including optimizations for parallel execution. The techniques in this paper can be directly applied to any (e.g., Oracle) machine learning product that supports feature selection. The scalable automatic search significantly improves the overall runtime. Especially in cloud applications, less runtime means less resource usage and cost. The approach can apply the following novel techniques to automatically select features.

[0016] With feature ranking, features are ranked using criteria such as variance, mutual information, or feature importance information obtained from a trained model. Then, instead of applying a threshold to the values, feature subsets are created based on the ranking order. For example, three features (1,2,3) might be ranked 2,3,1 in descending order of importance using a filtering method. Feature subsets (2), (2,3), (2,3,1) can be created based on the calculated rankings. This reduces the number of subsets considered from 2^n (i.e., exponential) to n (i.e., linear).

[0017] With ensemble ranking, sometimes a single ranking algorithm cannot fully evaluate the feature importance of all possible data sets. Therefore, multiple ranking algorithms are used to complement each other. Each ranking algorithm produces a feature ranking that defines a feature subset for evaluation according to the ranking order. In addition, there may be (one or more) ensemble rankings that combine multiple feature rankings into a new ranking. For example, the ensemble ranking can be based on the average of the feature scores of all other rankings.

[0018] In an embodiment, for each feature of the training data set, a computer calculates a relevance score based on a relevance scoring function and statistical information of the values ​​of the feature that appear in the training data set. A ranking based on the relevance score of the feature is calculated for each feature. Based on the ranking of the feature, a sequence of different subsets of the features is generated. For each different subset of the sequence of different subsets of features, a fitness score is generated based on training a machine learning (ML) model configured for the different subset.

[0019] With a scalable search, instead of evaluating all n feature subsets or applying a sequential search, an exponential subset selection function is applied. This function will select many small subsets of features and a few large subsets, all based on the feature ranking order. Here are a few reasons for taking this approach.

[0020] Small subsets can be evaluated faster than large subsets due to the reduced size of the dataset and the amount of processing required.

[0021] The relative size differences between small subsets are more noticeable than those of large subsets. For example, if you have a dataset with 1000 features, you may want to test a subset of 5 features even if you know that a subset of 10 features is good because it reduces the size by half. If you find that a subset of 900 features is good, then the potential reduction from a subset with 895 features is not enough to justify its evaluation.

[0022] By avoiding random steps, subset evaluation is based on ranking of normalized feature scores scaled between 0 and 1. Features are sorted by their scores. This will produce the best feature with a score of 1, which is the highest ranked, while the worst feature has a score of 0 and is ranked last. Sometimes, multiple features may have the same score, making the ranking order between these features naturally arbitrary. If the last element of the previous subset has the same score value as the last element of the current subset, then the embodiment counteracts this effect by not adding a subset for evaluation. In this case, the difference between the last subset and the current subset is based only on the random order of the features, not on a value-based ranking. Therefore, there is no need to further evaluate the current subset. This heuristic also applies to scores of 0 at the lower end of the ranked features. Despite this approach, the complete data set (a subset with all features) can be additionally selected for evaluation in the following cases:

[0023] The ranking does not reflect the true feature importance, or

[0024] All features are important for prediction.

[0025] With scalable search / evaluation, feature subsets are selected by examining given feature rankings without any prior stepwise evaluation. Therefore, composing subsets is fast (i.e., not computationally intensive). The time-consuming evaluation can then be parallelized by training and testing the target model using the selected subsets. Since the subsets are independent, parallelism can scale to the number of subsets evaluated. Even without high parallelism, this approach can facilitate load balancing through the inherent knowledge that smaller subsets can be evaluated faster than larger ones due to the reduction in dataset size.

[0026] With work stealing, given a cluster of computing nodes (e.g., cores in a CPU or hosts in a computer cluster), evaluation subsets should be scheduled on separate nodes so that the number of features on all subsets on each node is similar. This will reduce the overall runtime since runtime is generally proportional to the number of features. Additionally, embodiments may force larger subsets to be scheduled on different nodes so that subsets on each node are evaluated in descending order of subset size. This leaves smaller subsets to be evaluated in later stages. If for some reason the evaluation on one node takes longer, this descending order will facilitate efficient work stealing of smaller subsets by nodes that have completed their workload.

[0027] 1.0 Sample Computer

[0028] Figure 1 1 is a block diagram depicting an example computer 100 in an embodiment. The computer 100 ranks features individually and combines features based on their rankings to achieve an optimal combination of features that can accelerate training and / or inference and / or prevent overfitting of a machine learning (ML) model. The computer 100 can be one or more of a rack server such as a blade server, a personal computer, a mainframe, a virtual computer, or other computing device.

[0029] The computer 100 may store the ML model 110 in its memory. Depending on the embodiment, the ML model 110 is designed for clustering, classification, regression, anomaly detection, prediction, or dimensionality reduction (i.e., simplification). Examples of machine learning algorithms include decision trees, support vector machines (SVMs), Bayesian networks, random algorithms such as genetic algorithms (GAs), and connectionist topologies such as artificial neural networks (ANNs). Implementations of machine learning may rely on matrices, symbolic models, and hierarchical and / or associative data structures. Parameterized (i.e., configurable) implementations of the best hybrid machine learning algorithms can be found in open source libraries such as scikit-learn (sklearn), Google's TensorFlow for Python and C++, or Georgia Institute of Technology's MLPack for C++. Shogun is an open source C++ML library with adapters for several programming languages, including C#, Ruby, Lua, Java, MatLab, R, and Python.

[0030] The life cycle of the ML model 110 has two phases. The first phase is preparation and requires training, such as in a lab. The second phase requires reasoning in a production environment, such as with real-time and / or streaming data.

[0031] During inference, the ML model 110 is applied to (e.g., unfamiliar) samples, which may be injected as input into the ML model 110. This causes the ML model 110 to process the samples according to the internal mechanisms of the ML model 110, which are specifically configured according to reinforcement learning during previous training of the ML model 110. For example, if the ML model 110 is a classifier, the ML model 110 may select one of a plurality of mutually exclusive labels (i.e., classifications), such as hot and cold, for the samples.

[0032] Whether during training or inference, the ML model 110 can process samples that may have different values ​​for the same feature (such as AC). For example, feature A may be color, and feature B may be temperature. Each sample may have the same or different color and the same or different temperature.

[0033] Depending on how rich (i.e., multidimensional) the sample data is, there may be many more features (e.g., hundreds or thousands) than the AC shown. The ML model 110 may be one of many ML model types, and those model types generally have sufficient logical flexibility to accept and process very large numbers of features of the sample data. However, the consumption of resources such as time and / or space may be positively correlated with the number of complexities of the available features.

[0034] There may be some practical limits on how many features the ML model 110 can process within a given resource budget. Therefore, in practice, some features should be ignored. However, deciding how many features to ignore and which features to ignore may affect how accurate the ML model 110 is.

[0035] ML model accuracy is an objective measure that can explain the frequency and / or severity of misclassifications. For example, a binary classifier may observe separate frequencies of false positives and false negatives, which may have different semantic severity. For example, mistaking a red traffic light for green may be worse than mistaking a green light for red.

[0036] Some features may be noisy, or more or less relevant to inference or inference accuracy. Other features may be more or less relevant to the performance of the ML model 110. As described below, the computer 100 can detect how relevant features such as AC are.

[0037] Training the ML model 110 may require processing many samples, which are usually required for reinforcement learning. These samples together can form a training corpus, such as 120. As shown, the training corpus 120 has at least sample 0-1.

[0038] Sample 0-1 has many features, including AC as shown in training corpus 120. Computer 100 has one or more scoring functions 130, each of which calculates a relevance score for each of features AC, estimating how much impact the feature may have on the accuracy of ML model 110.

[0039] In an embodiment, the relevance score of the feature is merely an estimate and is not based on the actual performance of the ML model 110. Therefore, the scoring function 130 may score the features AC without actually operating the ML model 110. In such an embodiment, the scoring function 130 may operate regardless of whether the ML model 110 is trained, and may operate without the ML model 110.

[0040] Thus, depending on the embodiment, the relevance scores for features AC may be calculated by a scoring function 130 that has no other inputs besides the training corpus 120. For example, the scoring function 130 may calculate the relevance score for feature A based solely on the values ​​in column A of the training corpus 120. The scoring function 130 may calculate the relevance score based on statistics of the values ​​of the features, such as variance, such as statistics 140 of the values ​​of feature C.

[0041] The training in this paper is supervised. Therefore, each training sample also has a classification label, as shown in the training corpus 120, or other prediction targets that the ML model 110 can be trained to infer. Features can be more or less related to labels (i.e., predictive labels). For example, most blue samples can be labeled as positive.

[0042] Thus, blue may be somewhat predictive of positivity, and more generally, color may be somewhat predictive of the classification label. Scoring function 130 may calculate a relevance score based on the relevance of the feature to the label. In embodiments, as discussed later herein, relevance may be calculated based on statistics such as mutual information or an F-score. Features that are somewhat relevant to the classification may have a higher relevance score.

[0043] Although the ML model 110 may not be trained, the scoring function 130 may be based on the configuration of a different ML model (not shown) that has been trained. For example, some types of ML models naturally rank features according to relevance / importance. For example, in the case of a trained random forest ML model, each feature has a learned feature importance.

[0044] Likewise, for a trained logistic regression ML model, each feature has a learned feature coefficient. In an embodiment, the scoring function 130 may be sensitive to the learned importance or coefficient of the feature. For example, the scoring function 130 may use one learned coefficient when scoring feature A, and may use a different learned coefficient when scoring feature B, but both coefficients come from the same trained regressor (not shown).

[0045] Feature importances learned from a trained model not shown may improve efficiency, such as when the model not shown may be trained much faster than ML model 110. For example, the model not shown may have used a much smaller training corpus than 120 or a small subset of 120.

[0046] As described above, there are multiple (e.g., many) ways to calculate the relevance score of a feature. In an embodiment, the scoring function 130 integrates multiple statistics or inputs for the same feature. In an embodiment, each statistic can have its own scoring function, and the scores from the multiple functions are integrated to derive an aggregate (e.g., mean) relevance score for the feature.

[0047] Once all features AC have relevance scores, the features can be relatively ranked according to the relevance scores. For example, features AC can be ranked by descending relevance scores as shown, such that feature C is the most important (i.e., ranked 1), and feature B is less important (i.e., ranked 3). As discussed later in this document, features AC can be redundantly scored by different scoring functions such as 130 to produce different rankings.

[0048] When only one scoring function is available, there is no need to normalize the relevance scores. However, when multiple scoring functions are used, the scores they compute may need to be normalized (e.g., to unity).

[0049] In this example, there is only one rank, so that the features can be ordered by descending relevance (e.g., C, A, B). Ranking improves the efficiency and effectiveness of feature selection. Without ranking, feature selection may require the exploration of natural combinations, which may expand exponentially based on the count of available features. Therefore, the optimal feature selection may be difficult to compute without heuristics (natural, ranking).

[0050] The ranking indicates feature relevance. Less relevant features can be more or less ignored without affecting the accuracy of the ML model 110. Therefore, each row of the subset 150 (ie, feature subset) may lack many or most of the available features AC.

[0051] The generation of feature subsets within subset 150 may occur as follows. Each of rows 1-3 of feature subset 150 is a generated feature subset. The more relevant a feature is, the more subsets / rows that feature occurs in.

[0052] The first subset 1 has only the most relevant feature C. Each additional row adds additional features in descending order of relevance (i.e., ascending order of rank). Therefore, the most relevant feature C appears in all rows. Similarly, the least relevant feature B appears in only one row.

[0053] Thus, each row can be generated incrementally based on the previous row. The result of generating feature subsets in this manner is that feature subsets 150 have as many rows as features. Thus, feature selection scales linearly with feature count units. This is a huge improvement over exhaustive combinatorial generation of subsets, which has exponential complexity.

[0054] Adding additional features to each row of feature subset 150 makes each subsequent row imply more sample data (i.e., more feature columns AC of corpus 120). More sample data often means better training and more learning accuracy. Therefore, it may be tempting to select the last row 3 (i.e., the subset with the most features) as the best feature subset.

[0055] However, there may be a point of diminishing returns such that adding more features will eventually stop improving accuracy, or may even reduce accuracy. Therefore, it is actually necessary to empirically test subsets of features of the proposed subset 150 to detect which subset is indeed the best among the subsets 150. Therefore, as shown, each row of the subset 150 has a measured fitness score indicating that the training accuracy of the ML model 110 is actually achieved when only the features of the subset of that row are used for training.

[0056] Therefore, even though row 3 has more features, row 2 still has the best feature set because it actually achieves the highest accuracy (i.e., fitness score). Therefore, choosing the best (i.e., most accurate) feature subset depends on the empirically measured training accuracy. This means that each feature subset in subsets 150 should be used to actually configure and train ML model 110 to empirically measure training accuracy.

[0057] Therefore, selecting the best feature subset requires multiple trainings of the ML model 110, where each training uses a different row of the subset 150. Each individual training may be computationally time and / or space resource intensive. Therefore, as described above, subset training that scales linearly with the available feature count may be the basis for feasibility / tractability.

[0058] Better than linear scaling is possible if many rows of subset 150 are (e.g., selectively) skipped (i.e., not generated and not used for training). For example, computer 100 may selectively sample (i.e., generate and train) fewer rows than those shown in subset 150. As long as the last row (i.e., the row with all features) is sampled, no features are completely absent from subset 150, no matter how many other rows are skipped.

[0059] In an embodiment, individual rows of the subset 150 are sampled for training according to an exponential sequence such that larger and larger groups of consecutive rows of the subset 150 are skipped. Thus, most small subsets are sampled, and most large subsets are not sampled. Because the training time depends on the feature count, small subsets (e.g., no matter how many there are) are trained quickly.

[0060] Large subsets are less important for two reasons. First, subsets are constructed incrementally by concatenating features of decreasing relevance, so that a large subset can have many or most features with less relevance. Second, incremental feature concatenation means increasing feature overlap between consecutive subsets, so that two consecutive large subsets are almost identical, while the features of the two smallest subsets differ by half.

[0061] The exponential sequence of subset sizes (i.e., feature counts) may start at one and increase by a constant floating point factor greater than one as each subsequent subset is generated. Instead of adding one additional feature to each next subset, an embodiment may add an exponentially increasing number of additional features, thereby skipping an increasingly larger number of subsets between each previous subset generated and the next subset. For example, if the constant factor is 1.2, and the first size is 1, then the next size is 1.2x1, which is rounded to 2.

[0062] For another size, multiply 2 by 1.2, and so on. For a training dataset with 10,000 features and a factor of 1.2, this produces 45 subsets of the following sizes: 1, 2, 3, 4, 5, 6, 8, 10, 12, 15, 18, 22, 27, 33, 40, 48, 58, 70, 84, 101, 122, 147, 177, 213, 256, 308, 370, 444, 533, 640, 768, 922, 1107, 1329, 1595, 1914, 2297, 2757, 3309, 3971, 4766, 5720, 6864, 8237, and 9885. Thus, the counts of the sampled subsets are logarithmically proportional to the count of available features.

[0063] In an embodiment, the feature relevance scores are unit normalized so that the highest score is scaled to one. Completely irrelevant features have a score of zero. In an embodiment, features with a relevance score of zero are excluded from the subset 150. In an embodiment, even the excluded features are included in the last (i.e., largest) subset that includes all features.

[0064] Multiple features may have the same relevance score and are therefore ranked consecutively. In an embodiment, when some features are equal in relevance score, those equal features are all added to the next generated subset. For example, if four features are equal, then the next subset is generated, rather than incrementally generating four subsets in sequence to add these features.

[0065] Regardless of how many rows are generated and the fitness score for accuracy, embodiments can ultimately select the most accurate feature subset. After selecting the best row of subset 150, computer 100 or another computer can immediately or eventually configure ML model 110 to expect that feature subset. For example, ML model 110 can be configured to expect features A and C, but not feature B. In embodiments, unselected features such as B can be removed from training corpus 120. In any case, ML model 110 can then be fully trained with training corpus 120.

[0066] 2.0 Example feature selection process

[0067] Figure 2 is a flow chart depicting a computer 100 that ranks features individually and combines features based on their rankings to achieve an optimal combination of features that can accelerate training and / or inference and / or prevent overfitting of a machine learning (ML) model. In one embodiment. Figure 1 discuss Figure 2 .

[0068] Steps 202 and 204 are prepared and repeated for each feature (such as AC) of a training data set (such as 120). Step 202 calculates a relevance score for each feature, such as shown in feature C, as calculated by scoring function 130 and based on value statistics 140 as discussed above. In an embodiment, scores for multiple features AC are calculated simultaneously. In an embodiment, step 202 does not refer to ML model 110 or the configuration of ML model 110. In an embodiment, feature scores can be reused with ML models other than 110.

[0069] Step 204 ranks (eg, sorts) features AC by relevance score. Thus, step 204 distinguishes more or less important features from more or less unimportant features.

[0070] Steps 206 and 208 may occur simultaneously. For example, step 206 may repeatedly call step 208 as follows. Step 206 generates a sequence of different subsets of features. The first subset of the sequence consists of the highest ranked (i.e., most important) features (such as C). Each subsequent subset may be an incremental extension of the previously generated feature subset based on adding the most important features (such as A) that have not yet been added.

[0071] When step 206 generates the next feature subset in the sequence, step 208 may be repeated asynchronously and / or more or less immediately for that feature subset. Step 208 configures the ML model 110 according to the feature subset generated by step 206. Step 208 also trains the ML model 110 with the training corpus 120 and measures the fitness (e.g., accuracy) achieved by the training.

[0072] Due to computational intensity, the training during step 208 is Figure 2 The slowest part of step 208. Embodiments may distribute each occurrence of step 208 to a cluster of computers to speed it up through coarse-grained horizontal scaling. Pending occurrences of step 208 may be delayed (e.g., queued as a backlog) and eventually dispatched individually, such as to a cluster or thread pool. An implementation of the ML model 110 may utilize multiple cores, threads, or computers for a single training (e.g., occurrence of step 208). An implementation of the ML model 110 may exploit data parallelism, such as by employing single instruction multiple data (SIMD) instructions or other vector hardware.

[0073] Each occurrence of step 208 generates a fitness score for a different feature subset. Thus, after step 208 occurs for all generated feature subsets, each subset has a fitness score, and the highest fitness feature subset(s) can be identified. For example, Figure 2 The processing can select the best of the generated feature subsets.

[0074] 3.0 Multiple Ranks and Processors

[0075] Figure 3 is a block diagram depicting an example computer 300 in an embodiment. Computer 300 utilizes multiple feature scoring functions to improve accuracy and utilizes multiple computing processors (eg, CPU cores) for acceleration. Computer 300 may be an implementation of computer 100.

[0076] As discussed above, computer 300 may have multiple feature scoring functions (not shown), each of which implements a different ranking of features 1-10, such as AB, shown as feature subsets 311. For example, in ranking A, feature 2 is most important, and feature 7 is least important. While in ranking B, feature 8 is most important, and feature 6 is least important.

[0077] The feature subsets of subset 311 are initially segmented into rankings AB, according to which the scoring function is involved. At least those feature subsets can be further aggregated into combined subsets 312. The subsets can be merged in their entirety. In an embodiment, each subset ranked AB is included in subset 311 or not included in subset 311 according to criteria that may include relevance scores and / or rankings.

[0078] Computer 300 performs multiprocessing in an implementation-dependent manner. For example, computer 300 may be a standalone computer with a single CPU that includes multiple processing cores AD for symmetric multiprocessing (SMP). Computer 300 may have multiple CPUs and / or general-purpose coprocessors.

[0079] Computer 300 may be a plurality of interconnected computers (e.g., a cluster of) such that each of the core ADs is a separate computer. Computer 300 may be a combination of different such topologies. Computer 300 may be a computer cloud with elastic horizontal scaling.

[0080] After subset selection to obtain feature subsets 312, each subset is used to configure and train an ML model (not shown) to measure the training accuracy of each subset. For example, the best feature subset can be determined. As a multiprocessor, the computer 300 can use task (i.e., coarse-grained) parallelism to asynchronously train the ML model, where each of the core ADs trains a different feature subset.

[0081] In an embodiment, all subsets 312 reside in a central repository from which cores AD can independently and asynchronously fetch (i.e., extract) the next subset to be trained. When a core completes training on a subset, the core can fetch another pending subset from the central repository for training until the central repository is empty and all subsets 312 have been scored for fitness for accuracy. Depending on the embodiment, the central repository can have an unordered subset stack or an ordered queue.

[0082] In an embodiment, the queue is sorted by descending counts of features in each subset. For example, the combined subset 312 can be a queue where the core AD extracts the next subset from the head of the queue (i.e., the bottom shown) rather than the tail of the queue (i.e., the top shown). In an embodiment, ties are resolved by decreasing the relevance score of the least relevant feature of each subset.

[0083] In an embodiment, equal counts of subsets of all subsets 312 are initially assigned to cores AD. In an embodiment, subsets are initially assigned to cores AD in a round-robin fashion, in order of decreasing feature count (i.e., large subsets before small subsets). Thus, all cores AD have the same number of subsets, and each core has some large subsets and some small subsets.

[0084] In the illustrated embodiment, cores AD may have different numbers of subsets, but the same number of features. For example, as shown, each core has a total of 11 features. In a work stealing embodiment, an idle core that completes early may steal a subset to be processed from a busy core. In an embodiment, if the stolen core has no additional subsets to share, then the core may reject the attempted stealing. In an embodiment, the stolen subset is the largest (i.e., most feature-rich) additional subset whose processing has not yet begun on the stolen core. In an embodiment, when a core AD has no more subsets to process (e.g., no more subsets to steal), each of the cores AD reaches the same global synchronization barrier.

[0085] 4.0 Example Implementation

[0086] In an embodiment, the ranking may be based on some or all of the following named mathematical artifacts provided by the python sklearn library:

[0087] Mutual information (mutual_info_classif, mutual_info_regression)

[0088] Analysis of variance (ANOVA) F score (f_classif, f_regression)

[0089] feature_importances of the trained RandomForestClassifier or RandomForestRegressor

[0090] feature_importances of the trained AdaBoostClassifier or AdaBoostRegressor

[0091] For example, ML models can be used for classification to produce the probability that an input conforms to several mutually exclusive discrete classes (e.g., circle, rectangle), or for regression to produce a continuous output (e.g., home value based on comparable homes). The correlation between features and the predicted target (i.e., label) can be measured based on the ratio of their variances as an ANOVA F-score.

[0092] In addition to the target ML model, which can be opaque (i.e., black box), the transparent internals of another ML model may also provide feature importance. For example, the following ML models have feature transparency. The Random Forest model is an ensemble of different decision trees. Adaptive Boosting (AdaBoost) is another ensemble model in which the constituent models (e.g., decision trees) are each specialized differently for corresponding extreme examples that are particularly difficult for a single generalized model to infer accurately.

[0093] Hardware Overview

[0094] According to one embodiment, the technology described herein is implemented by one or more special-purpose computing devices. The special-purpose computing device can be hard-wired to perform these technologies, or can include digital electronic devices (such as one or more application-specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs) that are permanently programmed to perform these technologies), or can include one or more general-purpose hardware processors that are programmed to perform these technologies according to program instructions in firmware, memory, other storage devices, or combinations. Such special-purpose computing devices can also combine customized hard-wired logic, ASICs, or FPGAs with customized programming to implement these technologies. The special-purpose computing device can be a desktop computer system, a portable computer system, a handheld device, a networking device, or any other device that combines hard-wiring and / or program logic to implement these technologies.

[0095] For example, Figure 4 4 is a block diagram illustrating a computer system 400 upon which embodiments of the present invention may be implemented. Computer system 400 includes a bus 402 or other communication mechanism for communicating information, and a hardware processor 404 coupled with bus 402 for processing information. Hardware processor 404 may be, for example, a general purpose microprocessor.

[0096] The computer system 400 also includes a main memory 406, such as a random access memory (RAM) or other dynamic storage device, coupled to the bus 402 for storing information and instructions to be executed by the processor 404. The main memory 406 may also be used to store temporary variables or other intermediate information during the execution of instructions executed by the processor 404. When stored in a non-transitory storage medium accessible to the processor 404, these instructions cause the computer system 400 to become a special-purpose machine customized to perform the operations specified in the instructions.

[0097] Computer system 400 also includes a read only memory (ROM) 408 or other static storage device coupled to bus 402 for storing static information and instructions for processor 404. A storage device 410, such as a magnetic disk, optical disk, or solid state drive, is provided and coupled to bus 402 for storing information and instructions.

[0098] The computer system 400 may be coupled to a display 412, such as a cathode ray tube (CRT), via the bus 402 for displaying information to a computer user. An input device 414, including alphanumeric and other keys, is coupled to the bus 402 for communicating information and command selections to the processor 404. Another type of user input device is a cursor control 416, such as a mouse, trackball, or cursor direction keys, for communicating direction information and command selections to the processor 404 and for controlling cursor movement on the display 412. Such input devices typically have two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), which allows the device to specify a position in a plane.

[0099] Computer system 400 may implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic that is combined with a computer system to make or program computer system 400 into a special purpose machine. According to one embodiment, computer system 400 performs the described techniques in response to processor 404 executing one or more sequences of one or more instructions contained in main memory 406. These instructions may be read into main memory 406 from another storage medium, such as storage device 410. Execution of the sequences of instructions contained in main memory 406 causes processor 404 to perform the process steps described herein. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions.

[0100] The term "storage medium" as used herein refers to any non-transient medium that stores data and / or instructions that cause a machine to operate in a particular manner. Such storage media may include non-volatile media and / or volatile media. Non-volatile media include, for example, optical disks, magnetic disks, or solid state drives, such as storage device 410. Volatile media include dynamic memory, such as main memory 406. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a pattern of holes, RAM, PROMs and EPROMs, FLASH-EPROMs, NVRAMs, any other memory chips, or cassette tapes.

[0101] Storage media are distinct from but can be used in conjunction with transmission media. Transmission media participate in the transfer of information between storage media. For example, transmission media include coaxial cables, copper wire, and optical fiber, including the wires that comprise bus 402. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.

[0102] Various forms of media may be involved in transmitting one or more sequences of one or more instructions to the processor 404 for execution. For example, the instructions may initially be carried on a disk or solid-state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to the computer system 400 may receive the data on the telephone line and use an infrared transmitter to convert the data into an infrared signal. An infrared detector may receive the data carried in the infrared signal and appropriate circuitry may place the data on the bus 402. The bus 402 transmits the data to the main memory 406 from which the processor 404 retrieves and executes the instructions. The instructions received by the main memory 2406 may optionally be stored on the storage device 410 before or after execution by the processor 404.

[0103] Computer system 400 also includes a communication interface 418 coupled to bus 402. Communication interface 418 provides bidirectional data communication coupled to network link 420, wherein network link 420 is connected to local network 422. For example, communication interface 418 can be an integrated services digital network (ISDN) card, a cable modem, a satellite modem, or a modem that provides a data communication connection with a corresponding type of telephone line. As another example, communication interface 418 can be a local area network (LAN) card to provide a data communication connection with a compatible LAN. A wireless link can also be implemented. In any such implementation, communication interface 418 sends and receives electrical signals, electromagnetic signals, or optical signals that carry digital data streams representing various types of information.

[0104] The network link 420 typically provides data communication to other data devices through one or more networks. For example, the network link 420 can provide a connection to a host computer 424 or to data equipment operated by an Internet Service Provider (ISP) 426 through a local network 422. The ISP 426 in turn provides data communication services through a global packet data communication network, now commonly referred to as the "Internet" 428. Both the local network 422 and the Internet 428 use electrical, electromagnetic, or optical signals that carry digital data streams. The signals through the various networks and the signals on the network link 420 and through the communication interface 418 (which carry the digital data to and from the computer system 400) are example forms of transmission media.

[0105] Computer system 400 can send messages and receive data, including program code, through the network(s), network link 420, and communication interface 418. In the Internet example, server 430 can send the requested code for an application program through Internet 428, ISP 426, local network 422, and communication interface 418.

[0106] The received code may be executed by processor 404 as it is received, and / or stored in storage device 410 or other non-volatile storage for later execution.

[0107] Software Overview

[0108] Figure 5 is a block diagram of a basic software system 500 that may be used to control the operation of computing system 400. Software system 500 and its components, including their connections, relationships, and functions, are merely exemplary and are not meant to limit implementation of the example embodiment(s). Other software systems suitable for implementing the example embodiment(s) may have different components, including components with different connections, relationships, and functions.

[0109] Software system 500 is provided for directing the operation of computing system 400. Software system 500, which may be stored on system memory (RAM) 406 and fixed storage (eg, hard disk or flash memory) 410, includes a kernel or operating system (OS) 510.

[0110] OS 510 manages low-level aspects of computer operation, including managing the execution of processes, memory allocation, file input and output (I / O), and device I / O. One or more applications, represented as 502A, 502B, 502C ... 502N, may be "loaded" (e.g., transferred from fixed storage 410 into memory 406) for execution by system 500. Applications or other software intended for use on computer system 400 may also be stored as a downloadable set of computer-executable instructions, for example, for downloading and installation from an Internet location (e.g., a Web server, app store, or other online service).

[0111] The software system 500 includes a graphical user interface (GUI) 515 for receiving user commands and data in a graphical manner (e.g., "clicks" or "touch gestures"). In turn, these inputs can be operated by the system 500 according to instructions from the operating system 510 and / or (one or more) applications 502. The GUI 515 is also used to display the results of operations from the OS 510 and (one or more) applications 502, and the user can provide additional input or terminate the session (e.g., log out).

[0112] The OS 510 may execute directly on the bare hardware 520 (e.g., the processor(s) 404) of the computer system 400. Alternatively, a hypervisor or virtual machine monitor (VMM) 530 may be inserted between the bare hardware 520 and the OS 510. In this configuration, the VMM 530 acts as a software "buffer" or virtualization layer between the OS 510 and the bare hardware 520 of the computer system 400.

[0113] VMM 530 instantiates and runs one or more virtual machine instances ("guests"). Each guest includes a "guest" operating system (such as OS 510), and one or more applications (such as (one or more) applications 502) designed to execute on the guest operating system. VMM 530 presents a virtual operating platform to the guest operating system and manages the execution of the guest operating system.

[0114] In some instances, VMM 530 can allow a guest operating system (OS) to run as if it were running directly on bare hardware 520 of computer system 500. In these instances, the same version of the guest operating system that is configured to execute directly on bare hardware 520 can also execute on VMM 530 without modification or reconfiguration. In other words, VMM 530 can provide full hardware and CPU virtualization to guest operating systems in some cases.

[0115] In other instances, the guest operating system may be specifically designed or configured to execute on the VMM 530 to improve efficiency. In these instances, the guest operating system is "aware" that it is executing on a virtual machine monitor. In other words, the VMM 530 may provide paravirtualization to the guest operating system in certain circumstances.

[0116] A computer system process includes an allocation of hardware processor time, as well as an allocation of memory (physical and / or virtual), an allocation of memory for storing instructions executed by the hardware processor, for storing data generated by the execution of instructions by the hardware processor, and / or for storing hardware processor state (e.g., contents of registers) between allocations of hardware processor time when the computer system process is not running. A computer system process runs under the control of an operating system and may run under the control of other programs executing on the computer system.

[0117] cloud computing

[0118] The term "cloud computing" is used generally herein to describe a computing model that enables on-demand access to a shared pool of computing resources, such as computer networks, servers, software applications, and services, and allows resources to be rapidly provisioned and released with minimal management effort or service provider interaction.

[0119] A cloud computing environment (sometimes referred to as a cloud environment or just a cloud) can be implemented in a variety of different ways to best suit different requirements. For example, in a public cloud environment, the underlying computing infrastructure is owned by an organization that makes its cloud services available to other organizations or the public. In contrast, a private cloud environment is generally intended for use by or within a single organization. A community cloud is intended to be shared by several organizations within a community; while a hybrid cloud includes two or more types of clouds (e.g., private, community, or public) bound together by data and application portability.

[0120] In general, the cloud computing model enables some of those responsibilities that might previously have been provided by an organization's own information technology department to instead be delivered as a service layer within a cloud environment for use by consumers (either inside or outside the organization, depending on the public / private nature of the cloud). Depending on the specific implementation, the precise definition of the components or features provided by or within each cloud service layer may vary, but common examples include: Software as a Service (SaaS), in which consumers use software applications running on a cloud infrastructure, while the SaaS provider manages or controls the underlying cloud infrastructure and applications. Platform as a Service (PaaS), in which consumers can use software programming languages ​​and development tools supported by the supplier of the PaaS to develop, deploy, and otherwise control their own applications, while the PaaS provider manages or controls other aspects of the cloud environment (i.e., everything under the runtime execution environment). Infrastructure as a Service (IaaS), in which consumers can deploy and run arbitrary software applications, and / or provide processes, storage, networks, and other basic computing resources, while the IaaS provider manages or controls the underlying physical cloud infrastructure (i.e., everything below the operating system layer). Database as a Service (DBaaS), where the consumer uses a database server or database management system running on a cloud infrastructure, while the DbaaS provider manages or controls the underlying cloud infrastructure and applications.

[0121] The above basic computer hardware and software and cloud computing environment are given to illustrate the basic underlying computer components that can be used to implement (one or more) example embodiments. However, (one or more) example embodiments are not necessarily limited to any particular computing environment or computing device configuration. Instead, according to the present disclosure, (one or more) example embodiments can be implemented in any type of system architecture or processing environment that a person skilled in the art will understand in view of the present disclosure as being able to support the features and functions of (one or more) example embodiments given herein.

[0122] Machine Learning Models

[0123] A machine learning model is trained using a specific machine learning algorithm. Once trained, the input is applied to the machine learning model to make a prediction, which may also be referred to as a predicate output or output in this article. The attributes of the input may be referred to as features, and the values ​​of the features may be referred to as feature values ​​in this article.

[0124] A machine learning model includes a model data representation or model artifact. The model artifact includes parameter values, which may be referred to herein as theta values, and are applied to the input by the machine learning algorithm to generate a predicted output. Training a machine learning model requires determining theta values ​​of the model artifact. The structure and organization of theta values ​​depends on the machine learning algorithm.

[0125] In supervised training, training data is used by a supervised training algorithm to train a machine learning model. The training data includes inputs and "known" outputs. In an embodiment, the supervised training algorithm is an iterative process. In each iteration, the machine learning algorithm applies the model artifacts and inputs to generate a predicted output. The error or variance between the predicted output and the known output is calculated using an objective function. In practice, the output of the objective function indicates the accuracy of the machine learning model based on the specific state of the model artifacts in the iteration. By applying an optimization algorithm based on the objective function, the theta value of the model artifact can be adjusted. An example of an optimization algorithm is gradient descent. Iterations can be repeated until the desired accuracy is achieved or some other criterion is met.

[0126] In software implementations, when a machine learning model is referred to as receiving inputs, executing, and / or generating outputs or predicates, a computer system process executing a machine learning algorithm applies a model artifact to the inputs to generate a predicted output. The computer system process executes the machine learning algorithm by executing software configured to cause the algorithm to execute.

[0127] Classes of problems that machine learning (ML) excels at include clustering, classification, regression, anomaly detection, prediction, and dimensionality reduction (i.e., simplification). Examples of machine learning algorithms include decision trees, support vector machines (SVMs), Bayesian networks, random algorithms such as genetic algorithms (GAs), and connectionist topologies such as artificial neural networks (ANNs). Implementations of machine learning can rely on matrices, symbolic models, and hierarchical and / or associative data structures. Parameterized (i.e., configurable) implementations of the best machine learning algorithms can be found in open source libraries such as Google's TensorFlow for Python and C++ or Georgia Tech's MLPack for C++. Shogun is an open source C++ ML library with adapters for several programming languages, including C#, Ruby, Lua, Java, Matlab, R, and Python.

[0128] Artificial Neural Networks

[0129] An artificial neural network (ANN) is a machine learning model that models at a high level a system of neurons interconnected by directed edges. An overview of neural networks is described in the context of hierarchical feedforward neural networks. Other types of neural networks share characteristics of the neural networks described below.

[0130] In a layered feedforward network such as a multilayer perceptron (MLP), each layer comprises a set of neurons. A layered neural network comprises an input layer, an output layer, and one or more intermediate layers known as hidden layers.

[0131] The neurons in the input layer and the output layer are referred to as input neurons and output neurons, respectively. Neurons in the hidden layer or the output layer may be referred to as activation neurons in this article. Activation neurons are associated with activation functions. The input layer does not contain any activation neurons.

[0132] From each neuron in the input layer and hidden layer, there is one or more directed edges to the activated neuron in the subsequent hidden layer or output layer. Each edge is associated with a weight. The edge from a neuron to an activated neuron represents the input from the neuron to the activated neuron, as adjusted by the weight.

[0133] For a given input to a neural network, each neuron in the neural network has an activation value. For an input neuron, the activation value is simply the input value for that input. For an activation neuron, the activation value is the output of the corresponding activation function that activated the neuron.

[0134] Each edge from a particular neuron to an activated neuron represents that the activation value of the particular neuron is the input to the activated neuron, i.e., the input to the activation function of the activated neuron, as adjusted by the weight of the edge. Thus, an activated neuron in a subsequent layer represents that the activation value of the particular neuron is the input to the activation function of the activated neuron, as adjusted by the weight of the edge. An activated neuron may have multiple edges pointing to the activated neuron, each edge representing that the activation value from the source origin neuron (as adjusted by the weight of the edge) is the input to the activation function of the activated neuron.

[0135] Each activated neuron is associated with a bias. To generate an activation value for an activated neuron, the neuron's activation function is applied to the weighted activation value and the bias.

[0136] Illustrative Data Structures for Neural Networks

[0137] Artifacts of a neural network may include matrices of weights and biases. Training a neural network may iteratively adjust the matrices of weights and biases.

[0138] For layered feedforward networks, as well as other types of neural networks, the artifacts may include one or more matrices of edges W. Matrix W represents the edges from layer L-1 to layer L. Assuming the number of neurons in layers L-1 and L is N[L-1] and N[L] respectively, then the dimensions of matrix W are N[L-1] columns and N[L] rows.

[0139] The bias for a particular layer L may also be stored in a matrix B having N[L] rows and one column.

[0140] The matrices W and B may be stored in RAM memory as vectors or arrays, or as a comma-separated set of values ​​in memory. When the artifact is stored persistently in a persistent storage device, the matrices W and B may be stored as comma-separated values ​​in a compressed and / or serialized form or other suitable persistent form.

[0141] The specific input applied to the neural network includes the value of each input neuron. The specific input can be stored as a vector. The training data includes multiple inputs, each input is called a sample in the sample set. Each sample includes the value of each input neuron. The sample can be stored as a vector of input values, and the multiple samples can be stored as a matrix, where each row in the matrix is ​​a sample.

[0142] When input is applied to a neural network, activation values ​​are generated for the hidden and output layers. For each layer, the activation values ​​can be stored in a column of a matrix A that has a row for each neuron in the layer. In a vectorized approach for training, the activation values ​​can be stored in a matrix that has a column for each sample in the training data.

[0143] Training a neural network requires storing and processing additional matrices. The optimization algorithm generates matrices of derivative values ​​used to adjust the weights W and bias B matrices. Generating derivative values ​​can use and require matrices that store intermediate values ​​generated when calculating the activation values ​​of each layer.

[0144] The number of neurons and / or edges determines the size of the matrices required to implement a neural network. The fewer the number of neurons and edges in a neural network, the smaller the matrices and the memory required to store them. In addition, a smaller number of neurons and edges reduces the amount of computation required to apply or train a neural network. Fewer neurons means fewer activation values ​​that need to be calculated and / or fewer derivatives that need to be calculated during training.

[0145] The properties of the matrix used to implement the neural network correspond to neurons and edges. A cell in the matrix W represents a specific edge from the L-1 layer to a neuron in the L layer. An activated neuron represents the activation function of that layer, including the activation function. The activated neurons in the L layer correspond to the rows of the matrix W for the weights of the edge between the L layer and the L-1 layer and the columns of the matrix W for the weights of the edge between the L layer and the L+1 layer. During the execution of the neural network, the neurons also correspond to one or more activation values ​​stored in the matrix A for that layer and generated by the activation function.

[0146] ANN is suitable for vectorization of data parallelism, which can utilize vector hardware, such as single instruction multiple data (SIMD), such as utilizing a graphics processing unit (GPU). Matrix partitioning can achieve horizontal scaling, for example, utilizing symmetric multiprocessing (SMP), such as utilizing a multi-core central processing unit (CPU) and / or multiple coprocessors (such as GPU). Feedforward calculations in ANN can occur with only one step in each neural layer. The activation values ​​in one layer are calculated based on the weighted propagation of the activation values ​​of the previous layer, so that the values ​​are calculated in turn for each subsequent layer, such as sequentially calculating the values ​​for each subsequent layer, such as utilizing the corresponding iterations of a for loop. Layering imposes the ordering of non-parallelizable calculations. Therefore, the network depth (i.e., the number of layers) can cause computational delays. Deep learning requires giving multilayer perceptrons (MLPs) multiple layers. Each layer implements data abstraction, while complex (i.e., multidimensional with several inputs) abstractions require multiple layers to implement cascade processing. Implementations of ANNs based on reusable matrices and matrix operations for feed-forward processing are readily available and parallelizable in neural network libraries such as Google's TensorFlow for Python and C++, OpenNN for C++, and Fast Artificial Neural Networks (FANN) from the University of Copenhagen. These libraries also provide model training algorithms such as backpropagation.

[0147] Back Propagation

[0148] The output of an ANN can be more or less correct. For example, an ANN that recognizes letters may mistake an I for an L because the letters have similar features. The correct output may have a specific value(s), while the actual output may have a slightly different value. The arithmetic or geometric difference between the correct output and the actual output may be measured as an error according to a loss function, such that zero represents error-free (i.e., perfectly accurate) behavior. For any edge in any layer, the difference between the correct output and the actual output is the delta value.

[0149] Backpropagation requires distributing the error backward through the layers of the ANN to all connected edges within the ANN in varying amounts. The propagation of the error results in adjustments to the edge weights, which depend on the gradient of the error on each edge. The gradient of an edge is calculated by multiplying the error delta of the edge by the activation value of the upstream neuron. When the gradient is negative, the greater the error contribution of the edge to the network, the more the edge weight should be reduced, which is negative reinforcement. When the gradient is positive, positive reinforcement requires increasing the weight of the edge whose activation will reduce the error. Edge weights are adjusted based on a percentage of the edge's gradient. The steeper the gradient, the greater the adjustment. Not all edge weights are adjusted by the same amount. As model training continues with additional input samples, the error of the ANN should decrease. Training can stop when the error stabilizes (i.e., stops decreasing) or disappears below a threshold (i.e., approaches zero). Christopher M. Bishop teaches example mathematical formulas and techniques for feed-forward multilayer perceptrons (MLPs), including matrix operations and backpropagation, in the related reference "EXACT CALCULATION OF THE HESSIAN MATRIX FOR THE MULTI-LAYER PERCEPTRON".

[0150] Model training can be supervised, or unsupervised. For supervised training, the expected (i.e., correct) output is already known for each example in the training set. The training set is configured by pre-assigning classification labels to each example (e.g., by a human expert). For example, a training set for optical character recognition may have blurred photos of individual letters, and an expert may pre-label each photo according to which letter is shown. Error calculation and back-propagation occur as explained above.

[0151] Because the desired output needs to be discovered during training, more unsupervised model training is involved. Unsupervised training can be easier to adopt because no human experts are required to pre-label training examples. Therefore, unsupervised training saves manpower. The natural way to implement unsupervised training is to use an autoencoder, which is a type of ANN. The autoencoder acts as an encoder / decoder (codec) with two sets of layers. The first set of layers encodes the input example into a condensed code that needs to be learned during model training. The second set of layers decodes the condensed code to regenerate the original input example. The two sets of layers are trained together as a combined ANN. The error is defined as the difference between the original input and the input regenerated after decoding. After sufficient training, the decoder outputs the original input more or less exactly.

[0152] For each input example, the autoencoder relies on a condensed code as an intermediate format. The intermediate condensed code does not exist initially, but only emerges through model training, which may be counterintuitive. Unsupervised training can implement a vocabulary of intermediate encodings based on features and distinctions of unexpected correlations. For example, which examples and which labels are used during supervised training may depend on a human expert's understanding of the problem space that is somewhat unscientific (e.g., anecdotal) or incomplete. Unsupervised training, on the other hand, is more or less entirely based on statistical trends to discover a suitable intermediate vocabulary, which converges reliably to optimality with sufficient training due to internal feedback generated by regenerating decoding. The implementation and integration techniques of autoencoders are taught in the related U.S. patent application No. 14 / 558,700 entitled "AUTO-ENCODER ENHANCED SELF-DIAGNOSTIC COMPONENTS FOR MODEL MONITORING". This patent application promotes supervised or unsupervised ANN models to first-class objects, which are suitable for management techniques such as monitoring and governance during model development (such as during training).

[0153] Random Forest

[0154] A random forest or random decision forest is a family of learning methods that construct a collection of randomly generated nodes and decision trees during a training phase. The different decision trees of the forest are constructed such that each is randomly restricted to only a specific subset of the feature dimensions of the dataset, such as by feature bootstrapping aggregation (bagging). Thus, as the decision trees grow, they gain accuracy without being forced to overfit the training data, as would happen if the decision trees were forced to learn all feature dimensions of the dataset. The predictions can be calculated based on the average (or other integral, such as softmax) of the predictions from the different decision trees.

[0155] Random forest hyperparameters can include: number-of-trees-in-the-forest (the number of trees in the forest), maximum-number-of-features-considered-for-splitting-a-node (the maximum number of features considered for splitting a node), number-of-levels-in-each-decision-tree (the number of levels in each decision tree), minimum-number-of-data-points-on-a-leaf-node (the minimum number of data points on a leaf node), method-for-sampling-data-points (the method for sampling data points), and so on.

[0156] In the foregoing specification, embodiments of the present invention have been described with reference to numerous specific details, which may vary from implementation to implementation. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. The sole and exclusive indicator of the scope of the invention, and what the applicants intend as the scope of the invention, is the literal and equivalent range of the set of claims issuing from this application in the specific form in which such claims are issued, including any subsequent corrections.

Claims

1. A computer-implemented method comprising: For each of the multiple features of the training dataset, a feature relevance score is calculated based on: feature relevance scoring function, and Statistics of the values ​​of this feature that appear in the training dataset; For each feature of the plurality of features, calculating a ranking based on feature relevance scores of the plurality of features; generating a sequence of different subsets of the plurality of features based on the ranking of the plurality of features, wherein each different subset in the sequence of different subsets of the plurality of features has a different size, wherein the different sizes of the sequence of different subsets of the plurality of features comprise an exponential sequence of sizes; For each different subset in the sequence of different subsets of the plurality of features, the following operations are performed: configuring the machine learning (ML) model to accept the different subset of the plurality of features, training a ML model based on the configured ML model, and calculating a fitness score based on training the ML model; selecting the most accurate subset of features that provides the highest training accuracy for the sequence of different subsets of the plurality of features; configuring and training the ML model based on the most accurate subset of features; Wherein, for each different subset in the sequence of different subsets of the plurality of features, calculating the fitness score based on the trained ML model comprises: A corresponding plurality of different subsets in the sequence of different subsets are distributed to each processor of the plurality of processors, wherein the corresponding plurality of different subsets in the sequence of different subsets are evaluated by the processor in descending order of size of the different subsets.

2. The method of claim 1, wherein: For each feature of the plurality of features, the feature relevance scoring function associates the feature with a prediction target for which an ML model may be trained to perform inference.

3. The method of claim 2, wherein: For each feature of the plurality of features, associating the feature with a predicted target includes calculating mutual information between the feature and the predicted target.

4. The method of claim 2, wherein: For each feature of the plurality of features, the feature is associated with a prediction target based on the f-score.

5. The method of claim 2, wherein: The prediction target includes a classification label.

6. The method of claim 2, wherein: The prediction target includes regression.

7. The method of claim 1, wherein: The calculated relevance score is based on the impact of the feature on the accuracy of the ML model.

8. The method of claim 7, wherein: The described impact of features on the accuracy of the ML model is based on adaptive boosting.

9. The method of claim 1, wherein: The method also includes, for each feature in the plurality of features, calculating the following: a second relevance score based on a second objective scoring function, and a second ranking based on second relevance scores of the plurality of features; The sequence of generating different subsets of the plurality of features is also based on the second ranking of the plurality of features.

10. The method of claim 1, wherein: Each different subset in the sequence of different subsets of the plurality of features includes a previous different subset in the sequence of different subsets of the plurality of features.

11. The method of claim 1, wherein: Training of the ML model is restricted to the different subsets of the plurality of features of the training dataset.

12. The method of claim 1, wherein: The size of the sequences of different subsets of the plurality of features is the same as the size of the plurality of features.

13. The method of claim 1, wherein: A first processor of the plurality of processors steals a largest different subset to be processed from a second processor of the plurality of processors.

14. A computer-implemented method comprising: For each of the multiple features of the training dataset, based on the statistics of the values ​​of that feature that appear in the training dataset, calculate the following: a first feature relevance score based on the first feature relevance scoring function, and a second feature relevance score based on a second feature relevance scoring function; For each of the plurality of features, the following term is calculated: a first ranking based on first feature relevance scores of the plurality of features, and a second ranking based on second feature relevance scores of the plurality of features; generating a sequence of different subsets of the plurality of features based on the first ranking of the plurality of features and the second ranking of the plurality of features, wherein each different subset in the sequence of different subsets of the plurality of features has a different size; For each different subset in the sequence of different subsets of the plurality of features: configuring the machine learning (ML) model to accept the different subset of the plurality of features, training a ML model based on the configured ML model, and Calculating a fitness score based on the trained ML model; selecting the most accurate subset of features that provides the highest training accuracy for the sequence of different subsets of the plurality of features; configuring and training the ML model based on the most accurate subset of features; Wherein, for each different subset in the sequence of different subsets of the plurality of features, calculating the fitness score based on the trained ML model comprises: A corresponding plurality of different subsets in the sequence of different subsets are distributed to each processor of the plurality of processors, wherein the corresponding plurality of different subsets in the sequence of different subsets are evaluated by the processor in descending order of size of the different subsets.

15. One or more non-transitory computer-readable media storing one or more sequences of instructions that, when executed by one or more processors, cause the method of any one of claims 1-14 to be performed.

16. A computer device comprising: one or more processors; and A memory coupled to the one or more processors and comprising instructions stored thereon, which, when executed by the one or more processors, cause the method of any one of claims 1-14 to be performed.

Citation Information

Patent Citations

  • Auto-encoder enhanced self-diagnostic components for model monitoring

    US11836746B2

  • Data mining platform for bioinformatics and other knowledge discovery

    US20080097938A1