Systems and methods for operating automated feature engineering
The automated feature engineering system addresses inefficiencies in current tools by iteratively evaluating and determining importance coefficients, enhancing efficiency and accuracy in feature extraction for machine learning models.
Patent Information
- Application Number
- JP2023519186
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-30
- Filing Date
- 2021-09-16
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2041-09-16
AI Technical Summary
Current feature engineering tools are time-consuming, difficult to reuse, and require large amounts of data, making them inefficient for enterprise data processing needs.
A method and system for automated feature engineering that includes selecting primitives from a pool, synthesizing features, iteratively evaluating their usefulness, and determining importance coefficients, allowing for efficient feature extraction without requiring extensive data.
Enhances feature engineering efficiency by reducing the need for large data sets and enabling rapid adaptation to different datasets, improving the accuracy of machine learning models.
Smart Images

Figure 0007792956000001 
Figure 0007792956000002 
Figure 0007792956000003
Abstract
Description
[Technical Field]
[0001] The present invention relates generally to processing data streams, and more particularly to feature engineering useful for performing machine learning on data in streams. [Background technology]
[0002] This application claims priority to U.S. Non-Provisional Patent Application No. 17 / 039,428, filed September 30, 2020, which is incorporated by reference in its entirety.
[0003] Feature engineering is typically the process of identifying and extracting predictive features from complex data analyzed by businesses and other enterprises. Features are key to the accuracy of predictions made by machine learning models. Therefore, feature engineering is often the determining factor in the success of a data analytics project. Feature engineering is generally a time-consuming process. Currently available feature engineering tools make it difficult to reuse previous work, so an entirely new feature engineering pipeline must be built for every data analytics project. Additionally, currently available feature engineering tools typically require large amounts of data to achieve good predictive accuracy. Therefore, current feature engineering tools cannot efficiently address enterprise data processing needs. Summary of the Invention
[0004] These and other problems are addressed by a method for processing data blocks in a data analysis system, a computer-implemented data analysis system, and a computer-readable memory. One embodiment of the method includes receiving a dataset from a data source. The method further includes selecting primitives from a pool of primitives based on the received dataset. Each of the selected primitives is configured to be applied to at least a portion of the dataset to synthesize one or more features. The method further includes synthesizing the plurality of features by applying the selected primitives to the received dataset. The method further includes iteratively evaluating the plurality of features and removing some features from the plurality of features to obtain a subset of features. Each iteration includes the evaluating step, which includes evaluating the usefulness of at least some of the plurality of features by applying the evaluated features to a different portion of the dataset, and removing some of the evaluated features based on the usefulness of the evaluated features to generate a subset of features. The method also includes determining an importance coefficient for each feature in the subset of features. The method also includes generating a machine learning model based on the subset of features and the importance coefficient for each feature in the subset of features. The machine learning model is configured to be used to make predictions based on new data.
[0005] One embodiment of a computer-implemented data analysis system includes a computer processor for executing computer program instructions. The system also includes a non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations. The operations include receiving a dataset from a data source. The operations further include selecting primitives from a pool of primitives based on the received dataset. Each of the selected primitives is configured to be applied to at least a portion of the dataset to synthesize one or more features. The operations further include synthesizing the plurality of features by applying the selected primitives to the received dataset. The operations further include iteratively evaluating the plurality of features and removing some features from the plurality of features to obtain a subset of features. Each iteration includes evaluating the usefulness of at least some of the plurality of features by applying the evaluated features to different portions of the dataset, and removing some of the evaluated features based on the usefulness of the evaluated features to generate a subset of features. The operations also include determining an importance coefficient for each feature in the subset of features. The method also includes generating a machine learning model based on the subset of features and the importance coefficients of each feature in the subset of features, The machine learning model is configured to be used to make predictions based on new data.
[0006] An embodiment of a non-transitory computer-readable memory stores executable computer program instructions, the instructions being executable to perform operations. The operations include receiving a dataset from a data source. The operations further include selecting primitives from a pool of primitives based on the received dataset. Each of the selected primitives is configured to be applied to at least a portion of the dataset to synthesize one or more features. The operations further include synthesizing the plurality of features by applying the selected primitives to the received dataset. The operations further include iteratively evaluating the plurality of features and removing some features from the plurality of features to obtain a subset of features. Each iteration includes evaluating the usefulness of at least some of the plurality of features by applying the evaluated features to a different portion of the dataset, and removing some of the evaluated features based on the usefulness of the evaluated features to generate a subset of features. The operations also include determining an importance factor for each feature in the subset of features. The method also includes generating a machine learning model based on the subset of features and the importance factor for each feature in the subset of features. The machine learning model is configured to be used to make predictions based on new data. [Brief explanation of the drawings]
[0007] Further features and advantages of the present invention will become apparent from the following detailed description of the invention, which proceeds with reference to the accompanying drawings. [Figure 1] FIG. 1 is a block diagram illustrating a machine learning environment including a machine learning server, according to one embodiment. [Figure 2] FIG. 1 is a block diagram illustrating a more detailed view of the feature engineering application of the machine learning server, according to one embodiment. [Figure 3] FIG. 1 is a block diagram illustrating a more detailed view of the feature generation module of the feature engineering application, according to one embodiment. [Figure 4]1 is a flowchart illustrating a method for generating a machine learning model, according to one embodiment. [Figure 5] 1 is a flowchart illustrating a method for training a machine learning model and making predictions using the trained model, according to one embodiment. [Figure 6] FIG. 2 is a high-level block diagram illustrating a functional view of an exemplary computer system for use as the machine learning server of FIG. 1, according to one embodiment.
[0008] The drawings depict various embodiments for purposes of illustration only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods illustrated herein may be utilized without departing from the principles of the present invention as described herein. Like reference symbols and designations in the various drawings refer to like elements. DETAILED DESCRIPTION OF THE INVENTION
[0009] 1 is a block diagram illustrating a machine learning environment 100 including a machine learning server 110, according to one embodiment. The environment 100 further includes a variety of data sources 120 connected to the machine learning server 110 via a network 130. Although the illustrated environment 100 includes only one machine learning server 110 coupled to the variety of data sources 120, an embodiment can have a variety of machine learning servers and a single data source.
[0010] The data source 120 stores electronic data. Machine Learning Server110. Data sources 120 may be storage devices such as hard disk drives (HDDs) or solid state drives (SSDs), computers that manage and provide access to multiple storage devices, storage area networks (SANs), databases, or cloud storage systems. Data sources 120 may also be computer systems that can retrieve data from another source. Data sources 120 may be remote from machine learning server 110 and provide data over network 130. Additionally, some or all of data sources 120 may be directly coupled to the data analysis system and provide data without passing the data over network 130.
[0011] The data provided by the data source 120 may be organized into data records (e.g., rows). Each data record includes one or more values. For example, a data record provided by the data source 120 may include a series of comma-separated values. The data may be Machine Learning Server The data from data sources 120 describes information related to businesses using 110. For example, data from data sources 120 may describe computer-based interactions with content and / or applications accessible on a website (e.g., click tracking data). As another example, data from data sources 120 may describe customer transactions online and / or in-store. Businesses may belong to one or more of a variety of industries, such as manufacturing, retail, finance, banking, etc.
[0012] The machine learning server 110 is a computer-based system utilized to build machine learning models and serve them to make predictions based on data. Exemplary predictions include whether a customer will make a transaction within a certain period of time, whether a transaction is fraudulent, whether a user will perform a computer-based interaction, etc. Data is retrieved, collected, or accessed from one or more diverse data sources 120 via a network 130. The machine learning server 110 can implement scalable software tools and hardware resources used to access, prepare, blend, and analyze data from the diverse data sources 120. The machine learning server 110 can be a computing device used to implement machine learning functions, including the feature engineering and modeling techniques described herein.
[0013] 1 as feature engineering application 140 and modeling application 150. Feature engineering application 140 performs automated feature engineering to extract predictive variables, or features, from data (e.g., temporal and relational data sets) provided by data source 120. Each feature is a variable potentially relevant to a prediction (called a target prediction) made using the corresponding machine learning model.
[0014] In one embodiment, the feature engineering application 140 selects primitives from a pool of primitives based on the data. The pool of primitives is maintained by the feature engineering application 140. The primitives define individual calculations that can be applied to raw data in a dataset to create one or more new features with associated values. The selected primitives constrain the input and output data types, so they can be applied to different types of datasets and stacked to create new calculations. The feature engineering application 140 synthesizes features by applying the selected primitives to data provided by the data source. The feature engineering application 140 then evaluates the features to determine the importance of each feature through an iterative process in which different portions of the data are applied to the features in each iteration. The feature engineering application 140 removes some features in each iteration to obtain a subset of features that are more predictive than the removed features.
[0015] For each feature in the subset, the feature engineering application 140 determines an importance factor, for example, by using a random forest. The importance factor indicates how important / relevant the feature is to predicting the target. The features in the subset and their importance factors can be sent to the modeling application 150 to build a machine learning model.
[0016] One advantage of feature engineering application 140 is that the use of primitives makes the feature engineering process more efficient than traditional feature engineering processes in which features are extracted from raw data. Additionally, feature engineering application 140 can evaluate primitives based on feature ratings and importance coefficients generated from the primitives. Metadata describing the evaluation of the primitives can be generated and used to determine whether to select a primitive for different data or a different prediction problem. Traditional feature engineering processes can generate large numbers of features (e.g., millions) without providing any guidance or solutions for engineering features more quickly or appropriately. Another advantage of feature engineering application 140 is that it does not require large amounts of data to evaluate features. Rather, it applies an iterative method to evaluate features, using a different portion of the data in each iteration.
[0017] The feature engineering application 140 can provide a graphical user interface (GUI) that allows a user to contribute to the feature engineering process. As an example, the GUI associated with the feature engineering application 140 provides a feature selection tool that allows a user to edit features selected by the feature engineering application 140. The GUI can also provide the user with options to specify variables to consider and change feature characteristics, such as the maximum allowable feature depth, the maximum number of features generated, and the date range of included data (e.g., specified by a cutoff time). More details about the feature engineering application 140 are described in conjunction with FIGS. 2-4.
[0018] The modeling application 150 trains a machine learning model using the features and feature importance coefficients received from the feature engineering application 140. Different machine learning techniques, such as linear support vector machines (linear SVMs), boosting algorithms (e.g., AdaBoost), neural networks, logistic regression, naive Bayes, memory-based learning, random forests, bagged trees, decision trees, boosted trees, or boosted stumps, may be used in different embodiments. The generated machine learning model, when applied to features extracted from a new dataset (e.g., a dataset from the same or a different data source 120), makes a target prediction. The new dataset may be missing one or more features, but these features may still be included with null values. In some embodiments, the modeling application 150 applies dimensionality reduction (e.g., via linear discriminant analysis (LDA), principal component analysis (PCA), etc.) to reduce the amount of data in the features of the new dataset to a smaller, more representative dataset.
[0019] In some embodiments, the modeling application 150 validates predictions before deploying to new datasets. For example, the modeling application 150 applies the trained model to a validation dataset to quantify the model's accuracy. Common metrics applied to measure accuracy include precision = true positives (TP) / ((true positives (TP) + false positives (FP)) and recall = true positives (TP) / ((true positives (TP) + false negatives (FN))), where precision is the number of outcomes (TP or true positives) that the model correctly predicts out of the total predicted by the model (TP + TF or false positives), and recall is the number of outcomes that the model correctly predicts (TP) out of the total number that actually occurred (TP + FN or false negatives). The F-measure (F-measure = 2 * PR (precision * recall) / P + R (precision + recall)) combines precision and recall into a single measure. In one embodiment, the modeling application 150 iteratively retrains the machine learning model until a stopping condition occurs, such as an accuracy measurement indication that the machine learning model is sufficiently accurate or several training rounds have been performed.
[0020] In some embodiments, modeling application 150 tailors the machine learning model to specific business needs. For example, modeling application 150 builds a machine learning model for recognizing fraudulent financial transactions and tailors the model to emphasize more significant (e.g., high-value transactions) fraudulent transactions to reflect the needs of the business, e.g., by transforming the predicted probabilities in a manner that emphasizes the more significant transactions. More details about modeling application 150 are described in conjunction with FIG. 5.
[0021] Network 130 represents a communication path between machine learning server 110 and data source 120. In one embodiment, network 130 is the Internet and uses standard communication technologies and / or protocols. Thus, network 130 may include links using technologies such as Ethernet, 802.11, worldwide interoperability for microwave accesses (WiMAX), 3G, Long Term Evolution (LTE), Digital Subscriber Line (DSL), Asynchronous Transfer Mode (ATM), InfiniBand, PCI Express Advanced Switching, etc. Similarly, networking protocols used in network 130 may include Multiprotocol Label Switching (MPLS), Transmission Control Protocol / Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transport Protocol (HTTP), Simple Mail Transfer Protocol (SMTP), File Transfer Protocol (FTP), etc.
[0022] Data exchanged over network 130 may be represented using technologies and / or formats including HyperText Markup Language (HTML), Extensible Markup Language (XML), etc. Additionally, all or some links may be encrypted using conventional encryption technologies such as Secure Sockets Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. In alternative embodiments, entities may use custom and / or proprietary data communication technologies instead of or in addition to the above technologies.
[0023] 2 is a block diagram illustrating a feature engineering application 200, according to one embodiment. Feature engineering application 200 is one embodiment of feature engineering application 140 of FIG. 1. Feature engineering application 200 includes a primitive selection module 210, a feature generation module 220, Metadata The feature engineering application 200 includes a generation module 230, and a database 240. The feature engineering application 200 receives a dataset from the data source 120 and generates a machine learning model based on the dataset. Those skilled in the art will recognize that other embodiments may have different and / or other components than those described herein, and that functionality may be distributed among the components in a different manner.
[0024] The primitive selection module 210 selects one or more primitives from a pool of primitives maintained by the feature engineering application 200. The pool of primitives includes a large number of primitives, such as hundreds or thousands of primitives. Each primitive has an algorithm that, when applied to data, performs a calculation on the data and generates a feature having an associated value. A primitive is associated with one or more attributes. A primitive's attributes can be a description of the primitive (e.g., a natural language description specifying the calculation performed by the primitive when applied to data), an input type (i.e., the type of input data), a return type (i.e., the type of output data), primitive metadata indicating how useful the primitive was in a previous feature engineering process, or other attributes.
[0025] In some embodiments, the pool of primitives includes a variety of different types of primitives. One type of primitive is an aggregation primitive. When applied to a dataset, an aggregation primitive identifies related data in the dataset, performs a determination on the related data, and creates a value that summarizes and / or aggregates the determination. For example, the aggregation primitive "Count" identifies values in related rows in the dataset, determines whether each of the values is a non-null value, and returns (outputs) a count of the number of non-null values in the rows of the dataset. Another type of primitive is a transformation primitive. When applied to a dataset, a transformation primitive creates a new variable from one or more existing variables in the dataset. For example, the transformation primitive "Weekend" evaluates a timestamp in the dataset and returns a binary value (e.g., true or false) indicating whether the date indicated by the timestamp occurs on a weekend. Another exemplary transformation primitive evaluates a timestamp and returns a count indicating the number of days until a specified date (e.g., the number of days until a particular holiday).
[0026] The primitive selection module 210 selects a set of primitives based on a dataset received from a data source, such as one of the data sources 120 in FIG. 1 . In some embodiments, the primitive selection module 210 uses a skim view approach, a summary view approach, or both approaches to select primitives. In the skim view approach, the primitive selection module 210 identifies one or more semantic representations of the dataset. The semantic representation of the dataset describes characteristics of the dataset and may be obtained without performing calculations on the data in the dataset. Examples of semantic representations of a dataset include the presence of one or more particular variables (e.g., column names) in the dataset, the number of columns, the number of rows, the input type of the dataset, other attributes of the dataset, and some combination thereof. To select primitives using the skim view approach, the primitive selection module 210 determines whether the identified semantic representation of the dataset matches the attributes of the primitives in the pool. If there is a match, the primitive selection module 210 selects the primitives.
[0027] The skim view approach is a rule-based analysis. The determination of whether an identified semantic representation of a dataset matches an attribute of a primitive is based on rules maintained by the feature engineering application 200. The rules specify which semantic representation of a dataset matches which attribute of a primitive, for example, based on keyword matches between the semantic representation of the dataset and the attribute of the primitive. In one example, the semantic representation of the dataset is a column named "Date of Birth," and the primitive selection module 210 selects primitives whose input type is "Date of Birth" that matches the semantic representation of the dataset. In another example, the semantic representation of the dataset is a column named "Timestamp," and the primitive selection module 210 selects primitives that have an attribute indicating that the primitive is appropriate for use with data that indicates a timestamp.
[0028] In the summary view approach, the primitive selection module 210 generates a representative vector from the dataset. The representative vector encodes data describing the dataset, such as data indicating the number of tables in the dataset, the number of columns per table, the average number of each column, and the average number of each row. The representative vector thus serves as a fingerprint of the dataset. A fingerprint is a compact representation of the dataset and may be generated by applying one or more fingerprint functions, such as a hash function, Rabin's fingerprint algorithm, or other types of fingerprint functions, to the dataset.
[0029] The primitive selection module 210 selects primitives for a dataset based on the representative vectors. For example, the primitive selection module 210 inputs the representative vectors for the dataset into a machine learning model. The machine learning model outputs primitives for the dataset. The machine learning model is trained by the primitive selection module 210 to select primitives for the dataset based on the representative vectors, for example. It can be trained based on training data including a plurality of representative vectors for a plurality of training datasets and a set of primitives for each of the plurality of training datasets. The set of primitives for each of the plurality of training datasets is used to generate features determined to be useful for making predictions based on the corresponding training dataset. In some embodiments, the machine learning model is trained continuously. For example, the primitive selection module 210 can further train the machine learning model based on at least some of the representative vectors for the dataset and the selected primitives.
[0030] The primitive selection module 210 can provide the selected primitives for display to a user (e.g., a data analysis engineer) in a GUI supported by the feature engineering application 200. The GUI may also allow the user to edit the primitives, such as adding other primitives to the set of primitives, creating new primitives, removing selected primitives, other types of actions, or some combination thereof.
[0031] The feature generation module 220 generates a group and an importance factor for each feature in the group. In some embodiments, the feature generation module 220 synthesizes multiple features based on the selected primitives and the dataset. In some embodiments, the feature generation module 220 applies each of the selected primitives to at least a portion of the dataset to synthesize one or more features. For example, the feature generation module 220 applies the "weekend" primitive to a column named "timestamp" in the dataset to synthesize a feature indicating whether a date occurs on a weekend. The feature generation module 220 can synthesize a large number of features for a dataset, such as hundreds or millions of features.
[0032] The feature generation module 220 evaluates the features and removes some of the features based on the evaluation to obtain a group of features. In some embodiments, the feature generation module 220 evaluates the features through an iterative process. In each round of iteration, the feature generation module 220 applies the features not removed by the previous iteration (also referred to as "remaining features") to a different portion of the dataset and determines a utility score for each feature. The feature generation module 220 removes some of the features with the lowest utility scores from the remaining features. In some embodiments, the feature generation module 220 uses a random forest to determine the utility scores for the features.
[0033] After the iterations are performed and a group of features is obtained, the feature generation module 220 determines an importance coefficient for each feature in the group. The feature importance coefficient indicates how important the feature is for predicting the target variable. In some embodiments, the feature generation module 220 determines the importance coefficient by using a random forest, e.g., a forest constructed based on at least a portion of the dataset. In some embodiments, the feature generation module 220 adjusts the feature importance score by inputting the feature and different portions of the dataset into a machine learning model. The machine learning model outputs a second importance score for the feature. The feature generation module 220 compares the importance coefficient with the second importance score to determine whether to adjust the importance coefficient. For example, the feature generation module 220 can change the importance coefficient to the average of the first importance coefficient and the second importance coefficient.
[0034] The feature generation module 220 then sends the group of features and their importance factors to a modeling application, such as modeling application 150, to train a machine learning model.
[0035] In some embodiments, feature generation module 220 may generate additional features based on an incremental approach. For example, feature generation module 220 receives new primitives added by a user through primitive selection module 210, e.g., after a group of features has been generated and their importance factors determined. Feature generation module 220 generates the additional features, evaluates the additional features, and / or determines importance factors for the additional features based on the new primitives without changing the group of features that were generated and evaluated.
[0036] The metadata generation module 230 generates metadata associated with the primitives used to synthesize the features in the group. The primitive metadata indicates how useful the primitive is to the dataset. The metadata generation module 230 may generate the primitive metadata based on the usefulness score and / or importance coefficient of the features generated from the primitive. The metadata may be used by the primitive selection module 210 in subsequent feature engineering processes to select primitives for other datasets and / or different predictions. The metadata generation module 230 may search for representative vectors for the primitives used to synthesize the features in the group and feed the representative vectors and primitives back into a machine learning model used to select primitives based on the representative vectors to further train the machine learning model.
[0037] In some embodiments, the metadata generation module 230 generates natural language descriptions of the features in the group, including information describing attributes of the feature, such as the algorithm included in the feature, the results of applying the feature to data, and the function of the feature.
[0038] Database 240 stores data associated with feature engineering application 200, such as data received, used, and generated by feature engineering application 200. For example, database 240 stores data sets received from data sources, primitives, features, feature importance factors, random forests used to determine feature utility scores, machine learning models for selecting primitives and determining feature importance factors, metadata generated by metadata generation module 230, etc.
[0039] 3 is a block diagram illustrating a feature generation module 300, according to one embodiment. Feature generation module 300 is one embodiment of feature generation module 220 of FIG. 2. It generates features based on a dataset for training a machine learning model. Feature generation module 300 includes a synthesis module 310, an evaluation module 320, a ranking module 330, and a completion module 340. Those skilled in the art will recognize that other embodiments can have different and / or other components than those described herein, and that functionality can be distributed among the components in different ways.
[0040] The synthesis module 310 synthesizes multiple features based on the dataset and the primitives selected for the dataset. For each primitive, the synthesis module 310 identifies a portion of the dataset, such as one or more columns of the dataset. For example, for a primitive with an input type of date of birth, the synthesis module 310 identifies the data in the date of birth column in the dataset. The synthesis module 310 applies the primitives to the identified columns to generate features for each row of the column. The synthesis module 310 can generate a large number of features for a dataset, such as hundreds or millions.
[0041] The evaluation module 320 determines a utility score for the combined features. The utility score for a feature indicates how useful the feature is for predictions made based on the dataset. In some embodiments, the evaluation module 320 iteratively applies different portions of the dataset to the features to evaluate the utility of the features. For example, in the first iteration, the evaluation module 320 applies a predetermined percentage (e.g., 25%) of the dataset to the features to construct a first random forest. The first random forest includes a number of decision trees. Each decision tree includes multiple nodes. Every node corresponds to a feature and includes a condition that describes how to forward the tree through the node based on the value of the feature (e.g., if a date occurs on a weekend, take one branch, otherwise take another branch). The feature for each node is determined based on information gain or Gini impurity reduction. The feature that maximizes information gain or Gini impurity reduction is selected as the split feature. The evaluation module 320 determines individual utility scores for the features based on either the information gain or Gini impurity reduction by the feature across the decision trees. Each feature's individual utility score is specific to one decision tree. After determining the feature's individual utility scores for each decision tree in the random forest, the evaluation module 320 determines a first feature utility score by combining the feature's individual utility scores. In one example, the first feature utility score is the average of the feature's individual utility scores. The evaluation module 320 removes 20% of the features with the lowest first utility scores, leaving 80% of the features. These features are referred to as the first remaining features.
[0042] In the second iteration, the evaluation module 320 applies the first remaining features to a different portion of the dataset. The different portion of the dataset may be 25% of the dataset that is different from the portion of the dataset used in the first iteration, or 50% of the dataset that includes the portion of the dataset used in the first iteration. The evaluation module 320 constructs a second random forest using the different portion of the dataset and determines a second utility score for each of the remaining features by using the second random forest. The evaluation module 320 removes 20% of the first remaining features and the remainder of the first remaining features (i.e., 80% of the first remaining features form the second remaining features).
[0043] Similarly, in each subsequent iteration, the evaluation module 320 applies the remaining features from the previous round to a different portion of the dataset, determines a usefulness score for the remaining features from the previous round, and removes some of the remaining features to obtain a smaller group of features.
[0044] The evaluation module 320 can continue the iterative process until it determines that a condition is met. The condition can be that a threshold number of features remain, that the minimum utility score of the remaining features is above a threshold, that the entire dataset has been applied to the features, that a threshold number of rounds have been completed with iterations, other conditions, or some combination thereof. The remaining features from the last round, i.e., the features not removed by the evaluation module 320, are selected to train the machine learning model.
[0045] The ranking module 330 ranks the selected features and determines an importance score for each selected feature. In some embodiments, the ranking module 330 constructs a random forest based on the selected features and the dataset. The ranking module 330 determines individual ranking scores for the selected features based on each decision tree in the random forest and obtains an average of the individual ranking scores as the ranking score for the selected feature. The ranking module 330 determines an importance coefficient for the selected features based on the ranking scores. For example, the ranking module 330 ranks the selected features based on their ranking scores and determines that the importance score of the highest-ranked selected feature is 1. Then, the ranking module 330 determines the ratio of the ranking score of each of the remaining selected features to the ranking score of the highest-ranked selected feature as the importance coefficient for the corresponding selected feature.
[0046] The completion module 340 completes the selected features. In some embodiments, the completion module 340 re-ranks the selected features to determine a second ranking score for each of the selected features. In response to determining that a feature's second ranking score differs from its initial ranking score, the completion module 340 can remove the feature from the group, generate feature metadata indicating uncertainty in the feature's importance, and alert the end user to the discrepancy and uncertainty.
[0047] 4 is a flowchart illustrating a method 400 for generating a machine learning model, according to one embodiment. In some embodiments, the method is performed by feature engineering application 140, although some or all of the operations in the method may be performed by other entities in other embodiments. In some embodiments, the operations in the flowchart are performed in a different order and include different and / or additional steps.
[0048] The feature engineering application 140 receives 410 a data set from a data source, for example, one of the data sources 120 .
[0049] The feature engineering application 140 selects 420 primitives from a pool of primitives based on the received dataset. Each selected primitive is configured to be applied to at least a portion of the dataset to synthesize one or more features. In some embodiments, the feature engineering application 140 selects primitives by generating a semantic representation of the dataset and selecting primitives associated with attributes that match the semantic representation of the dataset. Additionally or alternatively, the feature engineering application 140 generates representative vectors for the dataset and inputs the representative vectors to a machine learning model. The machine learning model outputs the selected primitives based on the vectors.
[0050] The feature engineering application 140 synthesizes 430 multiple features based on the selected primitives and the received dataset. The feature engineering application 140 applies each of the selected primitives to an associated portion of the dataset to synthesize a feature. For example, for each selected primitive, the feature engineering application 140 identifies one or more variables in the dataset and applies the primitive to the variables to generate a feature.
[0051] The feature engineering application 140 iteratively evaluates 440 the plurality of features and removes some features from the plurality of features to obtain a subset of features. In each iteration, the feature engineering application 140 evaluates the usefulness of at least some of the plurality of features by applying a different portion of the dataset to the evaluated features and removes some of the evaluated features based on the usefulness of the evaluated features.
[0052] The feature engineering application 140 determines 450 an importance factor for each feature in the subset of features. In some embodiments, the feature engineering application 140 constructs a random forest based on the subset of features and at least a portion of the dataset to determine the importance factor for the subset of features.
[0053] The feature engineering application 140 generates 460 a machine learning model based on the subset of features and the importance coefficients of each feature in the subset of features. The machine learning model is configured to be used to make predictions based on new data.
[0054] 5 is a flowchart illustrating a method 500 for training a machine learning model and making predictions using the trained model, according to one embodiment. In some embodiments, the method is performed by feature engineering application 140, although some or all of the operations in the method may be performed by other entities in other embodiments. In some embodiments, the operations in the flowchart are performed in a different order and include different and / or additional steps.
[0055] The modeling application 150 trains 510 a model based on the features and feature importance coefficients. In some embodiments, the features and importance coefficients are generated by the feature engineering application 140, for example, by using the method 400 described above. The modeling application 150 may use different machine learning techniques in different embodiments. Examples of machine learning techniques include linear support vector machines (linear SVMs), boosting other algorithms (e.g., AdaBoost), neural networks, logistic regression, naive Bayes, memory-based learning, random forests, bag trees, decision trees, boosted trees, or boosted stumps, etc.
[0056] Modeling application 150 receives 520 a dataset from a data source (e.g., data source 120) associated with an enterprise. The enterprise can belong to one or more of a variety of industries, such as manufacturing, retail, finance, banking, etc. In some embodiments, modeling application 150 tailors the trained model to the needs of the particular industry. For example, if the trained model is to recognize fraudulent financial transactions, modeling application 150 tailors the trained model to emphasize fraudulent transactions that are more significant to reflect the needs of the enterprise (e.g., high-value transactions), for example, by transforming the predicted probabilities in a manner that emphasizes the more significant transactions.
[0057] The modeling application 150 obtains 530 values for the features from the received dataset. In some embodiments, the modeling application 150 retrieves 530 values for the features from the dataset, for example, in embodiments where the features are variables included in the dataset. In some embodiments, the modeling application 150 obtains the values for the features by applying the primitives used to synthesize the features to the dataset.
[0058] The modeling application 150 inputs 540 the feature values into the trained model, which outputs a prediction, which can be a prediction of whether a customer will make a transaction within a certain period of time, whether a transaction is fraudulent, whether a user will perform a computer-based interaction, etc.
[0059] FIG. 6 is a high-level block diagram illustrating a functional view of an exemplary computer system 600 for use as the machine learning server 110 of FIG. 1, according to one embodiment.
[0060] The illustrated computer system includes at least one processor 602 coupled to a chipset 604. The processor 602 may include multiple processor cores on the same die. The chipset 604 includes a memory controller hub 620 and an input / output (I / O) controller hub 622. The memory 606 and the graphics adapter 612 are coupled to the memory controller hub 620, and the display 618 is coupled to the graphics adapter 612. The storage device 608, the keyboard 610, the pointing device 614, and the network adapter 616 may be coupled to the I / O controller hub 622. In some alternative embodiments, the computer system 600 may have additional, fewer, or different components, and the components may be combined differently. For example, an embodiment of the computer system 600 may lack a display and / or a keyboard. Additionally, the computer system 600 may be instantiated as a rack-mounted blade server or as a cloud server instance in some embodiments.
[0061] The memory 606 holds instructions and data used by the processor 602. In some embodiments, the memory 606 is a random access memory. The storage device 608 is a non-transitory computer-readable storage medium. The storage device 608 may be an HDD, an SSD, or another type of non-transitory computer-readable storage medium. Data processed and analyzed by the machine learning server 110 may be stored in the memory 606 and / or the storage device 608.
[0062] Pointing device 614 may be a mouse, trackball, or other type of pointing device and is used in combination with keyboard 610 to input data into computer system 600. Graphics adapter 612 displays images and other information on display 618. In one embodiment, display 618 includes touch screen functionality for receiving user input and selections. Network adapter 616 couples computer system 600 to network 160.
[0063] The computer system 600 is adapted to execute computer modules to provide the functionality described herein. As used herein, the term "module" refers to computer program instructions and other logic for providing a particular function. A module can be implemented in hardware, firmware, and / or software. A module can include one or more processes and / or can be provided by only a portion of a process. Modules are typically stored in the storage device 608, loaded into the memory 606, and executed by the processor 602.
[0064] The particular naming of components, term capitalization, attributes, data structures, or other programming or structural aspects are not required or important, and mechanisms for implementing the described embodiments may have different names, formats, or protocols. Furthermore, the system may be implemented through a combination of hardware and software as described, or entirely with hardware elements. Also, the particular division of functionality among various system components described herein is merely exemplary and not required. Functions performed by a single system component may instead be performed by various components, and functions performed by various components may instead be performed by a single component.
[0065] Some portions of the above description are characterized in terms of algorithms and symbolic representations of operations of information. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. These operations, while described in functional or logical terms, will be understood to be implemented by computer programs. Further, without loss of generality, it is sometimes convenient to refer to the arrangement of these operations in terms of modules or functional names.
[0066] As is apparent from the above description, unless otherwise stated, throughout the description, descriptions utilizing terms such as "processing" or "computing" or "calculating" or "determining" or "displaying" relate to the actions and processes of a computer system or similar electronic computing device that operates on and transforms data represented as physical (electronic) quantities in the memory or registers of the computer system or other information storage, transmission or display device.
[0067] Certain embodiments described herein include process steps and instructions that are described in the form of algorithms. It should be noted that the process steps and instructions of the embodiments may be implemented in software, firmware, or hardware, and if implemented in software, may be downloaded to reside on and operate from different platforms used in real-time network operating systems.
[0068] Finally, it should be noted that the language used in the specification has been chosen primarily for ease of reading and descriptive purposes, and not to delineate or limit the subject matter of the present invention. Accordingly, the disclosure of embodiments is intended to be illustrative, but not limiting.
Claims
1. 1. A computer-implemented method comprising: receiving a data set from a data source; selecting primitives from a pool of primitives based on the received dataset, each of the selected primitives configured to be applied to at least a portion of the dataset to synthesize one or more features; synthesizing a plurality of features by applying the selected primitives to the received dataset; Iteratively evaluating the plurality of features and removing some features from the plurality of features to obtain a subset of features, each iteration comprising: evaluating the usefulness of at least some of the features by applying different portions of the dataset to the evaluated features; an evaluating step including removing some of the evaluated features based on the usefulness of the evaluated features to generate the subset of features; determining an importance factor for each feature of said subset of features; generating a machine learning model based on the subset of features and the importance coefficients of each feature in the subset of features, wherein the machine learning model is configured to be used to make predictions based on new data.
2. selecting the primitive from the plurality of primitives based on the received data set, generating a semantic representation of the received dataset; selecting primitives associated with attributes that match the semantic representation of the received data set; 10. The method of claim 1, comprising:
3. selecting the primitive from the plurality of primitives based on the received data set, generating a representative vector from the received data set; inputting the representative vectors into a machine learning model, the machine learning model outputting the selected primitives based on the representative vectors; The method of claim 1 further comprising:
4. The step of iteratively evaluating the plurality of features and removing some features from the plurality of features to obtain the subset of features includes: applying the plurality of features to a first portion of the dataset and determining a first utility score for each of the plurality of features; removing a portion of the plurality of features based on the first usefulness score for each of the plurality of features to obtain a preliminary subset of features; applying the preliminary subset of features to a second portion of the dataset and determining a second utility score for each of the preliminary subset of features; removing a portion of the preliminary subset of features from the preliminary subset of features based on a second utility score for each of the preliminary subset of features; 10. The method of claim 1, comprising:
5. Determining the importance factor for each of the subset of features comprises: ranking the subset of features by inputting the subset of features and a first portion of the dataset into a machine learning model, the machine learning model outputting a first ranking score for each of the subset of features; determining the importance factors of the subset of features based on their ranking scores; 10. The method of claim 1, comprising:
6. ranking the subset of features by inputting the subset of features and a second portion of the dataset into a machine learning model, the machine learning model outputting a second ranking score for each of the subset of features; determining a second importance factor for each of the subset of features based on the ranking scores of the features; adjusting the importance score of each of the subset of features based on a second importance factor of the feature; The method of claim 5 further comprising:
7. synthesizing the plurality of features based on the subset of primitives and the received dataset, For each primitive in the subset: identifying one or more variables in the dataset; applying the primitives to the one or more variables to generate one or more features of the plurality of features; 10. The method of claim 1, comprising:
8. 1. A system comprising: a computer processor for executing computer program instructions; a non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations; and the operation comprises: receiving a data set from a data source; selecting primitives from a pool of primitives based on the received dataset, each of the selected primitives configured to be applied to at least a portion of the dataset to synthesize one or more features; synthesizing a plurality of features by applying the selected primitives to the received dataset; and iteratively evaluating the plurality of features and removing some features from the plurality of features to obtain a subset of features, each iteration comprising: evaluating the usefulness of at least some of the features by applying different portions of the dataset to the evaluated features; removing some of the evaluated features based on the usefulness of the evaluated features to generate the subset of features; determining an importance factor for each feature of said subset of features; generating a machine learning model based on the subset of features and the importance coefficients for each feature in the subset of features, the machine learning model being configured to be used to make predictions based on new data. A system with.
9. Selecting the primitive from the plurality of primitives based on the received data set includes: generating a semantic representation of the received dataset; selecting primitives associated with attributes that match the semantic representation of the received data set; The system of claim 8 , comprising:
10. Selecting the subset of primitives from the plurality of primitives based on the received data set includes: generating a representative vector from the received data set; inputting the representative vector into a machine learning model, the machine learning model outputting the selected primitive based on the representative vector; The system of claim 8 further comprising:
11. Iteratively evaluating the plurality of features and removing some features from the plurality of features to obtain the subset of features includes: applying the plurality of features to a first portion of the dataset and determining a first utility score for each of the plurality of features; removing a portion of the plurality of features based on the first usefulness score for each of the plurality of features to obtain a preliminary subset of features; applying the preliminary subset of features to a second portion of the dataset and determining a second utility score for each of the preliminary subset of features; removing a portion of the preliminary subset of features from the preliminary subset of features based on a second utility score for each of the preliminary subset of features; The system of claim 8 , comprising:
12. Determining the importance factor for each of the subset of features includes: ranking the subset of features by inputting the subset of features and a first portion of the dataset into a machine learning model, the machine learning model outputting a first ranking score for each of the subset of features; determining the importance factors of the subset of features based on their ranking scores; The system of claim 8 , comprising:
13. The operation is ranking the subset of features by inputting the subset of features and a second portion of the dataset into a machine learning model, the machine learning model outputting a second ranking score for each of the subset of features; determining a second importance factor for each of the subset of features based on the ranking scores of the features; adjusting the importance score of each of the subset of features based on a second importance factor of the feature; The system of claim 12 further comprising:
14. synthesizing the plurality of features based on the subset of primitives and the received dataset, For each primitive in the subset: identifying one or more variables in the dataset; applying the primitives to the one or more variables to generate one or more features of the plurality of features; The system of claim 8 , comprising:
15. 1. A non-transitory computer-readable memory storing executable computer program instructions for processing data blocks in a data analysis system, the instructions comprising: receiving a data set from a data source; selecting primitives from a pool of primitives based on the received dataset, each of the selected primitives configured to be applied to at least a portion of the dataset to synthesize one or more features; synthesizing a plurality of features by applying the selected primitives to the received dataset; and iteratively evaluating the plurality of features and removing some features from the plurality of features to obtain a subset of features, each iteration comprising: evaluating the usefulness of at least some of the features by applying different portions of the dataset to the evaluated features; removing some of the evaluated features based on the usefulness of the evaluated features to generate the subset of features; determining an importance factor for each feature of said subset of features; generating a machine learning model based on the subset of features and the importance coefficients of each feature in the subset of features, the machine learning model being configured to be used to make predictions based on new data; a non-transitory computer-readable memory executable to perform operations including:
16. Selecting the primitive from the plurality of primitives based on the received data set includes: generating a semantic representation of the received dataset; selecting primitives associated with attributes that match the semantic representation of the received data set; 16. The non-transitory computer-readable memory of claim 15, comprising:
17. Selecting the primitive from the plurality of primitives based on the received data set includes: generating a representative vector from the received data set; inputting the representative vector into a machine learning model, the machine learning model outputting the selected primitive based on the representative vector; 16. The non-transitory computer-readable memory of claim 15, further comprising:
18. Iteratively evaluating the plurality of features and removing some features from the plurality of features to obtain the subset of features includes: applying the plurality of features to a first portion of the dataset and determining a first utility score for each of the plurality of features; removing a portion of the plurality of features based on the first usefulness score for each of the plurality of features to obtain a preliminary subset of features; applying the preliminary subset of features to a second portion of the dataset and determining a second utility score for each of the preliminary subset of features; 16. The non-transitory computer-readable memory of claim 15, further comprising removing a portion of the preliminary subset of features from the preliminary subset of features based on a second usefulness score for each of the preliminary subset of features.
19. Determining the importance factor for each of the subset of features includes: ranking the subset of features by inputting the subset of features and a first portion of the dataset into a machine learning model, the machine learning model outputting a first ranking score for each of the subset of features; determining the importance factors of the subset of features based on their ranking scores; 16. The non-transitory computer-readable memory of claim 15, comprising:
20. The operation is ranking the subset of features by inputting the subset of features and a second portion of the dataset into a machine learning model, the machine learning model outputting a second ranking score for each of the subset of features; determining a second importance factor for each of the subset of features based on the ranking scores of the features; adjusting the importance score of each of the subset of features based on a second importance factor of the feature; 20. The non-transitory computer-readable memory of claim 19, further comprising:
Citation Information
Patent Citations
Feature amount generation device and feature amount generation method
JP2020013511A
Methods and systems for sequential feature selection based on significance testing
US9189750B1