Integrated Future Engineering
The integrated feature engineering method addresses inefficiencies in current tools by generating primitives and synthesizing features through entity sets, enhancing data processing efficiency and user interaction.
Patent Information
- Application Number
- JP2023539955
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-12-30
- Filing Date
- 2021-12-16
- Publication Date
- 2025-07-28
- Estimated Expiration
- 2041-12-16
AI Technical Summary
Current feature engineering tools face architectural issues that make data processing inefficient and time-consuming, requiring significant data volumes and complicating interactions with the feature creation process.
A method and system for integrated feature engineering that includes generating primitives from data entities, applying time parameters, and synthesizing features through entity sets, enabling efficient data processing and automated feature creation without the need for extensive data movement.
Facilitates more efficient data processing by reducing complexity and time requirements, allowing for automated feature engineering and entity set creation, even with smaller data volumes, and improving user interaction.
Smart Images

Figure 0007714039000001 
Figure 0007714039000002 
Figure 0007714039000003
Abstract
Description
Technical Field
[0001] The described aspects generally relate to processing data streams, and more particularly to integrated feature engineering where automated feature engineering for performing machine learning on stream data is integrated with entity set creation.
Background Art
[0002] Feature engineering is a process related to identifying and extracting predictable features in complex data typically analyzed by companies and other enterprises. Features are important for the accuracy of predictions by machine learning models. Therefore, feature engineering is often a determining factor in the success of data analysis projects. Feature engineering is typically a time-consuming process and typically requires a significant amount of data to achieve good prediction accuracy. Often, the data used to create features comes from different sources and data combination is required before feature engineering. However, there are architectural issues with data movement between tools for combining data and tools for creating features, causing the process of creating features to be even more time-consuming. Additionally, the architectural issues make it more difficult for data analysis engineers to interact with the feature creation process. Therefore, current feature engineering tools cannot efficiently contribute to the data processing needs of enterprises.
Summary of the Invention
[0003] The foregoing and other problems are addressed by a method, a computer-implemented system, and a computer-readable memory. Aspects of the method include receiving a plurality of data entities from different data sources. The plurality of data entities are used to train a model that makes predictions based on new data. The method further includes generating primitives based on the plurality of data entities. Each of the primitives is configured to be applied to variables of the plurality of data entities to synthesize features. The method further includes receiving a time parameter from a client device associated with a user. The time parameter specifies a time value and is to be used to synthesize features from the plurality of data entities. The method further includes generating an entity set by aggregating the plurality of data entities after the primitives are generated and the time parameter is received. The method further includes synthesizing a plurality of features based on the entity set, the primitives, and the time parameter. Further, the method also includes training a model based on the plurality of features.
[0004] Aspects of a computer-implemented system include a computer processor for executing computer program instructions. Further, the system also includes a non-transitory computer-readable memory storing computer program instructions executable by the operating computer processor. The operations include receiving a plurality of data entities from different data sources. The plurality of data entities are used to train a model that makes predictions based on new data. The operations further include generating primitives based on the plurality of data entities. Each of the primitives is configured to be applied to variables of the plurality of data entities to synthesize features. The operations further include receiving a time parameter from a client device associated with a user. The time parameter specifies a time value and is to be used to synthesize features from the plurality of data entities. The operations further include generating an entity set by aggregating the plurality of data entities after the primitives are generated and the time parameter is received. The operations further include synthesizing a plurality of features based on the entity set, the primitives, and the time parameter. Further, the operations also include training a model based on the plurality of features.
[0005] Aspects of non-transitory computer-readable memories store executable computer program instructions. The instructions are executable to perform operations. The operations include receiving a plurality of data entities from different data sources. The plurality of data entities are used to train a model that makes predictions based on new data. The operations further include generating primitives based on the plurality of data entities. Each of the primitives is configured to be applied to variables of the plurality of data entities to synthesize features. The operations further include receiving a time parameter from a client device associated with a user. The time parameter specifies a time value and is to be used to synthesize features from the plurality of data entities. The operations further include generating an entity set by aggregating the plurality of data entities after the primitives are generated and the time parameter is received. The operations further include synthesizing a plurality of features based on the entity set, the primitives, and the time parameter. Further, the operations include training a model based on the plurality of features.
[0006] The drawings depict various aspects for illustrative purposes only. Those skilled in the art will immediately recognize that alternative aspects of the structures and methods illustrated herein may be utilized without departing from the principles of the aspects described herein. In the various drawings, like reference numerals and names indicate like elements.
Brief Description of the Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
DETAILED DESCRIPTION OF THE INVENTION
[0008] FIG. 1 is a block diagram illustrating a machine learning environment 100 including a machine learning server 110 according to one aspect. The environment 100 further includes a plurality of data sources 120 connected to the machine learning server 110 via a network 130. The illustrated environment 100 includes only one machine learning server 110 coupled to a plurality of data sources 120, but aspects are possible having a plurality of machine learning servers and a single data source.
[0009] The data source 120 is Machine learning serverProvide electronic data to 110. The data source 120 can be, for example, a storage device such as a hard disk drive (HDD) or a solid state drive (SSD), a computer that manages and provides access to multiple storage devices, a storage area network (SAN), a database, or a cloud storage system. Additionally, the data source 120 can be a computer system capable of retrieving data from another source. Different data sources may be associated with different users, different organizations, or different departments within the same organization. The data source 120 is remote from the machine learning server 110 and may provide data via the network 130. In addition, some or all of the data sources 120 may be directly connected to the data analysis system and provide data without passing the data through the network 130.
[0010] The data provided by the data source 120 may be organized into data records (e.g., rows). Each data record contains one or more values. For example, the data records provided by the data source 120 may contain a series of comma-separated values. The data Machine learning server describes information related to the enterprise that uses 110. For example, the data from the data source 120 can describe computer-based interactions with content accessible on a website and / or application (e.g., click tracking data). The enterprise is one or more of various industries such as, for example, computer technology and manufacturing.
[0011] The machine learning server 110 is a computer-based system utilized to provide a machine learning model that can be used to build a machine learning model and make predictions based on data. Exemplary predictions include application monitoring, network traffic data flow monitoring, user action prediction, and the like. Data is collected, gathered, or otherwise accessed from a plurality of data sources 120 via the network 130. The machine learning server 110 can implement scalable software tools and hardware resources employed when accessing, preparing, mixing, and analyzing data from a wide variety of data sources 120. The machine learning server 110 can be a computing device used to implement machine learning functions including the feature engineering and modeling techniques described herein.
[0012] In FIG. 1, the machine learning server 110 can be configured to support one or more software applications, exemplified as the feature engineering application 150 and the training application 160. The feature engineering application 150 performs integrated feature engineering where automated feature engineering is integrated with the creation of entity sets. The integrated feature engineering process begins by generating feature engineering algorithms and parameters from individual data entities, then proceeds to create entity sets by combining individual data entities, and further proceeds to extract predictor variables, i.e., features, from the data of the entity sets. Each feature is a variable that is potentially related to the prediction (referred to as the target prediction) that the corresponding machine learning model will make.
[0013] The Feature Engineering Application 150 generates primitives based on individual data entities. In one aspect, the Feature Engineering Application 150 selects a primitive from a pool of primitives based on individual data entities. The pool of primitives is maintained by the Feature Engineering Application 150. A primitive defines an individual calculation that can be applied to the raw data of a dataset to create one or more new features with associated values. The selected primitives can be applied and stacked across different types of data to create new calculations so as to constrain the input and output data types. The Feature Engineering Application 150 enables a user (such as a data analysis engineer) to provide time values that can later be used by the Feature Engineering Application 150 to create time-based features. A time-based feature is a feature extracted from data associated with a specific time or time period. A time-based feature may be used, for example, to train a model to make time-based predictions such as predictions for a specific time or time period.
[0014] After the Feature Engineering Application 150 generates primitives and receives time parameters from the user, it combines individual data entities to generate an entity set. In some aspects, the Feature Engineering Application 150 generates an entity set based on variables in individual data entities. For example, the Feature Engineering Application 150 identifies two individual data entities having a common variable, determines a parent-child relationship between them, and generates an intermediate data entity based on the parent-child relationship. The Feature Engineering Application 150 combines the intermediate data entities to generate an entity set.
[0015] After the entity set is generated, feature engineering application 150 synthesizes features by applying primitive and temporal parameters to the data of the entity set. Next, through an iterative process of applying different parts of the data to the features in each iteration, the features are evaluated to determine the importance of each feature. Feature engineering application 150 removes some of the features in each iteration to obtain a subset of features that are more useful for prediction than the removed features. For each feature in the subset, feature engineering application 150 determines an importance coefficient, for example, using a random forest. The importance coefficient indicates how important / how relevant a feature is to the target prediction. The features of the subset and their importance coefficients can be sent to training application 160 that builds a machine learning model.
[0016] Compared with conventional feature engineering tools, the architecture of feature engineering application 150 facilitates more efficient data processing and provides a less complex experience to users who can provide input to the feature engineering process. Feature engineering application 150 integrates automated feature engineering and entity set creation in a way that enables data input (from either data source 120 or the user) for both automated feature engineering and entity set creation at the start of integrated processing, preventing the need to request data during processing. In this way, feature engineering application 150 can operate efficiently because it does not need to wait for responses from other entities during operation. Additionally, all user inputs (including editing primitives and providing time parameters) at the start of the integrated feature engineering process also provide a better experience to the user. The user does not need to monitor the remaining processing. Therefore, it overcomes the challenges faced by conventional feature engineering tools.
[0017] Another advantage of the feature engineering application 150 is that the use of primitives makes the feature engineering process more efficient than the conventional feature engineering process in which features are extracted from raw data. Further, the feature engineering application 150 can evaluate primitives based on the evaluation and importance factors of the feature(s) generated from the primitives. It is possible to generate metadata that describes the evaluation of the primitives and use the metadata to determine whether to select primitives for different data or different prediction problems. The conventional feature engineering process could generate a huge number (e.g., millions, etc.) of features without providing any guidance or solution for engineering faster and better features. Yet another advantage of the feature engineering application 150 is that it does not require a large amount of data to evaluate features. Instead, an iterative approach for evaluating features is applied, using different portions of the data in each iteration.
[0018] The training application 160 trains a machine learning model based on the features received from the feature engineering application 150 and the importance factors of the features. Different machine learning techniques, such as linear support vector machine (linear SVM), boosting of other algorithms (e.g., AdaBoost), neural network, logistic regression, naive Bayes, memory-based learning, random forest, bagged tree, decision tree, boosted tree, or boosted stump, etc., may be used in different manners. The generated machine learning model performs target prediction when applied to features extracted from a new dataset (e.g., a dataset from the same or different data source 120).
[0019] In some aspects, the training application 160 validates the predictions before deploying the trained model to a new dataset. For example, the training application 160 applies the trained model to a validation dataset to quantify the accuracy of the model. Common metrics applied to accuracy measurements include Precision = TP / (TP + FP) and Recall = TP / (TP + FN), where precision is the number of results (TP or true positives) correctly predicted by the model out of the total number of predictions made by the model (TP + FP or false positives), and recall is the number of results (TP) correctly predicted by the model out of the total number that actually occurred (TP + FN or false negatives). The F-score (F-score = 2*PR / (P + R)) unifies precision and recall into a single measure. In one aspect, the training application 160 repeatedly retrains the machine learning model until the occurrence of a stopping condition, such as an indication of an accuracy measurement indicating that the machine learning model is sufficiently accurate, or the number of training rounds performed.
[0020] Network 130 represents the communication path between machine learning server 110 and data source 120. In one aspect, network 130 is the Internet and uses standard communication techniques and / or protocols. Thus, network 130 can include links using technologies such as Ethernet, 802.11, WiMAX (worldwide interoperability for microwave access), 3G, LTE (Long Term Evolution), digital subscriber line (DSL), asynchronous transfer mode (ATM), InfiniBand, PCI Express Advanced Switching, etc. Similarly, the networking protocols used in network 130 can include multi-protocol label switching (MPLS), TCP / IP (transmission control protocol / Internet protocol), UDP (User Datagram Protocol), HTTP (hypertext transport protocol), SMTP (simple mail transfer protocol), file transfer protocol (FTP), etc.
[0021] The data exchanged via network 130 can be represented using technologies and / or formats including HTML (hypertext markup language), XML (extensible markup language), etc. Additionally, all or part of the links can be encrypted using conventional encryption techniques such as SSL (secure sockets layer), TLS (transport layer security), virtual private network (VPN), IPsec (Internet Protocol security), etc. In another aspect, an entity can use custom and / or proprietary data communication technologies instead of or in addition to those described above.
[0022] FIG. 2 is a block diagram illustrating a feature engineering application 200 according to one aspect. The feature engineering application 200 receives a plurality of data entities from different data sources and synthesizes features from the data entities. The feature engineering application 200 is an aspect of the feature engineering application 150 of FIG. 1. The feature engineering application 200 includes a primitive generation module 210, a time parameter module 220, an entity feature module 230, a feature synthesis module 240, and a database 250. Those skilled in the art will recognize that other aspects may be different from and / or have other components than those described herein, and that functionality may be distributed among components in different ways.
[0023] The primitive generation module 210 generates a list of primitives based on the received data entities and enables the user to edit the list. In some aspects, the primitive generation module 210 generates a list of primitives by selecting primitives from a pool of primitives maintained by the feature engineering application 200 based on the data entities. In some aspects, the primitive generation module 210 obtains information indicating one or more relationships between the data entities before generating the list of primitives. The primitive generation module 210 selects the primitives of the list from the pool of primitives based on the data entities as well as the information indicating the relationships between the data entities. The information may indicate one or more parent-child relationships between the data entities. The primitive generation module 210 may obtain the information by receiving the information from the user or by determining one or more parent-child relationships, for example, using the methods described below in connection with the entity feature module 230.
[0024] A primitive pool contains a number of primitives, such as hundreds or thousands of primitives for example. Each primitive includes an algorithm that, when applied to data, performs a calculation on the data and generates a feature with an associated value. A primitive is associated with one or more attributes. The attributes of a primitive can be a description of the primitive (e.g., a natural language description specifying the calculation performed by the primitive when applied to data), an input type (i.e., the type of input data), a return type (i.e., the type of output data), metadata of the primitive indicating how useful the primitive was in previous feature engineering processes, or other attributes.
[0025] In some aspects, the primitive pool contains multiple different types of primitives. One type of primitive is an aggregation primitive. An aggregation primitive, when applied to a dataset, identifies relevant data of the dataset, makes a determination on the relevant data, and creates a value that summarizes and / or aggregates the determination. For example, the aggregation primitive "count" identifies the values of relevant rows of a dataset, determines whether each of the values is a non-NULL value, and returns (outputs) the number of non-NULL values in the rows of the dataset. Another type of primitive is a transformation primitive. A transformation primitive, when applied to a dataset, creates a new variable from one or more existing variables of the dataset. For example, the transformation primitive "weekend" evaluates the timestamps of a dataset and returns a binary value (e.g., true or false) indicating whether the date indicated by the timestamp is a weekend. Another exemplary transformation primitive evaluates a timestamp and returns a count indicating the number of days until a specified date (e.g., the number of days until a particular holiday).
[0026] In some embodiments, the primitive generation module 210 selects primitives based on the received data entities using a skim view approach, a summary view approach, or both approaches. In the skim view approach, the primitive generation module 210 identifies one or more semantic representations of each data entity. A semantic representation of a data entity describes the characteristics of the data entity and may be obtained without performing calculations on the data of the data entity. Examples of semantic representations of a data set include the presence of one or more specific variables (e.g., column names) of the data entity, the number of columns, the number of rows, the input type of the data entity, other attributes of the data entity, and combinations thereof. To select a primitive using the skim view approach, the primitive generation module 210 determines whether the identified semantic representation of the data entity matches the attributes of the primitives in the pool. If there is a match, the primitive generation module 210 selects the primitive.
[0027] The skim viewer approach is a rule-based analysis. The determination of whether the identified semantic representation of a data entity matches a primitive attribute is based on rules maintained by the feature engineering application 200. The rules specify, for example, which semantic representations of data entities match which primitive attributes based on the matching of keywords in the semantic representation of the data entity and the keywords of the primitive attributes. As an example, if the semantic representation of a data entity is the column name "Date of Birth", the primitive generation module 210 selects a primitive with an input type of "Date of Birth" that matches the semantic representation of the dataset. In another example, if the semantic representation of a dataset is the column name "Timestamp", the primitive generation module 210 selects a primitive that has an attribute indicating that the primitive is suitable for use with data indicating a timestamp.
[0028] In the summary viewer approach, the primitive generation module 210 generates a representative vector from the data entity. The representative vector encodes data that describes the data entity, such as, for example, the number of tables in the data entity, the number of columns per table, the average of each column, and data indicating the average of each row. Thus, the representative vector serves as a fingerprint of the dataset. The fingerprint is a compact representation of the dataset and may be generated by applying one or more fingerprint functions to the dataset, such as, for example, a hash function, Rabin's fingerprint algorithm, or other types of fingerprint functions.
[0029] The primitive generation module 210 selects a primitive for a data entity based on a representative vector. For example, the primitive generation module 210 inputs the representative vector of the data entity into a machine learning model. The machine learning model outputs the primitives of the data set. The machine learning model is trained, for example, by the primitive selection module 210, to select the primitives of the data set based on the representative vector. It may be trained based on training data including a plurality of representative vectors of a plurality of training data entities and a set of primitives for each of the plurality of training data entities. The set of primitives for each of the plurality of training data entities was used to generate features determined to be useful for making predictions based on the corresponding training data set. In some embodiments, the machine learning model is continuously trained. For example, the primitive generation module 210 can further train the machine learning model based on the representative vector of the data entity and at least some of the selected primitives.
[0030] In some embodiments, the primitive generation module 210 generates primitives based also on input from a user (e.g., a data analysis engineer). The primitive generation module 210 provides the user with a list of primitives selected from a pool of primitives based on the data entity for display in a user interface. The user interface enables the user to edit the primitives, such as adding other primitives to the list of primitives, creating new primitives, removing primitives, performing other types of actions, or combining, etc. The primitive generation module 210 updates the list of primitives based on the user's edits. Thus, the primitives generated by the primitive generation module 210 incorporate the user's input. Details regarding user information are described below in connection with FIG. 4.
[0031] The time parameter module 220 determines one or more cut-off times based on one or more time values. The cut-off time is the time at which the prediction is made. Data associated with timestamps before the cut-off time can be used to extract features for the label. However, data associated with timestamps after the cut-off time should not be used to extract features for the label. The cut-off time can be specific to a subset of the primitives generated by the primitive generation module 210 or can be global for all generated primitives. In some aspects, the time parameter module 220 receives time values from the user. The time value can be a timestamp or a time period. The time parameter module 220 enables the user to provide different time values for different primitives or the same time value for multiple primitives.
[0032] In some aspects, the primitive generation module 210 and the time parameter module 220, separately or together, provide a graphical user interface (GUI) that enables the user to edit primitives and enter time values. An example of the GUI provides tools that enable the user to view, search, edit the primitives generated by the primitive generation module 210, and provide one or more values of the time parameters of the primitive. In some aspects, the GUI enables the user to select the number of data entities to use for feature synthesis. The GUI provides the user with the option to control whether all or a portion of the data entities received by the feature engineering application 200 are used for feature synthesis. Details regarding the GUI are described in connection with FIG. 3.
[0033] The entity feature module 230 synthesizes features based on data entities received from the feature engineering application 200, primitives generated by the primitive generation module 210, and a cut-off time determined by the time parameter module 220. After the primitives are generated and the time parameters are received, the entity feature module 230 creates an entity set from the data entities received by the feature engineering application 200. In some aspects, the entity feature module 230 determines variables of the data entities and identifies two or more data entities sharing a common variable as a subset. The entity feature module 230 may identify subsets of multiple data entities from the data entities received from the feature engineering application 200.
[0034] For each subset, the entity feature module 230 determines the parent-child relationships of the subset. Using the data entities of the subset as parent entities, it identifies each of one or more other data entities of the subset as child entities. In some aspects, the entity feature module 230 determines the parent-child relationships based on the primary variables of the data entities. The primary variable of a data entity is a variable that has a value that uniquely identifies an entity (such as a user, an action, etc.). Examples of primary variables are variables associated with user identity information, action identity information, object identity information, or combinations thereof. The entity feature module 230 identifies the primary variable in each data entity and determines the hierarchy among the primary variables, for example, by performing rule-based analysis. The entity feature module 230 maintains rules that specify which variable has a higher hierarchy than other variables. For example, the rule specifies that the "user ID" variable has a higher position in the hierarchy than the "user action ID" variable.
[0035] In some embodiments, the entity feature module 230 determines the primary variables of the data entities based on user input. For example, the entity feature module 230 detects the variables of the data entities, provides the detected variables for display to the user, and provides the user with an opportunity to identify which variables are the primary variables. As another example, the entity feature module 230 determines the primary variables and provides the user with an opportunity to confirm or disapprove the determination. The entity feature module 230 combines a subset of data entities based on a parent-child relationship to generate intermediate data entities. The entity feature module 230 may support a GUI that facilitates user input. An example of the GUI is described below in connection with FIG. 5. In some embodiments, the entity feature module 230 determines the primary variables of the data entities without user input. For example, the entity feature module 230 compares the variables of a data entity with the variables of another data entity regarding the same data type. Examples of data types include numerical data types, categorical data types, time series data types, text data types, and the like. The entity feature module 230 determines a matching score for each variable of the data entity, where the matching score indicates the probability that the variable matches the variables of other data entities. The entity feature module 230 selects the variable of the data entity having the highest matching score as the primary variable of the data entity. Further, the entity feature module 230 combines the intermediate data entities to generate an entity set that incorporates all the data entities received by the feature engineering application 200.
[0036] In some aspects, the entity feature module 230 determines whether to normalize a data entity, for example, by determining whether there are multiple primary variables in the data entity. In one aspect, the entity feature module 230 determines whether there are duplicate values of different variables in the data entity. For example, if the "Colorado" value of the location variable always corresponds to the "Mountain" value of the region variable, the "Colorado" value and the "Mountain" value are duplicate values. In response to the determination to normalize the data entity, the entity feature module 230 normalizes the data entity by splitting the data entity into two new data entities. Each of the two new data entities includes a subset of the variables of the data entity. The process described above is called normalization. In some aspects of the normalization process, the entity feature module 230 identifies a first primary variable and a second primary variable from the variables of a given entity. Examples of primary variables of a data entity include, for example, user identity information, action identity information, object identity information, or combinations thereof. The entity feature module 230 classifies the variables of a given entity into a first variable group and a second variable group. The first variable group includes the first primary variable and one or more other variables of the given entity related to the first primary variable. The second variable group includes the second primary variable and one or more other variables of the given entity related to the second primary variable. The entity feature module 230 generates one of the two new entities from the variables of the first group and the values of the variables of the first group, and generates the other of the two new entities from the variables of the second group and the values of the variables of the first group. Next, the entity feature module 230 identifies a subset of the data entity from the two new data entities and the received data entity excluding the normalized data entity.
[0037] The entity feature module 230 applies primitives and cut-off times to the entity set to synthesize a group of features and importance coefficients for each feature in the group. The entity feature module 230 extracts data from the entity set based on the cut-off time, and applies primitives to the extracted data to synthesize a plurality of features. In some embodiments, the entity feature module 230 applies each of the selected primitives to at least a portion of the extracted data to synthesize one or more features. For example, the entity feature module 230 applies the "weekend" primitive to the column named "timestamp" in the entity set to synthesize a feature indicating whether the date is on a weekend. The entity feature module 230 can synthesize a large number of features for the entity set, such as hundreds or millions of features.
[0038] The entity feature module 230 evaluates the features and, based on the evaluation, removes some of the features to obtain a feature group. In some embodiments, the entity feature module 230 evaluates the features through an iterative process. In each round of the iteration, the entity feature module 230 applies the features that were not removed by the previous iteration (also referred to as "remaining features") to different parts of the extracted data to determine a usefulness score for each of the features. The entity feature module 230 removes some of the features with the lowest usefulness scores from the remaining features.
[0039] In some embodiments, the entity feature module 230 uses a random forest to determine a feature usefulness score. The feature usefulness score indicates how useful a feature is for making predictions based on an entity set. In some embodiments, the entity feature module 230 repeatedly applies different portions of the entity set to the feature to evaluate the feature's usefulness. For example, in a first iteration, the entity feature module 230 applies a predetermined percentage (e.g., 25%) of the entity set to the feature to build a first random forest. The first random forest includes a plurality of decision trees. Each decision tree includes a plurality of nodes. Each node corresponds to a feature and includes a condition that describes how to traverse the tree through the node based on the value of the feature (e.g., if the date is on a weekend, take one branch, otherwise take another branch). The feature of each node is determined based on information gain or Gini impurity reduction. The feature that maximizes the reduction of information gain or Gini impurity is selected as the splitting feature. Entity feature module 320 determines an individual usefulness score for a feature based on either information gain by the feature across the decision tree or reduction of Gini impurity. The individual usefulness score for a feature amount is specific to one decision tree. After determining the individual usefulness score for a feature for each of the decision trees of the random forest, Entity feature module 320 determines a first usefulness score for the feature by combining the individual usefulness scores for the feature. In one example, the first usefulness score for the feature is the average of the individual usefulness scores for the feature. The entity feature module 230 deletes 20% of the features with the lowest first usefulness scores so that 80% of the features remain. These features are referred to as the first remaining features.
[0040] In the second iteration, the entity feature module 230 applies the first remaining features to a different portion of the entity set. The different portion of the entity set can be 25% of the entity set that is different from the portion of the entity set used in the first iteration, or it can be 50% of the entity set that includes the portion of the entity set used in the first iteration. Entity feature module 320 constructs a second random forest using a different portion of the entity set and determines a second usefulness score for each of the remaining features using the second random forest. The entity feature module 230 removes 20% of the first remaining features and the remainder of the first remaining features (i.e., 80% of the first remaining features form the second remaining features).
[0041] Similarly, in each subsequent iteration, the entity feature module 230 applies the remaining features from the previous round to a different portion of the entity set, determines the usefulness scores of the remaining features from the previous round, and removes a portion of the remaining features to obtain a smaller group of features. The entity feature module 230 can continue the iterative process until it determines that the conditions are met. The conditions can be that a number of features below a threshold remain, the lowest usefulness score of the remaining features is above the threshold, the entire entity set has been applied to the features, a threshold number of rounds have been completed in the iteration, other conditions, or a combination thereof. The remaining features of the last round, i.e., the features not removed by the entity feature module 230, are selected for training the machine learning model.
[0042] After the iteration is completed and the feature group is obtained, the entity feature module 230 determines the importance coefficient of each feature in the group. The importance coefficient of a feature indicates how important the feature is for predicting the target variable. In some aspects, the entity feature module 230 determines the importance coefficient using a random forest, for example, one constructed based on at least a part of the entity set. In some aspects, the entity feature module 230 constructs a random forest based on the selected features and the entity set. The entity feature module 230 determines the individual ranking score of the selected features based on each decision tree of the random forest, and obtains the average of the individual ranking scores as the ranking score of the selected features. The entity feature module 230 determines the importance coefficient of the selected features based on the ranking score. For example, the entity feature module 230 ranks the selected features based on the ranking score, and determines that the importance score of the highest-ranked selected feature is 1. Next, the entity feature module 230 determines the ratio of the ranking score of each of the remaining selected features to the ranking score of the highest-ranked selected feature as the importance coefficient of the corresponding selected feature.
[0043] In some aspects, the entity feature module 230 adjusts the importance score of a feature by inputting the feature and a different part of the entity set into a machine learning model. The machine learning model outputs a second importance score of the feature. The entity feature module 230 compares the importance coefficient with the second importance score to determine whether to adjust the importance coefficient. For example, the entity feature module 230 can change the importance coefficient to the average of the importance coefficient and the second importance coefficient.
[0044] Figure 3 is a diagram showing a user interface 300 that enables user input for primitives and time parameters according to one aspect. The user interface 300 is provided by the primitive generation module 210 and the time parameter module 220 of the feature engineering application 200. The user interface 300 enables the user to view and interact with the primitives that would be used to synthesize features. Further, the user is also enabled to input values of the time parameters. The user interface 300 includes a search bar 310, a display unit 320, a training data unit 330, and a time parameter unit 340. Other aspects of the user interface 300 have more, fewer, or different components.
[0045] The search bar 310 enables the user to type in search terms to find primitives that the user wishes to view or interact with in other ways. The display unit 320 presents a list of primitives. It presents all the primitives selected by the primitive generation module 210 or the primitives that match the search terms entered by the user. In some embodiments, the primitives are listed in order, for example, based on the importance of the primitive to the prediction, the relevance of the primitive to the user's search terms, or other factors. For each primitive in the list, the display unit presents the name of the primitive, the category (aggregation or transformation), and a description. The name, category, and description of the primitive explain the function and algorithm of the primitive and help the user understand the primitive. Each primitive in the list is associated with a checkbox that the user may click to select the primitive. Although not shown in FIG. 3, the user interface 300 enables the user to remove selected primitives, add new primitives, change the order of the primitives, or perform other types of interaction with the primitives. The training data section 330 can select the maximum number of data entities that will be used to create features. What has been described above provides the user with an opportunity to control the amount of training data that will be used to create features. In the user interface, it is possible to specify a trade-off between the processing time to generate features and the quality of the features.
[0046] The time parameter section 340 provides options for the user to input time values, such as specifying a time, defining a time window, or both. Further, the time parameter section 340 also provides an option for the user to select whether to use the time value as a common time or as a case-specific time. The common time is applied to all the primitives used to create features, whereas the case-specific time is applied to a subset of the primitives.
[0047] FIG. 4 is a diagram showing a user interface 400 for entity set creation according to one aspect. The user interface is provided by the entity feature module 230 of the feature engineering application 200. The user interface 400 displays data entities 410, 420, and 430 received by the feature engineering application 200 to the user. Further, for example, for each data entity, information of the data entity such as the data entity name, the number of columns, and the primary key is also displayed. The data entity name helps the user identify the data entity. Since the columns of the data entity correspond to variables, the number of columns corresponds to the number of variables of the data entity, and the main column corresponds to the primary variable. The user interface 400 further displays the parent-child relationship between two data entities 410 and 420 and the common columns shared by the two data entities. The data entity 410 is shown as the parent data entity, and the data entity 420 is shown as the child data entity. The user interface 400 provides a drop-down icon 440 (individually referred to as the drop-down icon 440) for each main column and common column displayed in the user interface 400 to enable the user to change the columns. For example, in response to receiving a selection of the drop-down icon, the user interface 400 provides a list of candidate columns to the user and enables the user to select different columns. In some aspects, the entity feature module 230 detects the columns of the data entity and selects candidate columns from the detected columns. Further, the user interface 400 also enables the user to add new relationships. The user interface 400 shows three data entities and one parent-child relationship, but may include more data entities and more parent-child relationships. In one aspect, a data entity may be a parent in one parent but a child in a different subset.
[0048] FIG. 5 is a flowchart illustrating a method 500 for synthesizing features by using data entities received from different data sources according to one aspect. In some aspects, the method is performed by the machine learning server 110, but in other aspects, some or all of the operations in the method may be performed by other entities. In some aspects, the operations in the flowchart are performed in a different order and include different and / or additional steps.
[0049] The machine learning server 110 receives a plurality of data entities from different data sources 510. The different data sources may be associated with different users, different organizations, or different departments within the same organization. The plurality of data entities are used to train a model that makes predictions based on new data. A data entity is a set of data that includes one or more variables. Exemplary predictions include application monitoring, network traffic data flow monitoring, user action prediction, and the like.
[0050] The machine learning server 110 generates primitives based on the plurality of data entities 520. Each of the primitives is configured to be applied to the variables of the plurality of data sets to synthesize features. In some aspects, the machine learning server 110 selects primitives from a pool of primitives based on the plurality of data entities. The machine learning server 110 provides the selected primitives for display to the user and may enable the user to edit the selected primitives, such as adding, deleting, or changing the order of the primitives. The machine learning server 110 generates primitives based on the selection and user editing.
[0051] The machine learning server 110 receives a time value from a client device associated with a user 530. The time value is used to synthesize one or more time-based features from a plurality of data entities. In some embodiments, the machine learning server 110 determines one or more cutoff times based on the time value and extracts data from an entity set based on the one or more cutoff times. The data to be extracted will be used by the machine learning server 110 to synthesize one or more time-based features from the extracted data.
[0052] After a primitive is generated and the time value is received, the machine learning server 110 generates an entity set by aggregating a plurality of data entities 540. The machine learning server 110 identifies a subset of data entities from a plurality of data sets. Each subset includes two or more data entities that share a common variable. Next, the machine learning server 110 generates intermediate data entities by aggregating the data entities of each subset. In some embodiments, the machine learning server 110 determines the primary variable of each data entity in the subset and, based on the primary variable of the data entities in the subset, identifies the data entities in the subset as parent entities and each of one or more other data entities included in the parent as child entities. The machine learning server 110 aggregates the data entities in the subset based on the parent-child relationship. The machine learning server 110 generates an entity set by aggregating the intermediate data entities.
[0053] In some embodiments, the machine learning server 110 generates two new data entities from a given data entity among a plurality of data entities based on variables of the given data entity. Each of the two new data entities includes a subset of the variables of the given data entity. For example, the machine learning server 110 identifies a first primary variable and a second primary variable from the variables of a given entity. Next, the variables of the given entity are classified into a first variable group and a second variable group. The first variable group includes the first primary variable and one or more other variables of the given entity related to the first primary variable. The second variable group includes the second primary variable and one or more other variables of the given entity related to the second primary variable. The machine learning server 110 generates one of the two new entities from the variables of the first group and the values of the variables of the first group, and generates the other of the two new entities from the variables of the second group and the values of the variables of the first group. The machine learning server 110 identifies a subset of data entities from the two new data entities and the plurality of data entities excluding the given data entity.
[0054] The machine learning server 110 synthesizes a plurality of features by applying primitives and time values to an entity set 550. In some embodiments, the machine learning server 110 applies primitives to an entity set to generate a pool of features. Next, the machine learning server 110 iteratively evaluates the pool of features to remove some features from the pool of features and obtains a plurality of features. In each iteration, the machine learning server 110 evaluates the usefulness of at least some of the plurality of features by applying different parts of the entity set to the evaluated features, and generates a plurality of features by removing some of the evaluated features based on the usefulness of the evaluated features.
[0055] The plurality of features includes one or more time-based features. In some embodiments, the machine learning server 110 determines one or more cut-off times based on time values. Next, the machine learning server 110 extracts data from the entity set based on the one or more cut-off times and generates one or more time-based features from the extracted data.
[0056] The machine learning server 110 trains a model based on the plurality of features. 560 The machine learning server 110 may use different machine learning techniques in different embodiments. Exemplary machine learning techniques include, for example, linear support vector machine (linear SVM), boosting of other algorithms (e.g., AdaBoost), neural networks, logistic regression, naive Bayes, memory-based learning, random forest, bagging tree, decision tree, boosting tree, boosting stamp, and the like. Next, the trained model is used to make predictions from the perspective of a new dataset.
[0057] FIG. 6 is a high-level block diagram illustrating a functional view of a typical computer system 600 for use as the machine learning server 110 of FIG. 1 according to an embodiment.
[0058] The illustrated computer system includes at least one processor 602 coupled to a chipset 604. The processor 602 can include multiple processor cores on the same die. The chipset 604 includes a memory controller hub 620 and an input / output (I / O) controller hub 622. A memory 606 and a graphics adapter 612 are coupled to the memory controller hub 620, and a display 618 is coupled to the graphics adapter 612. A storage device 608, a keyboard 610, a pointing device 614, and a network adapter 616 may be coupled to the I / O controller hub 622. In some other embodiments, the computer system 600 can have additional, fewer, or different components, and the components may be coupled differently. For example, embodiments of the computer system 600 may be without a display and / or a keyboard. Additionally, in some embodiments, the computer system 600 may be instantiated as a rack-mounted blade server or as a cloud server instance.
[0059] The memory 606 holds instructions and data used by the processor 602. In some embodiments, the memory 606 is a random access memory. The storage device 608 is a non-transitory computer-readable recording medium. The storage device 608 can be an HDD, an SSD, or other types of non-transitory computer-readable recording media. Data processed and analyzed by the machine learning server 110 can be stored in the memory 606 and / or the storage device 608.
[0060] The pointing device 614 may be a mouse, trackball, or other type of pointing device and is used in combination with a keyboard 610 for entering data into the computer system 600. The graphics adapter 612 displays images and other information on the display 618. In some aspects, the display 618 includes touch screen capabilities for receiving user input and selections. The network adapter 616 connects the computer system 600 to a network 140 connection.
[0061] The computer system 600 is adapted to execute computer modules for providing the functionality described herein. As used herein, the term "module" refers to computer program instructions and other logic for providing a particular functionality. A module can be implemented in hardware, firmware, and / or software. A module can include one or more processes and / or can be provided by only a portion of a process. A module is typically stored in the storage device 608, loaded into the memory 606, and executed by the processor 602.
[0062] The particular naming of components, capitalization of terms, attributes, data structures, or other programming or structural aspects are not essential or important, and the mechanisms implementing the described aspects can have different names, formats, or protocols. Further, the system can be implemented via a combination of hardware and software as described, or may be implemented entirely with hardware elements. Additionally, the particular functional partitioning between the various system components described herein is merely exemplary and not essential, and functions performed by a single system component may instead be performed by multiple components, and functions performed by multiple components may instead be performed by a single component.
[0063] Some portions of the above description present features with respect to algorithms of operations of information and symbolic representations. The description and representation of the algorithms described above are used by those skilled in the art of data processing technology to most effectively convey the research content to other persons skilled in the art. It is understood that the operations described above are functionally or logically described while being implemented by a computer program. Furthermore, it has been shown that referring to the above-described arrangement of operations as a module, or by a function name, is sometimes convenient without sacrificing generality.
[0064] As is apparent from the above discussion, unless otherwise specified, throughout this specification, discussions using terms such as "processing" or "calculating" or "computing" or "determining" or "displaying" refer to the actions and processes of a computer system or a similar electronic computing device that manipulates and transforms data represented as physical (electronic) quantities within the memory or registers or other such information storage, transmission, or display devices of the computer system.
[0065] Certain aspects described in this specification include processing steps and instructions described in the form of algorithms. It should be noted that the processing steps and instructions of the aspects are embodied in software, firmware, or hardware, and when embodied in software, can be resident on and downloaded from different platforms used by a real-time network operating system and operated therefrom. Finally, in principle, the words used in the specification are selected for readability and for educational purposes and may not have been selected to demarcate or define the boundaries of the subject matter of the present invention. Accordingly, the disclosure of this aspect is intended to be illustrative and not limiting.
Claims
1. Receiving a plurality of data entities from different data sources, and generating primitives based on the plurality of data entities, each of the primitives including an algorithm for synthesizing features having associated values when applied to one or more of the plurality of data entities, and generating an entity set by aggregating the plurality of data entities, and generating a pool of features by applying the primitives to the entity set, updating the pool of features over a plurality of iterations, each of the plurality of iterations including determining a usefulness score for each feature in the pool of features using a portion of the entity set that is different from a portion of the entity set used in a different iteration, and removing at least one feature from the pool of features based on the usefulness score and outputting the updated pool of features as a plurality of features in response to a stop condition, thereby synthesizing the plurality of features, and generating the machine learning model configured to generate an output based on new data by training the machine learning model using the plurality of features A method characterized by comprising.
2. Generating the entity set by aggregating the plurality of data entities includes identifying a subset of data entities from the plurality of data entities, each subset of data entities including two or more data entities sharing a common variable, and generating intermediate data entities by aggregating the data entities in each subset of data entities, and generating the entity set by aggregating the intermediate data entities The method according to claim 1, characterized by comprising.
3. Generating the intermediate data entities by aggregating the data entities in each of the subsets of data entities includes Determining a primary variable for each data entity in the subset of data entities; Identifying, based on each of the primary variables determined for each data entity in the subset of data entities, a data entity in the subset of data entities as a parent entity and each of one or more other data entities in the subset of data entities as a child entity; The method according to claim 2, characterized by including the above.
4. Identifying the subset of data entities from the plurality of data entities comprises: Generating two new data entities from the given data entity among the plurality of data entities based on a variable defining the given data entity, each of the two new data entities including a subset of the variables defining the given data entity; Identifying the subset of data entities from the plurality of data entities excluding the given data entity and the two new data entities; The method according to claim 2, characterized by including the above.
5. Generating the two new data entities from the given data entity among the plurality of data entities based on a variable defining the given data entity comprises: Identifying a first primary variable and a second primary variable from the variables defining the given entity; Classifying the variables defining the given entity into a first group of variables and a second group of variables, wherein the first group of variables includes the first primary variable and one or more other variables defining the given entity related to the first primary variable, and the second group of variables includes the second primary variable and one or more other variables defining the given entity related to the second primary variable; Generating one of the two new entities with the first group of variables and respective values for the first group of variables; generating another one of the two new entities based on the second group of variables and respective values for the second group of variables The method according to claim 4, characterized by including this. [
6. ] Synthesizing the plurality of features by applying the primitive to the entity set is based on a time value, and determining one or more cut-off times based on the time value; and extracting data from the entity set based on the one or more cut-off times; and synthesizing the plurality of features from the extracted data The method according to claim 1, further characterized by further including this. [
7. ] a computer processor that executes computer program instructions; receiving a plurality of data entities from different data sources; generating a primitive based on the plurality of data entities, each of the primitives including an algorithm for synthesizing features having associated values when applied to one or more of the plurality of data entities; generating an entity set by aggregating the plurality of data entities; generating a pool of features by applying the primitive to the entity set; updating the pool of features over a plurality of iterations, each of the plurality of iterations including determining a usefulness score for each feature in the pool of features using a portion of the entity set different from a portion of the entity set used in a different iteration, and deleting at least one feature from the pool of features based on the usefulness score including this, and outputting the updated pool of features as a plurality of features in response to a stop condition to synthesize the plurality of features; and generating the machine learning model configured to generate an output based on new data by training the machine learning model using the plurality of features A non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations including A system characterized by comprising **Claim 8** Generating the entity set by aggregating the plurality of data entities comprises Identifying a subset of data entities from the plurality of data entities, each subset of data entities including two or more data entities sharing a common variable, and Generating intermediate data entities by aggregating the data entities in each subset of data entities, and Generating the entity set by aggregating the intermediate data entities The system according to claim 7, characterized by comprising **Claim 9** Generating the intermediate data entities by aggregating the data entities in each of the subsets of data entities comprises Determining a primary variable for each data entity in the subset of data entities, and Based on each of the primary variables determined for each data entity in the subset of data entities, identifying the data entities in the subset of data entities as parent entities and one or more other data entities in the subset of data entities as child entities The system according to claim 8, characterized by comprising **Claim 10** Identifying the subset of data entities from the plurality of data entities comprises Generating two new data entities from the given data entity among the plurality of data entities based on the variables defining the given data entity, each of the two new data entities including a subset of the variables defining the given data entity, and Identifying the subset of data entities from the plurality of data entities excluding the given data entity and the two new data entities The system according to claim 8, characterized by comprising
11. Generating the two new data entities from the given data entity among the plurality of data entities based on a variable that defines the given data entity, identifying a first primary variable and a second primary variable from the variable that defines the given entity; classifying the variable that defines the given entity into a first group of variables and a second group of variables, wherein the first group of variables includes the first primary variable and one or more other variables that define the given entity related to the first primary variable, and the second group of variables includes the second primary variable and one or more other variables that define the given entity related to the second primary variable; generating one of the two new entities based on the first group of variables and respective values for the first group of variables; generating the other of the two new entities based on the second group of variables and respective values for the second group of variables The system according to claim 10, characterized in that it includes the above.
12. Synthesizing the plurality of features by applying the primitive to the entity set is based on a time value, and determining one or more cut-off times based on the time value; extracting data from the entity set based on the one or more cut-off times; synthesizing the plurality of features from the extracted data The system according to claim 7, further characterized in that it further includes the above.
13. A non-transitory computer-readable memory storing computer program instructions executable for processing data blocks in a data analysis system, wherein the computer program instructions receive a plurality of data entities from different data sources; generating a primitive based on the plurality of data entities, each of the primitives including an algorithm for synthesizing features having related values when applied to one or more of the plurality of data entities; generating an entity set by aggregating the plurality of data entities; generating a pool of features by applying the primitive to the entity set; updating the pool of features over a plurality of iterations, each of the plurality of iterations determining a usefulness score for each feature in the pool of features using a different part of the entity set than the part of the entity set used in a different iteration, and removing at least one feature from the pool of features based on the usefulness score including; and responding to a stop condition by outputting the updated pool of features as a plurality of features; and combining the plurality of features; and generating the machine learning model configured to generate an output based on new data by training the machine learning model using the plurality of features A non-transitory computer-readable memory, characterized in that it is executable to perform operations including.
14. Generating the entity set by aggregating the plurality of data entities includes identifying a subset of data entities from the plurality of data entities, each subset of data entities including two or more data entities that share a common variable; and generating intermediate data entities by aggregating the data entities in each subset of data entities; and generating the entity set by aggregating the intermediate data entities The non-transitory computer-readable memory according to claim 13, characterized in that it includes.
15. Generating the intermediate data entities by aggregating the data entities in each of the subsets of data entities includes determining a primary variable for each data entity in the subset of data entities; Identifying, based on each of the primary variables determined for each data entity in the subset of data entities, a data entity in the subset of data entities as a parent entity and each of one or more other data entities in the subset of data entities as child entities The non-transitory computer-readable memory according to claim 14, comprising the above
16. Identifying the subset of data entities from the plurality of data entities is Generating, based on a variable defining a given data entity, two new data entities from the given data entity among the plurality of data entities, each of the two new data entities including a subset of the variables defining the given data entity Identifying the subset of data entities from the plurality of data entities excluding the given data entity and the two new data entities The non-transitory computer-readable memory according to claim 14, comprising the above
17. Synthesizing the plurality of features by applying the primitive to the entity set is based on a time value and Determining one or more cut-off times based on the time value Extracting data from the entity set based on the one or more cut-off times Synthesizing the plurality of features from the extracted data The non-transitory computer-readable memory according to claim 13, further comprising the above
18. During each of the plurality of iterations, determining the usefulness score for each feature in the pool of features is Constructing a random forest including a decision tree having nodes representing different features in the pool of features, and Applying a portion of the entity set that is different from the portion of the entity set used in the different iteration sets to the random forest The non-transitory computer-readable memory according to claim 13, characterized in that it is performed by
19. During each of the plurality of iterations, determining the usefulness score for each feature in the pool of features comprises constructing a random forest including a decision tree having nodes representing different features in the pool of features, and applying a portion of the entity set that is different from the portion of the entity set used in the different iteration sets to the random forest The method according to claim 1, characterized in that it is performed by
20. During each of the plurality of iterations, determining the usefulness score for each feature in the pool of features comprises constructing a random forest including a decision tree having nodes representing different features in the pool of features, and applying a portion of the entity set that is different from the portion of the entity set used in the different iteration sets to the random forest The system according to claim 7, characterized in that it is performed by
Citation Information
Patent Citations
Data processing device, method, and semiconductor manufacturing method
WO2020115943A1