Dataset feature type inference
By identifying and enhancing the feature types in training datasets using a large language model, the method addresses the issue of dataset quality in machine learning models, resulting in improved predictive accuracy.
Patent Information
- Application Number
- JP2024216078
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-12-11
- Publication Date
- 2025-07-03
AI Technical Summary
Existing machine learning models suffer from inaccuracies due to the quality of training datasets, which often lack diversity and representativeness, affecting their predictive capabilities.
A method involving feature type inference is employed, where a dataset is analyzed to identify and label different types of features, and a labeled dataset is adjusted using a large language model to enhance its scope and inclusiveness, followed by training a machine learning model on this adjusted dataset.
Improves the robustness and accuracy of machine learning models by ensuring they are trained on a more diverse and representative dataset, leading to better predictive performance.
Smart Images

Figure 2025100418000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to dataset feature type inference.
Background Art
[0002] A machine learning (ML) model is trained using a training dataset. The quality of the training dataset affects the accuracy and realism of the predictions made by the ML model. For example, the training dataset can define the prediction patterns of the ML model. A sufficiently diverse and representative training dataset containing various scenarios and features can enable the ML model to make reasonable predictions for different input data.
[0003] The subject matter claimed in the present disclosure is not limited to embodiments that solve problems or operate only in environments such as those described above. Rather, this background art description is provided only to illustrate an example of the technical area in which the embodiments described in the present disclosure may be implemented.
Summary of the Invention
[0004] According to one aspect of an embodiment, one or more operations may include accessing a dataset that includes a plurality of data subsets. Feature type candidates corresponding to the data subsets may be identified. The one or more operations may further include constructing a first machine learning model using different sets of feature type candidates. Each of the different sets of feature type candidates may be scored based on the respective accuracy of each of the first machine learning models corresponding to each different set of feature type candidates with respect to the dataset. Based on the scores of the different feature type sets, a final set of feature types may be selected from the different sets of feature type candidates. The operation may further include training a second machine learning model using the labeled dataset generated by applying the final set of feature types to the dataset.
[0005] The objectives and advantages of the embodiments will be realized and achieved by at least the elements, mechanisms, and combinations specifically pointed out in the claims. It should be understood that both the foregoing general description and the following detailed description are illustrative and not restrictive of the claimed invention.
Brief Description of the Drawings
[0006] The embodiments will be described and explained more specifically and in detail through the accompanying drawings including the following figures.
Figure 1
Figure 2
Figure 3
Figure 4
Modes for Carrying Out the Invention
[0007] A machine learning model can be trained using a training dataset for making predictions. The training dataset can include training instances or individual data points used to train the ML model. An individual data point can correspond to features and a target variable that the ML model can be designed to predict. Features can define the characteristics of the data that the ML model can use to make predictions. For example, the ML model can perform different types of processing or analysis of the data depending on different characteristics of the data defined by different feature types. Features can include various data types such as, among others, numerical, categorical, text-based, etc.
[0008] In some examples, the training dataset can be represented in different formats suitable for an ML model. For example, the training dataset can be represented in a table format with multiple columns and rows. In such an example, the columns can correspond to specific features with different feature types, and the rows can represent individual instances or data points of those features.
[0009] According to one or more embodiments of the present disclosure, the feature types of the training dataset can be identified. For example, a feature type inference can be performed on the training dataset. For example, different types of data within the training dataset can be identified and labeled with corresponding feature types. For example, in an example where the training dataset is represented as a table-formatted dataset, the different columns can represent different types of data. In such an example, the feature type inference process can determine and label each feature type that at least partially defines one or more characteristics of the data included in the corresponding column for each column or subset of the training dataset.
[0010] In some embodiments, the training dataset can be adjusted based on different types of data (indicated by the identified feature types). For example, in some embodiments, various large language model prompts can be generated and provided to a large language model. Using the responses from the large language model, the training dataset can be adjusted by improving the existing data and / or adding additional data. Adjusting the training dataset using a large language model can improve the scope and inclusiveness of the training dataset. As a result, the machine learning model generated using the training dataset can be improved. For example, the machine learning model can be more robust and can predict target features more accurately.
[0011] Embodiments of the present disclosure will be described with reference to the accompanying drawings.
[0012] FIG. 1 shows an example system 100 configured for machine learning training, according to one or more embodiments of the present disclosure. Generally, system 100 can be configured to train and / or generate an ML model 114. In some embodiments, system 100 can be configured to train an ML model 114 using a dataset 102 that can be adjusted or enhanced to improve the training of the ML model 114.
[0013] In some embodiments, system 100 can include a feature type inference (FTI) module 104 and a data conditioning module 108, which can generally be referred to as “modules.” In some embodiments, one or more of the modules can include code and routines configured to enable a computing system to perform one or more operations. Additionally, or alternatively, one or more of the modules can be implemented using hardware including one or more processors, CPUs, graphics processing units (GPUs), data processing units (DPUs), parallel processing units (PPUs), microprocessors (e.g., for performing or controlling one or more operations), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), accelerators (e.g., deep learning accelerators (DLAs)), and / or other processor types. In these and other embodiments, one or more of the modules can be implemented using a combination of hardware and software. In the present disclosure, operations described as being performed by a particular module can include operations that the particular module can direct a corresponding computing system to perform. In these and other embodiments, one or more of the modules can be implemented by one or more computing systems, such as those described in more detail with respect to FIG. 4, for example.
[0014] In some embodiments, the dataset 102 can be a training dataset used to train the ML model 114. The dataset 102 can be obtained from any source or can be constructed using any data compilation technique. The data can include numerical data, such as strings of characters including letters, symbols, or other characters, numbers, or combinations of numbers and characters. The data can also include data in other formats.
[0015] In some embodiments, the data within the dataset 102 can be organized into one or more data subsets. For example, data of the same category can be organized into the same data subset. For example, an example of data can be real estate data including address, land price, land size, and land value improvement. As an example, data representing land price can form part of a data subset.
[0016] In these and other embodiments, data of the same category grouped into a data subset can be referred to as a feature of the dataset 102. As an example, the dataset 102 can include data in a table format that can be arranged in columns and rows. In these and other embodiments, each column can represent a feature of the dataset 102, and each row can contain values in one or more of those columns. The values within one of those rows can be related to each other. For example, according to the previous example, the values of each column in a single row can be related to the same address.
[0017] In some embodiments, the data subsets of the dataset 102 can include various types of features. In such examples, multiple different types of features can be identified, and one or more data subsets can be labeled according to the type of related features.
[0018] The FTI module 104 can be configured to analyze data within one or more data subsets of the dataset 102 to identify different types of features included in the one or more data subsets. The one or more data subsets can accordingly be labeled to indicate the type of features included in the one or more data subsets. For example, the FTI module 104 can generate a labeled dataset 106 corresponding to the dataset 102, having feature labels corresponding to the feature types of the data subsets of the dataset 102. In some examples, different data types (e.g., indicated by the feature labels) can include, among other things, categorical variables (e.g., text features), identifier (ID)-like features (e.g., numerical IDs, alphanumeric codes, etc.). For example, the first data subset can include an address, and the second data subset can include an income value. The first data subset and the second data subset can be assumed to have no labels identifying the type of data before the dataset 102 is analyzed by the FTI module 104. The corresponding labeled dataset 106 (e.g., generated by the FTI module 104) can include a first data subset labeled as an address and a second data subset labeled as income.
[0019] In addition, or alternatively, the corresponding labeled dataset 106 can include feature type indications regarding the characteristics of the first data subset and the second data subset. For example, a feature type of "sentence" can be associated with the first data subset. In addition, or alternatively, a feature type of "currency" can be associated with the second data subset. In some embodiments, the FTI module 104 can perform one or more operations described in this disclosure with respect to FIGS. 2 and 3 as part of the feature type inference and generation of the labeled dataset 106.
[0020] In some embodiments, the labeled dataset 106 can be processed by the data adjustment module 108 to generate an adjusted dataset 110. In these and other embodiments, the data adjustment module 108 can execute one or more algorithms and / or operations to adjust the scope of the labeled dataset 106. As an example, the dataset 102 can be adjusted such that the scope of the dataset 102 can be expanded.
[0021] In some embodiments, the data adjustment module 108 can be configured to instruct a large language model (LLM) to generate one or more additional features for the labeled dataset 106. Using the one or more additional features, an adjusted dataset 110 can be generated. For example, the adjusted dataset 110 can include additional data in addition to the existing data of the dataset 102. In some embodiments, the one or more additional features generated by the LLM can vary based on the type of features corresponding to the one or more data subsets. For example, the one or more additional features can include enhancement of existing data and addition of external data determined based on the existing data. In some examples, the enhancement of existing data can include further grouping or splitting of the existing data such that at least one new feature is generated. The new feature can make a portion of the existing data more distinguishable within the dataset 102. In contrast, the external data can include new data that does not exist in the dataset 102 but can be related to at least one existing feature of the dataset 102. One or more examples of the data adjustment process are further explained and described in the U.S. patent application titled “Data Adjustment Using Large Language Model” by Lei Liu, Wei-Peng Chen, and Sou Hasegawa, filed on December 21, 2023 (Attorney Docket No. F1423.10580US01), which is hereby incorporated by reference in its entirety.
[0022] In some embodiments, the ML model 114 can be trained using the adjusted dataset 110. For example, the ML training module 112 can provide the adjusted ML model 114, and by learning the patterns and / or relationships between the features and the target feature in the adjusted dataset 110, the ML model 114 can be enabled to predict the value of the target feature when the values of other features are given. In some embodiments, the ML model 114 can be any type of ML model. For example, the ML model can be, inter alia, a supervised learning model (e.g., a regression model, a classification model), an unsupervised learning model (e.g., a clustering model), a deep learning model (e.g., a convolutional neural network, a recurrent neural network, a transformer model).
[0023] As shown, the ML model 114 can be used to predict the value of the target feature. For example, a dataset including one or more of the features of the dataset can be provided to the ML model 114. The ML model 114 can predict the value of the target feature based on the values of the provided one or more features. By providing the adjusted dataset 110 (e.g., a dataset having more features than the dataset 102), the ML model 114 can predict the value of the target feature more accurately. Therefore, the adjustment of the dataset 102 for generating the adjusted dataset 110 can improve the training of the ML model 114, and thus improve machine learning techniques.
[0024] Changes, additions, or omissions can be made to the system 100 without departing from the scope of the present disclosure. For example, in some embodiments, the system 100 may include any number of other components not explicitly illustrated or described. Also, the system 100 may perform any number of operations not explicitly described without departing from the scope of the present disclosure, and / or may not perform all of the operations explicitly described.
[0025] FIG. 2 shows an example process 200 for generating a labeled dataset 206 based on a dataset 202, according to one or more embodiments of the present disclosure. In some embodiments, other more operations of process 200 may be performed by the FTI module 104 of FIG. 1.
[0026] The dataset 202 may be the same as or similar to the dataset 102 of FIG. 1. Additionally or alternatively, the labeled dataset 206 may be the same as or similar to the labeled dataset 106 of FIG. 1.
[0027] In an embodiment, process 200 may include a feature type (FT) candidate generation operation 204 “FT candidate generation 204”. FT candidate generation 204 may include one or more operations that may be performed on the dataset 202 to identify one or more FT candidates 208.
[0028] In some embodiments, FT candidate generation 204 may include identifying different feature type candidates for a plurality of different subsets of the dataset 202. For example, the dataset 202 may include multiple columns of data, each of which may correspond to a particular feature. In these and other embodiments, FT candidate generation 204 may include identifying a set of feature type candidates for each of one or more of the columns, where each is a candidate feature type for the feature corresponding to the column.
[0029] In some embodiments, one or more operations of FT candidate generation 204 may be performed by a Feature Type Inference Model (FTI model). In some embodiments, the FTI model may include any suitable ML model configured to analyze dataset 202 and predict possible feature types (e.g., feature type candidates) for different subsets of dataset 202. For example, in some embodiments, the FTI model can predict different feature type candidates for one or more of the data subsets. Additionally, or alternatively, the FTI model can assign a probability value to each feature type candidate. The probability value may indicate the likelihood that the corresponding feature type candidate is the actual feature type of the corresponding data subset.
[0030] In some embodiments, for a data subset of dataset 202, respective groups of feature type candidates may be compiled. For example, the feature type candidates identified for each of one or more of the data subsets can be included in their respective groups corresponding to the data subsets.
[0031] In some embodiments, FT candidate generation 204 may include filtering out one or more feature type candidates from the group of feature type candidates. For example, in some embodiments, feature type candidates having a probability value that does not meet a specific probability threshold (e.g., 0.25) can be removed from one or more of the groups of feature candidates. For example, a specific group of feature type candidates corresponding to a specific data subset may include a first number of feature type candidates for that specific data subset. In these and other embodiments, feature type candidates that do not meet the probability threshold can be removed from the specific group of feature type candidates such that the specific group of feature type candidates can have a second number of feature type candidates that is less than the first number.
[0032] Additionally, or alternatively, one or more of the FT candidates 208 of the above groups may be filtered based on the ranking of the FT candidates 208 within each group. The ranking may be based on the corresponding probability values. For example, in some embodiments, a threshold number “m” of FT candidates 208 may be set such that none of the above groups can exceed “m” FT candidates 208. In these and other embodiments, the FT candidates 208 within each group may be ranked according to their respective probability values, and the top “m” FT candidates 208 may be maintained within the group and other FT candidates removed. In some embodiments, “m” may be a fixed number. Additionally, or alternatively, the value of “m” may be based on a certain percentage of the FT candidates 208. In these and other embodiments, the value of “m” may be based on computing resources (e.g., processing power, memory availability, etc.) that may be available for executing the process 200. The feature type candidate groups before or after filtering may each include one or more feature type candidates.
[0033] In some embodiments, the FTI model may be configured to predict feature type candidates from previously defined feature types. Training of the FTI model using a training data set having defined feature types may result in the FTI model being trained to predict which of such feature types may correspond to data subsets of the data set 202.
[0034] For example, in some embodiments, the FTI model may be trained using a training data set having feature types defined according to one or more of the following categories, namely, numerical, categorical, date / time, text, int, float, double timestamp, and string. Additionally, or alternatively, the training data set may not have a general “text” feature type, but rather may have more specific feature types for text, such as those shown in Table 1 below (which may correspond to defined “SortingHat” feature type categories).
Table 1
[0035] In some embodiments, one or more training data sets used to train the FTI model may include data subsets that have already been labeled according to previously defined feature types, such as those described above. Additionally, or alternatively, one or more training data sets used to train the FTI model may include one or more data subsets that are not labeled according to previously defined feature types and / or do not have features corresponding to previously defined feature types. In these and other embodiments, the training data sets used to train the FTI can be enhanced to include additional defined feature types.
[0036] For example, for data subsets (e.g., columns) that do not correspond to (e.g., do not overlap with) previously defined feature types (e.g., the feature types described above), one or more rules can be applied to obtain and / or generate sample values corresponding to such feature types and provide them to the FTI model to train the FTI model for such feature types.
[0037] The FT candidate 208 can include candidate feature types that can be respectively identified for the data subsets included in the data set 202. In some embodiments, the FT candidate 208 can be organized into respective groups of candidate feature types corresponding to different data subsets. As shown above, each group can include one or more candidate feature types for the data subset to which it corresponds. In some embodiments, one or more of the groups can be those remaining after performing filtering based on probability values as described above. Additionally, or alternatively, one or more of the groups can include candidate feature types predicted for the corresponding data subsets without any filtering being performed.
[0038] In these and other embodiments, the FT candidates 208 may each include a respective probability value associated with each feature type candidate and their corresponding data subsets. As shown above, each probability value may indicate the respective probability that the corresponding feature type candidate is the actual feature type of the corresponding data subset.
[0039] In some embodiments, the process 200 may include a feature type set generation operation 210 (FT set generation 210). The FT set generation 210 may include one or more operations corresponding to organizing the FT candidates 208 into one or more FT sets 212. In some embodiments, the FT set generation 210 may include identifying different combinations of the FT candidates 208 based on groups of FT candidates. In these and other embodiments, the FT set generation 210 may include identifying any different possible combinations of the FT candidates 208 based on any group of the FT candidates 208.
[0040] For example, the data set 202 can include "n" different data subsets, and thus, the FT candidates 208 can be organized into "n" different groups, one group for each data subset. Further, each of the "n" groups of FT candidates 208 can have a certain number of FT candidates 208 contained therein, and that number can vary from group to group. Each of the different combinations of the FT candidates 208 can include "n" FT candidates 208, and each FT candidate of a particular combination is selected from a different one of the groups of FT candidates 208. Additionally, alternatively, any possible combination of the "n" FT candidates 208 can be identified based on the different FT candidates included in the groups of FT candidates.
[0041] In these and other embodiments, each of the FT sets 212 may correspond to one of a plurality of combinations of the FT candidates 208. In some embodiments, the total number of combinations of FT candidates can be very large (e.g., thousands or millions of combinations). This number can depend on the number of different groups of FT candidates 208 (which can be determined by the number of data subsets), and on the number of FT candidates 208 included in each group of FT candidates 208. Thus, in some examples, the total number of FT sets 212 can be very large.
[0042] In some embodiments, the FT sets 212 may be filtered. For example, in some embodiments, a binding probability is determined for each FT set 212, and the FT sets 212 may be filtered based on the binding probability. For example, the FT sets 212 may be ranked according to their respective binding probabilities, and the threshold number of the highest-ranked FT sets 212 may be selected while the others may be filtered out. In these and other embodiments, the FT sets 212 may be filtered based on a threshold percentage of the FT sets 212, and the highest-ranked FT sets 212 within the threshold percentage may be selected while the (e.g., ranked) FT sets outside the threshold percentage may be filtered out.
[0043] In addition, or alternatively, the FT sets 212 may be filtered based on a binding probability threshold. For example, the FT sets 212 that meet the binding probability threshold may be selected, and the FT sets 212 that do not meet the binding probability threshold may be filtered out. In these and other embodiments, the thresholds that may be used to filter the FT sets 212 may be based on the computing resources (e.g., processing power, memory availability, etc.) available to execute the process 200.
[0044] The combination probability of the FT sets 212 may be determined according to any suitable technique. For example, in some embodiments, the probabilities of each feature type candidate included in each FT set 212 may be multiplied together to obtain the corresponding combination probability. For example, a particular FT set 212 may include four feature type candidates that may each have probability values “p1”, “p2”, “p3”, and “p4”. A particular combination probability “Cp” for the particular FT set 212 may be determined by the following equation: Cp = p1 * p2 * p3 * p4
[0045] In some embodiments, process 200 may include a machine learning model generation operation 214 (ML model generation 214). ML model generation 214 may include one or more operations used to generate a feature type ML model 216 (ML model 216) based on the FT sets 212 and the data set 202. In some embodiments, for each of the FT sets 212, a respective ML model 216 may be generated. In these and other embodiments, the ML model 216 may be generated for the FT sets 212 remaining after filtering out one or more FT sets 212 as described above, for example. In the present disclosure, the reference to a machine learning model being a “feature type” machine learning model is merely intended to distinguish the ML model generated using the FT sets 212 from other ML models described herein.
[0046] For example, in some embodiments, a particular FT set 212 may be provided to an automated machine learning model generator (ML generator). Additionally or alternatively, in some embodiments, the ML generator may include a rule-based ML generator configured to generate a particular machine learning pipeline (ML pipeline) based on the particular FT set 212.
[0047] For example, a particular ML pipeline may include data processing and modeling that can be used to generate a corresponding particular ML model. Additionally, or alternatively, different preprocessors that may be included in the ML pipeline may be better suited to analyze data having characteristics corresponding to some feature types than others. Thus, depending on the feature types included in a particular FT set 212 provided to the ML generator for generation of a particular ML pipeline, a particular preprocessor may be selected to include in the particular ML pipeline.
[0048] In some embodiments, the ML generator can apply rules to the feature types corresponding to different data subsets (such as those shown in a particular FT set 212) to determine which preprocessors can be used to analyze each data subset. In some embodiments, the process of providing the individual FT sets 212 to the ML generator is performed for each FT set 212, such that the ML generator can generate a corresponding ML pipeline for each FT set 212. In these and other embodiments, training data is provided to the individual ML pipelines, the individual ML pipelines can process the training data, and the processing of the training data by the ML pipelines can create the corresponding ML models 216. In some embodiments, each ML model 216 can thus correspond to one of the FT sets 212 and can be generated based on the corresponding FT set 212.
[0049] In some embodiments, the training data used to generate and train the ML models 216 may be the same for each ML model 216. Additionally, or alternatively, the training data used to generate and train two or more of the ML models 216 may be different.
[0050] In some embodiments, the training data may be sampled from the dataset 202. In these and other embodiments, the sampling strategy may vary depending on the ML type of the ML pipeline used to generate the corresponding ML model 216.
[0051] For example, in an ML pipeline corresponding to a classification task and operation, a specific number of instances of each corresponding data subset of the dataset 202 may be sampled for the training data. In these and other embodiments, those instances may be randomly sampled from the corresponding data subsets. Additionally, or alternatively, if a particular data subset does not contain a specific number of instances, all of those instances may be sampled as training data.
[0052] As another example, in an ML pipeline corresponding to a regression task and operation, the continuous values of the corresponding data subsets included in the dataset 202 may be converted to discrete values using any suitable technique. For example, in some embodiments, Doane's rule may be used to convert the continuous values to discrete values. In these and other embodiments, following the conversion, a specific number of instances of the discrete values may be sampled. In these and other embodiments, those instances may be randomly sampled from the corresponding data subsets. Additionally, or alternatively, if a particular data subset does not contain a specific number of instances, all of those instances may be sampled as training data. In some embodiments, the number of samples for classification may be the same as the number of samples for regression. In these and other embodiments, the number of samples for classification may be different from the number of samples for regression.
[0053] In some embodiments, process 200 may include an ML model evaluation operation 218 (ML model evaluation 218). ML model evaluation 218 may include one or more operations that can be used to determine the accuracy of ML model 216. In some embodiments, ML model evaluation 218 can be based on dataset 202.
[0054] For example, validation data sampled from dataset 202 can be provided as input data to each of ML models 216. In some embodiments, the validation data can be sampled in a manner similar to or similar to the training data. Additionally, alternatively, the validation data can be at least partially different from the training data. Also, the number of samples used for the validation data may be different from or the same as the number of samples used for the training data. For example, in some embodiments, the number of samples obtained for a particular data subset for the training data may be greater than the number of samples obtained for the particular data subset for the validation data.
[0055] ML model 216 can make one or more predictions based on the provided validation data and can output such predictions. The output can be verified using the data included in dataset 202, and for each ML model 216, a corresponding accuracy can be determined. For example, in some embodiments, each ML model 216 can be given a score based on its accuracy. In some embodiments, the determination of the accuracy for each of the respective ML models 216 can be included in the evaluation result 220. Additionally, alternatively, the FT set 212 corresponding to the ML model 216 can be scored according to the accuracy of the corresponding ML model 216. For example, a particular ML model 216 can be generated using a particular FT set 212. For the particular ML model 216, a particular accuracy can be determined. Additionally, alternatively, using the particular accuracy corresponding to the particular ML model 216, a score corresponding to the particular FT set 212 used to generate the particular ML model 216 can be determined.
[0056] In some embodiments, process 200 may include a feature type selection operation 222 (FT selection 222). The FT selection 222 may select a particular FT set 212 from the FT set 212 as the corresponding feature type for a data subset of the data set 202. In some embodiments, the FT selection 222 may be based on the evaluation result 220. For example, the FT selection 222 may identify which of the ML models 216 was the most accurate as identified by the ML model evaluation 218 based on the evaluation result 220. In these and other embodiments, the FT selection 222 may identify which of the FT sets 212 was used to generate the most accurate ML model 216. And the identified FT set 212 may be selected as the final feature type set 224 (final FT set 224).
[0057] In some embodiments, process 200 may include a data set labeling operation 226. The data set labeling operation 226 may be configured to generate a labeled data set 206 based on the data set 202 and the final FT set 224. For example, the data set labeling operation 226 may annotate a data subset of the data set 202 with the corresponding feature types included in the final FT set 224.
[0058] Accordingly, process 200 may be configured to perform feature type inference on the data set 202 to identify the feature types of the data set 202. As shown here, feature type identification can be used to adjust the training data set and improve machine learning model training. Additionally, or alternatively, the feature type inference described here provides a specific process that can be used to automate the feature type identification of the training data set, which improves the efficiency of generating training data for machine learning. Accordingly, improved training data efficiency improves the training efficiency of machine learning and, accordingly, the technology itself.
[0059] Figure 3 shows a flowchart of an example method 300 for performing feature type inference according to one or more embodiments of the present disclosure. Method 300 may be performed by any suitable system, apparatus, or device. For example, method 300 may be implemented using the system 100 of FIG. 1 or the computing system 400 of FIG. 4. Although shown as separate blocks, the steps and operations associated with one or more blocks of method 300 may be divided into additional blocks, combined into fewer blocks, or eliminated depending on the particular implementation. For example, one or more of the operations described above with respect to process 200 of FIG. 2 may be performed as part of method 300.
[0060] Method 300 may include block 302. At block 302, a data set including a plurality of data subsets may be accessed. The feature type corresponding to the data subset may not yet be specified. Data sets 102 and 202 described with respect to FIGS. 1 and 2 respectively may be examples of the data set accessed.
[0061] At block 304, a feature type candidate corresponding to the data subset may be identified. In some embodiments, the feature type candidate may be identified based on one or more of the operations described with respect to FT generation 204 of FIG. 2.
[0062] At block 306, a plurality of first machine learning models may be generated using a plurality of different sets of feature type candidates. In some embodiments, the different sets of feature type candidates may be identified based on one or more of the operations described with respect to FT set generation 210 of FIG. 2. Additionally, or alternatively, ML model 216 of FIG. 2 may be an example of a first machine learning model. In these and other embodiments, the first machine learning model may be generated based on one or more of the operations described with respect to ML model generation 214 of FIG. 2.
[0063] At block 308, each set of different feature type candidates can be scored. Additionally, or alternatively, the scoring can be based on the respective accuracy of the first machine learning model corresponding to each set of different feature type candidates with respect to the data set. In some embodiments, the accuracy of the first machine learning model can be determined based on one or more operations described with respect to the ML model evaluation 218 of FIG. 2. Additionally, or alternatively, the scoring of the different feature type candidate sets can be performed using one or more operations described with respect to the ML model evaluation 218 of FIG. 2.
[0064] At block 310, a final feature type set can be selected from the different feature type candidate sets. In some embodiments, the final feature type set can be selected based on one or more operations described with respect to the FT selection 222 of FIG. 2.
[0065] In some embodiments, the final feature type set can be applied to the data set to generate a labeled data set. For example, in some embodiments, the labeled data set can be generated based on one or more operations described with respect to the data set labeling 226 of FIG. 2.
[0066] At block 312, the second machine learning model can be trained using the labeled data set. Additionally, or alternatively, in some embodiments, the labeled data set can be adjusted, and the adjustment can use the feature type indications included in the labeled data set. Then, the adjusted data set can be used to train the second machine learning model.
[0067] Modifications, additions, or omissions can be made to method 300 without departing from the scope of the present disclosure. For example, one or more operations can be included or omitted.
[0068] Figure 4 shows a block diagram of an example computing system 400 according to at least one embodiment of the present disclosure. The computing system 400 can be configured to implement or direct one or more suitable operations described in the present disclosure. For example, the computing system 400 can be configured to execute one or more blocks of the process 200 of FIG. 2, or the method 300 of FIG. 3. Additionally, or alternatively, one or more of the modules of FIG. 1 can be implemented by, or can include, the computing system 400. The computing system 400 can include a processor 450, a memory 452, and a data storage 454. The processor 450, the memory 452, and the data storage 454 can be communicatively coupled.
[0069] Generally, the processor 450 can include any suitable dedicated or general-purpose computer, computing entity, or processing device, including various computer hardware or software modules, and can be configured to execute instructions stored on some applicable computer-readable medium. For example, the processor 450 can include a microprocessor, a microcontroller, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other digital or analog circuitry configured to interpret and / or execute program instructions and / or process data. Although shown as a single processor in FIG. 5, the processor 450 can include any number of processors configured to individually or collectively execute or direct the execution of any number of processes described in the present disclosure. Also, one or more of the processors can be present on one or more different electronic devices, such as different servers, for example.
[0070] In some embodiments, the processor 450 may be configured to interpret and / or execute program instructions and / or process data for program instructions and / or data stored in the memory 452, the data storage 454, or both the memory 452 and the data storage 454. In some embodiments, the processor 450 may fetch program instructions from the data storage 454 and load the program instructions into the memory 452. After the program instructions are loaded into the memory 452, the processor 450 may execute the program instructions.
[0071] The memory 452 and the data storage 454 may include a computer-readable storage medium for carrying or storing computer-executable instructions or data structures. By way of example and not limitation, such computer-readable storage media can include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory devices (e.g., solid state memory devices), or other non-transitory storage media that can be used to carry or store particular program code in the form of computer-executable instructions or data structures and that can be accessed by a general purpose or special purpose computer. The tangible or non-transitory computer-readable storage media may include the foregoing and combinations of the foregoing. The term "non-transitory" as used in this disclosure should be construed to exclude only those types of transitory media that have been found ineligible for patentable subject matter in the decision of the Federal Circuit in Nuijten, 500 F.3d 1346 (Fed. Cir. 2007).
[0072] Combinations of the foregoing may also be included within the scope of computer-readable storage media. Computer-executable instructions may include, for example, instructions and data configured to cause the processor 450 to perform a particular process or group of processes.
[0073] Without departing from the scope of the present disclosure, changes, additions, or omissions may be made to the computing system 400. For example, in some embodiments, the computing system 400 may include any number of other components that are not explicitly illustrated or described.
[0074] The foregoing disclosure is not intended to limit the present disclosure to the exact form disclosed or to a particular field of use. Accordingly, it is contemplated that various alternative embodiments and / or modifications to the present disclosure are possible, whether explicitly described or implied herein, in light of the present disclosure. Although embodiments of the present disclosure have been described as such, it can be appreciated that variations can be made in form and detail without departing from the scope of the present disclosure. Accordingly, the present disclosure is limited only by the claims.
[0075] In some embodiments, the various components, modules, engines, and services described herein may be implemented as a plurality of objects or processes executed on a computer system (e.g., as separate threads). Although some of the systems and methods described herein are generally described as being implemented in software (stored in and / or executed by general-purpose hardware), a specific hardware implementation, or an implementation combining software and specific hardware is also possible and contemplated.
[0076] In accordance with general convention, various features shown in the drawings may not be drawn to scale. The figures presented in this disclosure are not intended to be actual views of any particular apparatus (e.g., device, system, etc.) or method, but rather are merely idealized representations used to illustrate various embodiments of the present disclosure. Accordingly, the dimensions of various features may be arbitrarily enlarged or reduced for clarity. Additionally, some of the drawings may be simplified for clarity. Thus, the drawings may not show all of the components of a given apparatus (e.g., device) or all of the operations of a particular method.
[0077] As used herein and in particular in the appended claims (e.g., the body of the appended claims), the terms are generally intended to be “open” terms (e.g., the term “comprising” should be interpreted as “including, but not limited to,” the term “having” should be interpreted as “having at least,” the term “including” should be interpreted as “including, but not limited to,” and so forth).
[0078] Also, when the specific number of claim recitations to be introduced is intended, such intention shall be explicitly recited in the claims, and in the absence of such recitation, such intention does not exist. For example, for the purpose of assistance in understanding, the claims appended below may include the use of the introductory phrases "at least one" and "one or more" to introduce claim recitations. However, the use of such phrases shall not be construed to mean that the introduction of a claim recitation by an indefinite article such as "a" or "an" limits any particular claim containing such introduced claim recitation to an embodiment containing only such recitation (for example, "a" and / or "an" shall be construed to mean "at least one" or "one or more"). The same applies to the use of a definite article used to introduce a claim recitation.
[0079] Furthermore, even when the specific number of claim recitations to be introduced is explicitly recited, it should be understood that such recitation shall be construed to mean at least the recited number (for example, a literal recitation of "two things" without other modifying phrases means at least two things, or two or more things). Also, in those cases where traditional expressions similar to "at least one of A, B, and C" or "one or more of A, B, and C" are used, generally, such syntax is intended to include only A, only B, only C, A and B together, A and C together, B and C together, or A, B, and C together, etc. For example, the use of the term "and / or" is intended to be so construed.
[0080] Also, discrete terms or phrases presenting two or more different terms, whether in the description of the embodiments, the claims, or the drawings, should be understood as intending to include one of those terms, any of those terms, or both terms. For example, the phrase "A or B" should be understood as including the possibilities of "A" or "B" or "A and B".
[0081] Also, the use of terms such as "first", "second", "third", etc. is not necessarily used here to imply a particular order or number of elements. Generally, terms such as "first", "second", "third", etc. are used to distinguish between different elements as general identifiers. If terms such as "first", "second", "third", etc. do not indicate that they imply a particular order, they should not be understood as implying a particular order. Further, if terms such as "first", "second", "third", etc. do not indicate that they imply a particular number of elements, they should not be understood as implying a particular number of elements. For example, it may be described that the first widget has a first side and the second widget has a second side. The use of the term "second side" with respect to the second widget is for distinguishing such a side of the second widget from the "first side" of the first widget, and not for implying that the second widget has two sides.
[0082] All examples and conditional language described herein are intended for educational purposes to assist the reader in understanding the concepts provided by the inventors of this application for advancing the present invention and technology, and should not be construed as limitations to the specifically described examples and conditions. Although the embodiments of the present disclosure have been described in detail, it should be understood that various modifications, substitutions, and alterations can be made to these embodiments without departing from the spirit and scope of the present disclosure.
[0083] Regarding the above description, the following additional remarks are disclosed. (Appendix 1) Access a dataset including a plurality of data subsets, Identify a plurality of candidate feature types corresponding to the plurality of data subsets, Construct a plurality of first machine learning models using a plurality of different candidate feature type sets, Score each of the plurality of different candidate feature type sets based on the accuracy of each of the plurality of first machine learning models corresponding to each different candidate feature type set with respect to the dataset, Select a final feature type set from the plurality of different candidate feature type sets based on the scores of the plurality of different candidate feature type sets, Train a second machine learning model using the labeled dataset generated by applying the final feature type set to the dataset, A method having this. (Appendix 2) The candidate feature types of the plurality of candidate feature types are identified based on an inferential machine learning analysis of the dataset, the method according to Appendix 1. (Appendix 3) The method further includes filtering the candidate feature types based on respective probability values corresponding to the likelihood that the candidate feature types are the actual feature types corresponding to their respective data subsets, the method according to Appendix 2. (Appendix 4) The plurality of different candidate feature type sets are based on different combinations of candidate feature types corresponding to different data subsets, the method according to Appendix 1. (Appendix 5) The plurality of different candidate feature type sets are selected based on respective joint probability values for each candidate feature type set, the respective joint probability values are determined based on the respective individual probability values of the individual candidate feature types included in the corresponding candidate feature type set, and the respective individual probability values correspond to the likelihood that the corresponding candidate feature type is the actual feature type corresponding to their respective data subsets, the method according to Appendix 1. (Appendix 6) The method according to Appendix 1, wherein constructing the plurality of first machine learning models includes training the plurality of first machine learning models using data sampled from the dataset. (Appendix 7) The method according to Appendix 1, further comprising determining the accuracy of each of the plurality of first machine learning models based on data sampled from the dataset used as verification data for the plurality of first machine learning models. (Appendix 8) One or more non-transitory computer-readable media storing instructions that, in response to being executed by one or more processors, cause a system to perform operations, the operations including accessing a dataset including a plurality of data subsets, identifying a plurality of candidate feature types corresponding to the plurality of data subsets, constructing a plurality of first machine learning models using a plurality of different sets of candidate feature types, scoring each of the plurality of different sets of candidate feature types based on the accuracy of each of the plurality of first machine learning models corresponding to each different set of candidate feature types with respect to the dataset, selecting a final set of feature types from the plurality of different sets of candidate feature types based on the scores of the plurality of different sets of candidate feature types, training a second machine learning model using a labeled dataset generated by applying the final set of feature types to the dataset, One or more non-transitory computer-readable media having the above. (Appendix 9) The one or more non-transitory computer-readable media according to Appendix 8, wherein the candidate feature types of the plurality of candidate feature types are identified based on inferential machine learning analysis of the dataset. (Appendix 10) The one or more non-transitory computer-readable media according to Appendix 9, wherein the operation further includes filtering the candidate feature types based on respective probability values corresponding to the likelihood that the candidate feature types are the actual feature types corresponding to their respective data subsets. (Appendix 11) The one or more non-transitory computer-readable media according to Appendix 8, wherein the plurality of different candidate feature type sets are based on different combinations of candidate feature types corresponding to different data subsets. (Appendix 12) The one or more non-transitory computer-readable media according to Appendix 8, wherein the plurality of different candidate feature type sets are selected based on respective joint probability values for each candidate feature type set, the respective joint probability values being determined based on respective individual probability values of the individual candidate feature types included in the corresponding candidate feature type set, and the respective individual probability values corresponding to the likelihood that the corresponding candidate feature types are the actual feature types corresponding to their respective data subsets. (Appendix 13) The one or more non-transitory computer-readable media according to Appendix 8, wherein constructing the plurality of first machine learning models includes training the plurality of first machine learning models using data sampled from the data set. (Appendix 14) The one or more non-transitory computer-readable media according to Appendix 8, wherein the operation further includes determining the accuracy of each of the plurality of first machine learning models based on data sampled from the data set used as validation data for the plurality of first machine learning models. (Appendix 15) A system, one or more processors, one or more non-transitory computer-readable storage media configured to store instructions, wherein the instructions, in response to being executed, cause the system to perform an operation, the operation including accessing a data set including a plurality of data subsets, Identify a plurality of candidate feature types corresponding to the plurality of data subsets, Construct a plurality of first machine learning models using a plurality of different sets of candidate feature types, Score each of the plurality of different sets of candidate feature types based on the accuracy of each first machine learning model among the plurality of first machine learning models corresponding to each different set of candidate feature types with respect to the data set, Select a final set of feature types from the plurality of different sets of candidate feature types based on the scores of the plurality of different sets of candidate feature types, Train a second machine learning model using the labeled data set generated by applying the final set of feature types to the data set, A system having the above. (Appendix 16) The candidate feature types of the plurality of candidate feature types are identified based on an inferential machine learning analysis of the data set, The operation further includes filtering the candidate feature types based on respective probability values corresponding to the likelihood that the candidate feature types are the actual feature types corresponding to their respective data subsets, The system according to Appendix 15. (Appendix 17) The plurality of different sets of candidate feature types are based on different combinations of candidate feature types corresponding to different data subsets. The system according to Appendix 15. (Appendix 18) The plurality of different sets of candidate feature types are selected based on respective joint probability values for each set of candidate feature types. The respective joint probability values are determined based on the respective individual probability values of the individual candidate feature types included in the corresponding set of candidate feature types. The respective individual probability values correspond to the likelihood that the corresponding candidate feature types are the actual feature types corresponding to their respective data subsets. The system according to Appendix 15. (Appendix 19) The system according to Appendix 15, wherein constructing the plurality of first machine learning models includes training the plurality of first machine learning models using data sampled from the data set. (Appendix 20) The system according to Appendix 15, wherein the operation further includes determining the accuracy of each of the plurality of first machine learning models based on data sampled from the data set used as verification data for the plurality of first machine learning models.
Claims
1. Access a dataset that includes a plurality of data subsets, Identify a plurality of candidate feature types corresponding to the plurality of data subsets, Construct a plurality of first machine learning models using a plurality of different sets of candidate feature types, Score each of the plurality of different sets of candidate feature types based on the accuracy of each of the plurality of first machine learning models corresponding to each different set of candidate feature types with respect to the dataset, Select a final set of feature types from the plurality of different sets of candidate feature types based on the scores of the plurality of different sets of candidate feature types, Train a second machine learning model using the labeled dataset generated by applying the final set of feature types to the dataset, A method having this.
2. The method according to claim 1, wherein the candidate feature types of the plurality of candidate feature types are identified based on an inferential machine learning analysis of the dataset.
3. The method according to claim 2, wherein the method further includes filtering the candidate feature types based on respective probability values corresponding to the likelihood that the candidate feature types are the actual feature types corresponding to their respective data subsets.
4. The method according to claim 1, wherein the plurality of different sets of candidate feature types are based on different combinations of candidate feature types corresponding to different data subsets.
5. The plurality of different sets of candidate feature types are selected based on respective joint probability values for each set of candidate feature types, and the respective joint probability values are determined based on the respective individual probability values of the individual candidate feature types included in the corresponding set of candidate feature types, and the respective individual probability values correspond to the likelihood that the corresponding candidate feature types are the actual feature types corresponding to their respective data subsets. The method according to claim 1.
6. The method according to claim 1, wherein constructing the plurality of first machine learning models includes training the plurality of first machine learning models using data sampled from the dataset.
7. The method according to claim 1, further comprising determining the accuracy of each of the plurality of first machine learning models based on data sampled from the dataset used as validation data for the plurality of first machine learning models.
8. One or more non-transitory computer-readable media storing instructions that, in response to being executed by one or more processors, cause a system to perform operations, the operations comprising: Accessing a dataset including a plurality of data subsets; Identifying a plurality of candidate feature types corresponding to the plurality of data subsets; Constructing a plurality of first machine learning models using a plurality of different candidate feature type sets; Scoring each of the plurality of different candidate feature type sets based on the accuracy of each of the plurality of first machine learning models corresponding to each different candidate feature type set with respect to the dataset; Selecting a final feature type set from the plurality of different candidate feature type sets based on the scores of the plurality of different candidate feature type sets; Training a second machine learning model using a labeled dataset generated by applying the final feature type set to the dataset; One or more non-transitory computer-readable media having the above.
9. A system comprising: One or more processors; One or more non-transitory computer-readable storage media configured to store instructions that, in response to being executed, cause the system to perform operations, the operations comprising: Accessing a dataset including a plurality of data subsets; Identifying a plurality of candidate feature types corresponding to the plurality of data subsets; Constructing a plurality of first machine learning models using a plurality of different candidate feature type sets; Scoring each of the plurality of different candidate feature type sets based on the accuracy of each of the plurality of first machine learning models corresponding to each different candidate feature type set with respect to the dataset; Selecting a final feature type set from the plurality of different candidate feature type sets based on the scores of the plurality of different candidate feature type sets; Selecting a final feature type set from the plurality of different candidate feature type sets based on the scores of the plurality of different candidate feature type sets; Training a second machine learning model using the labeled dataset generated by applying the final feature type set to the dataset A system having this.