Data set encoding using generative artificial intelligence
Generative AI encoding of tabular datasets improves ML model training and prediction accuracy by generating embeddings and applying weights to pairwise feature comparisons, addressing the limitations of existing methods with tabular data.
Patent Information
- Application Number
- JP2025019689
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-27
- Filing Date
- 2025-02-10
- Publication Date
- 2025-09-08
AI Technical Summary
Existing machine learning models struggle with accurately predicting outcomes from tabular datasets due to lack of well-known representations and pre-trained models, issues like data sparsity, mixed feature types, and unknown dataset structures, leading to inefficiencies and reduced accuracy.
Employ generative artificial intelligence to encode tabular datasets by generating embeddings and applying weights to pairwise comparisons of features, enhancing dataset coverage and comprehensiveness, thereby improving ML model training and prediction accuracy.
Enhances the robustness and accuracy of ML models by increasing feature dimensionality and variability, allowing them to better handle non-descriptive features and maintain consistent performance across different datasets.
Smart Images

Figure 2025130695000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to using generative artificial intelligence to encode datasets for machine learning models. [Background technology]
[0002] Machine learning (ML) generally employs ML models trained with a training dataset to make predictions that become more accurate as training progresses. The size and dimensionality of the training dataset affect the accuracy and reliability of the predictions made by the ML model. ML may be used in a wide variety of applications, including, but not limited to, housing market forecasting, web search, online fraud detection, medical diagnosis, speech recognition, email filtering, image recognition, virtual personal assistants, and automatic translation.
[0003] The subject matter claimed in this disclosure is not limited to embodiments that operate only in environments such as those described above or that solve any disadvantages such as those described above. Rather, this background is provided only to illustrate, by way of example, the scope of technologies in which some embodiments described in this disclosure may be practiced. Summary of the Invention
[0004] According to an aspect of an embodiment, the operations may include identifying features corresponding to the dataset. A pre-trained generative artificial intelligence model may obtain an embedding (e.g., a feature embedding) for each of the features. A pairwise comparison of the embeddings may also be generated. An encoded dataset may be generated by applying weights calculated using the pre-trained generative artificial intelligence model to the pairwise comparison of the embeddings for each of the features identified in the dataset.
[0005] The object and advantages of the embodiments will be realized and attained at least by the elements, features, and combinations particularly pointed out in the claims. It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory and are not restrictive of the invention as claimed. [Brief explanation of the drawings]
[0006] [Figure 1] 1 illustrates an example process for encoding a data set. [Figure 2A] 1 illustrates an example system for encoding a data set. [Figure 2B] 1 represents an example of an encoded data set. [Figure 3] 1 illustrates an example process for encoding a dataset used to train a machine learning model. [Figure 4] 1 depicts a flowchart of an example method for encoding a dataset using generative artificial intelligence. [Figure 5] 1 is an example computing system. DETAILED DESCRIPTION OF THE INVENTION
[0007] All that is shown in the drawings is in accordance with one or more embodiments of the present disclosure. Through the use of the accompanying drawings, example embodiments will be described and explained with additional specificity and detail.
[0008] A machine learning (ML) model may be trained using a training dataset to make predictions. The training dataset may include training instances or individual data points used to train the ML model. The individual data points may correspond to features and one or more target variables that the ML model may be designed to predict. Features may define characteristics of the data that the ML model can use to make predictions. Features may include various data types, such as numeric, categorical, and / or text-based, among others. Features may be included in one or more headers related to the dataset, such as within the column headers of a tabular dataset.
[0009] In some embodiments, using generative artificial intelligence (AI) to encode a training dataset may improve the accuracy and / or efficiency of identifying features in a tabular dataset and / or generating pairwise comparisons of features compared to traditional approaches in which one or more humans (e.g., data scientists / feature engineers / domain experts) manually encode the dataset. For example, using generative AI to encode a training dataset may reduce human bias / error.
[0010] In some cases, encoding a training dataset may improve the training of an ML model and its ability to make accurate predictions, such as in situations where the training dataset is represented with a set of additional features suitable for the ML model. For example, a training dataset may be represented in a tabular format with multiple rows and columns. In such cases, the columns may represent features of the training dataset, and the rows may represent individual instances or data points of the features. ML models (e.g., deep ML models) generally perform better with unstructured data, such as image data, because a wide range of known representations and pre-trained models may be available. However, unlike image data, no well-known representations or pre-trained models exist for tabular data. Some issues that may cause this lack of well-known representations and / or pre-trained models may be challenges associated with tabular data, such as lack of locality, data sparsity, a mix of feature types, and / or lack of knowledge of the dataset structure. Encoding a training dataset that is a tabular dataset may improve the training of an ML model on the tabular dataset and / or the ability of the ML model to make accurate predictions of a target variable related to the tabular dataset.
[0011] The present disclosure may relate, inter alia, to systems and methods related to encoding a training dataset using generative AI. Encoding a training dataset may improve a training dataset used in training an ML model by enhancing the training dataset, for example, by adding additional features to the training dataset and by determining correlations between features within the training dataset.
[0012] In some embodiments, encoding the training dataset may include using generative AI to increase the dimensionality of features in the training dataset to improve variability in the training dataset through embeddings, such as word embeddings or contextual embeddings. Additionally, encoding the training dataset may include applying weights to pairwise comparisons of embeddings generated for each feature in the training dataset using generative AI.
[0013] Encoding a training dataset using generative AI (e.g., a pre-trained generative AI model) can improve the coverage and comprehensiveness of the training dataset. As a result, the ML model generated using the training dataset can be improved. For example, the ML model can be more robust and more accurately predict target variables.
[0014] Embodiments of the present disclosure will now be described with reference to the accompanying figures.
[0015] 1 illustrates an example process 100 configured to generate an encoded data set 114 using a data set 102 in accordance with one or more embodiments of the present disclosure. In some embodiments, the process 100 may include a feature identification process 104, an embedding process 108, and a data encoding process 112.
[0016] In some embodiments, the dataset 102 may include any suitable type of data to be analyzed. For example, the dataset 102 may include electronic data. In some embodiments, the dataset may be a training dataset used to train an ML model. The dataset 102 may be obtained from any source or constructed using any data compilation technique. The dataset 102 may include numeric data, strings of characters, such as phonograms, symbols, or other characters, or combinations of numbers and characters, and / or other types of data.
[0017] In some embodiments, the dataset 102 may be represented or stored as a tabular set of data. For example, the data included in the dataset 102 may be stored in columns and rows. In some cases, the columns may represent categories of information. For example, a column header may identify the category of information included in the corresponding column, where each column may contain different categories of information. In some embodiments, the columns may represent features 106 of the dataset 102, and each row may include one or more values of the column. In these and other embodiments, the columns may represent instances of data. For example, each row may include a single data entry that may be characterized by a corresponding value in the category of information represented by the column. In some embodiments, the values in a row may be associated together. For example, each value in a column within a single row may be associated with the same real estate listing.
[0018] In some embodiments, the dataset 102 may be stored in one or more databases. In some embodiments, the dataset 102 may include public data. For example, the dataset 102 may include housing market data. In some embodiments, the dataset 102 may include public data used in one or more existing ML projects. Additionally or alternatively, the dataset 102 may include private data. For example, the private data may include protected health information. In these and other embodiments, portions of the process 100 may be performed using an on-premise server of a customer that owns or controls the private data. For example, the on-premise server may enhance the confidentiality of private data included as part of the dataset 102.
[0019] In some embodiments, the dataset 102 may include one or more features 106. The features 106 may be characteristics of the dataset 102. For example, the features 106 may include one or more of the following: number / quantity of rows, number / quantity of columns, presence or absence of values, presence of missing values, presence of number categories, presence of string categories, presence of text, mean, median, mode, distribution, maximum, minimum, label or title of a category of information, and / or any other characteristic. In some embodiments, the features 106 may include any suitable categorical description of the data in the dataset 102. For example, the features 106 may be text-based data contained in column headers. In some embodiments, the features 106 may not be descriptive of the dataset 102. In these and other embodiments, nondescriptive may refer to not having significance or conveying meaningful content in a particular context. For example, features 106 extracted from column headers containing "Feature 1," "Feature 2," "Feature 3," etc. may not be descriptive of the dataset 102. Additionally or alternatively, the dataset 102 may be organized based on one or more features 106 that describe the dataset 102. For example, the dataset 102 may include real estate data corresponding to a city, square footage, and year built, where the city, square footage, and year built are each one of the features 106.
[0020] In some embodiments, the feature identification process 104 may include obtaining a dataset 102. Additionally or alternatively, the feature identification process 104 may include identifying features 106 contained in the dataset 102. For example, the dataset 102 may include tabular data, and the features 106 may be data contained in column headers of the dataset 102. In some embodiments, the feature identification process 104 may be performed by a feature identification module, such as feature identification module 204, described in more detail below in connection with FIG. 2A .
[0021] In some embodiments, the embedding process 108 may include generating embeddings 110 of the features 106, where the embeddings 110 are vector representations of the features 106 such that the similarity between each of the features 106 is expressed numerically. In some cases, the embeddings 110 may be generated as numerical representations of the features 106 such that an ML model may utilize the features 106 in ML model training and / or perform one or more operations on the features 106. In these and other embodiments, the embedding process 108 may include encoding the features 106 with additional information such that each of the features 106 is represented by multiple numerical elements (dimensions). In some embodiments, the embeddings 110 may be high-dimensional vector representations of the features 106. For example, the embeddings 110 may include 100, 1,000, or more dimensions for each of the features 106. In some embodiments, the additional information may include information regarding the semantic similarity of each feature 106 compared to each of the other features 106. Semantic similarity may indicate the relatedness of two or more words, sentences, and / or phrases included in the features based on whether the words, sentences, and / or phrases have similar meanings and / or frequently appear in similar contexts. In these and other embodiments, the additional information may be represented as an additional dimension included in the embedding 110. In some cases, the additional information may be generated using word embeddings and / or context embeddings. For example, word embedding techniques such as Word2Vec and / or Global Vectors for Word Representation (GloVe), and / or context embedding techniques such as Embeddings from Language Model (ELMo) and / or Bidirectional Encoder Representations from Transformers (BERT) may encode the features 106 with additional information indicative of semantic similarity. In some embodiments, the additional information may be represented in the embedding 110 as one or more numerical elements included in the embedding 110. In some embodiments, the embedding process 108 may be performed by an AI module, such as AI module 208, described in more detail below in connection with FIG. 2A.
[0022] In some embodiments, the data encoding process 112 may include generating an encoded data set 114 using the embeddings 110. In some cases, the data encoding process 112 may include using a generating AI to generate weights and apply the weights to pairwise comparisons of the embeddings 110, such as in system 200 described in more detail below in connection with FIG. 2A. For example, the data encoding process 112 may be performed by a data encoding module, such as data encoding module 218 described in more detail below in connection with FIG. 2A.
[0023] In some embodiments, the encoded dataset 114 may be a representation of the dataset 102 in which each of the features 106 in the dataset 102 is weighted based on a pairwise comparison of the embeddings 110. In some embodiments, the encoded dataset 114 may be a training dataset that may be used to train an ML model, such as the ML model 304 described in more detail below in connection with FIG. 3. For example, the encoded dataset 114 may include data suitable for training an ML model to perform one or more actions, such as the action execution process 310 described in more detail below in connection with FIG. 3.
[0024] Modifications, additions, or omissions may be made to process 100 without departing from the scope of the present disclosure. For example, in some embodiments, process 100 may include any number of other components that may not be explicitly illustrated or described.
[0025] 2A illustrates an example system 200 configured to generate an encoded data set 220 using a data set 202, in accordance with one or more embodiments of the present disclosure. In some embodiments, system 200 may include a feature identification module 204, an AI module 208, a paired comparison module 212, and a data encoding module 218, which may be generally referred to as "modules." Additionally or alternatively, in some embodiments, system 200 may perform one or more operations of process 100 described in connection with FIG. 1.
[0026] In some embodiments, one or more of the modules may include code and routines configured to enable a computing system to perform one or more operations. Additionally or alternatively, one or more of the modules may be implemented using hardware, including one or more processors, central processing units (CPUs), graphics processing units (GPUs), data processing units (DPUs), parallel processing units (PPUs), microprocessors (e.g., for performing or controlling the execution of one or more operations), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), accelerators (e.g., deep learning accelerators (DLAs)), and / or other processor types. In these and other embodiments, one or more of the modules may be implemented using a combination of hardware and software. In this disclosure, operations described as being performed by a particular module may include operations that the particular module may instruct a corresponding computing system to perform. In these and other embodiments, one or more of the modules may be implemented by one or more computing systems, such as those described in further detail in connection with FIG. 5.
[0027] In some embodiments, the dataset 202 may be the same as or similar to the dataset 102 described above in connection with Figure 1. Additionally or alternatively, in some embodiments, the features 206 may be the same as or similar to the features 106 described above in connection with Figure 1.
[0028] In some embodiments, features 206 of dataset 202 may be identified by feature identification module 204. For example, feature identification module 204 may be configured to identify that features 206 are column names of tabular data in dataset 202. In some cases, feature identification module 204 may also extract features 206 identified by feature identification module 204. In these and other embodiments, feature identification module 204 may classify features 206 based on the type of data included in each feature 206. For example, feature identification module 204 may classify features 206 based on whether the feature 206 includes a categorical variable (e.g., a text-based feature), an identifier (ID)-style feature (e.g., a numeric ID, an alphanumeric code, etc.), and / or other types of data. In some embodiments, feature identification module 204 may be configured to send one or more identified and / or extracted features 206 to AI module 208.
[0029] In some embodiments, the AI module 208 may be configured to instruct a generative AI model to generate embeddings 210 to enrich the features 206 with additional information generated based on the dataset 202. For example, the AI module 208 may be configured to instruct a large-scale language model (LLM) such that the AI module 208 can obtain (know) connections between words, sentences, and / or frames that may be useful for generating the embeddings 210.
[0030] In some embodiments, the AI module 208 may be configured to generate an embedding 210 that is n-dimensional in size, where n may be selected by a user and / or by the AI module 208 itself. In these and other embodiments, the dimensionality of the embedding 210 may be based on resource constraints, such as the limited data processing capabilities of the AI module 208, a user's time limit, and / or some other criteria.
[0031] In some embodiments, the paired comparison module 212 may be configured to obtain embeddings 210 and generate one or more paired comparisons 214 of the embeddings 210. For example, based on obtaining embeddings 210 of {[Emb(f1)], [Emb(f2)], [Emb(f3)]} representing features 206 of {f1, f2, f3}, the paired comparison module 212 may generate {[Emb(f1),Emb(f2)], [Emb(f1),Emb(f3)], [Emb(f2),Emb(f3)]} as pairs for the paired comparisons 214. In some embodiments, generating the paired comparisons 214 may include one or more calculations of correlations between the embeddings 210. For example, the paired comparisons 214 may include a calculation of correlation based on the distance between two embeddings 210, such as using cosine similarity or other measures of distance between vectors, such as Euclidean distance, Manhattan distance, or Word Mover's Distance.
[0032] In some embodiments, the paired comparisons 214 may be used by the AI module 208 to calculate weights 216 that may be applied to features 206 of the dataset 202. For example, the AI module 208 may use one or more generative AI models and / or statistical analysis to calculate the weights 216, such that the weights 216 indicate the strength of the correlation between the features 206. In some cases, the AI module 208 may use the paired comparisons 214 to generate the weights 216 so that, when the weights 216 are applied to the features 206, pairs of features 206 with the highest correlation are easily identifiable. For example, a weight 216 close to 1 may indicate a high correlation, while a weight 216 close to 0 may indicate a low correlation. In some embodiments, the set of paired comparisons 214 may include pairs in which particular features are compared to themselves, such that a value of 1 may be calculated as the weight for such paired comparisons 214.
[0033] In some embodiments, the data encoding module 218 may be configured to generate the encoded data set 220 by applying the weights 216 to the paired comparisons 214. For example, the data encoding module 218 may multiply the correlation values corresponding to the paired comparisons 214 by their respective weights 216 and / or may organize the weights 216 to correspond to their respective paired comparisons 214. In some embodiments, the encoded data set 220 may be an example of the encoded data set 114 in FIG. 1 above.
[0034] In some embodiments, the encoded dataset 220 may be a compilation of paired comparisons 214 organized in a tabular format such that the weights 216 for each of the paired comparisons 214 correspond to columns for first features represented in the paired comparison and rows for second features represented in the paired comparison. In these and other embodiments, the encoded dataset 220 may be formatted to be used to train an ML model. For example, the encoded dataset 220 may be a heatmap representation of the correlation between each of the features 206. For example, the heatmap representation may include a color gradient (e.g., grayscale shading) that can be used to indicate the level of correlation.
[0035] 2B depicts an exemplary heatmap representation 250 of a dataset related to real estate data, including features including "ScreenPorch," "Street," "1stFlrSF," etc. By way of example, each of the features labeled in the left row can be considered a first feature 252, each of the features labeled in the bottom column can be considered a second feature 254, and each intersection in the heatmap representation 250 can include a weight calculated from a pairwise comparison 256 of the corresponding first feature 252 with the corresponding second feature 254. As depicted in FIG. 2B , the highlighted pairwise comparison 256 of the first feature 252, “Street,” and the second feature 254, “Alley,” is shown to include a weight of “0.491,” with a color gradient shown as a medium gray further indicating that the features “Street” and “Alley” are highly correlated (e.g., compare the black (very low correlation) of the pairwise comparison 256 of “TotRmsAbvGrd” and “Alley” with the light gray (high correlation) of the pairwise comparison 256 of “TotalBsmtSF” and “BsmtFinSF2”).
[0036] Returning to FIG. 2A , the encoded dataset 220 may improve the training of an ML model by allowing weights 216 to be used in place of data corresponding to one or more features 206 in one or more additional datasets to which the ML model is applied. For example, using a dataset 202 in which weights 216 indicate correlations between features 206 may be used to help an ML model, such as an LLM, obtain consistent results even with different datasets (e.g., using different features that tend to classify the same / similar data, such as the features “Sodium,” “Salt,” “NaCl,” and “Electrolyte”). An ML model trained using the encoded dataset 220 can more efficiently process such different datasets with different features by knowing that both “Sodium” and “Salt” may be highly correlated with other features, such as “Sugar,” or a target variable, such as “Calories.” Additionally or alternatively, the encoded dataset 220 may be useful for training an ML model to make accurate predictions on datasets with non-descriptive features, because the ML model can use weights 216 to replace missing information that descriptive features may have provided. For example, the ML model can perform one or more downstream tasks, such as target variable prediction or regression calculations, based on the encoded dataset 220, even if the features are only "1," "2," "3," etc.
[0037] Modifications, additions, or deletions may be made to system 200 without departing from the scope of the present disclosure. For example, in some embodiments, system 200 may include any number of other components that may not be explicitly illustrated or described.
[0038] 3 illustrates an example process 300 configured to generate target variable predictions 312 using an encoded dataset 302 in accordance with one or more embodiments of the present disclosure. In some embodiments, the process 300 may include an ML model training process 306 and / or an action execution process 310.
[0039] In some embodiments, the encoded data set 302 may be the same as or similar to the encoded data set 114 in FIG. 1 above and / or the encoded data set 220 in FIG. 2A above.
[0040] In some embodiments, the ML model training process 306 may include obtaining the encoded dataset 302 and the ML model 304 to generate a trained ML model 308. For example, the ML model training process 306 may include the ML model 304 learning patterns and / or relationships between features in the encoded dataset 302 and one or more target variables. In some embodiments, the ML model 304 may be any type of ML model. For example, the ML model 304 may be a supervised learning model (e.g., a regression model, a classification model, a tree-based model), an unsupervised learning model (e.g., a clustering model), and / or a deep learning model (e.g., a convolutional neural network, a recurrent neural network, a transformer model), among others.
[0041] In some embodiments, the ML model 304 may include one or more parameters that may be adjusted by the ML model training process 306. The adjustments may be a product of the ML model 304 learning from data input as part of the ML model training process 306, such as the encoded dataset 302. For example, the ML model 304 may be trained until the parameters result in predictions that satisfy a threshold (e.g., the predictions are determined to simulate actual outcomes to a desired level). Additionally or alternatively, the ML model training process 306 may include one or more hyperparameters. In some embodiments, hyperparameters may be parameters of the ML model 304 that control the configuration settings and / or details of the ML model training process 306. For example, hyperparameters may control the model architecture, learning rate, and / or some other aspect of the ML model training process 306. In these and other embodiments, hyperparameters may be adjusted externally (e.g., by a user or by automated techniques such as grid search or random search) as compared to parameters that are adjusted internally as part of the ML model training process 306. In some embodiments, hyperparameters may be used to adjust the overall behavior of the ML model 304 during the ML model training process 306. For example, hyperparameters that determine the number of nodes and / or layers in a neural network may be used to adjust the model complexity of the ML model 304.
[0042] In some embodiments, at least one hyperparameter may be used for dropout regulation of the ML model 304, which may include ignoring one or more parameters during the ML model training process 306 to prevent overfitting (e.g., a model that provides accurate predictions on training data but not on new data). By way of example, the hyperparameter may be a masking weight threshold hyperparameter that determines the probability of dropping out a weight, such that a model is dropped out during the ML model training process 306 in response to a weight satisfying a preset condition, such as being less than or equal to 0.1. In some cases, the dropout may be setting a weight that satisfies the preset condition to have a value of zero. In some embodiments, adjusting the masking weight threshold hyperparameter may improve the predictive performance of the ML model 304. For example, decreasing the masking weight threshold hyperparameter may improve predictive performance for datasets with descriptive features.
[0043] In some embodiments, after generating the trained ML model 308 using the ML model training process 306, the trained ML model 308 may be used to perform one or more operations. For example, the trained ML model 308 may perform one or more classification operations, regression operations, recognition operations, processing operations, etc. In some embodiments, the trained ML model 308 may perform one or more operations in an operation execution process 310 to generate a target variable prediction 312. For example, the target variable prediction 312 may include generating a stock price prediction based on the trained ML model 308 performing one or more regressions on a dataset.
[0044] In some embodiments, the action execution process 310 using the trained ML model 308 may include obtaining at least one target variable prediction 312 using one or more inputs related to the dataset and a first threshold hyperparameter. In these and other embodiments, in response to the dataset including features that are not descriptive of the dataset, the action execution process 310 may further include, after obtaining the target variable prediction 312, using a second threshold hyperparameter lower than the first threshold hyperparameter to improve the target variable prediction 312. For example, the masking weight threshold hyperparameter may be adjusted from 0.5 to 0.3 in response to the trained ML model 308 obtaining a dataset with non-descriptive features after it is determined that the first initial target variable prediction 312 generated is too inaccurate.
[0045] Modifications, additions, or deletions may be made to process 300 without departing from the scope of the present disclosure. For example, in some embodiments, process 300 may include any number of other components that may not be explicitly illustrated or described.
[0046] 4 is a flowchart of an example method 400 for encoding a dataset using generative AI, in accordance with at least one embodiment described in this disclosure. Method 400 may be performed by any suitable system. For example, method 400 may be implemented using, and / or one or more operations of method 400 may be performed by, system 200 of FIG. 2A. Additionally or alternatively, one or more operations of method 400 may be performed by a computing system such as that described in connection with FIG. 5. Although represented as separate blocks, steps and operations associated with one or more blocks of method 400 may be separated into additional blocks, combined into fewer blocks, or eliminated, depending on the particular implementation.
[0047] Method 400 may include block 402. At block 402, a plurality of features corresponding to a dataset may be identified. Dataset 102 of FIG. 1 or dataset 202 of FIG. 2A may be examples of datasets. Features may be identified based on the location of the feature within the dataset and / or its data type. For example, features may include text contained in one or more headers associated with the dataset. Additionally or alternatively, metadata and / or some other characteristic may be used to identify the features. In some embodiments, one or more of the features may not be descriptive of the dataset. In some embodiments, one or more of the feature identification operations may include one or more of the operations described above with respect to feature identification module 204.
[0048] At block 404, an embedding for each of the features may be obtained using the generative AI model. In some embodiments, each embedding for each feature may include one or more of a word embedding and / or a context embedding. For example, the one or more embeddings may be word and / or context embeddings that represent semantic similarities between each feature as vector representations. In some embodiments, the one or more embedding operations may include one or more of the operations described above with respect to the AI module 208.
[0049] At block 406, a paired comparison for each embedding may be generated. In some embodiments, the paired comparisons may be organized such that a generative AI model can utilize the paired comparisons. For example, each paired comparison may be a numeric vector representation corresponding to the correlation between two particular features. In some embodiments, one or more operations of the paired comparison may include one or more operations described above with respect to the paired comparison module 212.
[0050] At block 408, an encoded data set may be generated by applying weights calculated using the generative AI model to the pairwise comparisons. For example, the generative AI model may be a pre-trained generative AI model. In some embodiments, the weights may be calculated using cosine similarity. In some embodiments, the one or more operations for generating the encoded data set may include one or more operations described above with respect to data encoding module 218.
[0051] Modifications, additions, or deletions may be made to method 400 without departing from the scope of the present disclosure. For example, the described steps and operations are provided merely as examples, and some of the steps and operations may be optional, combined into fewer steps and operations, or expanded into additional steps and operations without departing from the essence of the disclosed embodiments.
[0052] For example, method 400 may further include using the encoded dataset to train an ML model and performing one or more operations using the ML model. Additionally or alternatively, using the encoded dataset to train the ML model in method 400 may include using at least one hyperparameter to set a weight among the weights to zero in response to the weight satisfying a predetermined condition. In these and other embodiments of method 400, performing one or more operations using the ML model may include obtaining a prediction of one or more target variables using one or more inputs related to the dataset and a first threshold hyperparameter. Additionally or alternatively, in response to the plurality of features not describing the dataset, method 400 may further include, after obtaining a prediction of the one or more target variables, improving the prediction of the one or more target variables using a second threshold hyperparameter, the second threshold hyperparameter being lower than the first threshold hyperparameter. In these and other embodiments, the first threshold hyperparameter and / or the second threshold hyperparameter may be fine-tuned / optimized based on multiple features (e.g., the optimized hyperparameter value may be based on whether a corresponding feature in the multiple features is descriptive or non-descriptive of the dataset). For example, method 400 may include optimizing a first threshold hyperparameter that is descriptive of the dataset, such as "sq / ft."
[0053] 5 illustrates a block diagram of an exemplary computing system in accordance with one or more embodiments of the present disclosure. Computing system 500 may include a processor 502, a memory 504, data storage 506, and / or a communication unit 508, all of which may be communicatively coupled. For example, the modules of FIG. 2A may be implemented as a computing system consistent with computing system 500.
[0054] In general, processor 502 may include any suitable special-purpose or general-purpose computer, computing entity, or processing device, including various computer hardware or software modules, and may be configured to execute instructions stored on any suitable computer-readable storage medium. For example, processor 502 may include a microprocessor, microcontroller, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or any other digital or analog circuitry configured to interpret and / or execute program instructions and / or process data.
[0055] 5 as a single processor, processor 502 may include any number of processors distributed across any number of networks or physical locations configured to individually or collectively perform any number of operations described in this disclosure. In some embodiments, processor 502 may interpret and / or execute program instructions and / or process data stored in memory 504, data storage 506, or memory 504 and data storage 506. In some embodiments, processor 502 may fetch program instructions from data storage 506 and load program instructions into memory 504.
[0056] After the program instructions are loaded into memory 504, processor 502 may execute the program instructions, such as instructions to cause computing system 500 to perform some of the operations of method 400 of Figure 4. For example, computing system 500 may execute the program instructions to identify a plurality of features corresponding to a data set.
[0057] Memory 504 and data storage 506 may include one or more computer-readable storage media for storing computer-executable instructions or data structures. Such computer-readable storage media may be any available media that can be accessed by a general-purpose or special-purpose computer, such as processor 502. In some embodiments, computing system 500 may or may not include memory 504 and data storage 506.
[0058] By way of example, and not limitation, such computer-readable storage media may include non-transitory computer-readable storage media including random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory devices (e.g., solid-state memory devices), or any other storage medium that can be used to store desired program code in the form of computer-executable instructions or data structures and that can be accessed by a general-purpose or special-purpose computer. Combinations of the above may also be included within the scope of computer-readable storage media. Computer-executable instructions may include, for example, instructions and data configured to cause processor 502 to perform a certain operation or group of operations.
[0059] The communications unit 508 may include any component, device, system, or combination thereof configured to transmit or receive information over a network. In some embodiments, the communications unit 508 may communicate with other devices at other locations or the same location, or with other components within the same system. For example, the communications unit 508 may include a modem, a network card (wireless or wired), an optical communications device, an infrared communications device, a wireless communications device (e.g., an antenna), and / or a chipset (e.g., a Bluetooth® device, an 802.6 device (metropolitan area network (MAN)), a Wi-Fi device, a WiMax device, a cellular communications facility, or others), etc. The communications unit 508 may enable data to be exchanged with a network and / or any other device or system described in this disclosure. For example, the communications unit 508 may enable the computing system 500 to communicate with other systems, such as computing devices and / or other networks.
[0060] Those skilled in the art, after reviewing the present disclosure, may recognize that modifications, additions, or deletions may be made to computing system 500 without departing from the scope of the present disclosure. For example, computing system 500 may include more or fewer components than explicitly illustrated or described.
[0061] The above disclosure is not intended to limit the disclosure to the precise form or field of use disclosed. As such, various alternative embodiments and / or modifications to the disclosure, whether expressly described or implied herein, are contemplated as possible in light of the present disclosure. While embodiments of the present invention have been described above, it will be apparent that changes can be made in form and detail without departing from the scope of the invention. Accordingly, the present invention is limited only by the claims.
[0062] In some embodiments, the various components, modules, engines, and services described herein may be implemented as objects or processes that run on a computing system (e.g., as separate threads). Although some of the systems and processes described herein are generally described as being implemented in software (stored and / or executed by general-purpose hardware), specific hardware implementations or combinations of software and specific hardware implementations are also possible and contemplated.
[0063] Terms used in this disclosure, particularly in the appended claims (e.g., the body of the appended claims), are generally intended as "open" terms (e.g., the word "including" should be interpreted as meaning "including, but not limited to," "having" should be interpreted as "having at least," "includes" should be interpreted as "including, but not limited to," etc.).
[0064] Furthermore, when a specific number is intended in an introduced claim recitation, such intention is clearly recited in the claim; otherwise, no such intention exists. For example, to facilitate understanding, the following appended claims may use introductory phrases such as "at least one" and "one or more" to introduce claim recitations. However, the use of such phrases should not be interpreted as suggesting that a particular claim containing such introduced claim recitation is limited to instances containing only one of the recited item, even if the same claim contains both an introductory phrase such as "one or more" or "at least one" and an indefinite article such as "a" or "an" (e.g., "a" and / or "an" should be interpreted to mean "at least one" or "one or more"). The same applies when introducing claim recitations using definite articles.
[0065] Furthermore, even when a specific number is explicitly stated in an introduced claim, those skilled in the art will understand that such a statement should generally be interpreted to mean at least the recited number (e.g., a statement simply stating "two items," without any other modifiers, means at least two items, or more than two items). Furthermore, when phrases like "at least one of A, B, and C, etc." or "one or more of A, B, and C, etc." are used, such structure is generally intended to include A only, B only, C only, both A and B, both A and C, both B and C, and / or all of A, B, and C, etc. Furthermore, use of the term "and / or" is intended to be interpreted in this manner.
[0066] Furthermore, any disjunctive word and / or phrase presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to contemplate the possibility of including one of those terms, either of those terms, or both of those terms. For example, the phrase "A or B" should be understood to include the possibilities of "A or B" or "A and B," even if the term "and / or" is used elsewhere.
[0067] All examples and conditional language set forth in this disclosure are intended for educational purposes to aid the reader in understanding the concepts and inventions contributed by the inventors to the advancement of the art, and should not be construed as being limited to such specifically set forth examples and conditions. Although embodiments of the present disclosure have been described in detail, various changes, substitutions, and alterations may be made thereto without departing from the spirit and scope of the present disclosure.
[0068] In addition to the above embodiments, the following supplementary notes are disclosed. (Appendix 1) identifying a plurality of features corresponding to a dataset; obtaining a respective embedding for each feature of the plurality of features using a pre-trained generative artificial intelligence model; generating a pairwise comparison of each of the embeddings; generating an encoded dataset by applying weights calculated using the pre-trained generative artificial intelligence model to the paired comparisons, the weights indicating correlations between features in the paired comparisons; and A method having the following. (Appendix 2) the plurality of features includes text contained in one or more headers related to the dataset; The method described in Appendix 1. (Appendix 3) the respective embeddings for each of the plurality of features include one or more of word embeddings or context embeddings. The method described in Appendix 1. (Appendix 4) The weights are calculated using cosine similarity. The method described in Appendix 1. (Appendix 5) using the encoded dataset to train a machine learning (ML) model; performing one or more actions using the ML model; and 2. The method of claim 1, further comprising: (Appendix 6) Using the encoded dataset to train the ML model further includes using at least one hyperparameter that sets a weight among the weights to zero in response to the weight satisfying a preset condition. The method described in Appendix 5. (Appendix 7) performing one or more operations using the ML model includes obtaining predictions of one or more target variables using one or more inputs related to the dataset and a first threshold hyperparameter; The method described in Appendix 5. (Appendix 8) In response to the plurality of features not describing the dataset, the method includes: after obtaining the prediction of the one or more target variables, further comprising: improving the prediction of the one or more target variables using a second threshold hyperparameter; the second threshold hyperparameter is lower than the first threshold hyperparameter; The method described in Appendix 7. (Appendix 9) one or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the system to perform operations; The operation is identifying a plurality of features corresponding to a dataset; obtaining a respective embedding for each feature of the plurality of features using a pre-trained generative artificial intelligence model; generating a pairwise comparison of each of the embeddings; generating an encoded dataset by applying weights calculated using the pre-trained generative artificial intelligence model to the paired comparisons, the weights indicating correlations between features in the paired comparisons; and 1. One or more non-transitory computer-readable media having: (Appendix 10) the plurality of features includes text contained in one or more headers related to the dataset; 10. One or more non-transitory computer-readable media as described in Appendix 9. (Appendix 11) the respective embeddings for each of the plurality of features include one or more of word embeddings or context embeddings. 10. One or more non-transitory computer-readable media as described in Appendix 9. (Appendix 12) The weights are calculated using cosine similarity. 10. One or more non-transitory computer-readable media as described in Appendix 9. (Appendix 13) the plurality of features are determined not to describe the data set based on the weights satisfying a threshold. 10. One or more non-transitory computer-readable media as described in Appendix 9. (Appendix 14) The operation is using the encoded dataset to train a machine learning (ML) model; performing one or more actions using the ML model; and 10. The one or more non-transitory computer-readable media of claim 9, further comprising: (Appendix 15) performing one or more operations using the ML model includes obtaining predictions of one or more target variables using one or more inputs related to the dataset; 15. One or more non-transitory computer-readable media as described in Clause 14. (Appendix 16) 1. A system having one or more processors and one or more non-transitory computer-readable storage media storing instructions, The instructions, in response to being executed by the one or more processors, cause the system to: identifying a plurality of features corresponding to a dataset; obtaining a respective embedding for each feature of the plurality of features using a pre-trained generative artificial intelligence model; generating a pairwise comparison of each of the embeddings; generating an encoded dataset by applying weights calculated using the pre-trained generative artificial intelligence model to the paired comparisons, the weights indicating correlations between features in the paired comparisons; and A system that causes an operation having the following steps to be performed. (Appendix 17) the plurality of features includes text contained in one or more headers related to the dataset; 17. The system of claim 16. (Appendix 18) the plurality of features are determined not to describe the data set based on the weights satisfying a threshold. 17. One or more systems as described in Clause 16. (Appendix 19) The operation is using the encoded dataset to train a machine learning (ML) model; performing one or more actions using the ML model; and 17. The system of claim 16, further comprising: (Appendix 20) performing one or more operations using the ML model includes obtaining predictions of one or more target variables using one or more inputs related to the dataset; 19. The system of claim 19. [Explanation of symbols]
[0069] 102,202 datasets 104 Feature Identification Process 106,206 Features 108 Embedding Process 110,210 embedded 112 Data Encoding Process 114,220,302 coded dataset 200 systems 204 Feature Identification Module 208 AI Module 212 Paired Comparison Module 214 Paired Comparisons 216 Weight 218 Data Encoding Module 304 ML model 306 ML model training process 308 trained ML models 310 Action Execution Process 312 Target Variable Prediction 500 Computing Systems 502 processor 504 memory 506 Data Storage 508 Communication Unit
Claims
1. identifying a plurality of features corresponding to a dataset; obtaining a respective embedding for each feature of the plurality of features using a pre-trained generative artificial intelligence model; generating a pairwise comparison of each of the embeddings; generating an encoded dataset by applying weights calculated using the pre-trained generative artificial intelligence model to the paired comparisons, the weights indicating correlations between features in the paired comparisons; and A method having the following.
2. the plurality of features includes text contained in one or more headers related to the dataset; The method of claim 1.
3. the respective embeddings for each of the plurality of features include one or more of word embeddings or context embeddings. The method of claim 1.
4. The weights are calculated using cosine similarity. The method of claim 1.
5. using the encoded dataset to train a machine learning (ML) model; performing one or more operations using the ML model; and The method of claim 1 further comprising:
6. Using the encoded dataset to train the ML model further includes using at least one hyperparameter that sets a weight among the weights to zero in response to the weight satisfying a preset condition. The method of claim 5.
7. performing one or more operations using the ML model includes obtaining predictions of one or more target variables using one or more inputs related to the dataset and a first threshold hyperparameter; The method of claim 5.
8. In response to the plurality of features not describing the dataset, the method includes: after obtaining the prediction of the one or more target variables, further comprising: improving the prediction of the one or more target variables using a second threshold hyperparameter; the second threshold hyperparameter is lower than the first threshold hyperparameter; The method of claim 7.
9. one or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the system to perform operations; The operation is identifying a plurality of features corresponding to a dataset; obtaining a respective embedding for each feature of the plurality of features using a pre-trained generative artificial intelligence model; generating a pairwise comparison of each of the embeddings; generating an encoded dataset by applying weights calculated using the pre-trained generative artificial intelligence model to the paired comparisons, the weights indicating correlations between features in the paired comparisons; and 1. One or more non-transitory computer-readable media having:
10. the plurality of features includes text contained in one or more headers related to the dataset; 10. One or more non-transitory computer-readable media as recited in claim 9.
11. the respective embeddings for each of the plurality of features include one or more of word embeddings or context embeddings.
10. One or more non-transitory computer-readable media as recited in claim 9.
12. The weights are calculated using cosine similarity.
10. One or more non-transitory computer-readable media as recited in claim 9.
13. the plurality of features are determined not to describe the data set based on the weights satisfying a threshold.
10. One or more non-transitory computer-readable media as recited in claim 9.
14. The operation is using the encoded dataset to train a machine learning (ML) model; performing one or more operations using the ML model; and 10. The one or more non-transitory computer-readable media of claim 9, further comprising:
15. performing one or more operations using the ML model includes obtaining predictions of one or more target variables using one or more inputs related to the dataset; 15. One or more non-transitory computer-readable media as recited in claim 14.
16. 1. A system having one or more processors and one or more non-transitory computer-readable storage media storing instructions, The instructions, in response to being executed by the one or more processors, cause the system to: identifying a plurality of features corresponding to a dataset; obtaining a respective embedding for each feature of the plurality of features using a pre-trained generative artificial intelligence model; generating a pairwise comparison of each of the embeddings; generating an encoded dataset by applying weights calculated using the pre-trained generative artificial intelligence model to the paired comparisons, the weights indicating correlations between features in the paired comparisons; and A system that causes an operation having the following steps to be performed.
17. the plurality of features includes text contained in one or more headers related to the dataset; 17. The system of claim 16.
18. the plurality of features are determined not to describe the data set based on the weights satisfying a threshold.
17. One or more systems according to claim 16.
19. The operation is using the encoded dataset to train a machine learning (ML) model; performing one or more operations using the ML model; and The system of claim 16 further comprising:
20. performing one or more operations using the ML model includes obtaining predictions of one or more target variables using one or more inputs related to the dataset; 20. The system of claim 19.