Data adjustment using large language model

By using large language models to analyze and generate prompts for training datasets, the method addresses the issue of dataset diversity, resulting in more robust and accurate machine learning models.

JP2025100333APending Publication Date: 2025-07-03FUJITSU LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024168492
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-21
Filing Date
2024-09-27
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing machine learning models often suffer from inadequate training datasets that lack diversity and representation, leading to inaccurate and unrealistic predictions.

Method used

A method involving large language models is used to analyze and condition training datasets by generating prompts based on dataset characteristics, allowing for the creation of additional data subsets that enhance dataset diversity and inclusiveness, thereby improving the training of machine learning models.

Benefits of technology

The method enhances the robustness and accuracy of machine learning models by expanding the scope and inclusiveness of training datasets, enabling them to make more accurate predictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025100333000001_ABST
    Figure 2025100333000001_ABST
Patent Text Reader

Abstract

To provide a method for adjusting data for a machine learning model using a large language model, a system, and a non-transitory computer readable medium.SOLUTION: A method may include: accessing a dataset including multiple data subsets, each of the data subsets corresponding to a feature of the dataset; analyzing data in the data subsets to determine characteristics of the data; selecting a prompt template from prompt templates for the one of the data subsets on the basis of the determined characteristics of the data; generating prompts using the prompt template and the data from the one of the data subsets; and providing the prompts to a large language model (LLM). The prompts may command the LLM to perform one or more operations with respect to the data of the one of the data subsets. One or more additional data subsets may be created for the dataset on the basis of the response of the LLM.SELECTED DRAWING: Figure 1B
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to using large language models to condition data for machine learning models.

Background Art

[0002] Machine learning (ML) models are trained using training datasets. The quality of the training dataset affects the accuracy and realism of the predictions made by the ML model. For example, the training dataset may define the prediction patterns of the ML model. A sufficiently diverse and representative training dataset that includes various scenarios and features may allow the ML model to make valid predictions for different input data.

[0003] The subject matter claimed in the present disclosure is not limited to embodiments that solve any disadvantages or that operate only in the environments described above. Rather, this background is provided only to illustrate one exemplary technical field in which some embodiments described in the present disclosure may be practiced.

Summary of the Invention

Means for Solving the Problems

[0004] According to one aspect of an embodiment, the method may include accessing a dataset that includes a plurality of data subsets, where each data subset corresponds to a certain feature of the dataset. The data within the one or those data subsets can be analyzed to determine the characteristics of the data. Additionally, based on the determined characteristics of the data of one of the data subsets, a prompt template can be selected from a plurality of prompt templates for one of the data subsets. Using the prompt template and the data from one of the data subsets, a plurality of large language model prompts can be generated. The plurality of large language model prompts can be provided to a large language model. The plurality of large language model prompts can instruct the large language model to perform one or more operations regarding the data of one of the data subsets. One or more additional data subsets can be created for the dataset based on the response of the large language model. Each of the one or more additional data subsets may correspond to a new feature of the dataset.

[0005] The objectives and advantages of the embodiments are realized and achieved at least by the elements, features, and combinations particularly pointed out in the claims. It is understood that both the foregoing general description and the following detailed description are explanatory and not restrictive of the claimed invention.

Brief Description of the Drawings

[0006] Exemplary embodiments are described and explained with further specificity and detail through the accompanying drawings.

[0007]

Figure 1A

[0008]

Figure 1B

[0009]

Figure 1C

[0010]

Figure 2

[0011]

Figure 3

[0012]

Figure 4

[0013]

Figure 5

DETAILED DESCRIPTION OF THE INVENTION

[0014] A machine learning model can be trained using a training data set to make predictions. The training data set can include training instances or individual data points used to train the ML model. The individual data points can correspond to features and target variables that the ML model can be designed to predict. Features can define characteristics of the data that the ML model can use to make predictions. Features can include various data types such as, among other things, numerical, categorical, text-based, etc.

[0015] In some cases, the training dataset may be represented in different formats suitable for the ML model. For example, the training dataset may be represented in a tabular format with multiple columns and rows. In such cases, the columns may represent the features of the training dataset, and the rows may represent the individual instances or data points of the features. In some cases, the training dataset may be adjusted so that the diversity of the features can be improved. For example, the number of features or columns may be increased to improve the diversity of the training dataset.

[0016] Some techniques for adjusting the training dataset may include incorporating additional new information into the existing training dataset. For example, the existing data in the training dataset can be analyzed to determine new information related to the existing data, and that new information may be incorporated to expand the range of data included in the training dataset. A variety of different approaches have been used to determine the new information.

[0017] For example, in one approach, the new information may be determined based on metadata analysis. For example, the metadata describing the information included in the training dataset can be analyzed to determine external information related to the training dataset. Metadata can represent different characteristics of the training dataset, such as table names, column names, data types, etc. A directory containing available reference datasets may be searched using the characteristics determined from the metadata added to the training dataset.

[0018] In another approach, the training dataset can be compared to a database that includes a reference dataset to identify similarly relevant data. The comparison may be made based on a similarity metric. The similarity metric can be calculated based on the semantic similarity between two or more datasets. Such an approach can determine external data related to the training dataset to broaden the scope of the training dataset, but may be limited in the scope of new information that can be added. For example, the scope of external information may be limited to the information present in the reference dataset. Further, such an approach may only add new information to the training dataset rather than improving the existing data within the training dataset.

[0019] The present disclosure may relate, among other things, to systems and methods related to conditioning a training dataset. Conditioning the training dataset can enrich the training dataset, such as by adding additional features or by improving it.

[0020] In some embodiments, the adjustment of the training dataset may be performed using a large language model. In these and other embodiments, information generated by the large language model can be used to add external information to the training dataset and improve the data already present within the training dataset. In some embodiments, feature type inference may be performed with respect to the training dataset, e.g., different types of data within the training dataset can be identified and labeled. For example, if the training dataset is represented as a tabular dataset, different columns can represent different types of data. In such cases, the feature type inference process can determine the type of each column or subset of the training dataset and label it accordingly. Based on the different types of data, various large language model prompts can be generated and provided to the large language model. The responses from the large language model can be used to adjust the training dataset by improving the existing data and / or adding additional data. Adjusting the training dataset using a large language model can improve the scope and inclusiveness of the training dataset. As a result, the machine learning model generated using the training dataset can be improved. For example, the machine learning model can be more robust and can predict the target feature more accurately.

[0021] Embodiments of the present disclosure will be described with reference to the accompanying drawings.

[0022] FIG. 1A shows an exemplary system 100 configured for machine learning training according to one or more embodiments of the present disclosure. In some embodiments, system 100 can include a feature type inference (FTI) process 104, a data conditioning process 108, and a machine learning (ML) model generation process 112. Generally, system 100 may be configured to train an ML model 114. In some embodiments, system 100 may be configured to train an ML model 114 using a dataset 102 that can be adjusted or improved to enhance the training of the ML model 114.

[0023] In some embodiments, the FTI process 104 can obtain the dataset 102. In some embodiments, the dataset 102 may be a training dataset that can be used to train the ML model 114. For example, the dataset 102 may include data suitable for training the ML model 114 to perform one or more operations to generate predictions. For example, the dataset 102 may include data corresponding to features and target features that can be used to train the ML model 114.

[0024] In some embodiments, dataset 102 may include one or more data subsets that may correspond to one or more features. In these and other embodiments, the one or more features may include different types of features. The FTI process 104 can analyze the data in the one or more data subsets to identify the different types of features included in the one or more data subsets. The one or more data subsets may be appropriately labeled to indicate the type of features included in the one or more data subsets. For example, the FTI process 104 may be able to generate a labeled dataset 106 that may correspond to dataset 102 and has feature labels corresponding to the data subsets of dataset 102. In some cases, different data types can include, among other things, categorical variables (e.g., text features), identifier (ID) style features (e.g., numerical IDs, alphanumeric codes, etc.). For example, the first data subset may include an address, and the second data subset may include income. The first data subset and the second data subset may not have labels that identify the type of data. The FTI process labels the first data subset as an address and the second data subset as income.

[0025] In some embodiments, the labeled dataset 106 can be processed by a data conditioning process 108 to generate a conditioned dataset 110. In these and other embodiments, the data conditioning process 108 may include one or more algorithms and / or operations for conditioning the scope of the labeled dataset 106. As an example, dataset 102 may be conditioned such that the scope of dataset 102 can be expanded.

[0026] In some embodiments, the data conditioning process 108 may be configured to instruct a large language model (LLM) to generate one or more additional features with respect to the labeled dataset 106. The one or more additional features may be used to generate the conditioned dataset 110. For example, the conditioned dataset 110 can include additional data in addition to the existing data of the dataset 102. In some embodiments, the one or more additional features generated by the LLM may vary based on the type of features corresponding to one or more data subsets. For example, the one or more additional features can include enhancement of existing data and addition of external data determined based on the existing data. In some cases, the enhancement of the existing data can include additional grouping or partitioning of the existing data such that at least one new feature is generated. The new feature may accentuate a portion of the existing data within the dataset 102. In contrast, the external data can include new data that does not exist in the dataset 102 but may be related to at least one existing feature of the dataset 102.

[0027] In some embodiments, the conditioned dataset 110 may be used to train the ML model 114. For example, the ML model 114 may learn patterns and / or relationships between the features and the target feature in the conditioned dataset 110, whereby the ML model 114 can predict the value of the target feature when the values of other features are given. In some embodiments, the ML model 114 can be any type of ML model. For example, the ML model can be, inter alia, a supervised learning model (e.g., a regression model, a classification model), an unsupervised learning model (e.g., a clustering model), a deep learning model (e.g., a convolutional neural network, a recurrent neural network, a transformer model).

[0028] As shown, the ML model 114 can be used to predict the value of a target feature. For example, a dataset containing one or more of the features of a dataset can be provided to the ML model 114. The ML model 114 may predict the value of the target feature based on the values of the one or more provided features. By providing an adjusted dataset, for example, a dataset with more features, the ML model can more accurately predict the value of the target feature. Thus, the adjustment of the dataset 102 can improve the training of the ML model and improve machine learning techniques.

[0029] FIG. 1B shows an exemplary system 120 configured to adjust a dataset used to train a machine learning model according to one or more embodiments of the present disclosure. In some embodiments, the system 120 may include a prompt generator 124, an LLM 128, and a dataset adjustment process 132. In some embodiments, the prompt generator 124 may be configured to generate one or more prompts 126 for the LLM 128 based on the dataset 122.

[0030] In some embodiments, the dataset 122 may include data that can be used to train an ML model. The dataset 122 may be obtained from any source or constructed using any data integration technique. The data can include character strings containing characters such as numerical data, letters, symbols, or other characters, numbers, or combinations of numbers and characters. The data may also include data in other formats.

[0031] The data within dataset 122 may be organized into one or more data subsets. For example, data of the same category may be organized into the same data subset. For example, the data may include an address, a lot value, a lot size, and a lot improvement. As an example, the data representing the lot value may form part of a data subset. In these and other embodiments, data grouped into the same category of data and data subsets may be referred to as features of dataset 122. As an example, dataset 122 may include tabular data that can be arranged in columns and rows. In these and other embodiments, each of the columns may represent a feature of dataset 122, and each of the rows may have values in one or more of the columns. The values in one of the rows may be associated together. For example, following the previous example, the values for each of the columns in a single row may be associated with the same address.

[0032] In some embodiments, one or more data subsets of dataset 122 may include different types of features. In such cases, different types of features may be identified, and one or more data subsets may be labeled according to the associated type of feature. In some embodiments, one or more data subsets may be labeled by an FTI process. In some embodiments, the operation of the FTI process may be discussed in more detail with respect to the FIT process 104 of FIG. 1A. An example of the FTI process is further described in U.S. Patent Application "Data Set Feature Type Inference" by Sou Hasegawa, Lei Liu, Wei-Peng Chen, filed on December 21, 2023 (Attorney Docket No. F1423.10578US01). This application is hereby incorporated by reference in its entirety into this specification.

[0033] System 120 may be configured to adjust dataset 122 to generate an adjusted dataset 134. System 120 may adjust dataset 122 by using LLM 128 to determine how additional data may be added to dataset 122 or how the data within a feature may be adjusted to create additional features.

[0034] In some embodiments, prompt generator 124 may be configured to analyze dataset 122. Based on the analysis of dataset 122, prompt generator 124 may generate a prompt 126 that may be provided to LLM 128. Prompt 126 may be used by LLM 128 to determine how to adjust dataset 122. In some embodiments, LLM 128 may refer to a sophisticated artificial intelligence system trained on a vast amount of text data to understand and generate human-like language prompts and responses. The LLM model may be designed to process and understand the complexity of natural language, including syntax, semantic content, and context. LLM 128 may understand various forms of human language and generate human-like responses. Additionally, with relevant training methods, knowledge of facts can be retrieved from the LLM. For example, during pre-training, the LLM model may be exposed to a large amount of diverse text data from the Internet and other sources such as articles, books, websites, and various documents containing factual information. The factual information may be utilized to adjust the dataset in a reasonable manner.

[0035] In some embodiments, the prompt generator 124 can analyze the dataset 122 to determine one or more characteristics of each of the data subsets. For example, the analysis can determine a first characteristic of the data of the first data subset and a second characteristic of the data of the second data subset. Based on the characteristics of the data of the data subsets, the prompt generator 124 can select one or more data subsets for which a prompt 126 can be generated. In these and other embodiments, the prompt generator 124 may generate a prompt 126 based on the characteristics of the data of the selected data subsets. In these and other embodiments, different prompts 126 may be generated for different data subsets based on the characteristics of the data in each of the data subsets.

[0036] In some embodiments, each of the prompts 126 may further include one or more commands for the LLM 128. The commands can include operations to perform on the data of the selected data subset. In these and other embodiments, different commands can be provided to the LLM 128 for different characteristics of the data of the data subsets and / or different types of features associated with the data subsets. For example, the prompts 126 for text features and ID style features may include different commands. Text features are features represented using text in different languages. In some cases, text features can include categorical features that take a limited number of possible values representing different categories. The categories can be nominal (e.g., without a particular order) or ordinal (e.g., in a particular order). For example, nominal categorical features can include, among other things, various names of schools, countries, colors. Some examples of ordinal categorical features can include, among other things, size (e.g., small, medium, large), school rankings.

[0037] In some embodiments, each prompt 126 may be generated to include one or more individual values of a data subset of the dataset 122 and one or more commands. The commands may include one or more operations to be performed by the LLM 128 with respect to the individual values of the data subset. As an example, the certain portion may be values from multiple rows from a single column within the dataset 122.

[0038] In some embodiments, the prompt generator 124 may generate a prompt 126 for a set of data from a data subset. For example, the prompt generator 124 may select a random number of individual values from the data subset. Thus, the set of data may not include all of the individual values from the data subset. For example, if the prompt 126 includes a command related to determining how to split a character string into two or more substrings, only the set of data may be used to generate the prompt 126 instead of the entire data subset. Using only the set of data may reduce the processing time and resources required to determine how to split the character string. In some embodiments, in response to determining how to split a character string based on the prompt 126. In these and other embodiments, data splitting rules may be determined based on the response from the LLM resulting from the prompt. In response to determining the data splitting rules using the set of data, the data splitting rules may be applied to the entire data subset.

[0039] In some embodiments, the prompt generator 124 may generate a prompt 126 using a prompt template. The prompt template may include a prompt that may include one or more blank fields. The prompt may be a word sequence that conveys a command for the LLM 128. The blank fields may be completed using the data of the data subset and / or the names of the features associated with the data subset. For example, each prompt 126 for a data subset can include the same word sequence, but can include different values from the data subset in the blank fields. For example, the prompt template may be something like: divide [feature] [data value] into meaningful substrings [divide [feature] [data value] into meaningful substrings]. [Feature] and [data value] may be blank fields within the prompt template. The feature blank field may be filled with the name of the feature associated with the data subset. The data value blank field may be filled with an individual value from the data subset.

[0040] In some embodiments, the prompt 126 may instruct the LLM 128 to perform one or more operations related to the data of the data set 122. For example, the prompt 126 may include instructing the LLM 128 to generate additional data regarding specific data of the data set 122 provided to the LLM 128 using the prompt 126. For example, the prompt 126 may instruct the LLM 128 to split a value such as a string into multiple values. As a result of splitting each value of the first data subset into multiple values, the first data subset may be split into two or more individual data subsets.

[0041] In some embodiments, the prompt 126 may instruct the LLM 128 to cluster the individual values of a data subset into two or more groups. For example, the individual values may be clustered together into groups based on at least one similarity in the format or meaning of the individual values. Each group may be assigned a cluster ID, and a new data subset can include the cluster ID for each individual value of the data subset. In some embodiments, the improvement of existing data can be described in more detail with respect to FIGS. 2 and 3 of the present disclosure.

[0042] Additionally or alternatively, in some embodiments, the prompt 126 may instruct the LLM 128 to determine additional data related to one or more data subsets of the data set 122. In some cases, the additional data can include features not present in the data set 122. For example, different prompts 126 may be generated for text features and ID style features. For example, the prompt 126 for ID style features can include clustering or splitting commands, and the prompt 126 for text features can include commands for determining external data.

[0043] In some embodiments, the LLM 128 may provide a response 130 based on the prompt 126. For example, the LLM 128 may perform one or more operations included in the prompt 126 with respect to a portion of the data subset included in the prompt 126. For example, the LLM 128 may perform operations such as, among other things, splitting data, clustering data, generating new data, etc.

[0044] In some embodiments, the response 130 may be evaluated. Based on the evaluation of the response, additional prompts 126 may be generated. The additional prompt may include the response and may include instructions for the LLM 128 to provide a different response than the initial response.

[0045] In some embodiments, the response 130 may be used by the data adjustment process 132 to generate the adjusted data set 134. For example, the data adjustment process may generate one or more additional data subsets to be included in the data set 122, at least based on the response 130. For example, the response 130 may include additional data generated for one or more data subsets of the data set 122, based on one or more operations such as splitting, clustering, and generating external information. In such cases, the additional data may be used to construct additional data subsets that may be added to the data set 122 to generate the adjusted data set 134. In some embodiments, the additional data subsets may replace existing data subsets. For example, the first data subset may be split into a second data subset and a third data subset. In some cases, the first data subset may be removed from the data set 122, and the second data subset and the third data subset may be added to the data set 122.

[0046] Modifications, additions, or omissions may be made to the system 120 without departing from the scope of the present disclosure. For example, in some embodiments, the system 120 may include any number of other components not explicitly illustrated or described.

[0047] As another example, in some embodiments, the prompt 126 may be used to determine additional characteristics of the data. For example, the initial characteristics of the data may be determined and used to generate one or more prompts 126. The one or more prompts 126 may be provided to the LLM 128, and responses from the LLM 128 may be collected. The responses can assist in determining other characteristics of the data. The other characteristics may be used to generate additional prompts 126 that may be provided to the LLM 128.

[0048] As another example, in some embodiments, any number of LLMs may be used to generate the adjusted dataset 134. For example, FIG. 1C shows a system configured to adjust a dataset used to train a machine learning model according to one or more embodiments of the present disclosure. In some embodiments, the dataset 152 may be used to generate one or more prompts provided to the first LLM 154 and the second LLM 158. For example, the dataset 152 may be used to generate a first set of prompts for the first LLM 154 and a second set of prompts for the second LLM 158. In some embodiments, the first set of prompts and the second set of prompts may be generated using a process similar to the process performed to generate the prompt 126 of FIG. 1B.

[0049] In some embodiments, the prompts may be split into two or more groups. For example, the prompts may be split based on the type of features associated with the data subsets. For example, prompts generated based on text features may be grouped together into a first set of prompts, and prompts generated based on ID style features may be grouped together into a second set of prompts. In some embodiments, the first set of prompts may be provided to the first LLM 154, and the second set of prompts may be provided to the second LLM 158.

[0050] In these and other embodiments, the first adjusted dataset 156 and the second adjusted dataset 160 may be generated based on the responses of the first LLM 154 and the second LLM 158, respectively. For example, the first LLM 154 may generate a first response based on a first set of prompts, and the second LLM 158 may generate a second response based on a second set of prompts. The first response may be used to determine the first adjusted dataset 156, and the second response may be used to determine the second adjusted dataset 160.

[0051] In some embodiments, the first adjusted dataset 156 may include additional data (e.g., external information) related to text features, and the second adjusted dataset 160 may include additional data (e.g., segmented data) related to ID style features. For example, the first adjusted dataset 156 and the second adjusted dataset 160 may include additional data subsets for the dataset 122.

[0052] In some embodiments, the first LLM 154 and the second LLM 158 may perform one or more operations in parallel based on the first group and the second group, respectively. For example, the first LLM 154 and the second LLM 158 may determine the first adjusted dataset 156 and the second adjusted dataset 160 simultaneously. In other embodiments, the first LLM 154 and the second LLM 158 may operate in a sequential order. For example, the first LLM 154 may perform an operation before the second LLM 158. Although FIG. 1C shows two LLMs, any suitable number of LLMs may be used.

[0053] In some embodiments, the first adjusted data set 156 and the second adjusted data set 160 may be combined with the data set 152 to generate an adjusted data set 162. The adjusted data set 162 may include the data set 122 and additional data subsets included in the first adjusted data set 156 and the second adjusted data set 160.

[0054] FIG. 2 shows a flowchart of an exemplary method 200 that includes operations performed by a computing system to adjust a data set with respect to ID style values, according to one or more embodiments of the present disclosure. The method 200 may be performed by any suitable system, apparatus, or device. For example, the method 200 may be implemented using the system 100 of FIG. 1A or the system 120 of FIG. 1B. Although shown in discrete blocks, the steps and operations associated with one or more blocks of the method 200 may, depending on the specific implementation, be divided into additional blocks, combined into fewer blocks, or removed.

[0055] Method 200 may include block 202. In block 202, one or more data subsets of a dataset may be identified. The data subsets may be identified based on an analysis of the dataset. The dataset may be an example of dataset 102 in FIG. 1. In these and other embodiments, those data subsets may be identified in response to each of the data subsets including a percentage of values within the data subset that meet a threshold of values that are considered unique. A value is considered unique if it is not repeated within a given data subset. For example, the data subset may include different identifiers for classes within a school, such as PSY101, PSY102, PSY103, and PSY104, and no identifier is assigned to more than one class. In these and other embodiments, a data subset may be considered unique in response to a threshold percentage of values within the data subset being unique. In some cases, the threshold percentage may be determined as a predetermined number. For example, the threshold percentage may be predetermined as 80% or 90%. In other cases, the threshold percentage may be determined using one or more algorithms. For example, an ML model may be trained to determine which data subsets may be considered unique to contain unique values.

[0056] In some embodiments, the one or more data subsets having a high percentage of unique values may be represented in an ID style value. The ID style value may be considered an ID style in which values within the same data subset share a consistent format or structure. For example, continuing with the example of identifiers for different classes, the values within the data subset may be represented in a [subject][level] format (e.g., [PSY]

[0101] , where PSY represents psychology and 101 represents the lowest level). In some cases, the ID style value may include a numeric ID and / or an alphanumeric ID in which the value includes at least one numeric value.

[0057] In block 204, one or more data subsets may be selected from the identified data subsets. For example, the identified data subsets may be analyzed to determine whether the identified data subsets contain independent and identically distributed (IID) values. Independent values may be characterized as those where the occurrence of one individual value or value does not affect the occurrence of another individual value or value. Values are considered to be of the same distribution if they follow the same probability distribution. IID values may be values that are independent of each other and drawn from the same distribution probability. For example, test scores from multiple classes may be IID. For example, the test scores may be determined independently of each other, and the scores for each class may be drawn from the same distribution.

[0058] In some embodiments, in response to determining that the one or more data subsets are not IID, the one or more data subsets may be retained to adjust the data set. In block 208, the data of the selected data subset may be analyzed to determine whether the selected data subset contains data that is semantically significant. In these and other embodiments, semantically significant data can refer to data that is meaningful or conveys meaningful content in a particular context. For example, semantically significant data can be interpreted using the actual meaning of the words, phrases, and / or terms included in the data. For example, an individual value of a data subset may include the character string "HIST101", and since "HIST" can represent a word or character string that has a meaning understood in the context, namely "history", this individual value may be semantically significant. Further, if the column name "Course ID" is considered, the individual value "HIST101" may be more likely to be determined to be semantically significant.

[0059] In some embodiments, the determination of whether individual data subsets of a selected data subset are semantically significant may be determined using an LLM. For example, one or more LLM prompts may be generated. The LLM prompt may instruct the LLM to determine whether one or more individual values of an individual data subset are semantically significant. In some embodiments, instead of all of the individual values of an individual data subset, a sample set of individual values may be provided to the LLM in the LLM prompt. Providing a sample set of individual values may reduce the time taken to determine whether an individual data subset is semantically significant. In some embodiments, column names may also be carried in the prompt to allow the LLM to make better predictions.

[0060] In these and other embodiments, the sample set of individual values may be randomly selected from the individual values of an individual data subset. In some embodiments, the values in the sample set may be inspected to remove any duplicate values. For example, if there is an individual value that is included in the sample set more than once, it may be detected and removed from the sample set. In response to determining that there are no duplicates in the sample set, the individual values of the sample set may be provided to the LLM as part of one or more prompts to determine whether the individual values are semantically significant.

[0061] In some embodiments, the response from the LLM may provide a first list of individual values that are semantically significant and a second list of individual values that are not semantically significant. In these and other embodiments, the number of individual values in the first and second lists may be determined. In some embodiments, the LLM may be prompted to provide the number of individual values in the first and second lists along with the first and second lists.

[0062] In some embodiments, to determine whether an individual data subset is semantically significant as a whole, the number of semantically significant values (e.g., the number of individual values in a first list) may be compared to the number of individual values that are not semantically significant (e.g., the number of individual values in a second list). In some embodiments, an individual data subset may be semantically significant if there are more (at least one more) individual values that are semantically significant than there are individual values that are not semantically significant in the sample set. In some cases, the number of individual values in the sample set may be set as an odd number other than one. This is to eliminate possible cases having the same number of individual values that are semantically significant and individual values that are not semantically significant. In other embodiments, for an individual data subset to be semantically significant, a threshold percentage of the individual values may need to be determined to be semantically significant.

[0063] In response to determining that the one or more data subsets are semantically significant, in block 210, LLM clustering may be performed on the one or more data subsets. LLM clustering may include grouping one or more individual values of an individual data subset into one or more clusters that may share one or more characteristics. In such a case, each individual value of an individual data subset may be embedded using an LLM such that it is represented as an embedding or a vector. For example, the embedding may include a numerical representation of a word, a sentence, and / or a phrase. The embedding may be created by encoding text information into a high-dimensional vector in a continuous space. The individual values may be clustered based on the similarity between the embeddings, and each cluster may be assigned a unique category ID. As an example, the individual values may be associated with course IDs for different courses (e.g., courses at a university). Each course may be associated with a course ID that represents the course. The course ID may include two parts, namely, a first part indicating the field of study (e.g., "PSY" for psychology, "HIST" for history, etc.) and a second part indicating the level of the course (e.g., 101 for the lowest level, 102 for a higher level, etc.). In such a case, individual course IDs in the same field of study (e.g., course IDs starting with "PSY") may be grouped into a cluster. Further, since semantic information is considered by the LLM for clustering, similar courses may be grouped together into a cluster. For example, courses "ENG" and "ESL", which are English courses and English as a second language courses, respectively, may be grouped into a cluster. In such a case, a new data subset may be generated that indicates the cluster ID corresponding to the individual values.

[0064] In some cases where the one or more data subsets are not semantically significant, in block 212, a data splitting rule for the individual values of the one or more data subsets may be determined using an LLM. For example, the individual values may be split into two or more split values. In some embodiments, the process of determining the splitting rule using an LLM can be described in more detail with respect to FIG. 3 of the present disclosure.

[0065] In these and other embodiments, syntactic clustering for the split values can be performed in block 214. For example, similarly split values may be grouped into a new data subset. For example, the first value of the first data subset may be "B4064600", and the second value of the first data subset may be "B4064900". The splitting rule can split the first value into "B" and "4064600", and the second value into "B" and "4064900". In such a case, the two "B"s may be clustered together into a second data subset, and "4064600" and "4064900" may be clustered together into a third data subset. In these and other embodiments, the second and third data subsets may not have been part of the data set.

[0066] In block 216, new data subsets may be added to the data set. For example, the second and third data subsets may be added to the data set. As another example, new data subsets generated using LLM clustering in block 210 may also be added to the data set. In some embodiments, the first data subset used to create the second and third data subsets may be replaced by the second and third data subsets. In other embodiments, the second and third data subsets may be added in addition to the first data subset.

[0067] Without departing from the scope of the present disclosure, modifications, additions, or omissions may be made to method 200. For example, the steps and operations outlined are provided as examples only, and some of the steps and operations may be optional and may be combined into fewer steps and operations or expanded into additional steps and operations without compromising the essence of the disclosed embodiments.

[0068] For example, a new data subset may be analyzed to detect any exceptions. For example, a new data subset may contain individual values that do not conform to the format or structure of the data subset. For example, the individual values of a new data subset may correspond to numbers. In some cases, certain individual values may contain exceptional values that are different from other individual values. For example, a particular individual value may sometimes include special characters along with a number (e.g., 10+). In such cases, that particular individual value may be handled such that the format of the exceptional value is adjusted to the same format as other individual values within the data subset (e.g., 10). In some embodiments, the adjustment of exceptional values may be performed using an LLM. For example, a prompt may be generated for each exceptional value to instruct the LLM to adjust the format of the exceptional value. The response of the LLM may be used to replace the exceptional value such that the particular individual value conforms to the format of the data subset.

[0069] Figure 3 shows a flowchart of an exemplary method 300 for determining a data splitting rule according to one or more embodiments of the present disclosure. In some embodiments, method 300 may represent various steps of block 212 in FIG. 2. Method 300 may be executed by any suitable system, apparatus, or device. For example, method 300 may be implemented using system 100 in FIG. 1A or system 120 in FIG. 1B. Although shown as discrete blocks, the steps and operations associated with one or more blocks of method 300 may be divided into additional blocks, combined into fewer blocks, or deleted, depending on the specific implementation. Method 300 may be executed with respect to one or more data subsets of a dataset that includes semantically insignificant data, such as the data subset determined with respect to block 208 in FIG. 2.

[0070] Method 300 may include block 302. In block 302, one or more values of a data subset may be provided to an LLM as part of an LLM prompt. The LLM prompt may instruct the LLM to split the individual values of the data subset. For example, the LLM prompt may instruct the LLM to split each individual value of one or more values into two or more significant substrings. In some embodiments, instead of the entire data subset, a random number of individual values may be provided to the LLM for splitting, which may improve the time spent by the LLM.

[0071] In block 304, the data splitting rule may be determined based on the response of the LLM. For example, the response of the LLM may provide two or more substrings for each individual value provided to the LLM in the LLM prompt. In these and other embodiments, the splitting pattern can be determined based at least on how the substrings were generated. For example, the response may be analyzed to determine the splitting pattern between the substrings provided by the LLM. Some examples of splitting patterns can include alphabet / number splitting (e.g., from "XLY101" to "XLY" and "101"), special character splitting (e.g., from "Kepler-10a" to "Kepler" and "10a"), and specific index splitting (e.g., from "B10c3748" to "B" and "10c3748", and from "C67d7922" to "C" and "67d67922"). In some embodiments, the splitting pattern of the data subset may be determined based on the fact that a threshold number of the split individual values share the same splitting pattern. For example, the majority (e.g., more than half) of the split individual values share the same splitting pattern.

[0072] In block 306, the determined data splitting rule may be applied to all individual values of the individual data subsets. For example, the same splitting pattern may be applied to each of the individual values, such that the entire data subset is split according to the same pattern. In some cases, the substrings of the individual values split according to the same pattern may be grouped together into a set of substrings.

[0073] In block 308, it may be determined whether the quality of the determined data splitting rule is sufficient to train the ML model. For example, sets of substrings can be analyzed to determine whether each set of substrings can be significant when training the ML model. For example, each set of substrings can be regarded as a new data subset and analyzed to determine whether that set of substrings is relevant for the ML model. A set of substrings may be relevant if the substrings are complete (e.g., contain sufficient information), accurate (e.g., have no errors, missing values, or outliers), diverse (e.g., cover diverse scenarios and situations), etc. For example, a set of substrings may not be considered significant if the substrings within the set are constant (e.g., one value for all substrings), contain a large proportion of missing values, and / or are highly correlated with other substrings within the set. In such cases, the set of substrings or the new data subset may not provide additional information to the ML model such that splitting the data subset is significant.

[0074] In response to determining that the data splitting rule does not provide sufficient data to train the ML model, in block 310, the determined splitting rule may be discarded and the LLM may be prompted to split the individual values of the data subset in a different way again. In these and other embodiments, returning to block 304, new data splitting rules can be determined using new substrings. In some embodiments, the loop of discarding the determined data splitting rule and determining a new data splitting rule (e.g., the loop of blocks 304, 306, 308, and 310) can be repeated until a sufficient data splitting rule is determined or until a threshold number of loop sequential iterations is met. For example, the threshold number can be set for the number of times the loop can be repeated to limit processing time and / or resource usage.

[0075] In response to determining that the quality of the data splitting rule is sufficient to train the ML model, one or more additional data subsets may be created for the data set. For example, the additional data subset may incorporate substrings. For example, using the example of the specific index splitting of "B10c3748" and "C67d7922", "B" and "C" may be grouped together in a first additional data subset, and "10c3748" and "67d7922" may be grouped in a second additional data subset.

[0076] Modifications, additions, or omissions may be made to method 300 without departing from the scope of the present disclosure. For example, the steps and operations outlined are provided as examples only, and some of the steps and operations may be optional without detracting from the essence of the disclosed embodiments, may be combined into fewer steps and operations, or may be expanded into additional steps and operations.

[0077] FIG. 4 shows a flowchart of another exemplary method 400 for adjusting one or more data subsets of a data set, according to one or more embodiments of the present disclosure. Method 400 may be performed by any suitable system, apparatus, or device. For example, method 400 may be implemented using system 100 of FIG. 1A or system 120 of FIG. 1B. Although shown in discrete blocks, the steps and operations associated with one or more blocks of method 400 may, depending on the specific implementation, be divided into additional blocks, combined into fewer blocks, or deleted. In some embodiments, method 400 may be performed with respect to one or more data subsets of a data set that includes categorical features or text features.

[0078] Method 400 may include block 402. In block 402, one or more data subsets of a dataset including text features may be obtained. For example, the features of the dataset or data subsets may be analyzed to determine a data subset including text features as described with respect to prompt 126 of FIG. 1B. In such a case, individual data subsets of the one or more data subsets may include multiple individual phrases. In some embodiments, the one or more data subsets may be analyzed to define one or more additional features or data subsets.

[0079] In block 404, sub - phrases of the individual phrases of the one or more data subsets may be identified. In some embodiments, a sub - phrase may include a part of an individual phrase. For example, an individual phrase may be divided into distinct parts or sub - phrases. Sub - phrases may be identified such that each sub - phrase has an independent meaning. For example, it may be an individual phrase "ABC University of XYZ", and in such a case, the sub - phrases may include "ABC", "University", and "XYZ", and each sub - phrase has an independent meaning. In response to determining two or more sub - phrases, the sub - phrases may be compared with the corresponding individual phrases to determine the similarity between the sub - phrases and the individual phrases. For example, by comparing the contextual meaning of the sub - phrases with the contextual meaning of the individual phrases, it can be determined how well the sub - phrases represent the individual phrases. Sub - phrases that sufficiently represent the meaning of the corresponding individual phrases may be identified as key phrases associated with the individual phrases.

[0080] In some embodiments, the comparison between a subphrase and the corresponding individual phrases may be performed using the embeddings of the subphrase and the individual phrases. For example, in block 406, the individual phrases and subphrases may be embedded using an LLM. In such cases, the embeddings generated by the LLM may include numerical representations of words, sentences, and / or phrases. The embeddings including numerical representations may be used to compare the subphrase with the corresponding individual phrases.

[0081] In block 408, key phrases corresponding to the individual phrases may be determined based on the embeddings. In some embodiments, the key phrases may be determined based on a comparison between the embeddings corresponding to the individual phrases and the corresponding subphrases. For example, the individual phrases may be compared with each subphrase to identify one or more subphrases that appropriately represent the individual phrases.

[0082] In some cases, the similarity may be determined using cosine similarity. For example, the cosine similarity between the embedding corresponding to an individual phrase and each of the sub - phrases may be determined. In some cases, the sub - phrase having the highest cosine similarity may be identified as the key phrase. For example, a number of sub - phrases in order of decreasing cosine similarity may be selected as key phrases. In some cases, the number may be pre - determined. For example, the top 10 sub - phrases may be selected as key phrases. Continuing the above example as an example of comparing a sub - phrase with the corresponding individual phrase, "ABC" may be embedded, "university" may be embedded, and "ABC University of XYZ" may be embedded. The cosine similarity between the embeddings corresponding to "ABC" and "ABC University of XYZ", between "university" and "ABC University of XYZ", and between "XYZ" and "ABC University of XYZ" may be calculated. In some cases, the cosine similarity between the embeddings corresponding to "ABC" and "ABC University of XYZ" and between "university" and "ABC University of XYZ" may meet the threshold, while the cosine similarity between "XYZ" and "ABC University of XYZ" may not meet the threshold, indicating that "ABC" and "university" represent "ABC University of XYZ" better than "XYZ".

[0083] In some embodiments, the key phrases may be used directly to generate additional data subsets. For example, at block 416, additional data subsets corresponding to a set of subphrases may be determined. For example, continuing with the above example, at block 416, "ABC" may be considered to be in the first additional data subset, "University" may be placed in the second additional data subset, and "XYZ" may be placed in the third additional data subset. Other individual phrases within the data subset may be split into subphrases according to the same splitting pattern (e.g., [school name], [school type], and [location]). In some cases, an individual phrase may not include all three parts. For example, a subphrase may be something like "DEF University" without a location. In such a case, a part of the third additional data subset corresponding to "DEF University" may be left empty or missing.

[0084] In block 410, the LLM may be prompted to improve at least one data subset based on the key phrases. For example, the key phrases may be used to generate one or more prompts for the LLM. In some embodiments, one or more prompts may be generated using a prompt template. In some embodiments, one or more prompts may instruct the LLM to generate a response that includes external data. External data may be defined as data that was not previously included in the data set or that can be derived from the data set. For example, the LLM may be prompted to generate new data that is related to but not included in the data of the data set. For example, continuing the above example, "ABC" and "university" or combinations thereof (e.g., "ABC University") may be used to generate a prompt. As an example, a prompt template selected from a plurality of prompt templates may be "What type of school is [key phrase]". The LLM may use external data not included in the data set to determine that "ABC University is a private school". Such prompts may be generated for each of the individual phrases included in the data subset.

[0085] In block 416, the response from the LLM (e.g., "ABC University is a private school.") may be used to generate one or more data subsets. For example, the response may indicate the school types of different schools included in a particular data subset. In such a case, the response may be used to generate a new data subset representing the school types.

[0086] In some embodiments, the individual phrases may be used to generate one or more prompts for the LLM without determining key phrases. For example, in block 412, the LLM may be prompted to generate new data from the individual phrases based on external data known to the LLM. For example, the individual phrases may be plugged into a prompt template to generate new data. For example, instead of a prompt such as "What type of school is ABC University?", the prompt may be "What school is XYZ's ABC University?". The response from the LLM may be used in block 416 to generate one or more additional data subsets.

[0087] In some embodiments, clustering techniques may be used to define at least one additional data subset from one or more data subsets. In some embodiments, the clustering technique may be similar to the clustering technique performed on the semantically significant data in block 210 of FIG. 2. For example, one or more individual phrases of a particular data subset may be grouped into one or more clusters that share one or more characteristics. In some embodiments, the clustering may be performed using the LLM. For example, one or more prompts for the LLM may be generated to instruct the LLM to group one or more individual phrases into one or more clusters based on the shared characteristics. For example, the individual phrases of the data subset may each correspond to a school name. In some cases, the individual phrases may be grouped alphabetically.

[0088] In some embodiments, the clustering technique may be implemented with respect to the embeddings. For example, in response to determining the embeddings at block 406, the LLM may be prompted at block 414 to cluster individual phrases based on the embeddings corresponding to the individual phrases, similar to the process performed at block 210 of FIG. 2.

[0089] In some embodiments, the clustering technique at block 414 may be implemented with respect to the individual phrases instead of the embeddings. For example, in response to obtaining a data subset of the dataset at block 402, the individual phrases of the data subset may be clustered at block 414.

[0090] In some embodiments, the response from the LLM may be used to generate an additional data subset. For example, one or more clusters to which individual phrases are assigned may be used to generate an additional data subset at block 416.

[0091] For example, in some embodiments, the individual phrases may be grouped based on the location of the school. For example, the LLM may determine the location of the school corresponding to the individual phrases of the data subset and be prompted to group the individual phrases into one or more clusters based on the location (e.g., state, region, etc.). In some embodiments, the individual phrases may be grouped based on any other shared characteristic.

[0092] Modifications, additions, or omissions may be made to method 400 without departing from the scope of the present disclosure. For example, the steps and operations outlined are provided by way of example only, and some of the steps and operations may be optional, may be combined into fewer steps and operations, or may be expanded into additional steps and operations without detracting from the essence of the disclosed embodiments.

[0093] FIG. 5 shows a block diagram of an exemplary computing system 500 according to at least one embodiment of the present disclosure. The computing system 500 may be configured to implement or direct one or more suitable operations described in the present disclosure. For example, the computing system 500 may be configured to execute one or more blocks of the method 200 of FIG. 2, the method 300 of FIG. 3, and the method 400 of FIG. 4. For example, the computing system 500 may be configured to generate prompts for an LLM. The computing system 500 may include a processor 550, a memory 552, and a data storage 554. The processor 550, the memory 552, and the data storage 554 may be communicatively coupled.

[0094] Generally, the processor 550 may include any suitable dedicated or general-purpose computer, computing entity, or processing device that includes various computer hardware or software modules, and may be configured to execute instructions stored on any applicable computer-readable storage medium. For example, the processor 550 may include a microprocessor, a microcontroller, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or any other digital or analog circuit configured to interpret and / or execute program instructions and / or process data. Although shown as a single processor in FIG. 5, the processor 550 may include any number of processors configured to individually or collectively execute or direct the execution of any number of operations described in the present disclosure. Additionally, one or more of the processors may be present on one or more different electronic devices such as different servers.

[0095] In some embodiments, the processor 550 may be configured to interpret and / or execute program instructions stored in the memory 552, the data storage 554, or both the memory 552 and the data storage 554, and / or to process data. In some embodiments, the processor 550 may fetch program instructions from the data storage 554 and load the program instructions into the memory 552. After the program instructions are loaded into the memory 552, the processor 550 may execute the program instructions.

[0096] The memory 552 and the data storage 554 may include a computer-readable storage medium that carries or stores computer-executable instructions or data structures. For example, such a computer-readable storage medium may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory devices (e.g., solid state memory devices), or any other non-transitory storage medium that can be used to carry or store specific program code in the form of computer-executable instructions or data structures and that can be accessed by a general purpose or special purpose computer, including a tangible or non-transitory computer-readable storage medium. In these and other embodiments, the term "non-transitory" as used in this disclosure should be construed to exclude only those types of transitory media that are considered outside the scope of patentable subject matter in the Federal Circuit decision of In re Nuijten, 500 F.3d 1346 (Fed. Cir. 2007).

[0097] The above combinations may also be included within the scope of the computer-readable storage medium. The computer-executable instructions may include, for example, instructions and data configured to cause the processor 550 to perform certain operations or groups of operations.

[0098] Modifications, additions, or omissions may be made to computing system 500 without departing from the scope of the present disclosure. For example, in some embodiments, computing system 500 may include any number of other components not explicitly illustrated or described.

[0099] The above disclosure is not intended to limit the present disclosure to the exact form disclosed or to a particular field of use. Thus, various alternative embodiments and / or modifications to the present disclosure are contemplated in light of the present disclosure, whether explicitly described or implied herein. Although embodiments of the present disclosure have been described in this manner, it can be recognized that changes can be made in form and detail without departing from the scope of the present disclosure. Accordingly, the present disclosure is limited only by the claims.

[0100] In some embodiments, the different components, modules, engines, and services described herein may be implemented as objects or processes (e.g., as separate threads) running on a computing system. Although some of the systems and methods described herein are generally described as being implemented in software (stored in and / or executed by general-purpose hardware), specific hardware implementations or combinations of specific software and specific hardware implementations are also possible and contemplated.

[0101] In accordance with general practice, the various features shown in the drawings may not be drawn to scale. The figures presented in this disclosure are not intended to be actual views of any particular apparatus (e.g., device, system, etc.) or method, but are merely idealized representations used to illustrate various embodiments of the disclosure. Thus, the dimensions of the various features may be arbitrarily enlarged or reduced for clarity. Additionally, some of the drawings may be simplified for clarity. Thus, the drawings do not necessarily show all of the components of a given apparatus (e.g., device) or all of the operations of a particular method.

[0102] As used herein, particularly in the appended claims (e.g., the body of the appended claims), terms are generally intended to be open terms (e.g., the term "comprising" should be interpreted as "comprising, but not limited to", the term "having" should be interpreted as "having at least", the term "including" should be interpreted as "including, but not limited to", etc.).

[0103] Furthermore, if a specific number of claim recitations is intended, such intent is explicitly recited in the claim, and if there is no such recitation, such intent does not exist. For example, for purposes of illustration, the following appended claims may include the use of introductory phrases "at least one" and "one or more" to introduce claim recitations. However, the use of such phrases should not be construed as implying that the introduction of a claim recitation by the indefinite article "a" or "an" limits any particular claim that includes such introduced claim recitation to an embodiment that includes only one such recitation, even if the same claim includes both an introductory phrase such as "one or more" or "at least one" and an indefinite article such as "a" or "an" (e.g., "a" and / or "an" should be interpreted as meaning "at least one" or "one or more"). The same applies to the use of the definite article to introduce a claim recitation.

[0104] In addition, even if a specific number recited in a claim being introduced is explicitly recited, it is understood that such recitation should be interpreted to mean at least that recited number (e.g., a recitation of only "two recitations" without other modifiers means at least two recitations, or two or more recitations). Further, when conventional expressions similar to "at least one of A, B, and C" or "one or more of A, B, and C" are used, generally, such syntax is intended to include only A, only B, only C, both A and B, both A and C, both B and C, or A, B, and C, etc. For example, the use of the term "and / or" is intended to be interpreted in this way.

[0105] Furthermore, any disjunctive phrase presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to contemplate the possibility of including one of the terms, any of the terms, or both terms. For example, the phrase "A or B" should be understood to include the possibilities of "A" or "B" or "A and B".

[0106] Furthermore, the use of terms such as "first," "second," "third," etc. is not necessarily used in this specification to imply a particular order or number of elements. In general, terms such as "first," "second," "third," etc. are used as general identifiers to distinguish different elements. When terms such as "first," "second," "third," etc. do not indicate an implied particular order, these terms should not be understood to imply a particular order. Furthermore, when terms such as "first," "second," "third," etc. do not indicate an implied number of elements, these terms should not be understood to imply a particular number of elements. For example, a first widget may be described as having a first side, and a second widget may be described as having a second side. The use of the term "second side" with respect to the second widget may be for the purpose of distinguishing such a side of the second widget from the "first side" of the first widget, and not for implying that the second widget has two sides.

[0107] All examples and conditional language recited in this disclosure are intended for educational purposes to assist the reader in understanding the disclosure and the concepts contributed by the inventor to further the art, and are to be construed as not being limited to such specifically recited examples and conditions. Although embodiments of the disclosure have been described in detail, various changes, substitutions, and alterations can be made without departing from the spirit and scope of the disclosure.

[0108] Regarding embodiments including the above embodiments, the following appendices are further disclosed. (Appendix 1) Accessing a data set including a plurality of data subsets, each data subset corresponding to a certain feature of the data set; Analyzing data in one of the data subsets to determine characteristics of the data; Based on the determined characteristics of the data of the one data subset among the data subsets, selecting a prompt template from a plurality of prompt templates for the one data subset among the data subsets; Using the prompt template and the data from the one data subset among the data subsets to generate a plurality of large language model prompts; Providing the plurality of large language model prompts to a large language model, wherein the plurality of large language model prompts instruct the large language model to perform one or more operations regarding the data of the one data subset among the data subsets; Generating one or more additional data subsets for the data set based on the response of the large language model, wherein each of the one or more additional data subsets corresponds to a new feature of the data set, including the steps of; Method. (Appendix 2) Training a machine learning (ML) model using the data set; Performing one or more operations using the ML model The method according to Appendix 1, further comprising. (Appendix 3) The data from the one data subset among the data subsets includes character strings, the plurality of large language model prompts include splitting each of the character strings into two or more substrings, and the one or more additional data subsets are generated using the two or more substrings, the method according to Appendix 1. (Appendix 4) A part of the data of the one data subset among the data subsets is provided to the large language model, and the method includes: Determining a string splitting rule based on the response of the large language model; Using the string splitting rule, splitting the remaining part of the data of the one data subset among the data subsets into two or more sub - strings The method according to Appendix 3, further comprising. (Appendix 5) Determining whether the quality of the string splitting rule is sufficient for the machine learning model; Determining a new string splitting rule in response to determining that the string splitting rule is not sufficient for the machine learning model The method according to Appendix 4, further comprising. (Appendix 6) The plurality of large - language model prompts include clustering the data of the one data subset among the data subsets, and the method further includes assigning an identifier for each cluster identified by the response of the large - language model, and the one or more additional data subsets include information about the clusters. The method according to Appendix 1. (Appendix 7) The plurality of large - language model prompts are provided in parallel to the large - language model. The method according to Appendix 1. (Appendix 8) The method according to Appendix 1, further comprising replacing the one data subset among the data subsets with one of the one or more additional data subsets. (Appendix 9) The data from the one data subset among the data subsets is text data, the plurality of large - language model prompts include requesting additional information about the text, and the one or more additional data subsets include the additional information from the large - language model. The method according to Appendix 1. (Appendix 10) Data from the one data subset of the data subsets is text data, and the one or more additional data subsets include information regarding sub - phrases of individual phrases in the data subset, the method according to Appendix 1. (Appendix 11) Identifying an exception in the one data subset of the data subsets, the exception including individual data values that are in a different format from other data values in the one data subset of the data subsets; Selecting a second prompt template from among the plurality of prompt templates for the exception based on one or more characteristics of the exception; Generating a second large - language model prompt using the second prompt template and the exception; Providing the second large - language model prompt to the large - language model to generate a second response; Replacing the exception in the one data subset of the data subsets based on the second response. The method according to Appendix 1, further comprising. (Appendix 12) One or more non - transitory computer - readable media storing instructions that, when executed by one or more processors, cause a system to perform operations, the operations being: Accessing a data set including a plurality of data subsets, each data subset corresponding to a certain characteristic of the data set; Analyzing data in one data subset of the data subsets to determine characteristics of the data; Selecting a prompt template from a plurality of prompt templates for the one data subset of the data subsets based on the determined characteristics of the data in the one data subset of the data subsets; Generating a plurality of large language model prompts using the prompt template and the data from the one data subset of the data subsets; Providing the plurality of large language model prompts to a large language model, wherein the plurality of large language model prompts instruct the large language model to perform one or more operations regarding the data of the one data subset of the data subsets; Generating one or more additional data subsets for the data set based on the response of the large language model, wherein each of the one or more additional data subsets corresponds to new features of the data set; One or more non-transitory computer-readable media comprising. (Appendix 13) The operations are: Training a machine learning (ML) model using the data set; Performing one or more operations using the ML model; The one or more non-transitory computer-readable media according to Appendix 12, further comprising. (Appendix 14) The data from the one data subset of the data subsets includes character strings, the plurality of large language model prompts includes splitting each of the character strings into two or more substrings, and the one or more additional data subsets are generated using the two or more substrings. The one or more non-transitory computer-readable media according to Appendix 12. (Appendix 15) A part of the data of the one data subset of the data subsets is provided to the large language model, and the operations are: Determining a string splitting rule based on the response of the large language model; Dividing the remaining part of the data of the one data subset among the data subsets using the string splitting rule into two or more substrings One or more non-transitory computer-readable media according to appendix 14, further comprising (Appendix 16) The operations are: Determining whether the quality of the string splitting rule is sufficient for the machine learning model; Determining a new string splitting rule in response to determining that the string splitting rule is not sufficient for the machine learning model One or more non-transitory computer-readable media according to appendix 15, further comprising (Appendix 17) The plurality of large language model prompts includes clustering the data of the one data subset among the data subsets, and the operations further include assigning an identifier for each cluster identified by the response of the large language model, and the one or more additional data subsets include information about the clusters. One or more non-transitory computer-readable media according to appendix 12 (Appendix 18) One or more non-transitory computer-readable media according to appendix 12, wherein the plurality of large language model prompts are provided in parallel to the large language model (Appendix 19) One or more non-transitory computer-readable media according to appendix 12, wherein the operations further include replacing the one data subset among the data subsets with one of the one or more additional data subsets (Appendix 20) One or more processors; One or more non-transitory computer-readable storage media configured to store instructions A system having the following, wherein the instructions, in response to being executed, cause the system to perform operations, and the operations are: Obtaining a first waveform profile corresponding to a first optical signal received at a first optical receiver via an optical link; Obtaining a second waveform profile corresponding to a second optical signal received at a second optical receiver via the optical link; Combining the first waveform profile and the second waveform profile to form a combined waveform profile; Obtaining a first reconstructed waveform profile that is an estimation of the first waveform profile and a second reconstructed waveform profile that is an estimation of the second waveform profile; Combining the first reconstructed waveform profile and the second reconstructed waveform profile to form a combined reconstructed waveform profile; Determining a power profile estimation corresponding to the optical link based on a comparison between the combined waveform profile and the combined reconstructed waveform profile; Adjusting one or more aspects of optical transmission through the optical link based on the determined power profile estimation. System.

Description of Reference Numerals

[0109] 102 Dataset 104 Feature Type Inference Process 106 Labeled Dataset 108 Data Adjustment Process 110 Adjusted Dataset 112 ML Model Generation Process 114 ML Model 122 Dataset 124 Prompt Generator 126 Prompt 128 Large Language Model 130 Response 132 Dataset adjustment process 134 Adjusted dataset 152 Dataset 154 First LLM 156 First adjusted dataset 158 Second LLM 160 Second adjusted dataset 162 Adjusted dataset 202 Identify data subsets based on unique values 204 Select data subsets for adjustment from the identified data subsets 208 Does the data subset contain semantically significant data? 210 LLM clustering 212 Determine data splitting rules using LLM 214 Syntax clustering 216 Add new data to the dataset 302 Provide a part of the data of the data subset to the large language model as part of the LLM prompt for splitting the individual values of the data subset 304 Determine data splitting rules based on the response of the large language model 306 Use the data splitting rules to split the remaining part of the data of one subset of the data subset into two or more substrings 308 Determine whether the quality of the data splitting rules is sufficient for the ML model 310 Prompt the LLM to split the individual values of the data subset in different ways 312 Generate one or more additional data subsets for the dataset 402 Obtain data subsets of the dataset containing categorical features 404 Identify subphrases corresponding to the individual phrases of the data subset Embed 406 individual phrases and corresponding sub - phrases using an LLM Identify key phrases corresponding to individual phrases in a 408 data subset Prompt the LLM to improve the data subset based on the key phrases Prompt the LLM to generate external data Prompt the LLM to cluster individual phrases Generate one or more additional data subsets 500 System 550 Process 552 Memory 554 Data storage

Claims

1. Accessing a dataset that includes a plurality of data subsets, each data subset corresponding to a certain feature of the dataset; Analyzing data in one of the data subsets to determine characteristics of the data; Selecting a prompt template from a plurality of prompt templates for the one data subset among the data subsets based on the determined characteristics of the data in the one data subset; Using the prompt template and the data from the one data subset among the data subsets to generate a plurality of large language model prompts; Providing the plurality of large language model prompts to a large language model, the plurality of large language model prompts instructing the large language model to perform one or more operations regarding the data in the one data subset among the data subsets; Generating one or more additional data subsets for the dataset based on the response of the large language model, each of the one or more additional data subsets corresponding to a new feature of the dataset; A method.

2. Training a machine learning (ML) model using the dataset; Performing one or more operations using the ML model; The method according to claim 1, further comprising.

3. The data from the one data subset among the data subsets includes character strings, the plurality of large language model prompts includes splitting each of the character strings into two or more substrings, and the one or more additional data subsets are generated using the two or more substrings. The method according to claim 1.

4. A part of the data in the one data subset among the data subsets is provided to the large language model, and the method includes: Determining a string splitting rule based on the response of the large language model; Dividing the remaining part of the data of the one data subset among the data subsets using the string splitting rule into two or more substrings The method according to claim 3, further comprising.

5. Determining whether the quality of the string splitting rule is sufficient for the machine learning model; Determining a new string splitting rule in response to determining that the string splitting rule is not sufficient for the machine learning model The method according to claim 4, further comprising.

6. The plurality of large language model prompts include clustering the data of the one data subset among the data subsets, and the method further includes assigning an identifier for each cluster identified by the response of the large language model, and the one or more additional data subsets include information regarding the clusters, the method according to claim 1.

7. The plurality of large language model prompts are provided in parallel to the large language model, the method according to claim 1.

8. The method according to claim 1, further comprising replacing one of the data subsets with one of the one or more additional data subsets.

9. The data from the one data subset among the data subsets is text data, the plurality of large language model prompts include requesting additional information regarding the text, and the one or more additional data subsets include the additional information from the large language model, the method according to claim 1.

10. The data from the one data subset among the data subsets is text data, and the one or more additional data subsets include information regarding sub-phrases of individual phrases in the data subset, the method according to claim 1.

11. Identifying an exception in the one data subset among the data subsets, the exception including individual data values that are in a different format from other data values in the one data subset among the data subsets; Based on one or more characteristics of the exception, selecting a second prompt template from among the plurality of prompt templates for the exception; Generating a second large language model prompt using the second prompt template and the exception; Providing the second large language model prompt to the large language model to generate a second response; Replacing the exception in the one data subset of the data subsets based on the second response. The method according to claim 1, further comprising.

12. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause a system to perform operations, the operations comprising: Accessing a data set including a plurality of data subsets, each data subset corresponding to a certain feature of the data set; Analyzing data in one data subset of the data subsets to determine characteristics of the data; Selecting a prompt template from a plurality of prompt templates for the one data subset of the data subsets based on the determined characteristics of the data in the one data subset of the data subsets; Generating a plurality of large language model prompts using the prompt template and the data from the one data subset of the data subsets; Providing the plurality of large language model prompts to a large language model, the plurality of large language model prompts instructing the large language model to perform one or more operations regarding the data in the one data subset of the data subsets; Generating one or more additional data subsets for the data set based on the response of the large language model, each of the one or more additional data subsets corresponding to a new feature of the data set. One or more non-transitory computer-readable media comprising.

13. The operations comprising: Training a machine learning (ML) model using the data set; executing one or more operations using the ML model; The one or more non - transitory computer - readable media of claim 12, further comprising. **Claim 14** The data from the one data subset of the data subsets includes a character string, the plurality of large - language model prompts includes splitting each of the character strings into two or more substrings, and the one or more additional data subsets are generated using the two or more substrings. The one or more non - transitory computer - readable media of claim 12. **Claim 15** A part of the data of the one data subset of the data subsets is provided to the large - language model, and the operations are: determining a string - splitting rule based on the response of the large - language model; using the string - splitting rule to split the remaining part of the data of the one data subset of the data subsets into two or more substrings The one or more non - transitory computer - readable media of claim 14, further comprising. **Claim 16** The operations are: determining whether the quality of the string - splitting rule is sufficient for a machine - learning model; determining a new string - splitting rule in response to determining that the string - splitting rule is not sufficient for the machine - learning model The one or more non - transitory computer - readable media of claim 15, further comprising. **Claim 17** The plurality of large - language model prompts includes clustering the data of the one data subset of the data subsets, and the operations further include assigning an identifier for each cluster identified by the response of the large - language model, and the one or more additional data subsets include information about the clusters. The one or more non - transitory computer - readable media of claim 12. **Claim 18** The plurality of large - language model prompts are provided in parallel to the large - language model. The one or more non - transitory computer - readable media of claim 12. **Claim 19** The one or more non-transitory computer-readable media of claim 12, wherein the operation further comprises replacing the one data subset of the data subsets with one of the one or more additional data subsets.

20. One or more processors; One or more non-transitory computer-readable storage media configured to store instructions A system having, wherein the instructions, in response to being executed, cause the system to perform operations, the operations comprising: Obtaining a first waveform profile corresponding to a first optical signal received at a first optical receiver via an optical link; Obtaining a second waveform profile corresponding to a second optical signal received at a second optical receiver via the optical link; Combining the first waveform profile and the second waveform profile to form a combined waveform profile; Obtaining a first reconstructed waveform profile that is an estimation of the first waveform profile and a second reconstructed waveform profile that is an estimation of the second waveform profile; Combining the first reconstructed waveform profile and the second reconstructed waveform profile to form a combined reconstructed waveform profile; Determining a power profile estimate corresponding to the optical link based on a comparison between the combined waveform profile and the combined reconstructed waveform profile; Adjusting one or more aspects of optical transmission through the optical link based on the determined power profile estimate. System.