Machine learning feature generation method for multi-agent agents in database
Through the machine learning feature generation method of multi-agent agents in the database, using large language models and iterative optimization technology, the accuracy and efficiency of machine learning tasks in the database are solved, and efficient feature generation and model training are achieved.
Patent Information
- Application Number
- CN202411650247.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-19
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-11-19
AI Technical Summary
In the prior art, it is difficult to execute machine learning tasks in the database and low accuracy, mainly because feature extraction operations encounter challenges in SQL expression, and it is difficult for the database management system to seamlessly support structured queries and vector computing.
A machine learning feature generation method in the database based on multi-agent agent is adopted to determine feature prompts through historical feature sets and preset machine learning tasks, and a large language model is used to generate new features, and a feature set is iteratively optimized until the matching feature set is decomposed to improve the effectiveness of the feature and the performance of the model.
It improves the accuracy and efficiency of machine learning tasks in the database, solves the problem of feature extraction and computing mode compatibility, and realizes more efficient feature generation and model training.
Smart Images

Figure CN119151016B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine learning, and in particular to a method for generating machine learning features within a database of a multi-agent agent. Background Art
[0002] In the field of structured data, there is an increasing demand for data analysis using machine learning techniques, including many typical scenarios such as stock prediction, fraud detection, and sales forecasting. Therefore, some database management systems integrate ML (Machine Learning) models directly in their database engines to simplify the model training and deployment process. Typically, these database management systems provide structured query language extensions that enable users to interact with these models seamlessly.
[0003] Currently, users still prefer to extract data from DBMS (Database Management System) into Python environment, and then perform subsequent feature engineering and model training, so as to instantly design and customize their own models. This preference stems from some key challenges faced by these DBMS: First, the success of ML tasks depends largely on the effectiveness of feature extraction rather than the model itself. Some feature extraction operations are challenging or even impossible when expressed in SQL (Structured Query Language), which makes it difficult to extract features for machine learning, and the execution results obtained after performing machine learning tasks based on the extracted features are not accurate enough. Second, DBMS is designed to efficiently handle relational algebra operations, while machine learning tasks usually involve vector calculations. It is still challenging to seamlessly support these two computing modes in a single system, making it difficult to perform machine learning tasks within the database management system. Third, unlike typical SQL analytical queries, accurately capturing users' real needs for ML tasks is an additional challenge in declarative languages without a lot of human intervention for tasks such as data exploration, data cleaning, and model testing.
[0004] There is currently no effective solution to the problem that it is difficult and inaccurate to perform machine learning tasks within a database in related technologies. Summary of the invention
[0005] Based on this, it is necessary to provide a method for generating machine learning features in a database of a multi-agent agent that can solve the problem of difficulty and low accuracy in performing machine learning tasks in the database in response to the above technical problems.
[0006] In a first aspect, a method for generating machine learning features in a database based on a multi-agent agent is provided in this embodiment, and the method comprises:
[0007] Determining a first feature set and a feature description of the first feature set based on performance indicators of the historical feature set in the machine learning model in the database;
[0008] Obtaining feature prompts corresponding to the first feature set according to a preset machine learning task and the historical feature set;
[0009] Acquire new features generated by a large language model in a database according to the first feature set, the feature description and the feature prompt, and obtain a second feature set by combining the first feature set and the new features;
[0010] Determining a third feature set according to performance indicators of the historical feature set and the second feature set in the machine learning model;
[0011] Decompose the third feature set until the decomposed feature set matches the third feature set, and obtain a fourth feature set required to perform the machine learning task according to the decomposition result.
[0012] In some embodiments, the feature prompt includes a static prompt and a dynamic prompt, and obtaining the feature prompt corresponding to the first feature set according to the preset machine learning task and the historical feature set includes:
[0013] Obtaining a static prompt according to the machine learning task;
[0014] A feature set whose similarity with the first feature set meets a preset degree is obtained from the historical feature set, and the dynamic prompt is obtained according to the feature set whose similarity meets the preset degree.
[0015] In some embodiments, decomposing the third feature set until the decomposed feature set matches the third feature set, and obtaining the fourth feature set required to perform the machine learning task according to the decomposition result includes:
[0016] Decomposing the third feature set to obtain a plurality of decomposed features;
[0017] Determining whether the credibility of the decomposed feature set meets the first preset condition;
[0018] If the credibility satisfies the first preset condition, it is determined that the feature set obtained by decomposition matches the third feature set, the decomposition is stopped and the fourth feature set is obtained; and / or,
[0019] Decomposing the third feature set to obtain a generated code of the decomposed feature set;
[0020] Determining whether the complexity of the generated code of the decomposed feature set meets the second preset condition;
[0021] If the complexity satisfies the second preset condition, it is determined that the feature set obtained by decomposition matches the third feature set, and the decomposition is stopped to obtain the fourth feature set.
[0022] In some embodiments, decomposing the third feature set includes:
[0023] According to the third feature set, obtaining a feature description of the third feature set;
[0024] Decomposing the feature description of the third feature set into a plurality of sub-descriptions;
[0025] The large language model is called to generate a decomposed feature set according to the sub-description.
[0026] In some embodiments, after obtaining new features generated by a large language model in a database according to the first feature set, the feature description and the feature prompt, and combining the first feature set and the new features to obtain a second feature set, the method further includes:
[0027] Determining whether the number of iterations of the second feature set reaches a preset number;
[0028] If not, the first feature set and the second feature set are used as the iterated historical feature set; according to the performance indicators of the iterated historical feature set in the machine learning model in the database, a new first feature set and a feature description of the new first feature set are determined; according to the preset machine learning task and the iterated historical feature set, feature prompts corresponding to the new first feature set are obtained; new features generated by the large language model according to the new first feature set, the feature description of the new first feature set and the feature prompts corresponding to the new first feature set are obtained, and the iterated second feature set is obtained by combining the new first feature set and the new features;
[0029] If so, the step of determining a third feature set based on the performance indicators of the historical feature set and the second feature set in the machine learning model is performed.
[0030] In some embodiments, after decomposing the third feature set until the feature set obtained by the decomposition matches the third feature set, and obtaining a fourth feature set required for performing the machine learning task in the database according to the decomposition result, the method further includes:
[0031] Obtaining a generated code for the fourth feature set;
[0032] If there is a code matching a preset pattern in the generated code of the fourth feature set, converting the matching code into a structured query statement of the database;
[0033] If there is a code that does not match the preset pattern in the generated code of the fourth feature set, the unmatched code is converted into a user-defined function.
[0034] In some embodiments, the method further comprises:
[0035] Combining the structured query statements according to operators in the structured query statements; and / or,
[0036] The structured query statement is assembled according to the common table expression.
[0037] In some embodiments, there is code that does not match the preset pattern in the generated code of the fourth feature set, and after converting the unmatched code into a user-defined function, the method further includes:
[0038] Obtaining a first execution cost of an operator of the user-defined function;
[0039] After splitting the user-defined function according to the attribute of the user-defined function, obtaining a second execution cost of the operator of the split user-defined function;
[0040] Whether to split the user-defined function is determined according to a comparison result of the first execution cost and the second execution cost.
[0041] In some embodiments, determining the first feature set and the feature description of the first feature set according to the performance indicators of the historical feature set in the machine learning model in the database includes:
[0042] Obtain the historical feature set in each node of the feature transformation tree;
[0043] Determine the node to be visited according to the performance parameter of the historical feature set in the machine learning model and the number of times each of the nodes has been visited;
[0044] Acquire a first feature set in the node to be visited and obtain a feature description of the first feature set.
[0045] In some embodiments, determining the third feature set according to the performance indicators of the historical feature set and the second feature set in the machine learning model in the database includes:
[0046] Adding the second feature set as a child node to the feature conversion tree;
[0047] Determine a new node to be visited according to the performance parameters of the feature set in each node of the feature conversion tree in the machine learning model and the number of times each node is visited;
[0048] Obtain a third feature set in the new node to be visited.
[0049] In a second aspect, a computer device is provided in this embodiment, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the machine learning feature generation method in a database based on multi-agent agents as described in the first aspect above.
[0050] According to a third aspect, a computer program product is provided in this embodiment, comprising a computer program, which, when executed by a processor, implements the method for generating machine learning features in a database based on multi-agent agents as described in the first aspect above.
[0051] The machine learning feature generation method in the database of the above-mentioned multi-agent agent provides suggestions for feature generation of a large language model by generating feature descriptions of feature sets, so that the generated features are effective features; after selecting the feature set according to performance indicators, the features are decomposed to avoid the problem of being unable to perform machine learning tasks due to complex generated features, thereby achieving the effect of improving the accuracy of machine learning in the database. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 A flowchart of a method for generating machine learning features in a database of a multi-agent agent in one embodiment;
[0053] Figure 2 A flowchart of a method for generating machine learning features in a database of a multi-agent agent in another embodiment;
[0054] Figure 3 is a schematic diagram of feature decomposition in one embodiment;
[0055] Figure 4 A structural block diagram of a machine learning feature generation device in a database of a multi-agent agent in one embodiment;
[0056] Figure 5 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0058] In one embodiment, Figure 1 As shown, a method for generating machine learning features in a database of a multi-agent agent is provided. This embodiment is illustrated by applying the method to a terminal, wherein the terminal contains a database of a multi-agent agent. It is understandable that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0059] Step 102, determining a first feature set and a feature description of the first feature set according to performance indicators of the historical feature set in the machine learning model in the database.
[0060] Among them, the historical feature set is one or more feature sets required to execute the machine learning tasks preset in the database. Performance indicators include but are not limited to performance parameters such as the F1 score and accuracy of the machine learning model. The historical feature set may include the initial feature set and one or more new features generated by the large language model based on the initial feature set; or, the historical feature set may also be one or more feature sets acquired in advance. The feature description is a description of the data attributes in the first feature set. The feature description can be set according to the machine learning task. For example, the feature description may include text annotation, part-of-speech tagging, position encoding, etc. of the feature.
[0061] Optionally, based on the performance indicators of the historical feature sets according to the first agent in the database, a feature set with the best performance indicators is determined as the first feature set. It can be understood that if the F1 score of the machine learning model is high, the performance indicator is better, and if the F1 score of the machine learning model is low, the performance indicator is poor; if the accuracy of the machine learning model is high, the performance indicator is better, and if the accuracy of the machine learning model is low, the performance indicator is poor. In addition, according to application requirements, the historical feature sets can also be sorted based on performance indicators, and multiple feature sets with better performance indicators can be used as multiple first feature sets.
[0062] Step 104, obtaining feature prompts corresponding to the first feature set according to the preset machine learning task and the historical feature set.
[0063] Among them, feature prompts are used to limit the feature output of the large language model in step 106. By obtaining feature prompts based on the preset machine learning task, the output second feature set can meet the requirements of the machine learning task; by obtaining feature prompts through the historical feature set, the historical feature output can be reflected on, so that the large language model outputs new features that are conducive to the realization of the machine learning task.
[0064] Optionally, the first intelligent agent obtains partial feature prompts based on the description of a preset machine learning task; the first intelligent agent may also obtain partial feature prompts based on one or more factors such as historical feature sets, generation paths of historical feature sets, performance of historical feature sets on machine learning model training data sets, etc.
[0065] Step 106, obtaining new features generated by the large language model in the database according to the first feature set, feature descriptions and feature prompts, and combining the first feature set and the new features to obtain a second feature set.
[0066] The second feature set is the union of the new feature and the first feature set. The large language model can generate multiple new features based on the first feature set, feature descriptions, and feature prompts, and respectively combine the multiple new features with the first feature set to obtain multiple second feature sets.
[0067] Optionally, the first agent generates an executable code, and based on the executable code implements a process of using the first feature set, feature descriptions and feature prompts as inputs to a large language model, obtaining new features output by the large language model, and combining the new features with the first feature set to obtain a second feature set.
[0068] Step 108, determining a third feature set based on the performance indicators of the historical feature set and the second feature set in the machine learning model in the database.
[0069] Optionally, the first agent in the database determines one or more feature sets with the best performance indicators as the third feature set based on the performance indicators of the historical feature set and the second feature set. The performance indicators may be performance parameters such as the F1 score and accuracy of the machine learning model, and may also include the number of times the feature set is selected as the first feature set.
[0070] Furthermore, before executing step 108, steps 102 to 106 may be repeatedly executed: after each second feature set is generated according to step 106, the second feature set is used as a historical feature set, and step 102 is executed again until the number of iterations reaches a preset number. At this time, the performance indicator in step 102 may also include the number of times that the multiple historical feature sets are determined as the first feature set according to the F1 score and accuracy of the machine learning model. Among them, the more times the historical feature set is determined as the first feature set, the more beneficial it is to guide the nodes selected in the future, and the better the performance indicator.
[0071] Step 210, decompose the third feature set until the decomposed feature set matches the third feature set, and obtain a fourth feature set required to perform the machine learning task in the database according to the decomposition result.
[0072] Among them, the feature set obtained by decomposition is a sub-feature, and the third feature set is a parent feature. The matching of the feature set obtained by decomposition and the third feature set means that the sub-features obtained by decomposition and their parent features maintain a correct association. Optionally, the second agent uses the ability of the large language model to decompose the feature description of the third feature set to obtain multiple sub-features, and generates simpler features after decomposition based on the sub-features. When it is verified that the values of the parent feature and the sub-feature are the same, it can be judged that the sub-feature matches its parent feature.
[0073] Optionally, the third feature set is decomposed multiple times, and when the feature set obtained by the decomposition matches the third feature set, the features obtained by the last decomposition are used as the fourth feature set.
[0074] In the method for generating machine learning features in the database of the above-mentioned multi-agent agent, the generation of machine learning features is realized in the database by integrating a large language model and a machine learning model in the database. Based on the historical feature set and the machine learning task, feature prompts are generated to limit the feature output of the large language model. While capturing the user's needs for machine learning tasks, the effectiveness of the new features output by the large language model can be improved by reflecting on the historical features, so that the feature set obtained based on the new features can meet the needs of the machine learning task. After determining the third feature set based on the performance of the machine learning model, the complexity of the machine learning features is reduced by decomposing the third feature set, so that the preset machine learning tasks can be smoothly and efficiently implemented in the database based on the machine learning features, solving the problem of difficulty and low accuracy in executing machine learning tasks in the database.
[0075] In one embodiment, after decomposing the third feature set until the decomposed feature set matches the third feature set, and obtaining the fourth feature set required to execute the machine learning task in the database based on the decomposition result, the method also includes: obtaining the generated code of the fourth feature set; if there is code matching a preset pattern in the generated code of the fourth feature set, converting the matching code into a structured query statement of the database; if there is code not matching the preset pattern in the generated code of the fourth feature set, converting the not matching code into a user-defined function.
[0076] Wherein, the generated code is an executable code for generating a feature set. The generated code can be a Python statement or a regular language. Optionally, when the generated code is a Python statement, the generated code of the fourth feature set is abstracted into a Python abstract syntax tree, and the matching between the generated code and the preset mode is implemented at the Python abstract syntax tree level. When the preset mode is a regular language, the matching between the generated code and the preset mode can be directly implemented.
[0077] The preset mode is a data structure of a preset code. Optionally, the preset mode is used to indicate the correspondence between the operator of the data frame in the generated code and the pure SQL statement, and / or the SQL statement with a user-defined function. It is determined whether the operator of the data frame in the generated code of the fourth feature set matches the operator in the preset mode. If there is a matching code in the generated code, the operator type in the matching code is obtained, and the corresponding structured query language is generated for different types of operators based on the corresponding relationship indicated by the preset mode. If there is an unmatched code in the generated code, the unmatched code is packaged into a user-defined function. For ease of understanding, taking the generated code as Python as an example, the unmatched code can be simply packaged by the specified general operator to generate a UDF implemented by Python. Optionally, the UDF (user-defined function) generated based on the unmatched code is a custom function using a linear algebra library. The UDF will be compiled using JIT technology and registered in the database during execution, thereby avoiding the limited flexibility of SQL as a structured language when processing complex linear operations or complex algorithms.
[0078] Optionally, before converting the generated code into a structured query language and / or a user-defined function, the generated code of the fourth feature set may be split into several code blocks, so as to ensure that each code block generates only one new feature in the DataFrame. Each code block is matched with a preset pattern respectively.
[0079] In this embodiment, by performing model matching on the generated code, each operator of the generated code is converted into a structured query language and / or a user-defined function within the database, thereby achieving the effect of efficiently generating the features required for machine learning within the database.
[0080] Furthermore, in one embodiment, the method further comprises: combining the structured query statements according to operators in the structured query statements; and / or combining the structured query statements according to common table expressions.
[0081] Among them, classification can be implemented based on the attributes of the operators, and multiple operators that appear continuously can be combined based on the classification results to obtain an SQL statement. Optionally, simple operators implemented based on pure SQL in the structured query statement are obtained, and the simple operators that appear continuously can be combined. UDF operators that appear continuously can also be combined to obtain an SQL statement.
[0082] Among them, the common expression is a temporary result set used in the SQL query. Using CTE (common expression) to implement SQL combination can avoid caching intermediate calculation results, thereby reducing unnecessary storage of calculation results and reducing the implementation cost required to implement feature-based machine learning tasks.
[0083] Optionally, the structured query statements are first combined according to the operators in the structured query statements, and then combined according to the common table expressions. The structured query statements can also be combined in the opposite order.
[0084] In this embodiment, by combining structured query statements, and / or combining structured query statements and user-defined functions to simplify SQL, redundant operations can be eliminated, and the cost required for repeated unpacking and packaging when transmitting data for multiple UDF operators can be reduced, thereby achieving the effect of accelerating SQL execution speed.
[0085] In one embodiment, there is code that does not match the preset pattern in the generated code of the fourth feature set. After converting the non-matching code into a user-defined function, the method also includes: obtaining a first execution cost of the operator of the user-defined function; after splitting the user-defined function according to the attributes of the user-defined function, obtaining a second execution cost of the operator of the split user-defined function; and judging whether to split the user-defined function based on the comparison result of the first execution cost and the second execution cost.
[0086] The first execution cost is the total cost required for executing each operator of the user-defined function before splitting, and the second execution cost is the total cost required for executing each operator after splitting the user-defined function. Optionally, a cost evaluation model can be established, and the hyperparameters of the cost of each operator can be defined through the cost evaluation model to obtain the first execution cost and the second execution cost of the user-defined function.
[0087] Optionally, the attributes involved in the user-defined function may be split to obtain an independent path, and the independent path obtained by the splitting is connected with the main branch in the user-defined function based on the connection operator to obtain the split user-defined function.
[0088] Since there is additional overhead when executing the newly added connection operator after the split, the splitting of the user-defined function is not necessarily a forward optimization. Therefore, it is necessary to judge whether the splitting operation is necessary based on the execution cost before and after the splitting. Optionally, when the first execution cost is greater than or equal to the second execution cost, it is judged that the user-defined function does not need to be split. When the first execution cost is less than the second execution cost, it is judged that the user-defined function needs to be split. Alternatively, when the first execution cost is less than the second execution cost, and the cost difference between the two is greater than the specified value of the preset setting, it is judged that the user-defined function needs to be split.
[0089] Since both the packing and unpacking of user-defined functions will generate additional overhead, in this embodiment, the judgment of splitting the user-defined function is implemented according to the operator cost, which can reduce the cost of executing the user-defined function.
[0090] In one embodiment, determining a first feature set and a feature description of the first feature set based on performance indicators of historical feature sets in a machine learning model within a database includes: obtaining historical feature sets in each node of a feature conversion tree; determining a node to be visited based on performance parameters of the historical feature sets in the machine learning model and the number of times each node is visited; obtaining the first feature set in the node to be visited, and obtaining a feature description of the first feature set.
[0091] Each historical feature set is a feature set of each node in the feature conversion tree. The root node of the feature conversion tree is a preset initial feature set. The feature conversion tree is constructed from scratch based on the large language model and the initial feature set.
[0092] Optionally, the first agent obtains a performance index of the feature conversion tree based on the performance parameters of the historical feature set in the machine learning model and the number of times each node is visited, which is called node utility. Among them, when the total number of times each node is visited is fixed, among the nodes with the same performance parameters, the utility of the node with more visits is relatively high; otherwise, the utility of the node is relatively low. When the number of times a certain node is visited is fixed, the lower the total number of times each node is visited, the higher the utility of the node. The higher the performance parameters such as the accuracy or F1 score of the feature set corresponding to the current node in the machine learning model, the higher the node utility; otherwise, the node utility is relatively high. Optionally, the first value can be calculated based on the total number of times each node is visited and the number of times the current node is visited. The second value is obtained based on one or more performance parameters of the current node. The first value and the second value are combined by a mathematical processing method of weighted addition and multiplication to obtain the node utility. The node utilities of multiple nodes are sorted, and one or more nodes with the highest node utility are selected as the nodes to be visited. The first feature set is obtained by visiting the nodes to be visited. A feature description of the first feature set is obtained based on the attributes of the first feature set.
[0093] In this embodiment, in the feature transformation tree, the nodes to be visited are selected based on the performance parameters of the historical feature sets in the machine learning model and the number of times each node is visited, so that the first feature set selected in the nodes to be visited can not only improve the accuracy of the downstream machine learning model when performing machine learning tasks, but also guide the selection of the node where the third feature set is located in the future, thereby achieving a combination of utilization and exploration.
[0094] Furthermore, in one embodiment, determining the third feature set based on the performance indicators of the historical feature set and the second feature set in the machine learning model in the database includes: adding the second feature set as a child node to the feature conversion tree; determining a new node to be visited based on the performance parameters of the feature set in each node of the feature conversion tree in the machine learning model and the number of times each node is visited to obtain the third feature set in the new node to be visited.
[0095] Among them, the method of selecting nodes corresponding to the third feature set is the same as the method of selecting nodes corresponding to the first feature set, that is, the first intelligent agent implements the sorting of node selection priority through node utility, and takes the feature set in one or more nodes with the highest priority as the third feature set.
[0096] In this embodiment, in the feature conversion tree, the nodes to be visited are selected based on the performance parameters of the historical feature set in the machine learning model and the number of times each node is visited. The utilization and exploration of the nodes are integrated to select the third feature set required for the machine learning task.
[0097] In one embodiment, feature prompts include static prompts and dynamic prompts. According to a preset machine learning task and a historical feature set, obtaining feature prompts corresponding to a first feature set includes: obtaining static prompts according to the machine learning task; obtaining a feature set in the historical feature set whose similarity with the first feature set meets a preset degree, and obtaining dynamic prompts according to the feature set whose similarity meets the preset degree.
[0098] The static prompt remains unchanged during the feature generation process of the large language model. Optionally, the static prompt can be obtained according to the description of the machine learning task. Furthermore, the static prompt can also be obtained according to the predefined feature generation operation. The dynamic prompt is based on the historical experience generated by the historical response when the large language model generates features.
[0099] Among them, the preset degree is used to indicate that the current feature set belongs to one or more feature sets in each historical feature set that are most similar to the first feature set. Optionally, the similarity between the two feature sets can be evaluated by the semantic similarity and / or structural similarity of the features in the historical feature set and the first feature set. Optionally, the similarity of the historical feature sets is sorted; a dynamic prompt is obtained based on one or more feature sets with the highest semantic similarity, or a dynamic prompt is obtained based on one or more feature sets with the highest structural similarity; or, semantic similarity and structural similarity are combined by multiplication, weighted addition, etc., and one or more feature sets with the highest comprehensive similarity are selected to obtain a dynamic prompt.
[0100] In this embodiment, dynamic prompts and static prompts are provided as input to the large language model to provide suggestions for feature generation of the large language model, thereby improving the effectiveness of new features generated by the large language model when performing machine learning tasks.
[0101] In one embodiment, decomposing the third feature set includes: obtaining a feature description of the third feature set according to the third feature set; decomposing the feature description of the third feature set into multiple sub-descriptions; and calling the large language model to generate a decomposed feature set according to the sub-descriptions.
[0102] Among them, the feature description of the third feature set can be obtained by the attributes of the third feature set. Optionally, the second agent can realize the decomposition of the feature description of the third feature set through an optimized proxy model with decomposition features, and decompose the feature description of the third feature set into a set of different sub-descriptions. The executable code for generating sub-descriptions is obtained through the first agent, and the large language model is called based on the executable code for generating sub-descriptions, and the sub-descriptions are input into the large language model to obtain the decomposed feature set output by the large language model. Furthermore, based on the method of obtaining the feature prompts corresponding to the first feature set, the feature prompts of the sub-descriptions can be obtained, the large language model generates the decomposed feature set, and the feature prompts of the sub-descriptions are synchronized as the input of the large language model. In this embodiment, by decomposing the third feature set, the complexity of the features is reduced, and the problem that the first agent cannot generate correct code based on the features due to the high complexity and logic of the features is avoided.
[0103] Furthermore, multiple decompositions can be achieved recursively, so that complex features are continuously decomposed into simple features. In order to ensure that the features obtained after each decomposition are less complex than the features before decomposition, in one embodiment, the third feature set is decomposed until the decomposed feature set matches the third feature set, and the fourth feature set required for executing the machine learning task is obtained according to the decomposition result, including: decomposing the third feature set to obtain multiple decomposed features; judging whether the credibility of the decomposed feature set meets the first preset condition; if the credibility meets the first preset condition, judging that the decomposed feature set matches the third feature set, stopping the decomposition and obtaining the fourth feature set; and / or, decomposing the third feature set to obtain the generated code of the decomposed feature set; judging whether the complexity of the generated code of the decomposed feature set meets the second preset condition; if the complexity meets the second preset condition, judging that the decomposed feature set matches the third feature set, stopping the decomposition and obtaining the fourth feature set.
[0104] Among them, credibility refers to the consistency between the feature set after decomposition and the feature set before decomposition. The first preset condition is used to indicate that the consistency between the feature set after decomposition and the feature set before decomposition is high enough. For example, among the multiple features obtained after decomposition, if there are features with a specified proportion whose generated code response results are consistent, the credibility is high; otherwise, the credibility is low. Therefore, the first preset condition can be set to that the response result of the feature generation code is consistent with the generated code response results of other features greater than or equal to the specified proportion. For ease of understanding, the specified proportion is set to 2 / 3. If there are k features whose generated code c={c 1 , c 2 , …c K}, after executing these codes c, select the code with the most number of execution results that are the same as the response results of other generated codes If the code The number of other codes with the same response results exceeds 2 / 3 of the total number of generated codes k, so it can be considered that code The specified ratio can also be modified according to application requirements.
[0105] Optionally, if the credibility meets the first preset condition, the sub-features finally decomposed are used as the fourth feature set; if the credibility does not meet the first preset condition, the sub-features decomposed are unreliable and further decomposed based on the current decomposed features.
[0106] The complexity refers to the complexity of the execution logic of the decomposed feature set. The second preset condition is used for the low complexity of the execution logic of the decomposed feature set. Optionally, the complexity of the generated code of the decomposed feature is calculated. The second preset condition is set to the complexity of the generated code of the feature is less than or equal to cp bound , where cp bound is a preset value. If the complexity meets the second preset condition, the sub-features obtained by the final decomposition are used as the fourth feature set; if the complexity does not meet the second preset condition, the complexity of the sub-features obtained by the decomposition is too high, and further decomposition is performed based on the current decomposition features.
[0107] Furthermore, in the decomposition process, there is a situation where the complexity of the feature code after decomposition is higher than the feature before decomposition. Therefore, the second preset condition can also be set as: the complexity is less than or equal to cp bound , or the complexity of the generated code of the feature before decomposition is greater than or equal to the complexity of the generated code of the feature after decomposition, then the second preset condition is met. bound If the complexity of the generated code of the feature before decomposition is less than the complexity of the generated code of the feature after decomposition, further decomposition is performed on the current decomposed feature.
[0108] Optionally, it can also be determined whether to further perform feature decomposition based on both complexity and credibility, that is, when the feature set obtained by decomposition meets the first preset condition or meets the second preset condition, the decomposition is stopped.
[0109] Furthermore, after decomposing the third feature set, the method further includes: if the column value of the third feature set is equal to the column value of the decomposed feature set, then stop decomposing the third feature set. The column value of the third feature set is the numerical value of each feature, and the column value can reflect the semantic information in the feature set. If the column value of the third feature set is equal to the column value of the decomposed feature set, it indicates that the complexity of the third feature set is low and does not need to be decomposed.
[0110] In this embodiment, by introducing the step of column value comparison, the decomposition can be terminated in advance according to the column value comparison result, thereby improving the efficiency of feature generation.
[0111] In one embodiment, after obtaining new features generated by a large language model in a database according to a first feature set, feature descriptions and feature prompts, and combining the first feature set and the new features to obtain a second feature set, the method also includes: determining whether the number of iterations of the second feature set reaches a preset number; if not, using the first feature set and the second feature set as historical feature sets after iteration; determining a new first feature set and a feature description of the new first feature set according to performance indicators of the iterated historical feature set in a machine learning model in the database; obtaining feature prompts corresponding to the new first feature set according to preset machine learning tasks and the iterated historical feature set; obtaining new features generated by the large language model according to the new first feature set, the feature description of the new first feature set and the feature prompts corresponding to the new first feature set, and combining the new first feature set and the new features to obtain an iterated second feature set; if so, executing the step of determining a third feature set according to the performance indicators of the historical feature set and the second feature set in the machine learning model.
[0112] The number of iterations is set and modified according to the needs. The performance indicators of the iterated historical feature set in the machine learning model in the database include the performance parameters of the machine learning model, the number of times each historical feature set is selected as the first feature set in each iteration round, and the number of current iterations. When the number of iterations is fixed, the more times the historical feature set is selected as the first feature set, the better the performance indicators of the historical feature set.
[0113] Wherein, when the number of iterations of the second feature set reaches a preset number, the second feature set when the step of determining the third feature set according to the performance indicators of the historical feature set and the second feature set in the machine learning model is performed is the second feature set obtained in the last iteration. Optionally, in each iteration, the large language model generates multiple new features according to the new first feature set, the feature description of the new first feature set, and the feature prompts corresponding to the new first feature set; and multiple second feature sets are obtained by combining the multiple new features with the first feature set.
[0114] Optionally, after obtaining the historical feature sets after the iteration, a performance indicator of each historical feature set is obtained according to the performance parameters of the machine learning model, and another performance indicator of the historical feature set is obtained according to the number of times each historical feature set is selected as the first feature set in each iteration round and the number of current iterations. After assigning preset weights to the two performance indicators, the final performance indicator of each historical feature set is obtained, and one or more historical feature sets with the best performance indicators are used as the first feature set.
[0115] Optionally, a feature conversion tree is constructed based on the historical feature set, and each node in the feature conversion tree identifies a feature set. The root node of the tree is the initial feature set, and the historical feature set at the first iteration is the initial feature set; the child nodes of the feature conversion tree are the second feature sets obtained by combining the first feature set and the new features. In each iteration, a node will be selected as the first feature set for expansion, and one or more second feature sets will be obtained by expansion.
[0116] In one embodiment, Figure 2 A flowchart of another method for generating machine learning features in a database based on multi-agent agents is provided, such as Figure 2 As shown, a two-stage processing method is adopted for machine learning features.
[0117] The first stage is multi-agent-based feature generation. Specifically, feature generation for ML tasks is based on large language models, with the main goal of dynamically identifying new features required for the task and using LLM (Large Language Model) agents to generate the corresponding code. Agents are carefully designed to interact with LLM and the external environment to complete specific tasks.
[0118] In the first stage of this embodiment, two types of LLM agents are established: the main agent and the optimization agent. Among them, the main agent is the first agent in the above embodiment, and the optimization agent is the second agent in the above embodiment. The main agent is responsible for proposing new feature suggestions and then realizing the generation of new features. The optimization agent is responsible for correcting the feature code of a complex feature and improving the correctness of the generated features by decomposing the complex features into simpler features. After obtaining the sampled data based on the original data sampling, the two agents use the capabilities of LLM and collaborate in an iterative manner to build a feature transformation tree from scratch, generate and improve new features, and combine the features and machine learning task information to obtain the optimal feature set of the feature transformation tree. The specific implementation is shown in Algorithm 1.
[0119] For ease of understanding, let's first explain the formal definition problem in Algorithm 1:
[0120] set up is a data set, where F is an initial feature set constructed by a set of preset initial features, , n is a natural number; Y is the target attribute for prediction or classification. Considering the downstream machine learning model L, the goal is to find a feature set with the best performance index , to maximize the performance of the machine learning model L. The features in are generated by the master agent and the optimization agent using the power of large language models to incorporate domain knowledge into feature generation. Given a dataset and a downstream machine learning model L, the feature engineering problem is defined as follows:
[0121]
[0122] in, is a performance indicator of the machine learning model L, including but not limited to the accuracy and F1 score of the machine learning model L; It is a new feature set generated by transforming the initial feature set F. By transforming the feature set, the newly generated feature set can maximize the performance index; F is an infinite feature set, which contains all features generated from the initial feature set F through one or more rounds of operations. Different from the related art that limits each operation to a fixed operator to narrow the search space, this embodiment does not limit the operator to ensure a larger search space for the feature.
[0123] Algorithm 1 is as follows:
[0124] 1: Agent-based feature generation ;
[0125] 2: ;
[0126] 3: for i∈{1,···,s} do:
[0127] 4: ;
[0128] 5: for j∈{1,···,k} do:
[0129] 6: F j ,c j ,d j ← Main Agent (F,Q);
[0130] 7: F j ,c j ←Optimize Agent (F j ,c j ,d j );
[0131] 8: ;
[0132] 9: ;
[0133] 10: ;
[0134] 11: ;
[0135] 12: for (F,c)∈Q do:
[0136] 13: ;
[0137] 14: POP_MAX_ And return the optimal feature set and corresponding code;
[0138] The following is an explanation of Algorithm 1:
[0139] Algorithm 1 mainly performs iterative search on the FTT tree (feature transformation tree), generating child nodes by using two agents to construct new features. The newly generated features will be incorporated into the continuous feature construction, thereby enhancing the existing feature set and gradually constructing a feature transformation tree. Each node in the feature transformation tree represents a feature set, starting from the parent node F par To child node F child Each directed edge of F par Generate new feature F new and add it to F child ;Right now, The root node of the tree is the initial feature set F. Specifically, each iteration generates k child nodes, and a total of s iterations are performed, where k and s are natural numbers and can be set and modified based on application requirements. The feature set in the iteratively generated k child nodes is the second feature set in the above embodiment.
[0140] Each iteration will select a node for expansion, that is, expansion based on the first feature set in the node to be visited in the above embodiment. In order to strike a balance between utilization and exploration, it is necessary to select nodes that can bring high performance to the downstream machine learning model to meet the "utilization" requirement, and also select nodes with less information to guide future selections to meet the "exploration" requirement. Therefore, in Algorithm 1, by defining the node utility U, the effect of considering both utilization and exploration is achieved to guide node selection during the search process. The calculation method of node utility U(F, T, L) is as follows:
[0141]
[0142] in, is the performance index of L; v 0 is the sum of all node visit times; v F is the number of visits to the current node; w is a weight factor used to achieve a trade-off between exploitation and exploration. Optionally, w is set to , w can also be set to other values. In addition, a triple is matched for each node, including its feature set F, the code c that generates new features in F, and its utility U.
[0143] Therefore, the specific workflow of the first stage described in Algorithm 1 is as follows: Given a dataset and machine learning model L. First, use the root triple (F, ,U(F,Y,L)) initializes a priority queue Q, the initial feature set is F, and the initial code is empty. The higher the node utility, the higher the priority of the node in Q. Then, Algorithm 1 iterates s steps, and selects the node with the highest node utility (the node to be visited) in each iteration, that is, POP_MAX_ , and use the large language model to expand the node with the highest node utility to obtain k child nodes. When expanding the child nodes, Algorithm 1 calls Main_Agent(F, Q) to generate the child feature F j and its corresponding code c j and feature description d j , and then call Optimization_Agent(F j, c j, d j )}, decompose complex logical features into simpler features to optimize the generated features F j , and combine the parent feature set F and code c with the newly generated feature set F j and code c j The newly generated child nodes will be added to Q. After expanding k child nodes, Algorithm 1 updates the parent node by recalculating the utility, because v 0 and v F Finally, select and return the largest value in Q without considering the exploration item. After executing s iterations, Algorithm 1 returns an optimal feature set based on Monte Carlo tree search. , and for generating Corresponding codes of new features Among them, the optimal feature set That is, the third feature set in the above embodiment.
[0144] In the first stage, the main agent is used to generate executable code for the new feature, based on which a new feature attribute can be generated.
[0145] refer to Figure 2 After LLM generates a new feature set based on the historical feature set, the main agent first analyzes the new feature set, thereby selecting features based on feature importance to support feature scalability. Among them, the main agent aims to discover new potential features by exploring the data set. Optionally, the main agent selects the newly generated Figure 2 The feature set 1 in the above example is analyzed, including: querying the top-k nodes in FTT that are most similar to the feature set of the currently selected node through the RAG program. The value of k in top-k depends on the length of the context that the LLM can accept. These similar nodes are provided to the main agent as historical experience, so as to achieve the effect of generating better features by reflecting on the historical feature generation records, so as to improve the quality of new features.
[0146] Optionally, the similarity between two feature sets is evaluated based on the semantic similarity and structural similarity of the linear combination, and the similarity is calculated as follows:
[0147] Sim(A1, A2) = Sim sem (A1, A2)×Sim struct (A1, A2)
[0148] Among them, Sim sem is the semantic similarity between two features. Existing embedding models such as the BGE model (BERT-based Generative Encoder) and the GTE model (Generative Text Embedding) can be used to encode the feature description d of each feature into a high-dimensional vector, and the similarity of the two texts is calculated by calculating the cosine value between the two representation vectors. Collect relevant positive and negative samples based on the preset machine learning tasks, and use contrastive learning to further fine-tune the embedding model, using infoNCE as the loss function. In contrastive learning, each data point is in the format of [q, k + , (k 1 , k 2 , …k N )], where q is the original query, i.e., the textual description of the current node information, and q is used to query the most similar nodes in history. p is a positive sample, k 1 , k 2 ,…k N is N negative samples. The goal of contrastive learning is to make the embedding of q The embeddings of 1 , k 2 , …k N) are as dissimilar as possible. Therefore, infoNCE is used as the loss function for further training. The similarity between feature sets A1 and A2 is represented as Sim sem (A1, A2), Sim sem The value range of (A1, A2) is [0,1]:
[0149]
[0150] Among them, Sim struct (A1, A2) is structural similarity: Considering that all features except the initial feature set are generated by combining several previous features, a feature set can be represented as a tree structure based on the dependency between features. For ease of understanding, assume that the initial feature set is [a, b, c, d, e], and the following five new features are generated based on the initial feature set, resulting in the above feature set 1:
[0151] A = a + b
[0152] B = b + c
[0153] D = d + e
[0154] X = A + B
[0155] Y = B + D
[0156] The displayed tree structure can be used to represent feature set 1. Optionally, the similarity between two feature sets is measured by using the edit distance and normalized to the [0,1] interval using the sigmoid function. Furthermore, the directed graph is converted to a directed tree by copying the shared nodes, thereby reducing the computational complexity of the structural similarity.
[0157] After the master agent analyzes the feature set, it cleans the new feature set. Optionally, it performs two key tasks, data normalization and noise elimination, based on the new features; the master agent identifies "null" values, incorrect data types, and other abnormal values in the new feature data, and corrects the erroneous data and abnormal data.
[0158] After the main agent performs data cleaning on the feature set, it integrates all previous responses of LLM as context and uses LLM to generate Python code for the entire feature generation process.
[0159] Among them, in the process of LLM generating feature set 2 based on feature set 1, the main agent will interact with the large language model by providing two types of prompts, including static prompts and dynamic prompts. The static prompts remain unchanged throughout the feature generation process, including a description of the task information of the machine learning task and predefined feature generation operations. Dynamic prompts are generated by querying the feature generation tree, including historical feature sets and the tree paths used to generate these feature sets. In order to evaluate the performance of the new potential features on the training data set, the main agent calculates relevant indicators and provides them as additional inputs to the LLM. Among them, the relevant indicators include the cross-validated prediction accuracy of the current node after training on the downstream specified model. The downstream specified model can be XGBOOST (eXtreme Gradient Boosting) or other models.
[0160] In the first stage, the optimization agent is used to further improve the correctness of the generated code. This is because features usually play a key role in improving the performance of downstream machine learning tasks, and the main agent tends to generate more complex features. When the complex logic of the generated features exceeds the ability of the LLM agent to generate correct code, it will cause abnormal code generation and regeneration. The solution of the related art is to re-call the main agent to generate features multiple times until the code is correct. However, this solution simply bypasses complex logic features and selects simple logic features, which will cause the performance of downstream machine learning tasks to degrade. Therefore, this embodiment sets an optimization agent to generate correct code features without affecting the performance of downstream machine learning tasks.
[0161] In order to ensure the correctness of the generated code and the performance of downstream machine learning tasks, the optimization agent adopts a decomposition method to decompose complex logical features into multiple simple logical features. Specifically, the optimization agent recursively uses the decomposition agent to decompose the complex feature generation task into a series of simple tasks until the generated child feature matches its corresponding parent feature, indicating that further decomposition is unnecessary. Furthermore, in order to improve efficiency, the optimization agent introduces a rollback mechanism, which checks the code complexity required to generate the feature and stops the decomposition when further decomposing the feature will increase the code complexity, so as to stop the decomposition in advance when the complexity increases. The specific implementation of feature decomposition is shown in Algorithm 2.
[0162] Algorithm 2:
[0163] 1: Optimize the agent’s feature decomposition process (F, c, d);
[0164] 2: F = {F};
[0165] 3: while True do:
[0166] 4: F next = ;
[0167] 5: for F∈ F do:
[0168] 6: F child ← ;
[0169] 7: if DecompEval(F) then:
[0170] 8: F child ← DecompAgent(F);
[0171] 9: if Rollback(F child , F) or StopDecomp(F child , F) then;
[0172] 10: F child = {F};
[0173] 11: ;
[0174] 12: if StopDecomp(F, F next ) then:
[0175] 13: break;
[0176] 14: F = F next ;
[0177] 15: Return F generated by iterative decomposition;
[0178] The following is an explanation of Algorithm 2:
[0179] In Algorithm 2, a feature F and its corresponding code c and feature description d are received, and a set of decomposed features F is returned. The received feature F is a newly generated feature by LLM, and the code c is the generated code of feature F. First, the returned feature set F is initialized with the input feature F, and then the complex features are recursively decomposed into a set of simple features. In each iteration, the DecompEval function is first used to evaluate whether feature F needs to be decomposed. If it is decided to decompose F, the DecompAgent function is called to use the ability of LLM to decompose feature F into a set of simple features F child When the generated feature set matches the initial feature set, Algorithm 2 terminates the decomposition. Return F generated by iterative decomposition, which is the fourth feature set in the above embodiment.
[0180] Among them, the DecompEval function in Algorithm 2 is used to evaluate whether a feature needs further decomposition based on the specified criteria.
[0181] Alternatively, the standard evaluation can be the credibility of the code obtained by self-consistency evaluation of the code generated by the master agent. The self-consistency theory in LLM indicates that the more consistent the answers of the agent in multiple responses, the higher the credibility of the final majority result. Therefore, for a feature F, k response codes c={c 1 , c 2 , …c K}, execute these codes c and compare their execution results. Select the code with the most identical execution results among the response codes c. If the maximum number of identical results exceeds a specified ratio of the total number of generated codes k, the code can be considered is reliable. Otherwise, the code Unreliable. When the code is unreliable, feature F is decomposed.
[0182] Exemplarily, the specified ratio of the total number of generated codes k is set to 2 / 3, and for four codes c that generate the same characteristics 1 , c 2 , c 3 , c 4 , if c 1 , c 2 , c 3 After executing on the original table, there are the same column values, while executing c 4 The results are different, so it can be considered that it meets self-consistency and meets the first preset condition in the above embodiment. This is because the ratio of identical results is 3 / 4 ≥ 2 / 3. The ratio of 2 / 3 comes from the Byzantine consensus theory, which points out that at least 3p+1 responses are required in a single round of communication to tolerate p erroneous responses. In actual application, the specified ratio of the total number of generated codes k can be modified to other values, which is not limited here.
[0183] Optionally, the standard evaluation may be the McCabe loop complexity M of the code c that generates the feature F. Loop complexity is a metric in software engineering that measures the complexity of the logic loop of the code by calculating the number of linearly independent paths in the code. In this embodiment, the complexity of the logic in the code c is evaluated by the loop complexity M. When the loop complexity M exceeds a specified value, it is determined that the second preset condition in the above embodiment is not met, and the feature F continues to be decomposed.
[0184] Exemplarily, the control flow graph of code c is analyzed, and the loop complexity of code c is calculated by the formula M = E – N + 2, where E represents the number of edges in the CFG (control flow graph) and N represents the number of nodes. For example, the loop complexity of a GROUPBY function that divides rows of data into two categories is 2, where E=4 and N=4. Among them, the code generated by the main agent can use Pandas library functions to operate on DataFrame: In order to facilitate the calculation of the loop complexity of features, the predefined loop complexity in the Pandas library function is introduced (for example, M(GROUPBY)=2, M(AVG)=2). For ease of understanding, when using AVG(GROUPBY(df[`sex'])) to calculate the features of the average salary of each gender, its loop complexity M=M(GROUPBY)+M(AVG)-1=2+2-1=3, where 1 is the loop complexity reduced by merging the CFG of the GROUPBY function with the CFG of the AVG function.
[0185] Optionally, the self-consistency and loop complexity of the code c of feature F can be combined to determine whether feature F needs to be decomposed: when the code c of feature F is unreliable or the loop complexity of the code exceeds cp bound When , the feature F is decomposed by Algorithm 2.
[0186] DecompAgent in Algorithm 2 can decompose the complex parent feature F into a set of simple sub-features F based on the agent. child . Divide the feature description d of the parent feature into a set of sub-descriptions d=d 1 ,d 2 ,…,d k , and then use the main agent to generate executable code c=c based on these sub-descriptions 1 ,c 2 ,…,c k , thereby generating k simple sub-features F child =F 1 ,F 2 ,…,F k For ease of understanding, Figure 3 A schematic diagram of eigendecomposition is provided. Figure 3As shown, for a complex feature Financial_Educational_Score, the optimization agent can decompose the generated code of the complex feature based on self-consistency and cyclomatic complexity, that is, the above-mentioned loop complexity, and decompose the complex feature into simpler features FD1, FD2 and FD3 based on the decomposed codes 1, 2, 3 and 4. Among them, FD1: Financial_Score is half of the normalized average profit of capital assets in each job category; FD2: Education_Score is twice the value of the number of years of education; FD3: Financial_Educational_Score is the weighted sum of Financial_Score and Education_Score, with weights of 0.5 and 2 respectively.
[0187] The Rollback function in Algorithm 2 is used to execute the above rollback mechanism. This setting is because some logically simple features have ambiguous feature descriptions. For example, the new feature "average area" can be calculated by dividing the area by the number of people, but there may be multiple columns in the original table containing information about the number of people, such as "number of men", "number of women", "number of children", etc. Although the code of these logically simple features shows a low cyclomatic complexity, it can easily become unreliable due to the ambiguity of the feature description. According to the DecompEval function mentioned above, these features need to be decomposed. However, decomposing these features will increase the cyclomatic complexity of the code for generating sub-features, resulting in increased logical complexity and prone to errors. In order to check the complexity of the sub-features generated after decomposition. Therefore, the Rollback mechanism is used to stop feature decomposition in advance. Specifically, the Rollback mechanism checks the cyclomatic complexity of the code required to generate the feature. After decomposing feature F and generating its corresponding sub-feature F child Afterwards, if F is generated child If the cyclomatic complexity of the code for F is equal to or greater than the cyclomatic complexity of the code that generated F, it will roll back to F and stop decomposing F. For example, if decomposing the feature "Average Area" leads to the generation of a sub-feature "Total Number of People" that needs to calculate the sum of the number of people based on age and gender, then this sub-feature will have a higher cyclomatic complexity.
[0188] Furthermore, in addition to the aforementioned Rollback mechanism for early stopping of decomposition, a StopDecomp function can also be introduced to terminate the iterative decomposition in Algorithm 2. The StopDecomp function is used to separate the column values generated by feature F from its sub-feature F child Specifically, if the column value of executing F is equal to the value of executing F in sequence childThe column values of all sub-features in indicate that further decomposition is unnecessary and the iterative decomposition in Algorithm 2 will be stopped.
[0189] Continue to see Figure 2 In the second stage, this embodiment is used to realize feature generation in the database. In the second stage, all processing tasks are integrated in the database engine, the workflow based on is converted into optimized SQL statements, features are efficiently generated in the database, and a large number of input and output (IO) operations required to export / import data between the database and the machine learning system are eliminated, thereby improving the efficiency of model generation in the database. The second stage includes two sub-stages: SQL generation through pattern matching and SQL SQL optimization. The feature generation in the database is specifically shown in Algorithm 3.
[0190] Algorithm 3:
[0191] 1: SQL generation in the database;
[0192] 2: Step 1: SQL generation through pattern matching;
[0193] 3: SQLs ← ;
[0194] 4: C 1 ,C 2 ,…,C k ← SplitFeatures(C);
[0195] 5: for i∈ 1,2,…k do:
[0196] 6: ast_i←AST(C i );
[0197] 7: SQLs ← SQLs + PatternMatching(ast i ,P);
[0198] 8: SQLs ←AppendTrain / Predict(SQLs);
[0199] 9: Step 2: SQL optimization;
[0200] 10: SQLs ← CombineOperator(SQLs);
[0201] 11: SQLs ← SplitPath(SQLs);
[0202] 12: return CTECombine(SQLs);
[0203] In step 1, the most common operations for generating new attributes are defined based on the code in Kaggle, and the operators for operating the data frame in pandas are obtained. Inspired by the Madlib system, each operator corresponds to a pure SQL statement or a SQL statement with a user-defined function (UDF). Considering that the traditional pure SQL interface is an efficient and portable native language suitable for writing data-intensive programs, when both pure SQL and UDF can implement the operator, pure SQL is preferred to improve efficiency. However, since SQL as a structured language has limited flexibility in handling complex linear operations or complex algorithms, it is necessary to use UDFs in linear algebra libraries such as Eigen and Armadillo, which can be compiled by JIT technology and registered in the database during execution.
[0204] Lines 2 to 7 of Algorithm 3 demonstrate the process of generating executable SQL from the initial Python code C. First, the feature code in the optimal feature set searched in the first stage is split into code blocks C by using the SplitFeatures function. 1 ,C 2 ,…,C k , and ensure that each code block only generates one new feature in the DataFrame. Then, the corresponding SQL is generated for each code block through the PatternMatching function in the AST (Abstract Syntax Tree), and the Train or Prediction operator is attached as required.
[0205] Line 7 of Algorithm 3 demonstrates the process of matching a code sequence with the Python statements provided in and converting it into SQL statements and UDF statements through the PatternMatching function. Similar to regular expressions, the code is parsed into a Python abstract syntax tree (AST) and pattern matching is performed at the AST level. For the matched code, the corresponding pure SQL statements and / or UDF statements are generated according to the operator type. If all the current codes cannot match any operators, they will fall back to a general operator called Fallback, which generates a UDF implemented in Python by simply wrapping the original code. Although Python UDF may slow down SQL execution, Fallback ensures minimal errors in the SQL generation process.
[0206] Furthermore, in order to implement subsequent machine learning model predictions, the model can be serialized into bytes or strings and stored in a database table through the AppendTrain function or the Predict function for further deserialization during prediction. Optionally, considering that the sample data may be inaccurate, the iterative verification process in the machine learning workflow can be followed. When the number of selected feature sets T>1, D train Divide into sub-datasets D' train and D validate , to evaluate the performance of each feature set on the full dataset (non-sample dataset). Then, the feature set with the highest performance is selected to train the final model.
[0207] In step 2, the SQL generated in step 1 is optimized again, aiming to improve performance by actively rewriting the query outside the database query optimizer.
[0208] Line 9 of Algorithm 3 demonstrates the process of rule-based operator merging. Considering that as the number of abstract syntax tree nodes increases during SQL parsing, the query optimizer of the database often cannot effectively handle query rewriting, which leads to a problem of slowing down the execution plan. To address this problem, this embodiment classifies simple operations for a small number of attributes and implements operators through pure SQL as simple operators, where pure SQL includes but is not limited to multi-element operators, simple unary operators, deletion operators, and missing value filling operators. The SQL operations of multiple simple operators that appear continuously in the pipeline are combined into one SQL statement. This can eliminate some redundant operations and simplify the original SQL. At the same time, Python / C++ UDF operators that appear continuously are combined together to reduce the repeated unpacking and packaging overhead when data is transferred in multiple UDF operators, thereby speeding up the final execution speed.
[0209] Line 10 in Algorithm 3 merges SQL through CTE (common table expressions): After generating the corresponding SQL and / or UDF for each operator, common table expressions can be used in SQL to combine SQL and / or UDF into a larger SQL expression, which avoids caching intermediate calculation results and allows the Volcano model to be used in the query engine to achieve a better streaming calculation process. Compared with the related art Madlib, which specifies the source table and target table for each operation and materializes the intermediate results after each operator, the SQL merging method in this embodiment eliminates unnecessary storage and computing overhead.
[0210] Line 11 of Algorithm 3 implements the UDF decomposition based on the value model. Through experiments, it can be observed that the call of UDF is the most time-consuming part of SQL execution. For sequential operator execution, each operator accepts all necessary attributes output by the previous operator, performs calculations and outputs the attributes to the next operator. While pure SQL statements can use the Volcano model to filter unnecessary data flows at runtime, the execution of UDFs may be completely different. Since it is impossible to understand the internal data access pattern and execution behavior of each UDF, the SQL engine treats it as a black box and needs to pay the full cost of unpacking and wrapping each input and output feature between the database format and the C++ raw data, even if the current operator does not need a certain feature.
[0211] Optionally, to improve the efficiency of SQL execution, the attributes involved in the current UDF calculation are split into an independent path, and the join operator is executed with the main branch after the calculation is completed. It should be noted that this rule-based optimization is not always beneficial, considering the additional overhead of the join operator. Therefore, the execution cost can be estimated before and after the split to determine whether to split. For example, if the original UDF is executed on table t, with attribute attrs, and is split into l attrs (UDF required attribute) and r attrs , then the new cost C UDF (t,l attrs )+C Join (t, l attrs , r attrs ) is less than the original cost C UDF (t,attrs).
[0212] Among them, you can use microbenchmarks for cost evaluation and calculate CPU and memory costs at startup for different table and attribute sizes. Optionally, execute pre-designed microbenchmarks on sampled data to calculate time ratios .in, Indicates the time it takes to package (Write) and unpack (Read) data of a unit size through UDF. Indicates the time to directly read (Read) and write (Write) data through SQL; thus the ratio of the two time ratio γ represents the additional overhead incurred by the packaging and unpacking of data in UDF compared to direct reading and writing of data in SQL during further execution. At the same time, hyperparameters for calculating the cost of each operator are defined, which are user-configurable. For example, hyperparameters include: cpu_operator_cost (COC), the cost of performing arithmetic operations on a row of data; cpu_tuple_cost (CTC), the cost of returning a row of data; io_byte_cost (IBC), the cost of reading one byte of data. Based on the configured hyperparameters, the cost of the connection operator and the UDF operator can be estimated by the following formula, where the function bytes(attr) represents the number of bytes required to store the attribute attr in a row of data. By estimating the cost of each operator, you can choose whether to split the path.
[0213]
[0214]
[0215] It is understandable that other hyperparameters may be configured, and the hyperparameter configuration is not limited herein.
[0216] In the second phase, Python code is seamlessly converted into native SQL queries based on semantic analysis and code execution. Given the inherent limitations of SQL in handling complex linear calculations, C++ user-defined functions are introduced to cleverly deal with these complexities. Specifically, a rule-based strategy is adopted to generate SQL queries at a fine-grained level; to further improve efficiency, common table expressions (CTEs) are used to integrate multiple SQL queries to avoid intermediate materialization. In addition, a comprehensive cost model is designed to guide the merging of SQL queries to ensure optimal performance. The final result of this process is the execution of optimized SQL queries in a database environment, thereby achieving seamless model generation without losing efficiency.
[0217] The third stage in this embodiment is model generation in the database. The execution engine executes the features generated in the second stage to generate a machine learning model and implement model training, or the model predicts based on the data in the database to obtain a prediction result.
[0218] It should be understood that the implementation steps involved in the embodiments described above are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed in turn or alternation with other steps or at least part of the steps or stages in other steps. For example, the second stage can first merge SQL through CTE, and then decompose UDF based on the value model as in Algorithm 3; it can also be as follows. Figure 2 As shown in the figure, the UDF is first decomposed based on the value model, and then the SQL is merged through CTE.
[0219] Based on the same inventive concept, the embodiment of the present application also provides a machine learning feature generation device in a database of a multi-agent agent for implementing the machine learning feature generation method in a database of a multi-agent agent involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in the embodiments of the machine learning feature generation device in the database of one or more multi-agent agents provided below can be referred to the limitations of the machine learning feature generation method in the database of a multi-agent agent above, and will not be repeated here.
[0220] In one embodiment, Figure 4 As shown, a machine learning feature generation device in a database of a multi-agent agent is provided, comprising: a first agent processing module and a second agent processing module, wherein:
[0221] The first agent processing module is used to determine the first feature set and the feature description of the first feature set according to the performance indicators of the historical feature set in the machine learning model in the database; obtain the feature prompt corresponding to the first feature set according to the preset machine learning task and the historical feature set; obtain the new features generated by the large language model in the database according to the first feature set, the feature description and the feature prompt, and obtain the second feature set by combining the first feature set and the new features; determine the third feature set according to the performance indicators of the historical feature set and the second feature set in the machine learning model;
[0222] The second intelligent agent processing module is used to decompose the third feature set until the decomposed feature set matches the third feature set, and obtain the fourth feature set required to perform the machine learning task according to the decomposition result.
[0223] In some of the embodiments, the feature prompts include static prompts and dynamic prompts. The first intelligent agent processing module obtains feature prompts corresponding to the first feature set based on the preset machine learning task and the historical feature set, including: obtaining static prompts based on the machine learning task; obtaining a feature set in the historical feature set whose similarity with the first feature set meets the preset degree, and obtaining dynamic prompts based on the feature set whose similarity meets the preset degree.
[0224] In some of the embodiments, the second agent processing module decomposes the third feature set until the decomposed feature set matches the third feature set, and obtains the fourth feature set required to perform the machine learning task according to the decomposition result, including: decomposing the third feature set to obtain multiple decomposed features; judging whether the credibility of the decomposed feature set meets a first preset condition; if the credibility meets the first preset condition, judging that the decomposed feature set matches the third feature set, stopping the decomposition and obtaining the fourth feature set; and / or, decomposing the third feature set to obtain a generated code for the decomposed feature set; judging whether the complexity of the generated code for the decomposed feature set meets a second preset condition; if the complexity meets the second preset condition, judging that the decomposed feature set matches the third feature set, stopping the decomposition and obtaining the fourth feature set.
[0225] In some of the embodiments, the second agent processing module decomposes the third feature set including: obtaining a feature description of the third feature set based on the third feature set; decomposing the feature description of the third feature set into multiple sub-descriptions; and calling the large language model to generate the decomposed feature set based on the sub-descriptions.
[0226] Furthermore, after the second agent processing module decomposes the third feature set, the execution method further includes: if the column value of the third feature set is equal to the column value of the decomposed feature set, then stopping the decomposition of the third feature set.
[0227] In some of the embodiments, the machine learning feature generation device in the database of the multi-agent agent also includes a code generation module, which is used to decompose the third feature set until the decomposed feature set matches the third feature set, and after obtaining the fourth feature set required to perform the machine learning task in the database according to the decomposition result, obtain the generated code of the fourth feature set; if there is a code matching a preset pattern in the generated code of the fourth feature set, the matching code is converted into a structured query statement of the database; if there is a code not matching the preset pattern in the generated code of the fourth feature set, the not matching code is converted into a user-defined function.
[0228] In some of the embodiments, the code generation module execution method further includes: combining the structured query statements according to operators in the structured query statements; and / or combining the structured query statements according to common table expressions.
[0229] In some of the embodiments, the code generation module contains code that does not match the preset pattern in the generated code of the fourth feature set, and after converting the non-matching code into a user-defined function, the execution method also includes: obtaining a first execution cost of the operator of the user-defined function; after splitting the user-defined function according to the attributes of the user-defined function, obtaining a second execution cost of the operator of the split user-defined function; and judging whether to split the user-defined function based on a comparison result of the first execution cost and the second execution cost.
[0230] In some of the embodiments, the first agent processing module determines the first feature set and the feature description of the first feature set based on the performance indicators of the historical feature set in the machine learning model in the database, including: obtaining the historical feature set in each node of the feature transformation tree; determining the node to be visited based on the performance parameters of the historical feature set in the machine learning model and the number of times each node is visited; obtaining the first feature set in the node to be visited and obtaining the feature description of the first feature set.
[0231] In some of the embodiments, the first agent processing module determines the third feature set based on the performance indicators of the historical feature set and the second feature set in the machine learning model in the database, including: adding the second feature set as a child node to the feature transformation tree; determining a new node to be visited based on the performance parameters of the feature set in each node of the feature transformation tree in the machine learning model and the number of times each node is visited; and obtaining the third feature set in the new node to be visited.
[0232] Each module in the machine learning feature generation device in the database of the multi-agent agent can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.
[0233] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. While storing data, the database of the computer device is also provided with a large language model and a machine learning model. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for generating machine learning features in a database of a multi-agent agent is implemented.
[0234] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0235] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.
[0236] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0237] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0238] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.
[0239] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0240] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A machine learning feature generation method in a database based on a multi-agent agent, characterized in that: The method comprises: Determine a first feature set and a feature description of the first feature set according to a performance indicator of the historical feature set in a machine learning model in a database; the feature description is a description of data attributes in the first feature set; Obtaining feature prompts corresponding to the first feature set according to a preset machine learning task and the historical feature set; Acquire new features generated by a large language model in a database according to the first feature set, the feature description and the feature prompt, and obtain a second feature set by combining the first feature set and the new features; Determining a third feature set according to performance indicators of the historical feature set and the second feature set in the machine learning model; Decomposing the third feature set until the feature set obtained by the decomposition matches the third feature set, and obtaining a fourth feature set required for performing the machine learning task according to the decomposition result; The feature prompt includes a static prompt and a dynamic prompt, and obtaining the feature prompt corresponding to the first feature set according to the preset machine learning task and the historical feature set includes: obtaining the static prompt according to the machine learning task; obtaining a feature set whose similarity with the first feature set meets a preset degree from the historical feature set, and obtaining the dynamic prompt according to the feature set whose similarity meets the preset degree; The decomposing the third feature set includes: obtaining a feature description of the third feature set according to the third feature set; decomposing the feature description of the third feature set into a plurality of sub-descriptions; and calling the large language model to generate a decomposed feature set according to the sub-descriptions.
2. The method according to claim 1, characterized in that Decomposing the third feature set until the feature set obtained by the decomposition matches the third feature set, and obtaining a fourth feature set required for performing the machine learning task according to the decomposition result includes: Decomposing the third feature set to obtain a plurality of decomposed features; Determining whether the credibility of the decomposed feature set meets the first preset condition; If the credibility satisfies the first preset condition, it is determined that the feature set obtained by decomposition matches the third feature set, the decomposition is stopped and the fourth feature set is obtained; and / or, Decomposing the third feature set to obtain a generated code of the decomposed feature set; Determining whether the complexity of the generated code of the decomposed feature set meets the second preset condition; If the complexity satisfies the second preset condition, it is determined that the feature set obtained by decomposition matches the third feature set, and the decomposition is stopped to obtain the fourth feature set.
3. The method according to claim 1, characterized in that After obtaining new features generated by the large language model in the database according to the first feature set, the feature description and the feature prompt, and combining the first feature set and the new features to obtain a second feature set, the method further includes: Determining whether the number of iterations of the second feature set reaches a preset number; If not, taking the first feature set and the second feature set as the iterated historical feature set; Determine a new first feature set and a feature description of the new first feature set according to the performance indicators of the iterated historical feature set in the machine learning model in the database; obtain feature prompts corresponding to the new first feature set according to the preset machine learning task and the iterated historical feature set; obtain new features generated by the large language model according to the new first feature set, the feature description of the new first feature set and the feature prompts corresponding to the new first feature set, and obtain the iterated second feature set in combination with the new first feature set and the new features; If so, the step of determining a third feature set based on the performance indicators of the historical feature set and the second feature set in the machine learning model is performed.
4. The method according to claim 1, characterized in that: After decomposing the third feature set until the feature set obtained by the decomposition matches the third feature set, and obtaining a fourth feature set required for performing the machine learning task in the database according to the decomposition result, the method further includes: Obtaining a generated code for the fourth feature set; If there is a code matching a preset pattern in the generated code of the fourth feature set, converting the matching code into a structured query statement of the database; If there is a code that does not match the preset pattern in the generated code of the fourth feature set, the unmatched code is converted into a user-defined function.
5. The method according to claim 4, characterized in that The method further comprises: Combining the structured query statements according to operators in the structured query statements; and / or, The structured query statement is assembled according to the common table expression.
6. The method according to claim 4, characterized in that If there is a code that does not match the preset pattern in the generated code of the fourth feature set, after converting the unmatched code into a user-defined function, the method further includes: Obtaining a first execution cost of an operator of the user-defined function; After splitting the user-defined function according to the attribute of the user-defined function, obtaining a second execution cost of the operator of the split user-defined function; Whether to split the user-defined function is determined according to a comparison result of the first execution cost and the second execution cost.
7. The method according to claim 1, characterized in that Determining the first feature set and the feature description of the first feature set according to the performance indicator of the historical feature set in the machine learning model in the database includes: Obtain the historical feature set in each node of the feature transformation tree; Determine the node to be visited according to the performance parameter of the historical feature set in the machine learning model and the number of times each of the nodes has been visited; Acquire a first feature set in the node to be visited and obtain a feature description of the first feature set.
8. The method according to claim 7, characterized in that Determining the third feature set according to the performance indicators of the historical feature set and the second feature set in the machine learning model in the database includes: Adding the second feature set as a child node to the feature conversion tree; Determine a new node to be visited according to the performance parameters of the feature set in each node of the feature conversion tree in the machine learning model and the number of times each node is visited; Obtain a third feature set in the new node to be visited.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Autonomous learning platform for novel feature discovery
CN110612535A
Method and system for realizing machine learning algorithm based on SQL (Structured Query Language)
CN117312357A