Method and system for processing data tables and automatically training machine learning models
By obtaining and utilizing the table relation configuration information of the data table, multiple data tables are spliced into basic sample tables and generating derivative features, the problem of handling multiple data tables in the prior art is solved, and automated machine learning sample generation and model training are realized.
Patent Information
- Application Number
- CN202011205070.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-02
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2040-11-02
AI Technical Summary
The prior art is difficult to process multiple data tables easily and effectively to obtain machine learning samples, and cannot perform machine learning tasks automatically.
By obtaining the table relationship configuration information of multiple data tables, the data table is spliced into a basic sample table, and derivative features are generated based on the table, and finally a sample table including multiple machine learning samples is formed.
It realizes convenient and efficient processing of multiple data tables, generates effective machine learning samples, improves the effectiveness of machine learning models, and automates the execution process of machine learning.
Smart Images

Figure CN114443639B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of data processing, and more specifically, to a method and system for processing data tables, and a method and system for automatically training a machine learning model. Background Art
[0002] In real business environments such as online advertising, recommendation systems, financial market analysis, and healthcare, data sources are extensive and are often stored in different data tables. At the same time, data such as user behavior or commodity trading volume changes over time, so there is a large amount of time-series relational data.
[0003] In machine learning applications, experienced scientists in modeling need to continuously try and make mistakes to build valuable features based on multiple related data tables to improve the performance of machine learning models. Summary of the Invention
[0004] Exemplary embodiments of the present disclosure are provided to provide a method and system for processing data tables to solve the problem in the prior art that multiple data tables cannot be conveniently and effectively processed to obtain machine learning samples. In addition, exemplary embodiments of the present disclosure also provide a method and system for automatically training a machine learning model to solve the problem in the prior art that machine learning cannot be automatically executed effectively starting from data splicing.
[0005] According to an exemplary embodiment of the present disclosure, a method for processing a data table is provided, including: obtaining table relationship configuration information about a plurality of data tables, where the table relationship configuration information includes: association relationships between pairwise data tables; based on the table relationship configuration information, splicing the plurality of data tables into a basic sample table; generating derivative features about the fields based on the fields in the basic sample table, and incorporating the generated derivative features into the basic sample table to form a sample table including multiple machine learning samples.
[0006] Optionally, the step of splicing the plurality of data tables into a basic sample table includes: splicing the fields in one of the two data tables with an association relationship based on an association field into the other data table according to a splicing order until the specified data table is spliced, and splicing the fields in the data tables other than the specified data table into the specified data table to form a basic sample table.
[0007] Optionally, the step of splicing the fields in the data tables other than the specified data table into the specified data table according to the splicing order to form the basic sample table includes: splicing the fields in the data tables other than the specified data table into the specified data table directly or after respectively performing various aggregation processes according to the splicing order; for each field that is spliced into the specified data table only after respectively performing various aggregation processes, screening out the aggregation fields with relatively low feature importance from the respective aggregation fields spliced into the specified data table after respectively performing various aggregation processes on this field; deleting the screened aggregation fields from the spliced specified data table to obtain the basic sample table.
[0008] Optionally, in the step of splicing the fields in the data tables other than the specified data table into the specified data table directly or after respectively performing various aggregation processes according to the splicing order, when the field in any data table other than the specified data table can be spliced into the specified data table directly from its initial data table according to the splicing order without aggregation processing, the field of this data table is directly spliced into the specified data table from its initial data table according to the splicing order; when the field in any data table other than the specified data table can be spliced into the specified data table from its initial data table according to the splicing order only after aggregation processing, in the process of splicing this field from its initial data table into the specified data table according to the splicing order, whenever it is necessary to perform aggregation processing to splice the field of this data table or the aggregation field of the field of this data table into the next data table, the respective aggregation fields obtained after respectively performing various aggregation processes on the field of this data table or the aggregation field of the field of this data table are spliced into the next data table, where the various aggregation processes include various time-series aggregation processes based on each time window and / or various non-time-series aggregation processes.
[0009] Optionally, the step of splicing the fields in the data tables other than the specified data table into the specified data table directly or after respectively performing various aggregation processes according to the splicing order includes: generating a splicing path for each field in each data table other than the specified data table except for the associated field between it and the data table it is to be spliced into, for splicing this field directly or after respectively performing various aggregation processes into the specified data table according to the splicing order; for each of the above fields, splicing this field into the specified data table directly or after respectively performing aggregation processes according to the splicing path of this field.
[0010] Optionally, the step of screening out the aggregation fields with relatively low feature importance from the aggregation fields spliced into the specified data table obtained after performing various aggregation processes on this field respectively includes: for each aggregation field in the aggregation fields spliced into the specified data table obtained after performing various aggregation processes on this field respectively, training a corresponding machine learning model based on the fields in the specified data table that have not undergone aggregation processing and this aggregation field; screening out from the aggregation fields spliced into the specified data table obtained after performing various aggregation processes on this field respectively: the aggregation fields with relatively poor effects of the corresponding machine learning models as the aggregation fields with relatively low feature importance.
[0011] Optionally, the table relationship configuration information includes: the maximum number of splicings. Among them, the step of splicing the fields in the data tables other than the specified data table into the specified data table in accordance with the splicing order to form a basic sample table includes: determining whether there is a data table among the multiple data tables that needs to be spliced into the specified data table in accordance with the splicing order and the number of splicings exceeds the maximum number of splicings; when it is determined that there is such a data table, splicing the fields in the data tables other than the specified data table and the determined data table among the multiple data tables into the specified data table in accordance with the splicing order to form a basic sample table; when it is determined that there is no such data table, splicing the fields in the data tables other than the specified data table among the multiple data tables into the specified data table in accordance with the splicing order to form a basic sample table.
[0012] Optionally, the step of generating derivative features about this field based on the fields in the basic sample table and incorporating the generated derivative features into the basic sample table includes: (a) performing the i-th round of derivation on the features in the current feature search space, and screening out the derivative features with relatively high feature importance from the derivative features generated in the i-th round, where the initial value of i is 1, and the initial value of the feature search space is the first predetermined number of fields with the highest feature importance in the basic sample table; (b) when i is less than the preset threshold, updating the feature search space to the first predetermined number of fields with the highest feature importance in the basic sample table except those that have been used as the feature search space, setting i = i + 1, and returning to execute step (a); (c) when i is greater than or equal to the preset threshold, incorporating the derivative features screened out in the previous i rounds into the basic sample table.
[0013] Optionally, the steps of performing the i-th round of derivation on the features in the current feature search space include: performing various first-order processes on each feature in the current feature search space to generate respective first-order derived features; and / or, performing various second-order processes on each pair of features in the current feature search space to generate respective second-order derived features; and / or, performing various third-order processes on each triple of features in the current feature search space to generate respective third-order derived features, where the first-order process is a process that takes only a single feature as the processing object; the second-order process is a process based on two features that performs processing on at least one of the two features; and the third-order process is a process based on three features that performs processing on at least one of the three features.
[0014] Optionally, the steps of screening out the derived features with relatively high feature importance from the derived features generated in the i-th round include: for each derived feature in the derived features generated in the i-th round, training a corresponding machine learning model based on the features in the current feature search space and the derived feature; screening out from the derived features generated in the i-th round: the derived features for which the effect of the corresponding machine learning model meets a preset condition.
[0015] Optionally, the time series aggregation processing includes at least one of the following items: summation, averaging, taking the maximum value, taking the minimum value, calculating the number of different field values, calculating the number of field values, calculating the variance, calculating the mean square deviation, finding the preset number of field values with the highest occurrence frequency, taking the previous field value, taking the previous non-empty field value; the non-time series aggregation processing includes at least one of the following items: summation, averaging, taking the maximum value, taking the minimum value, calculating the number of different field values, calculating the number of field values, calculating the variance, calculating the mean square deviation, finding the preset number of field values with the highest occurrence frequency.
[0016] Optionally, the steps of obtaining the table relationship configuration information regarding multiple data tables include: obtaining the table relationship configuration information regarding the multiple data tables according to the input operations performed by the user on the screen, where the input operations include: input operations for specifying an association relationship between two data tables.
[0017] According to another exemplary embodiment of the present disclosure, there is provided a method for automatically training a machine learning model, including: a sample table including multiple machine learning samples obtained by performing the steps of the method as described above; based on the sample table, respectively training machine learning models using different machine learning algorithms and different hyperparameters; determining the machine learning model with the best effect from the trained machine learning models as the finally trained machine learning model.
[0018] According to another exemplary embodiment of the present disclosure, a system for processing data tables is provided, including: a configuration information acquisition device adapted to acquire table relationship configuration information about a plurality of data tables, wherein the table relationship configuration information includes: association relationships between pairwise data tables; a splicing device adapted to splice the plurality of data tables into a basic sample table based on the table relationship configuration information; a sample table generation device adapted to generate derivative features about the fields based on the fields in the basic sample table and incorporate the generated derivative features into the basic sample table to form a sample table including multiple machine learning samples.
[0019] Optionally, the splicing device is adapted to splice the fields in one of the two data tables having an association relationship to the other data table based on the association fields according to a splicing order until the fields in the data tables other than the specified data table are spliced to the specified data table to form a basic sample table.
[0020] Optionally, the splicing device includes: a splicing unit adapted to directly or respectively perform various aggregation processes on the fields in the data tables other than the specified data table according to the splicing order and then splice them to the specified data table; a screening unit adapted to, for each field that is spliced to the specified data table only after performing various aggregation processes respectively, screen out the aggregation fields with relatively low feature importance from the respective aggregation fields spliced to the specified data table obtained by performing various aggregation processes on this field respectively; a basic sample table generation unit adapted to delete the screened-out aggregation fields from the spliced specified data table to obtain a basic sample table.
[0021] Optionally, when the fields in any data table other than the specified data table can be directly spliced to the specified data table from their initial data tables according to the splicing order without aggregation processes, the splicing unit is adapted to directly splice the fields of this data table from their initial data tables to the specified data table according to the splicing order; when the fields in any data table other than the specified data table can be spliced to the specified data table from their initial data tables according to the splicing order only after aggregation processes, during the process of splicing the fields of this data table from their initial data tables to the specified data table according to the splicing order, whenever it is necessary to perform aggregation processes to splice the fields of this data table or the aggregation fields of the fields of this data table to the next data table, the respective aggregation fields obtained by performing various aggregation processes on the fields of this data table or the aggregation fields of the fields of this data table are spliced to the next data table, wherein the various aggregation processes include various time-series aggregation processes based on respective time windows and / or various non-time-series aggregation processes.
[0022] Optionally, the splicing unit is adapted to generate a splicing path for each field in each data table other than the specified data table, except for the association fields between the data table to be spliced and the data table to which it is to be spliced, and splice the field into the specified data table after directly or separately performing various aggregation processes on the field in accordance with the splicing order; and for each field, directly or separately perform an aggregation process on the field in accordance with the splicing path of the field and splice the field into the specified data table.
[0023] Optionally, the screening unit is adapted to train a corresponding machine learning model for each aggregated field in the aggregated fields spliced into the specified data table obtained by separately performing various aggregation processes on the field, based on the unaggregated fields in the specified data table after splicing and the aggregated field; and screen out from the aggregated fields spliced into the specified data table obtained by separately performing various aggregation processes on the field: the aggregated fields with relatively poor effects of the corresponding machine learning models as the aggregated fields with relatively low feature importance.
[0024] Optionally, the table relationship configuration information includes: a maximum number of splices. Among them, the splicing device is adapted to determine whether there is a data table among the multiple data tables that needs to be spliced more than the maximum number of splices to be finally spliced into the specified data table in accordance with the splicing order; where when it is determined that there is such a data table, fields in the data tables among the multiple data tables other than the specified data table and the determined data table are spliced into the specified data table in accordance with the splicing order to form a basic sample table; when it is determined that there is no such data table, fields in the data tables among the multiple data tables other than the specified data table are spliced into the specified data table in accordance with the splicing order to form a basic sample table.
[0025] Optionally, the sample table generation device is adapted to perform the following processing: (a) perform an i-th round of derivation on the features in the current feature search space, and screen out the derived features with relatively high feature importance from the derived features generated in the i-th round, where the initial value of i is 1, and the initial value of the feature search space is the first predetermined number of fields with the highest feature importance in the basic sample table; (b) when i is less than a preset threshold, update the feature search space to the first predetermined number of fields with the highest feature importance in the basic sample table except those that have been used as the feature search space, set i = i + 1, and return to perform processing (a); (c) when i is greater than or equal to the preset threshold, incorporate the derived features screened out in the previous i rounds into the basic sample table.
[0026] Optionally, the sample table generation device is adapted to perform the following processes: performing various first-order processes on each feature in the current feature search space respectively to generate respective first-order derivative features; and / or, performing various second-order processes on every two features in the current feature search space respectively to generate respective second-order derivative features; and / or, performing various third-order processes on every three features in the current feature search space respectively to generate respective third-order derivative features, wherein the first-order process is a process that takes only a single feature as the processing object; the second-order process is a process based on two features that performs processing on at least one of the two features; the third-order process is a process based on three features that performs processing on at least one of the three features.
[0027] Optionally, for each derivative feature generated in the i-th round, the sample table generation device is adapted to train a corresponding machine learning model based on the features in the current feature search space and the derivative feature; and screen out from the derivative features generated in the i-th round: the derivative features for which the effect of the corresponding machine learning model meets a preset condition.
[0028] Optionally, the time series aggregation process includes at least one of the following items: summation, averaging, taking the maximum value, taking the minimum value, calculating the number of different field values, calculating the number of field values, calculating variance, calculating mean square deviation, obtaining the preset number of field values with the highest occurrence frequency, taking the previous field value, taking the previous non-null field value; the non-time series aggregation process includes at least one of the following items: summation, averaging, taking the maximum value, taking the minimum value, calculating the number of different field values, calculating the number of field values, calculating variance, calculating mean square deviation, obtaining the preset number of field values with the highest occurrence frequency.
[0029] Optionally, the configuration information acquisition device is adapted to obtain table relationship configuration information about the multiple data tables according to the input operations performed by the user on the screen, wherein the input operations include: input operations for specifying an association relationship between two data tables.
[0030] According to another exemplary embodiment of the present disclosure, there is provided a system for automatically training a machine learning model, including: the system for processing data tables as described above; a training device adapted to train machine learning models using different machine learning algorithms and different hyperparameters respectively based on a sample table including multiple machine learning samples obtained by the system for processing data tables; and a determination device adapted to determine the machine learning model with the best effect from the trained machine learning models as the finally trained machine learning model.
[0031] According to another exemplary embodiment of the present disclosure, a system including at least one computing device and at least one storage device storing instructions is provided, wherein, when the instructions are run by the at least one computing device, the at least one computing device is caused to execute the method for processing a data table as described above or the method for automatically training a machine learning model as described above.
[0032] According to another exemplary embodiment of the present disclosure, a computer-readable storage medium storing instructions is provided, wherein, when the instructions are run by at least one computing device, the at least one computing device is caused to execute the method for processing a data table as described above or the method for automatically training a machine learning model as described above.
[0033] The method and system for processing a data table according to an exemplary embodiment of the present disclosure provide a convenient and effective way to process a data table, which not only improves the processing efficiency and reduces the usage threshold of feature engineering, but also facilitates extracting effective features to form machine learning samples to improve the effect of a machine learning model.
[0034] In addition, the method and system for automatically training a machine learning model according to an exemplary embodiment of the present disclosure can not only automatically train a machine learning model that meets the requirements, greatly reducing the threshold of machine learning, and further, since the obtained machine learning samples include effective feature information, the effect of the corresponding trained machine learning model can be further improved.
[0035] Additional aspects and / or advantages of the general concept of the present disclosure will be set forth in part in the description that follows, and in part will be obvious from the description, or may be learned by practice of the general concept of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Through the following description with reference to the drawings of exemplary embodiments, the above and other objects and features of the exemplary embodiments of the present disclosure will become more apparent, wherein:
[0037] Figure 1 A flowchart showing a method for processing a data table according to an exemplary embodiment of the present disclosure;
[0038] Figure 2 A flowchart showing a method for splicing a plurality of data tables into a basic sample table according to an exemplary embodiment of the present disclosure;
[0039] Figure 3 A flowchart showing a method for generating derivative features and incorporating the generated derivative features into a basic sample table according to an exemplary embodiment of the present disclosure;
[0040] Figure 4 A flowchart showing a method for automatically training a machine learning model according to an exemplary embodiment of the present disclosure;
[0041] Figure 5 Block diagram showing a system for processing a data table according to an exemplary embodiment of the present disclosure;
[0042] Figure 6 Block diagram showing a splicing device according to an exemplary embodiment of the present disclosure;
[0043] Figure 7 Block diagram showing a system for automatically training a machine learning model according to an exemplary embodiment of the present disclosure. Detailed implementation manners
[0044] Reference will now be made in detail to the embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings, wherein the same reference numerals always refer to the same components. The embodiments will be described below with reference to the accompanying drawings to explain the present disclosure.
[0045] Figure 1 Flowchart showing a method for processing a data table according to an exemplary embodiment of the present disclosure. Here, as an example, the method can be executed by a computer program, or by a dedicated hardware device or an aggregate of software and hardware resources for performing machine learning, big data computing, or data analysis, etc. For example, the method can be executed by a data warehouse software for data storage and management, a machine learning platform for implementing machine learning-related services, etc.
[0046] Refer to Figure 1 , in step S10, obtain table relationship configuration information about a plurality of data tables.
[0047] The table relationship configuration information includes: the association relationship between two data tables. For example, there is an association relationship between data table A and data table B. Further, as an example, the table relationship configuration information may further include at least one of the following items: the association fields based on which the association relationship between two data tables is established, the specific type of the association relationship between two data tables, the maximum number of splicing times, the type of at least one data table among the plurality of data tables, and the splicing end point. It should be understood that the table relationship configuration information may further include other appropriate information for configuring the splicing method between a plurality of data tables, and the present disclosure places no limitation thereon.
[0048] As an example, according to the input operations performed by the user on the screen, table relationship configuration information about the multiple data tables can be obtained. As an example, the input operations may include: input operations for specifying an association relationship between two data tables. In addition, as an example, the input operations may further include at least one of the following items: input operations for specifying the type of the association relationship between two data tables, input operations for specifying the association fields based on which the association relationship between two data tables is established, input operations for specifying the type of at least one data table among the multiple data tables, input operations for specifying the maximum number of splicing times allowed, and input operations for specifying the data table as the splicing end point.
[0049] Here, each data record in the data table can be regarded as a description of an event or an object, corresponding to an example or a sample. In the data record, there is attribute information reflecting the performance or nature of the event or object in a certain aspect, that is, a field. For example, a row of the data table corresponds to a data record, and a column of the data table corresponds to a field. In two data tables having an association relationship based on association fields, the meaning of the corresponding association field in one data table is the same as the meaning of the corresponding association field in the other data table, so that the data records in the two data tables can be corresponded based on these two association fields to achieve splicing. For example, when splicing, the data records with the same field values of these two association fields in the two data tables can be spliced together. It should be understood that the field names of these two association fields can be the same or different. For example, one association field can be the "ID" field, and its corresponding association field can be the "UserID" field. Although the field names are different, the business information they describe is essentially the same, both for describing the ID number of the user. For example, in two data tables having an association relationship based on association fields, the corresponding association field in one data table can be the primary key of this data table, and the corresponding association field in the other data table can be the foreign key of this primary key.
[0050] As an example, the types of association relationships between pairwise data tables may include: one-to-one, one-to-many, many-to-one, and many-to-many. Specifically, if there is an association relationship between data table A and data table B based on association field C, and the same field value of association field C in data table A can only appear in one data record, and the same field value of association field C in data table B can only appear in one data record, then the type of the association relationship between data table A and data table B is: one-to-one; if the same field value of association field C in data table A can only appear in one data record, and the same field value of association field C in data table B may appear in multiple data records, then the type of the association relationship between data table A and data table B is: one-to-many; if the same field value of association field C in data table A may appear in multiple data records, and the same field value of association field C in data table B can only appear in one data record, then the type of the association relationship between data table A and data table B is: many-to-one; if the same field value of association field C in data table A may appear in multiple data records, and the same field value of association field C in data table B may appear in multiple data records, then the type of the association relationship between data table A and data table B is: many-to-many.
[0051] As an example, the types of data tables may include but are not limited to at least one of the following items: static table, time series table, slice table. It should be understood that the types of data tables may also include other appropriate types, and the present disclosure does not limit this.
[0052] In step S20, based on the table relationship configuration information, the multiple data tables are spliced into a basic sample table.
[0053] As an example, the fields in the data tables other than the specified data table can be spliced into the specified data table to form a basic sample table according to the splicing order of splicing the fields in one of the two data tables with an association relationship to the other data table based on the association field between them until the specified data table is spliced.
[0054] Specifically, based on the association relationships between every two data tables among the multiple data tables, the splicing order of the multiple data tables can be determined. Then, according to this splicing order, the fields in the data tables other than the specified data table are spliced into the specified data table to form a basic sample table. For example, if there is an association relationship between data table A and data table B, an association relationship between data table B and data table C, an association relationship between data table C and data table D, an association relationship between data table D and data table E, an association relationship between data table E and data table F, and data table E is the specified data table (i.e., the splicing end point mentioned above), then the splicing order can be: data table A → data table B → data table C → data table D → data table E, and data table F → data table E. For example, the fields of data table A need to be spliced into data table B first, and then from data table B spliced into data table C until spliced into data table E; the fields of data table B need to be spliced into data table C first, then from data table C spliced into data table D, and then from data table D spliced into data table E; the fields of data table F are spliced into data table E.
[0055] In addition, as an example, it can be determined whether there is a data table among the multiple data tables that needs to be spliced into the specified data table according to the splicing order and the number of splicing times exceeds the maximum splicing times; when it is determined that there is such a data table, according to the splicing order, the fields in the data tables other than the specified data table and the determined data table among the multiple data tables are spliced into the specified data table to form a basic sample table, that is, the data table that needs to be spliced into the specified data table according to the splicing order and the number of splicing times exceeds the maximum splicing times does not participate in the splicing; when it is determined that there is no such data table, according to the splicing order, the fields in the data tables other than the specified data table among the multiple data tables are spliced into the specified data table to form a basic sample table.
[0056] This will be combined with Figure 2 to describe an exemplary embodiment of step S20.
[0057] In step S30, derivative features regarding the fields are generated based on the fields in the basic sample table, and the generated derivative features are incorporated into the basic sample table to form a sample table including multiple machine learning samples.
[0058] As an example, a derivative feature can be generated based on at least one field.
[0059] This will be combined with Figure 3 to describe an exemplary embodiment of step S30.
[0060] Figure 2 A flowchart showing a method for splicing multiple data tables into a basic sample table according to an exemplary embodiment of the present disclosure.
[0061] As Figure 2As shown, in step S201, the fields in the data tables other than the specified data table are spliced directly (i.e., without aggregating the field values) or after various aggregation processes are performed separately according to the splicing order to the specified data table.
[0062] As an example, when the fields in any data table other than the specified data table can be spliced directly from their original data tables to the specified data table according to the splicing order without the need for aggregation processing, the fields of this data table can be directly spliced from their original data tables to the specified data table according to the splicing order.
[0063] As an example, when the fields in any data table other than the specified data table can be spliced from their original data tables to the specified data table according to the splicing order only after aggregation processing, during the process of splicing this field from its original data table to the specified data table according to the splicing order, whenever aggregation processing is required to splice the field of this data table or the aggregated field of the field of this data table to the next data table, the respective aggregated fields obtained after performing various aggregation processes on the field of this data table or the aggregated field of the field of this data table can be spliced to the next data table. In addition, when the field of this data table or the aggregated field of the field of this data table can be spliced directly to the next data table without the need for aggregation processing, the field of this data table or the aggregated field of the field of this data table can be directly spliced to the next data table. That the fields in the data table can be spliced from their original data tables to the specified data table according to the splicing order only after aggregation processing can be understood as: during the process of splicing the fields in the data table from their original data tables to the specified data table according to the splicing order, there is at least one time when only aggregation processing can continue the splicing downward.
[0064] For example, when a data table (i.e., the data table to be spliced) needs to be spliced into another data table (i.e., the data table to be spliced into), if the association relationship between the data table to be spliced and the data table to be spliced into is: one-to-one or one-to-many, the fields of the data table to be spliced can be directly spliced into the data table to be spliced into; if the association relationship between the data table to be spliced and the data table to be spliced into is: many-to-one or many-to-many, the fields of the data table to be spliced need to be respectively subjected to various aggregation processes and then spliced into the data table to be spliced into. Here, the fields of the data table to be spliced that are respectively subjected to various aggregation processes and then spliced into the fields of the data table to be spliced into are the aggregated fields of the fields of the data table to be spliced. It should be understood that performing n kinds of aggregation processes on a field respectively will result in n aggregated fields. For example, data table 1 needs to be spliced into data table 2, data table 2 needs to be spliced into data table 3, and the fields of data table 1 need to be aggregated before being spliced into data table 2, and the fields of data table 2 need to be aggregated before being spliced into data table 3. For field a of data table 1 other than the associated field with data table 2, n kinds of aggregation processes can be respectively performed on field a, and the n aggregated fields of field a can be spliced into data table 2. When splicing data table 2 into data table 3, n kinds of aggregation processes can be respectively performed on the n aggregated fields of field a again, and the aggregated result of n*n aggregated fields can be spliced into data table 3. As an example, when the type of the data table to be spliced is a slice table, the data table to be spliced can be spliced into the data table to be spliced into in a last-join manner.
[0065] As an example, the various aggregation processes may include various time-series aggregation processes and / or various non-time-series aggregation processes based on each time window. As an example, the table relationship configuration information may further include at least one of the following items: the types of the various time-series aggregation processes, the types of the various non-time-series aggregation processes, the sizes of the respective time windows. As an example, the types of the various time-series aggregation processes, the types of the various non-time-series aggregation processes, and the sizes of the respective time windows can be specified by the user. It should be understood that the same time-series aggregation process under different time windows also belongs to different aggregation processes; performing the same aggregation process based on different aggregation reference fields also belongs to different aggregation processes.
[0066] As an example, the fields can be aggregated according to the types of the fields. As an example, the types of the fields in the data table can be specified by the user. As an example, the types of the fields in the data table can be set according to the input operations performed by the user in the graphical interface displayed on the screen for setting the types of the fields in the data table.
[0067] As an example, the types of fields may include at least one of the following items: singlestring (single string type), arraystring (array string type), kvstring (key-value string type), continuenum (continuous value type), time (timestamp type).
[0068] As an example, the types of time-series aggregation processing may include at least one of the following items: summation, averaging, taking the maximum value, taking the minimum value, calculating the number of different field values, calculating the number of field values, calculating variance, calculating mean square deviation, finding the preset number of field values with the highest occurrence frequency, taking the previous field value, taking the previous non-null field value.
[0069] As an example, the types of non-time-series aggregation processing may include at least one of the following items: summation, averaging, taking the maximum value, taking the minimum value, calculating the number of different field values, calculating the number of field values, calculating variance, calculating mean square deviation, finding the preset number of field values with the highest occurrence frequency.
[0070] In addition, as an example, for each field in each data table except the specified data table, except for the association fields between the data table to be spliced and the data table to which it is to be spliced, a splicing path for directly or separately performing various aggregation processes on the field and then splicing it to the specified data table according to the splicing order may be generated; for each field, the field is spliced to the specified data table after directly or separately performing aggregation processing according to the splicing path of the field. Thus, the underlying computing engine can process the fields in each data table except the specified data table according to the splicing path and splice them to the specified data table.
[0071] In step S202, for each field that is spliced to the specified data table only after performing various aggregation processes, the aggregation fields with relatively low feature importance are filtered out from the various aggregation fields obtained after performing various aggregation processes on the field and then spliced to the specified data table.
[0072] In step S203, the filtered aggregation fields are deleted from the specified data table after splicing to obtain a basic sample table.
[0073] In other words, the fields in the specified data table obtained after splicing that are not obtained through aggregation processing are directly retained, while among the fields that have been obtained through aggregation processing (i.e., aggregation fields), only the aggregation fields with relatively high feature importance are retained.
[0074] Various appropriate methods can be used to screen out the aggregated fields with relatively low feature importance. As an example, for each field that is spliced into the specified data table after various aggregation processes are performed separately, for each aggregated field among the aggregated fields spliced into the specified data table obtained after various aggregation processes are performed on this field separately, a corresponding machine learning model can be trained based on the unaggregated fields in the specified data table after splicing and this aggregated field; from the aggregated fields obtained after various aggregation processes are performed on this field separately and spliced into the specified data table, screen out: the aggregated fields with relatively poor performance of the corresponding machine learning model as the aggregated fields with relatively low feature importance.
[0075] Since the performance of the corresponding machine learning model can reflect the feature importance (e.g., predictive power) of the candidate aggregated field, the candidate aggregated fields can be screened by measuring the performance of the machine learning model corresponding to each candidate aggregated field. For example, the better the performance of the corresponding machine learning model, the higher the feature importance. As an example, a specified model evaluation metric can be used to evaluate the performance of the machine learning model corresponding to each candidate aggregated field. As an example, the model evaluation metric can be AUC (Area Under ROC Curve, the area under the ROC (Receiver Operating Characteristic) curve), MAE (Mean Absolute Error), or logloss (logarithmic loss function), etc.
[0076] As an example, for each field that is spliced into the specified data table after various aggregation processes are performed separately, a certain number of aggregated fields with the lowest feature importance can be screened out from the aggregated fields obtained after various aggregation processes are performed on this field separately and spliced into the specified data table, and this number can be determined based on the total number of fields in the data table where this field was initially located.
[0077] According to the method of splicing multiple data tables into a basic sample table in the present disclosure, effective feature information of the associated table can be automatically mined to form a basic sample table.
[0078] Figure 3 A flowchart showing a method for generating derivative features and incorporating the generated derivative features into a basic sample table according to an exemplary embodiment of the present disclosure.
[0079] As Figure 3 shown, in step S301, perform the i-th round of derivation on the features in the current feature search space, and screen out the derivative features with relatively high feature importance from the derivative features generated in the i-th round, where the initial value of i is 1.
[0080] As an example, the initial value of the feature search space can be the first predetermined number of fields with the highest feature importance in the basic sample table. In addition, other appropriate methods can also be used to determine the initial value of the feature search space. For example, meta-learning can be used to determine the initial value of the feature search space.
[0081] As an example, the steps of performing the i-th round of derivation on the features in the current feature search space may include:
[0082] Performing various first-order processes on each feature in the current feature search space respectively to generate respective first-order derived features;
[0083] And / or, performing various second-order processes on every two features in the current feature search space respectively to generate respective second-order derived features;
[0084] And / or, performing various third-order processes on every three features in the current feature search space respectively to generate respective third-order derived features. Thus, at least one of the generated first-order derived features and / or second-order derived features and / or third-order derived features can be used as the derived features of the i-th round of derivation.
[0085] As an example, the first-order process can be a process that takes only a single feature as the processing object; the second-order process can be a process based on two features for at least one of the two features; the third-order process can be a process based on three features for at least one of the three features.
[0086] As an example, the user can specify the ways of the first-order process, the second-order process, and the third-order process. For example, the ways of the first-order process may include: discretizing continuous features. The ways of the second-order process may include: combining and / or aggregating two features, and / or aggregating another feature with one feature as the aggregation benchmark. For example, the ways of the third-order process may include: combining and / or aggregating three features, and / or aggregating two other features with one feature as the aggregation benchmark, and / or aggregating another feature with two features as the aggregation benchmark.
[0087] As an example, the steps of screening out the derived features with relatively high feature importance from the derived features generated in the i-th round may include: for each derived feature in the derived features generated in the i-th round, training a corresponding machine learning model based on the features in the current feature search space and the derived feature; screening out from the derived features generated in the i-th round: the derived features for which the effect of the corresponding machine learning model meets the preset conditions.
[0088] In step S302, determine whether i is less than the preset threshold.
[0089] In step S303, when i is less than a preset threshold, update the feature search space to the first predetermined number of fields with the highest feature importance in the basic sample table except those that have already been used as the feature search space, set i = i + 1, and return to execute step S301.
[0090] In step S304, when i is greater than or equal to the preset threshold, incorporate the derivative features screened out in the previous i rounds into the basic sample table.
[0091] As an example, when performing each round of derivation, only part of the data records in the basic sample table can be used. A suitable number of data records can be selected in each round, and the number selected in different rounds can be the same or different, which can be reasonably set according to calculation requirements (considering aspects such as computing configuration and computing time required). Also, since it is necessary to verify the effect of the corresponding machine learning model, verification samples are required. The part of the data records in the basic sample table used in each round can be split into the training samples and verification samples of this round according to time series. Finally, derivative features can be generated for all the data records in the basic sample table according to the specific forms of the derivative features screened out in the previous i rounds and incorporated into the basic sample table.
[0092] In addition, as an example, the number of derivative features selected in each round can be the same or different; the number of fields in the feature search space used in each round can be the same or different.
[0093] According to the exemplary embodiments of the present disclosure, valuable feature information can be automatically derived based on the fields in the basic sample table to improve the comprehensiveness and effectiveness of the features in the sample table.
[0094] According to the exemplary embodiments of the present disclosure, the user only needs to establish an association relationship between any two data tables to obtain a sample table including relatively comprehensive valuable feature information spliced based on multiple data tables. The user operation is simple, intuitive, and easy to understand, so that users without experience in feature engineering can also construct a sample table with good effects for machine learning model training.
[0095] Figure 4 A flowchart showing a method for automatically training a machine learning model according to an exemplary embodiment of the present disclosure is shown. Here, as an example, the method can be executed by a computer program, or by a dedicated hardware device or an aggregate of software and hardware resources for performing machine learning, big data computing, or data analysis, etc. For example, the method can be executed by a machine learning platform for implementing machine learning-related services.
[0096] As Figure 4 shown, in step S10, obtain table relationship configuration information regarding multiple data tables, where the table relationship configuration information includes: association relationships between pairwise data tables.
[0097] In step S20, based on the table relationship configuration information, the multiple data tables are spliced into a basic sample table.
[0098] In step S30, derivative features about the fields are generated based on the fields in the basic sample table, and the generated derivative features are incorporated into the basic sample table to form a sample table including multiple machine learning samples. It should be understood that steps S10 to S30 can be implemented with reference to the specific implementation manners described above Figures 1 to 3 and will not be elaborated herein.
[0099] In step S40, machine learning models using different machine learning algorithms and different hyperparameters are respectively trained based on the sample table.
[0100] As an example, appropriate algorithms such as random search, grid search, Bayesian optimization, and the hyperparameter optimization algorithm hyperband can be used to respectively train machine learning models using different machine learning algorithms and different hyperparameters based on the sample table.
[0101] In step S50, the machine learning model with the best effect is determined from the trained machine learning models as the finally trained machine learning model.
[0102] According to the exemplary embodiment of the present disclosure, the user only needs to perform an input operation for establishing an association relationship between any two data tables that is easy to operate and intuitive to understand, and then a machine learning model that meets the requirements can be trained. Thus, business personnel who do not have professional capabilities related to machine learning can also independently complete the modeling work, greatly reducing the threshold of machine learning, and also freeing modeling engineers from learning about the business in the target field and enabling them to be engaged in more professional production work.
[0103] Figure 5 A block diagram of a system for processing data tables according to an exemplary embodiment of the present disclosure is shown.
[0104] As Figure 5 shown, the system for processing data tables according to the exemplary embodiment of the present disclosure includes: a configuration information acquisition device 10, a splicing device 20, and a sample table generation device 30.
[0105] Specifically, the configuration information acquisition device 10 is adapted to acquire table relationship configuration information about multiple data tables, where the table relationship configuration information includes: the association relationships between pairwise data tables.
[0106] The splicing device 20 is adapted to splice the multiple data tables into a basic sample table based on the table relationship configuration information.
[0107] The sample table generation device 30 is adapted to generate derivative features regarding the fields based on the fields in the base sample table, and incorporate the generated derivative features into the base sample table to form a sample table including multiple machine learning samples.
[0108] As an example, the splicing device 20 may be adapted to splice the fields in one of the two data tables having an association relationship to the other data table based on the association field until reaching the splicing order of the specified data table, and splice the fields in the data tables other than the specified data table to the specified data table to form a base sample table.
[0109] Figure 6 The block diagram of the splicing device according to an exemplary embodiment of the present disclosure is shown.
[0110] As Figure 6 shown, the splicing device 20 may include: a splicing unit 201, a screening unit 202, and a base sample table generation unit 203.
[0111] Specifically, the splicing unit 201 is adapted to splice the fields in the data tables other than the specified data table to the specified data table directly or after performing various aggregation processes respectively according to the splicing order.
[0112] The screening unit 202 is adapted to screen out the aggregation fields with relatively low feature importance from the respective aggregation fields spliced to the specified data table obtained after performing various aggregation processes respectively for each field that is spliced to the specified data table only after performing various aggregation processes.
[0113] The base sample table generation unit 203 is adapted to delete the screened aggregation fields from the spliced specified data table to obtain a base sample table.
[0114] As an example, when the fields in any data table other than the specified data table can be spliced to the specified data table directly from their original data tables according to the splicing order without the need for aggregation processing, the splicing unit 201 may be adapted to splice the fields of the data table directly from their original data tables to the specified data table according to the splicing order.
[0115] As an example, the splicing unit 201 may be adapted such that when a field in any data table other than the specified data table can only be spliced into the specified data table in accordance with the splicing order from the data table where it is initially located after aggregation processing, during the process of splicing the field from the data table where it is initially located into the specified data table in accordance with the splicing order, whenever it is necessary to perform aggregation processing to splice the field of the data table or the aggregated field of the field of the data table into the next data table, the respective aggregated fields obtained by performing various aggregation processes on the field of the data table or the aggregated field of the field of the data table are spliced into the next data table, where the various aggregation processes include various time-series aggregation processes and / or various non-time-series aggregation processes based on respective time windows.
[0116] As an example, the splicing unit 201 may be adapted to generate a splicing path for each field in each data table other than the specified data table, except for the association field between the data table to which it is to be spliced, to splice the field into the specified data table directly or after performing various aggregation processes in accordance with the splicing order; and for each such field, splice the field into the specified data table after directly or separately performing aggregation processing in accordance with the splicing path of the field.
[0117] As an example, the screening unit 202 may be adapted to train a corresponding machine learning model for each aggregated field spliced into the specified data table obtained by performing various aggregation processes on the field separately, based on the unaggregated fields in the specified data table after splicing and the aggregated field; and screen out from the respective aggregated fields spliced into the specified data table obtained by performing various aggregation processes on the field separately: the aggregated fields with relatively poor performance of the corresponding machine learning model as the aggregated fields with relatively low feature importance.
[0118] As an example, the table relationship configuration information includes: a maximum number of splices, where the splicing device 20 may be adapted to determine whether there is a data table among the multiple data tables that needs to be spliced more than the maximum number of splices to be finally spliced into the specified data table in accordance with the splicing order; where, when it is determined that there is such a data table, the fields in the data tables among the multiple data tables except the specified data table and the determined data table are spliced into the specified data table in accordance with the splicing order to form a basic sample table; when it is determined that there is no such data table, the fields in the data tables among the multiple data tables except the specified data table are spliced into the specified data table in accordance with the splicing order to form a basic sample table.
[0119] As an example, the sample table generation device 30 may be adapted to perform the following processing: (a) perform the i-th round of derivation on the features in the current feature search space, and screen out the derived features with relatively high feature importance from the derived features generated in the i-th round, where the initial value of i is 1, and the initial value of the feature search space is the first predetermined number of fields with the highest feature importance in the basic sample table; (b) when i is less than the preset threshold, update the feature search space to the first predetermined number of fields with the highest feature importance in the basic sample table except those that have been used as the feature search space, let i = i + 1, and return to perform processing (a); (c) when i is greater than or equal to the preset threshold, incorporate the derived features screened out in the previous i rounds into the basic sample table.
[0120] As an example, the sample table generation device 30 may be adapted to perform the following processing: perform various first-order processes on each feature in the current feature search space to generate respective first-order derived features; and / or, perform various second-order processes on every two features in the current feature search space to generate respective second-order derived features; and / or, perform various third-order processes on every three features in the current feature search space to generate respective third-order derived features, where the first-order process is a process that takes only a single feature as the processing object; the second-order process is a process based on two features that performs processing on at least one of the two features; the third-order process is a process based on three features that performs processing on at least one of the three features.
[0121] As an example, for each of the derived features generated in the i-th round, the sample table generation device 30 may be adapted to train a corresponding machine learning model based on the features in the current feature search space and the derived feature; and screen out from the derived features generated in the i-th round: the derived features for which the effect of the corresponding machine learning model meets the preset conditions.
[0122] As an example, the time series aggregation processing may include at least one of the following items: summation, averaging, taking the maximum value, taking the minimum value, calculating the number of different field values, calculating the number of field values, calculating the variance, calculating the mean square deviation, finding the preset number of field values with the highest occurrence frequency, taking the previous field value, taking the previous non-empty field value; the non-time series aggregation processing may include at least one of the following items: summation, averaging, taking the maximum value, taking the minimum value, calculating the number of different field values, calculating the number of field values, calculating the variance, calculating the mean square deviation, finding the preset number of field values with the highest occurrence frequency.
[0123] As an example, the configuration information acquisition device 10 may be adapted to obtain the table relationship configuration information about the multiple data tables according to the input operations performed by the user on the screen, where the input operations include: input operations for specifying an association relationship between two data tables.
[0124] Figure 7 A block diagram showing a system for automatically training a machine learning model according to an exemplary embodiment of the present disclosure.
[0125] As Figure 7 shown, the system for automatically training a machine learning model according to an exemplary embodiment of the present disclosure includes: a configuration information acquisition device 10, a splicing device 20, a sample table generation device 30, a training device 40, and a determination device 50.
[0126] Specifically, the configuration information acquisition device 10 is adapted to acquire table relationship configuration information regarding a plurality of data tables, where the table relationship configuration information includes: association relationships between pairwise data tables.
[0127] The splicing device 20 is adapted to splice the plurality of data tables into a basic sample table based on the table relationship configuration information.
[0128] The sample table generation device 30 is adapted to generate derivative features regarding the fields based on the fields in the basic sample table, and incorporate the generated derivative features into the basic sample table to form a sample table including multiple machine learning samples.
[0129] The training device 40 is adapted to respectively train machine learning models using different machine learning algorithms and different hyperparameters based on the sample table including multiple machine learning samples obtained by the system for processing data tables.
[0130] The determination device 50 is adapted to determine the machine learning model with the best effect from the trained machine learning models as the finally trained machine learning model.
[0131] It should be understood that the specific implementation manners of the system for processing data tables and the system for automatically training a machine learning model according to the exemplary embodiments of the present disclosure may be implemented with reference to the relevant specific implementation manners described in conjunction with Figures 1 to 4 and will not be elaborated herein.
[0132] The devices included in the system for processing data tables and the system for automatically training a machine learning model according to the exemplary embodiments of the present disclosure may be respectively configured as software, hardware, firmware, or any combination of the above items for performing specific functions. For example, these devices may correspond to dedicated integrated circuits, may also correspond to pure software code, or may also correspond to modules combining software and hardware. In addition, one or more functions implemented by these devices may also be uniformly executed by components in a physical entity device (such as a processor, a client, or a server, etc.).
[0133] It should be understood that the method for processing data tables according to an exemplary embodiment of the present disclosure can be implemented by a program recorded on a computer-readable medium. For example, according to an exemplary embodiment of the present disclosure, a computer-readable medium for processing data tables can be provided, wherein a computer program for performing the following method steps is recorded on the computer-readable medium: obtaining table relationship configuration information about a plurality of data tables, wherein the table relationship configuration information includes: the association relationship between every two data tables; based on the table relationship configuration information, splicing the plurality of data tables into a basic sample table; generating derivative features about the fields based on the fields in the basic sample table, and incorporating the generated derivative features into the basic sample table to form a sample table including multiple machine learning samples.
[0134] It should be understood that the method for automatically training a machine learning model according to an exemplary embodiment of the present disclosure can be implemented by a program recorded on a computer-readable medium. For example, according to an exemplary embodiment of the present disclosure, a computer-readable medium for automatically training a machine learning model can be provided, wherein a computer program for performing the following method steps is recorded on the computer-readable medium: obtaining table relationship configuration information about a plurality of data tables, wherein the table relationship configuration information includes: the association relationship between every two data tables; based on the table relationship configuration information, splicing the plurality of data tables into a basic sample table; generating derivative features about the fields based on the fields in the basic sample table, and incorporating the generated derivative features into the basic sample table to form a sample table including multiple machine learning samples; based on the sample table, respectively training machine learning models using different machine learning algorithms and different hyperparameters; determining the machine learning model with the best effect from the trained machine learning models as the finally trained machine learning model.
[0135] The computer program in the above computer-readable medium can run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. It should be noted that the computer program can also be used to execute additional steps other than the above steps or perform more specific processing when executing the above steps. The content of these additional steps and further processing has been described with reference to Figures 1 to 4 and will not be repeated here to avoid redundancy.
[0136] It should be noted that the systems for processing data tables and automatically training machine learning models according to the exemplary embodiments of the present disclosure can completely rely on the running of computer programs to implement corresponding functions, that is, each device corresponds to each step in the functional architecture of the computer program, so that the entire system is called through a special software package (for example, lib library) to implement the corresponding functions.
[0137] On the other hand, each device included in the system for processing data tables and the system for automatically training a machine learning model according to an exemplary embodiment of the present disclosure may also be implemented by hardware, software, firmware, middleware, microcode, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments for performing the corresponding operations may be stored in a computer-readable medium such as a storage medium, so that the processor can execute the corresponding operations by reading and running the corresponding program code or code segments.
[0138] For example, an exemplary embodiment of the present disclosure may also be implemented as a computing device, which includes a storage component and a processor. A set of computer-executable instructions is stored in the storage component. When the set of computer-executable instructions is executed by the processor, a method for processing data tables or a method for automatically training a machine learning model is executed.
[0139] Specifically, the computing device may be deployed in a server or a client, or may also be deployed on a node device in a distributed network environment. In addition, the computing device may be a PC computer, a tablet device, a personal digital assistant, a smart phone, a web application, or other devices capable of executing the above instruction set.
[0140] Here, the computing device does not have to be a single computing device, but may also be an aggregate of any devices or circuits capable of executing the above instructions (or instruction sets) individually or jointly. The computing device may also be a part of an integrated control system or a system manager, or may be configured to be interconnected with a portable electronic device locally or remotely (e.g., via wireless transmission).
[0141] In the computing device, the processor may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0142] Certain operations described in the method for processing data tables and the method for automatically training a machine learning model according to an exemplary embodiment of the present disclosure may be implemented in a software manner, certain operations may be implemented in a hardware manner, and in addition, these operations may also be implemented in a combination of software and hardware.
[0143] The processor may run the instructions or code stored in one of the storage components, where the storage component may also store data. The instructions and data may also be sent and received via a network interface device through a network, where the network interface device may adopt any known transmission protocol.
[0144] The storage component can be integrated with the processor. For example, RAM or flash memory can be arranged within an integrated circuit microprocessor, etc. In addition, the storage component can include independent devices, such as external disk drives, storage arrays, or other storage devices that can be used by any database system. The storage component and the processor can be operatively coupled or can communicate with each other, for example, through I / O ports, network connections, etc., such that the processor can read the files stored in the storage component.
[0145] In addition, the computing device can further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the computing device can be connected to each other via a bus and / or a network.
[0146] The operations involved in the method of processing data tables and the method of automatically training a machine learning model according to an exemplary embodiment of the present disclosure can be described as various interconnected or coupled functional blocks or functional diagrams. However, these functional blocks or functional diagrams can be equally integrated into a single logical device or operate with non-exact boundaries.
[0147] For example, as described above, a computing device according to an exemplary embodiment of the present disclosure for processing data tables can include a storage component and a processor, where a set of computer-executable instructions is stored in the storage component. When the set of computer-executable instructions is executed by the processor, the following steps are performed: obtaining table relationship configuration information regarding a plurality of data tables, where the table relationship configuration information includes: the association relationships between pairwise data tables; based on the table relationship configuration information, splicing the plurality of data tables into a basic sample table; generating derivative features regarding the fields based on the fields in the basic sample table, and incorporating the generated derivative features into the basic sample table to form a sample table including multiple machine learning samples.
[0148] For example, as described above, a computing device according to an exemplary embodiment of the present disclosure for automatically training a machine learning model can include a storage component and a processor, where a set of computer-executable instructions is stored in the storage component. When the set of computer-executable instructions is executed by the processor, the following steps are performed: obtaining table relationship configuration information regarding a plurality of data tables, where the table relationship configuration information includes: the association relationships between pairwise data tables; based on the table relationship configuration information, splicing the plurality of data tables into a basic sample table; generating derivative features regarding the fields based on the fields in the basic sample table, and incorporating the generated derivative features into the basic sample table to form a sample table including multiple machine learning samples; based on the sample table, training machine learning models using different machine learning algorithms and different hyperparameters respectively; determining the machine learning model with the best effect from the trained machine learning models as the finally trained machine learning model.
[0149] The foregoing describes various exemplary embodiments of the present disclosure. It should be understood that the above description is merely exemplary and not exhaustive, and the present disclosure is not limited to the disclosed exemplary embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the present disclosure. Therefore, the scope of protection of the present disclosure should be determined by the scope of the claims.
Claims
1. A method for processing data tables, comprising: obtaining table relationship configuration information about multiple data tables, wherein the table relationship configuration information includes: the association relationships between pairwise data tables; based on the table relationship configuration information, splicing the multiple data tables into a basic sample table; generating derivative features about the fields based on the fields in the basic sample table, and incorporating the generated derivative features into the basic sample table to form a sample table including multiple machine learning samples; wherein the step of generating derivative features about the fields based on the fields in the basic sample table and incorporating the generated derivative features into the basic sample table includes: (a) performing the i-th round of derivation on the features in the current feature search space, and screening out the derivative features with relatively high feature importance from the derivative features generated in the i-th round, wherein the initial value of i is 1, and the initial value of the feature search space is the first predetermined number of fields with the highest feature importance in the basic sample table; (b) when i is less than a preset threshold, updating the feature search space to the first predetermined number of fields with the highest feature importance in the basic sample table except those that have been used as the feature search space, setting i = i + 1, and returning to execute step (a); (c) when i is greater than or equal to the preset threshold, incorporating the derivative features screened out in the previous i rounds into the basic sample table; wherein the step of screening out the derivative features with relatively high feature importance from the derivative features generated in the i-th round includes: for each derivative feature in the derivative features generated in the i-th round, training a corresponding machine learning model based on the features in the current feature search space and this derivative feature; screening out from the derivative features generated in the i-th round: the derivative features whose corresponding machine learning model effects meet preset conditions.
2. The method according to claim 1, wherein, the step of splicing the multiple data tables into a basic sample table includes: splicing the fields in one of the two data tables with an association relationship into the other data table based on the association fields according to the splicing order until the specified data table is spliced, and splicing the fields in the data tables other than the specified data table into the specified data table to form a basic sample table.
3. The method according to claim 2, wherein, the step of splicing the fields in the data tables other than the specified data table into the specified data table according to the splicing order to form a basic sample table includes: splicing the fields in the data tables other than the specified data table into the specified data table directly or after respectively performing various aggregation processes according to the splicing order; for each field that is spliced into the specified data table after respectively performing various aggregation processes, screening out the aggregation fields with relatively low feature importance from the respective aggregation fields spliced into the specified data table obtained by respectively performing various aggregation processes on this field; deleting the screened-out aggregation fields from the spliced specified data table to obtain a basic sample table.
4. The method according to claim 3, wherein, in the step of splicing the fields in the data tables other than the specified data table into the specified data table directly or after respectively performing various aggregation processes according to the splicing order, When a field in any data table other than the specified data table can be spliced into the specified data table directly from the data table where it is initially located in accordance with the splicing order without the need for aggregation processing, the field of this data table is directly spliced into the specified data table from the data table where it is initially located in accordance with the splicing order; When a field in any data table other than the specified data table can be spliced into the specified data table from the data table where it is initially located in accordance with the splicing order only through aggregation processing, during the process of splicing this field from the data table where it is initially located into the specified data table in accordance with the splicing order, whenever aggregation processing is required to splice the field of this data table or the aggregated field of the field of this data table into the next data table, the respective aggregated fields obtained after performing various aggregation processes on the field of this data table or the aggregated field of the field of this data table are spliced into the next data table. Among them, the various aggregation processes include various time-series aggregation processes and / or various non-time-series aggregation processes based on each time window.
5. The method according to claim 3 or 4, wherein, The step of splicing the fields in the data tables other than the specified data table into the specified data table directly or respectively after performing various aggregation processes in accordance with the splicing order includes: For each field in each data table other than the specified data table except for the associated field between it and the data table it is to be spliced into, generating a splicing path for splicing this field directly or respectively after performing various aggregation processes into the specified data table in accordance with the splicing order; For each such field, splicing this field into the specified data table directly or respectively after performing aggregation processing in accordance with the splicing path of this field.
6. The method according to claim 3, wherein, The step of screening out the aggregated fields with relatively low feature importance from the respective aggregated fields spliced into the specified data table obtained after performing various aggregation processes on this field respectively includes: For each aggregated field among the respective aggregated fields spliced into the specified data table obtained after performing various aggregation processes on this field, training a corresponding machine learning model based on the fields in the specified data table that have not undergone aggregation processing after splicing and this aggregated field; Screening out from the respective aggregated fields spliced into the specified data table obtained after performing various aggregation processes on this field: the aggregated fields with relatively poor effects of the corresponding machine learning models as the aggregated fields with relatively low feature importance.
7. The method according to claim 2, wherein, The table relationship configuration information includes: the maximum number of splices, Among them, the step of splicing the fields in the data tables other than the specified data table into the specified data table in accordance with the splicing order to form a basic sample table includes: Determining whether there is a data table among the multiple data tables that needs to be spliced more times than the maximum number of splices when finally spliced into the specified data table in accordance with the splicing order; When it is determined that there is such a data table, splicing the fields in the data tables among the multiple data tables other than the specified data table and the determined data table into the specified data table in accordance with the splicing order to form a basic sample table; When it is determined that there is none, the fields in the data tables other than the specified data table among the multiple data tables are spliced into the specified data table in the splicing order to form a basic sample table.
8. The method according to claim 1, wherein, the step of performing the i-th round of derivation on the features in the current feature search space includes: performing various first-order processes on each feature in the current feature search space to generate respective first-order derived features; and / or, performing various second-order processes on every two features in the current feature search space to generate respective second-order derived features; and / or, performing various third-order processes on every three features in the current feature search space to generate respective third-order derived features, wherein, the first-order process is a process that takes only a single feature as the processing object; the second-order process is a process based on two features that performs processing on at least one of the two features; the third-order process is a process based on three features that performs processing on at least one of the three features.
9. The method according to claim 4, wherein, the time-series aggregation processing includes at least one of the following items: summation, averaging, taking the maximum value, taking the minimum value, calculating the number of different field values, calculating the number of field values, calculating the variance, calculating the mean square deviation, finding the preset number of field values with the highest occurrence frequency, taking the previous field value, taking the previous non-null field value; the non-time-series aggregation processing includes at least one of the following items: summation, averaging, taking the maximum value, taking the minimum value, calculating the number of different field values, calculating the number of field values, calculating the variance, calculating the mean square deviation, finding the preset number of field values with the highest occurrence frequency.
10. The method according to claim 1, wherein, the step of obtaining the table relationship configuration information about multiple data tables includes: obtaining the table relationship configuration information about the multiple data tables according to the input operation performed by the user on the screen, wherein, the input operation includes: an input operation for specifying an association relationship between two data tables.
11. A method for automatically training a machine learning model, including: a sample table including multiple machine learning samples obtained by performing the steps of the method according to any one of claims 1 to 10; based on the sample table, respectively training machine learning models using different machine learning algorithms and different hyperparameters; determining the machine learning model with the best effect from the trained machine learning models as the finally trained machine learning model.
12. A system for processing data tables, including: a configuration information acquisition device adapted to acquire the table relationship configuration information about multiple data tables, wherein the table relationship configuration information includes: the association relationship between two data tables; a splicing device adapted to splice the multiple data tables into a basic sample table based on the table relationship configuration information; a sample table generation device adapted to generate derived features about the fields based on the fields in the basic sample table and incorporate the generated derived features into the basic sample table to form a sample table including multiple machine learning samples; wherein, the sample table generation device is adapted to perform the following processing: (a) Perform the i-th round of derivation on the features in the current feature search space, and screen out the derived features with relatively high feature importance from the derived features generated in the i-th round. Here, the initial value of i is 1, and the initial value of the feature search space is the first predetermined number of fields with the highest feature importance in the basic sample table; (b) When i is less than the preset threshold, update the feature search space to the first predetermined number of fields with the highest feature importance in the basic sample table except those that have been used as the feature search space, let i = i + 1, and return to execute process (a); (c) When i is greater than or equal to the preset threshold, incorporate the derived features screened out in the previous i rounds into the basic sample table; Among them, the sample table generation device is adapted to, for each derived feature in the derived features generated in the i-th round, train a corresponding machine learning model based on the features in the current feature search space and this derived feature; and screen out from the derived features generated in the i-th round: the derived features whose effects of the corresponding machine learning model meet the preset conditions.
13. The system according to claim 12, wherein, The splicing device is adapted to splice the fields in one of the two data tables with an association relationship to another data table based on the association field according to the splicing order until it is spliced to the specified data table, and splice the fields in the data tables other than the specified data table to the specified data table to form the basic sample table.
14. The system according to claim 13, wherein, The splicing device includes: A splicing unit, adapted to splice the fields in the data tables other than the specified data table to the specified data table directly or after respectively performing various aggregation processes according to the splicing order; A screening unit, adapted to screen out the aggregation fields with relatively low feature importance from the respective aggregation fields spliced to the specified data table obtained by respectively performing various aggregation processes on each field that is spliced to the specified data table only after various aggregation processes; A basic sample table generation unit, adapted to delete the screened aggregation fields from the spliced specified data table to obtain the basic sample table.
15. The system according to claim 14, wherein, The splicing unit is adapted to, when the fields in any data table other than the specified data table can be directly spliced to the specified data table from their initial data tables according to the splicing order without aggregation processing, directly splice the fields of this data table from their initial data tables to the specified data table according to the splicing order; The splicing unit is adapted to, when the fields in any data table other than the specified data table can be spliced to the specified data table from their initial data tables according to the splicing order only after aggregation processing, in the process of splicing this field from its initial data table to the specified data table according to the splicing order, whenever it is necessary to perform aggregation processing to splice the field of this data table or the aggregation field of the field of this data table to the next data table, splice the respective aggregation fields obtained by respectively performing various aggregation processes on the field of this data table or the aggregation field of the field of this data table to the next data table, Among them, the various aggregation processes include various time-series aggregation processes and / or various non-time-series aggregation processes based on each time window.
16. The system according to claim 14 or 15, wherein, The splicing unit is adapted to generate a splicing path for each field in each data table except the specified data table, except for the associated field between the data table to be spliced and the data table to be spliced to, and splice the field to the specified data table after directly or separately performing various aggregation processes on the field according to the splicing order; and for each field, splice the field to the specified data table after directly or separately performing aggregation processes according to the splicing path of the field.
17. The system according to claim 14, wherein, The screening unit is adapted to train a corresponding machine learning model for each aggregated field obtained by separately performing various aggregation processes on the field and spliced to the specified data table, based on the unaggregated field in the specified data table after splicing and the aggregated field; And screen out from the aggregated fields obtained by separately performing various aggregation processes on the field and spliced to the specified data table: the aggregated fields with relatively poor effects of the corresponding machine learning models as the aggregated fields with relatively low feature importance.
18. The system according to claim 13, wherein, The table relationship configuration information includes: the maximum number of splices, wherein, the splicing device is adapted to determine whether there is a data table among the multiple data tables that needs to be spliced more than the maximum number of splices to be finally spliced to the specified data table according to the splicing order; wherein, when it is determined that there is, splice the fields in the data tables among the multiple data tables except the specified data table and the determined data table to the specified data table according to the splicing order to form a basic sample table; when it is determined that there is no, splice the fields in the data tables among the multiple data tables except the specified data table to the specified data table according to the splicing order to form a basic sample table.
19. The system according to claim 12, wherein, The sample table generation device is adapted to perform the following processes: Perform various first-order processes on each feature in the current feature search space respectively to generate respective first-order derivative features; and / or, perform various second-order processes on every two features in the current feature search space respectively to generate respective second-order derivative features; and / or, perform various third-order processes on every three features in the current feature search space respectively to generate respective third-order derivative features, wherein, the first-order process is a process that takes only a single feature as the processing object; The second-order process is a process based on two features and performed on at least one of the two features; The third-order process is a process based on three features and performed on at least one of the three features.
20. The system according to claim 15, wherein, The time-series aggregation process includes at least one of the following items: summation, averaging, taking the maximum value, taking the minimum value, calculating the number of different field values, calculating the number of field values, calculating the variance, calculating the mean square deviation, finding the preset number of field values with the highest occurrence frequency, taking the previous field value, taking the previous non-null field value; Non-temporal aggregation processing includes at least one of the following items: summation, averaging, taking the maximum value, taking the minimum value, calculating the number of different field values, calculating the number of field values, calculating variance, calculating mean square deviation, and obtaining the preset number of field values with the highest occurrence frequency.
21. The system according to claim 12, wherein, the configuration information acquisition device is adapted to obtain table relationship configuration information about the plurality of data tables according to an input operation performed by a user on a screen, wherein the input operation includes: an input operation for specifying an association relationship between two data tables.
22. A system for automatically training a machine learning model, comprising: the system for processing data tables according to any one of claims 12 to 21; a training device, adapted to respectively train machine learning models using different machine learning algorithms and different hyperparameters based on a sample table including a plurality of machine learning samples obtained by the system for processing data tables; a determination device, adapted to determine the machine learning model with the best effect from the trained machine learning models as the finally trained machine learning model.
23. A system including at least one computing device and at least one storage device storing instructions, wherein, when the instructions are run by the at least one computing device, the at least one computing device is caused to execute the method for processing data tables according to any one of claims 1 to 10 or the method for automatically training a machine learning model according to claim 11.
24. A computer-readable storage medium storing instructions, wherein, when the instructions are run by at least one computing device, the at least one computing device is caused to execute the method for processing data tables according to any one of claims 1 to 10 or the method for automatically training a machine learning model according to claim 11.
Citation Information
Patent Citations
A method and a system for realizing data table splicing and automatic training of a machine learning model
CN109739855A