Method and System for Generating Combined Features of Machine Learning Samples

By iteratively generating and filtering the combined features of machine learning samples, the problem of automated combined features is solved, the model prediction effect is improved and the computing efficiency is optimized.

CN114298323BActive Publication Date: 2025-07-08THE FOURTH PARADIGM BEIJING TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111615354.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2017-09-08
Publication Date
2025-07-08
Estimated Expiration
2037-09-08

AI Technical Summary

Technical Problem

The prior art is difficult to automate the characteristics of combinatorial machine learning samples, resulting in poor prediction results and inefficient computing.

Method used

By obtaining historical data records, iteratively generate candidate combination features, and using pre-sorting and resorting methods to filter out high-important features as combination features of machine learning samples, including using binning operations, machine learning model evaluation and verification processes.

Benefits of technology

With the reduction of computing resources, feature combination is effectively realized and the prediction effect of machine learning models is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114298323B_ABST
    Figure CN114298323B_ABST
Patent Text Reader

Abstract

A method and system for generating combined features of machine learning samples are provided. The method includes: obtaining historical data records, where the historical data records include multiple attribute information; iteratively performing feature combination among at least one discrete feature according to a search strategy to generate candidate combined features, and selecting target combined features as the combined features of machine learning samples. For each round of iteration, pre-rank the importance of each candidate combined feature in the candidate combined feature set; screen out a part of the candidate combined features according to the pre-ranking result to form a candidate combined feature pool; re-rank the importance of each candidate combined feature in the candidate combined feature pool; select at least one candidate combined feature with higher importance as the target combined feature according to the re-ranking result. According to the method and system, automatic feature combination can be effectively realized with less computing resources, and the effect of the model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application with the application date of September 8, 2017, application number 201710803886.7, and title "Method and System for Generating Combined Features of Machine Learning Samples". Technical Field

[0002] The present invention generally relates to the field of artificial intelligence, and more specifically, to a method and system for generating combined features of machine learning samples. Background Art

[0003] With the emergence of massive data, artificial intelligence technology has developed rapidly. In order to extract value from a large amount of data, it is necessary to generate samples suitable for machine learning based on data records.

[0004] Here, each data record can be regarded as a description of an event or object, corresponding to an example or sample. In a data record, there are various items reflecting the performance or nature of the event or object in a certain aspect, and these items can be called "attributes".

[0005] How to convert the attributes of the original data record into the features of the machine learning sample will have a great impact on the effect of the machine learning model. In fact, the prediction effect of the machine learning model is related to the model selection, available data, and feature extraction, etc. That is to say, the prediction effect of the model can be improved by improving the feature extraction method. On the contrary, if the feature extraction is inappropriate, it will lead to the deterioration of the prediction effect.

[0006] However, in the process of determining the feature extraction method, it often requires technicians to not only master the knowledge of machine learning but also have an in-depth understanding of the actual prediction problem. And the prediction problem is often combined with different practical experiences in different industries, resulting in it being difficult to achieve a satisfactory effect. In particular, when combining different features, on the one hand, it is difficult to grasp which features to combine from the perspective of the prediction effect, and on the other hand, considering the operation efficiency, it is also difficult to effectively screen out a specific combination method. To sum up, it is difficult to automatically combine features in the prior art. Summary of the Invention

[0007] The exemplary embodiments of the present invention aim to overcome the defect in the prior art that it is difficult to automatically combine the features of machine learning samples.

[0008] According to an exemplary embodiment of the present invention, a method for generating combined features of machine learning samples is provided, including: (A) obtaining historical data records, where the historical data records include multiple attribute information; and (B) iteratively performing feature combination among at least one discrete feature generated based on the multiple attribute information according to a search strategy to generate candidate combined features, and selecting target combined features from the generated candidate combined features as the combined features of machine learning samples, where for each round of iteration, pre-rank the importance of each candidate combined feature in the candidate combined feature set; screen out a part of the candidate combined features from the candidate combined feature set according to the pre-rank result to form a candidate combined feature pool; re-rank the importance of each candidate combined feature in the candidate combined feature pool; and select at least one candidate combined feature with higher importance from the candidate combined feature pool as the target combined feature according to the re-rank result.

[0009] Optionally, in the method, the pre-rank is performed based on a first quantity of historical data records, the re-rank is performed based on a second quantity of historical data records, and the second quantity is not less than the first quantity.

[0010] Optionally, in the method, screen out candidate combined features with higher importance from the candidate combined feature set according to the pre-rank result to form a candidate combined feature pool.

[0011] Optionally, in the method, the candidate combined feature set includes the candidate combined features generated in the current round of iteration; or the candidate combined feature set includes the candidate combined features generated in the current round of iteration and the candidate combined features generated in the previous round of iteration that have not been selected as target combined features.

[0012] Optionally, in the method, generate the candidate combined features for the next round of iteration by combining the target combined feature selected in the current round of iteration with the at least one discrete feature; or generate the candidate combined features for the next round of iteration by pairwise combining the target combined features selected in the current round of iteration and the previous round of iteration.

[0013] Optionally, in the method, the at least one discrete feature includes discrete features converted from continuous features generated based on the multiple attribute information through the following processing: for each continuous feature, perform at least one binning operation to generate discrete features composed of at least one binned feature, where each binning operation corresponds to one binned feature.

[0014] Optionally, in the method, the at least one binning operation is selected from a predetermined number of binning operations for each iteration or for all iterations, wherein the importance of the binning features corresponding to the selected binning operation is not lower than that of the binning features corresponding to the unselected binning operation.

[0015] Optionally, in the method, the at least one binning operation is selected by the following process: for each binning feature among the binning features corresponding to the predetermined number of binning operations, a binning single-feature machine learning model is obtained, the importance of each binning feature is determined based on the effects of the respective binning single-feature machine learning models, and the at least one binning operation is selected based on the importance of each binning feature, wherein the binning single-feature machine learning model corresponds to each binning feature.

[0016] Optionally, in the method, the at least one binning operation is selected by the following process: for each binning feature among the binning features corresponding to the predetermined number of binning operations, a binning overall machine learning model is obtained, the importance of each binning feature is determined based on the effects of the respective binning overall machine learning models, and the at least one binning operation is selected based on the importance of each binning feature, wherein the binning overall machine learning model corresponds to the binning basic feature subset and each binning feature.

[0017] Optionally, in the method, the at least one binning operation is selected by the following process: for each binning feature among the binning features corresponding to the predetermined number of binning operations, a binning composite machine learning model is obtained, the importance of each binning feature is determined based on the effects of the respective binning composite machine learning models, and the at least one binning operation is selected based on the importance of each binning feature, wherein the binning composite machine learning model includes a binning basic sub-model and a binning additional sub-model based on a boosting framework, wherein the binning basic sub-model corresponds to the binning basic feature subset, and the binning additional sub-model corresponds to each binning feature.

[0018] Optionally, in the method, the binning basic feature subset includes the target combination features selected before the current iteration.

[0019] Optionally, in the method, the pre-ranking is performed by the following process: for each candidate combination feature in the candidate combination feature set, a pre-ranking single-feature machine learning model is obtained, and the importance of each candidate combination feature is determined based on the effects of the respective pre-ranking single-feature machine learning models, wherein the pre-ranking single-feature machine learning model corresponds to each candidate combination feature.

[0020] Optionally, in the method, pre-ranking is performed through the following process: for each candidate combined feature in the candidate combined feature set, a pre-ranking overall machine learning model is obtained, and the importance of each candidate combined feature is determined based on the effects of the respective pre-ranking overall machine learning models, where the pre-ranking overall machine learning model corresponds to a pre-ranking basic feature subset and each of the candidate combined features.

[0021] Optionally, in the method, pre-ranking is performed through the following process: for each candidate combined feature in the candidate combined feature set, a pre-ranking composite machine learning model is obtained, and the importance of each candidate combined feature is determined based on the effects of the respective pre-ranking composite machine learning models, where the pre-ranking composite machine learning model includes a pre-ranking basic sub-model and a pre-ranking additional sub-model based on a boosting framework, where the pre-ranking basic sub-model corresponds to a pre-ranking basic feature subset, and the pre-ranking additional sub-model corresponds to each of the candidate combined features.

[0022] Optionally, in the method, the pre-ranking basic feature subset includes target combined features selected before the current round of iteration.

[0023] Optionally, in the method, re-ranking is performed through the following process: for each candidate combined feature in the candidate combined feature pool, a re-ranking single-feature machine learning model is obtained, and the importance of each candidate combined feature is determined based on the effects of the respective re-ranking single-feature machine learning models, where the re-ranking single-feature machine learning model corresponds to each of the candidate combined features.

[0024] Optionally, in the method, re-ranking is performed through the following process: for each candidate combined feature in the candidate combined feature pool, a re-ranking overall machine learning model is obtained, and the importance of each candidate combined feature is determined based on the effects of the respective re-ranking overall machine learning models, where the re-ranking composite machine learning model corresponds to a re-ranking basic feature subset and each of the candidate combined features.

[0025] Optionally, in the method, re-ranking is performed through the following process: for each candidate combined feature in the candidate combined feature pool, a re-ranking composite machine learning model is obtained, and the importance of each candidate combined feature is determined based on the effects of the respective re-ranking composite machine learning models, where the re-ranking composite machine learning model includes a re-ranking basic sub-model and a re-ranking additional sub-model based on a boosting framework, where the re-ranking basic sub-model corresponds to a re-ranking basic feature subset, and the re-ranking additional sub-model corresponds to each of the candidate combined features.

[0026] Optionally, in the method, the re-ranking basic feature subset includes target combined features selected before the current round of iteration.

[0027] Optionally, in the method, step (B) further includes: for each round of iteration, checking whether the selected target combined feature is suitable as the combined feature of the machine learning sample.

[0028] Optionally, in the method, in step (B), the effect change of the machine learning model based on the target combined feature that has passed the check after introducing the selected target combined feature is used to check whether the selected target combined feature is suitable as the combined feature of the machine learning sample.

[0029] Optionally, in the method, when the check result is that the selected target combined feature is suitable as the combined feature of the machine learning sample, the selected target combined feature is used as the combined feature of the machine learning sample, and the next round of iteration is executed; when the check result is that the selected target combined feature is not suitable as the combined feature of the machine learning sample, another part of candidate combined features is screened out from the candidate combined feature set according to the pre-sorting result to form a new candidate combined feature pool.

[0030] According to another exemplary embodiment of the present invention, a computer-readable medium for generating combined features of machine learning samples is provided, wherein a computer program for executing the method described above is recorded on the computer-readable medium.

[0031] According to another exemplary embodiment of the present invention, a computing device for generating combined features of machine learning samples is provided, including a storage component and a processor, wherein a set of computer-executable instructions is stored in the storage component, and when the set of computer-executable instructions is executed by the processor, the method described above is executed.

[0032] According to another exemplary implementation of the present invention, a system for generating combined features of machine learning samples is provided, including: a data record acquisition device for acquiring historical data records, where the historical data records include multiple attribute information; and a feature combination device for iteratively performing feature combination among at least one discrete feature generated based on the multiple attribute information according to a search strategy to generate candidate combined features, and selecting a target combined feature from the generated candidate combined features as the combined feature of the machine learning sample. For each round of iteration, the feature combination device performs pre-sorting of the importance of each candidate combined feature in the candidate combined feature set, screens out a part of candidate combined features from the candidate combined feature set according to the pre-sorting result to form a candidate combined feature pool, performs re-sorting of the importance of each candidate combined feature in the candidate combined feature pool, and selects at least one candidate combined feature with higher importance from the candidate combined feature pool as the target combined feature.

[0033] Optionally, in the system, the feature combination device performs pre-sorting based on a first quantity of historical data records and performs re-sorting based on a second quantity of historical data records, and the second quantity is not less than the first quantity.

[0034] Optionally, in the system, the feature combination device screens out candidate combined features with relatively high importance from the candidate combined feature set according to the pre-sorting result to form a candidate combined feature pool.

[0035] Optionally, in the system, the candidate combined feature set includes candidate combined features generated in the current round of iteration; or, the candidate combined feature set includes candidate combined features generated in the current round of iteration and candidate combined features generated in previous rounds of iteration that have not been selected as target combined features.

[0036] Optionally, in the system, the feature combination device generates candidate combined features for the next round of iteration by combining the target combined features selected in the current round of iteration with the at least one discrete feature; or, the feature combination device generates candidate combined features for the next round of iteration by pairwise combining the target combined features selected in the current round of iteration and the previous round of iteration.

[0037] Optionally, in the system, the at least one discrete feature includes discrete features converted from continuous features generated based on the multiple attribute information through the following processing: for each continuous feature, at least one binning operation is performed to generate discrete features composed of at least one binned feature, where each binning operation corresponds to one binned feature.

[0038] Optionally, in the system, the at least one binning operation is selected from a predetermined number of binning operations for each round of iteration or for all rounds of iteration, where the importance of the binned features corresponding to the selected binning operation is not lower than the importance of the binned features corresponding to the unselected binning operation.

[0039] Optionally, in the system, the feature combination device selects the at least one binning operation through the following processing: for each of the binned features corresponding to the predetermined number of binning operations, a binned single-feature machine learning model is obtained, the importance of each binned feature is determined based on the effects of the respective binned single-feature machine learning models, and the at least one binning operation is selected based on the importance of each binned feature, where the binned single-feature machine learning model corresponds to each of the binned features.

[0040] Optionally, in the system, the feature combination device selects the at least one binning operation through the following process: for each binning feature among the binning features corresponding to the predetermined number of binning operations, obtain a binning overall machine learning model, determine the importance of each binning feature based on the effects of the respective binning overall machine learning models, and select the at least one binning operation based on the importance of each binning feature, where the binning overall machine learning model corresponds to a binning basic feature subset and each of the binning features.

[0041] Optionally, in the system, the feature combination device selects the at least one binning operation through the following process: for each binning feature among the binning features corresponding to the predetermined number of binning operations, obtain a binning composite machine learning model, determine the importance of each binning feature based on the effects of the respective binning composite machine learning models, and select the at least one binning operation based on the importance of each binning feature, where the binning composite machine learning model includes a binning basic sub-model and a binning additional sub-model based on a boosting framework, where the binning basic sub-model corresponds to a binning basic feature subset, and the binning additional sub-model corresponds to each of the binning features.

[0042] Optionally, in the system, the binning basic feature subset includes target combination features selected before the current round of iteration.

[0043] Optionally, in the system, the feature combination device performs pre-ranking through the following process: for each candidate combination feature in the candidate combination feature set, obtain a pre-ranking single feature machine learning model, and determine the importance of each candidate combination feature based on the effects of the respective pre-ranking single feature machine learning models, where the pre-ranking single feature machine learning model corresponds to each of the candidate combination features.

[0044] Optionally, in the system, the feature combination device performs pre-ranking through the following process: for each candidate combination feature in the candidate combination feature set, obtain a pre-ranking overall machine learning model, and determine the importance of each candidate combination feature based on the effects of the respective pre-ranking overall machine learning models, where the pre-ranking overall machine learning model corresponds to a pre-ranking basic feature subset and each of the candidate combination features.

[0045] Optionally, in the system, the feature combination device performs pre-sorting through the following process: for each candidate combined feature in the candidate combined feature set, a pre-sorting composite machine learning model is obtained, and the importance of each candidate combined feature is determined based on the effects of the respective pre-sorting composite machine learning models. Among them, the pre-sorting composite machine learning model includes a pre-sorting basic sub-model and a pre-sorting additional sub-model based on the boosting framework. Among them, the pre-sorting basic sub-model corresponds to the pre-sorting basic feature subset, and the pre-sorting additional sub-model corresponds to each candidate combined feature.

[0046] Optionally, in the system, the pre-sorting basic feature subset includes the target combined features selected before the current round of iteration.

[0047] Optionally, in the system, the feature combination device performs re-sorting through the following process: for each candidate combined feature in the candidate combined feature pool, a re-sorting single-feature machine learning model is obtained, and the importance of each candidate combined feature is determined based on the effects of the respective re-sorting single-feature machine learning models. Among them, the re-sorting single-feature machine learning model corresponds to each candidate combined feature.

[0048] Optionally, in the system, the feature combination device performs re-sorting through the following process: for each candidate combined feature in the candidate combined feature pool, a re-sorting overall machine learning model is obtained, and the importance of each candidate combined feature is determined based on the effects of the respective re-sorting overall machine learning models. Among them, the re-sorting composite machine learning model corresponds to the re-sorting basic feature subset and each candidate combined feature.

[0049] Optionally, in the system, the feature combination device performs re-sorting through the following process: for each candidate combined feature in the candidate combined feature pool, a re-sorting composite machine learning model is obtained, and the importance of each candidate combined feature is determined based on the effects of the respective re-sorting composite machine learning models. Among them, the re-sorting composite machine learning model includes a re-sorting basic sub-model and a re-sorting additional sub-model based on the boosting framework. Among them, the re-sorting basic sub-model corresponds to the re-sorting basic feature subset, and the re-sorting additional sub-model corresponds to each candidate combined feature.

[0050] Optionally, in the system, the re-sorting basic feature subset includes the target combined features selected before the current round of iteration.

[0051] Optionally, in the system, the feature combination device also checks for each round of iteration whether the selected target combined features are suitable as combined features of machine learning samples.

[0052] Optionally, in the system, the feature combination device uses the change in the effect after introducing the selected target combined feature of a machine learning model based on the target combined feature that has passed the inspection to inspect whether the selected target combined feature is suitable as the combined feature of the machine learning sample.

[0053] Optionally, in the system, when the inspection result is that the selected target combined feature is suitable as the combined feature of the machine learning sample, the feature combination device uses the selected target combined feature as the combined feature of the machine learning sample and performs the next round of iteration; when the inspection result is that the selected target combined feature is not suitable as the combined feature of the machine learning sample, the feature combination device screens out another part of the candidate combined features from the candidate combined feature set according to the pre-sorting result to form a new candidate combined feature pool.

[0054] In the method and system for generating the combined features of machine learning samples according to the exemplary embodiments of the present invention, a part is screened out from the combined features generated in each round of iteration through pre-sorting and re-sorting in a specific manner to finally form a combined feature set of the machine learning sample, so that automatic feature combination can be effectively achieved with less computing resources, and the effect of the machine learning model can be improved. Description of the Drawings

[0055] These and / or other aspects and advantages of the present invention will become clearer and easier to understand from the following detailed description of the embodiments of the present invention in conjunction with the drawings, in which:

[0056] Figure 1 A block diagram showing a system for generating combined features of machine learning samples according to an exemplary embodiment of the present invention;

[0057] Figure 2 A block diagram showing a feature combination device according to an exemplary embodiment of the present invention;

[0058] Figure 3 A block diagram showing a feature combination device according to another exemplary embodiment of the present invention;

[0059] Figure 4 A block diagram showing a training system of a machine learning model according to an exemplary embodiment of the present invention;

[0060] Figure 5 A flowchart showing a method for generating combined features of machine learning samples according to an exemplary embodiment of the present invention;

[0061] Figure 6 An example showing a search tree for iteratively generating combined features according to an exemplary embodiment of the present invention;

[0062] Figure 7A flowchart showing a method for training a machine learning model according to an exemplary embodiment of the present invention; and

[0063] Figure 8 A flowchart showing a method for generating combined features of machine learning samples according to another exemplary embodiment of the present invention. Detailed Description of the Invention

[0064] To enable those skilled in the art to better understand the present invention, the following further describes in detail the exemplary embodiments of the present invention with reference to the accompanying drawings and specific embodiments.

[0065] In the exemplary embodiments of the present invention, automatic feature combination is performed in the following manner: discrete features that can be combined are generated based on the attribute information of data records, and discrete feature combinations serving as candidate combined features are generated in an iterative manner. In each round of iteration, a part of the target combined features are screened out from the candidate combined features through pre-sorting and re-sorting in a specific manner to form a combined feature set of machine learning samples.

[0066] Here, machine learning is an inevitable product of the development of artificial intelligence research to a certain stage, which is committed to improving the performance of the system itself by means of computing and using experience. In a computer system, "experience" usually exists in the form of "data". Through machine learning algorithms, a "model" can be generated from the data. That is to say, by providing empirical data to machine learning algorithms, a model can be generated based on these empirical data. When facing new situations, the model will provide corresponding judgments, that is, prediction results. Whether training a machine learning model or using a trained machine learning model for prediction, the data needs to be converted into machine learning samples including various features. Machine learning can be implemented in the form of "supervised learning", "unsupervised learning" or "semi-supervised learning". It should be noted that the exemplary embodiments of the present invention do not specifically limit specific machine learning algorithms. In addition, it should also be noted that other means such as statistical algorithms can be combined during the process of training and applying the model.

[0067] Figure 1 A block diagram showing a system for generating combined features of machine learning samples according to an exemplary embodiment of the present invention. Figure 1 The shown system includes a data record acquisition device 100 and a feature combination device 200.

[0068] Specifically, the data record acquisition device 100 is used to acquire historical data records, where the historical data records include multiple attribute information. Here, as an example, the data record acquisition device 100 can acquire historical data records that have been labeled for supervised machine learning.

[0069] The above historical data records can be data generated online, pre-generated and stored data, or data received from the outside through an input device or a transmission medium. These data can relate to the attribute information of individuals, enterprises or organizations, such as identity, education background, occupation, assets, contact information, liabilities, income, profit, tax, etc. Alternatively, these data can also relate to the attribute information of business-related items, such as information about the transaction amount, the two parties to the transaction, the subject matter, the transaction location, etc. of a sales contract. It should be noted that the content of the attribute information mentioned in the exemplary embodiments of the present invention can relate to the performance or nature of any object or matter in a certain aspect, and is not limited to the limitation or description of individuals, objects, organizations, units, institutions, projects, events, etc.

[0070] The data record acquisition device 100 can acquire structured or unstructured data from different sources, such as text data or numerical data, etc. The acquired data records can be used to form machine learning samples and participate in the training / test process of machine learning models. These data can come from inside the entity that expects to obtain the model prediction result, such as a bank, an enterprise, a school, etc. that expects to obtain the prediction result; these data can also come from outside the above entity, such as a data provider, the Internet (such as a social networking site), a mobile operator, an APP operator, a courier company, a credit institution, etc. Optionally, the above internal data and external data can be combined for use to form machine learning samples carrying more information.

[0071] The above data can be input into the data record acquisition device 100 through an input device, or automatically generated by the data record acquisition device 100 according to the existing data, or can be obtained by the data record acquisition device 100 from the network (such as a storage medium on the network (such as a data warehouse)). In addition, an intermediate data exchange device such as a server can help the data record acquisition device 100 obtain corresponding data from an external data source. Here, the acquired data can be converted into an easily processable format by a data conversion module such as a text analysis module in the data record acquisition device 100.

[0072] The feature combination device 200 is used to iteratively perform feature combination among at least one discrete feature generated based on the multiple attribute information according to a search strategy to generate candidate combined features, and select a target combined feature from the generated candidate combined features as the combined feature of the machine learning sample. Specifically, for each round of iteration, the feature combination device 200 pre-orders the importance of each candidate combined feature in the candidate combined feature set, screens out a part of the candidate combined features from the candidate combined feature set according to the pre-order result to form a candidate combined feature pool, re-orders the importance of each candidate combined feature in the candidate combined feature pool, and selects at least one candidate combined feature with higher importance from the candidate combined feature pool as the target combined feature according to the re-order result.

[0073] Here, the feature combination device 200 can first generate discrete features that can be combined based on the multiple attribute information recorded in the historical data (the discrete feature can be regarded as the smallest unit capable of feature combination). In this process, the feature combination device 200 can discretize continuous features as needed to obtain discrete features that are convenient for mutual combination. Here, a single discrete feature can be regarded as a first-order feature. According to the exemplary embodiments of the present invention, higher-order feature combinations such as second-order and third-order can be performed to generate corresponding candidate combined features, where "order" represents the number of single discrete features participating in the combination.

[0074] As an example, according to the search strategy for combined features, the combined features of the machine learning sample can be generated in an iterative manner. For this purpose, the feature combination device 200 can gradually generate candidate combined features in an iterative manner. As an example, in each round of iteration, the feature combination device 200 can combine new candidate combined features according to the preset search strategy. Here, the feature combination device 200 can pre-order the importance of each candidate combined feature currently participating in the screening, and screen out a part of the candidate combined features according to the pre-order result. For example, a part of the candidate combined features with higher importance (for example, among 100 candidate combined features, the first to tenth most important features can be screened out), a part of the candidate combined features with spaced importance arrangements (for example, among 100 candidate combined features, the first most important feature, the eleventh most important feature, the twenty-first most important feature..., the ninety-first most important feature), etc. The screened candidate combined features can form a candidate combined feature pool, so that the feature combination device 200 can further screen out the target combined features with higher importance through re-ordering. After determining the target combined feature as described above, the next round of iteration can be executed to further obtain new target combined features.

[0075] It should be noted that the recording acquisition device 100 and the feature combination device 200 can be configured as individual units composed of software, hardware, and / or firmware, and some or all of these units can be integrated or cooperate together to complete specific functions.

[0076] As an example, the following will refer to Figure 2 to describe the block diagram of the feature combination device 200 according to an exemplary embodiment of the present invention. Referring to Figure 2 , the feature combination device 200 may include a candidate combination feature generation unit 210, a pre-sorting unit 220, and a re-sorting unit 230.

[0077] Here, the candidate combination feature generation unit 210 is used to generate candidate combination features for each iteration according to a search strategy.

[0078] The candidate combination feature generation unit 210 may first generate discrete features that can be combined based on the attribute information recorded in the historical data in the first iteration. Here, the corresponding discrete features can be obtained by discretizing continuous features (for example, the continuous value attribute information itself). Preferably, when discretizing continuous features, the candidate combination feature generation unit 210 may perform at least one binning operation for each continuous feature to generate discrete features composed of at least one binned feature, where each binning operation corresponds to one binned feature.

[0079] Specifically, for at least a part of the attribute information recorded in the historical data, corresponding continuous features can be generated. Here, continuous features are a type of feature opposite to discrete features (for example, categorical features), and their values can be numerical values with a certain continuity, such as distance, age, amount, etc. In contrast, as an example, the values of discrete features do not have continuity. For example, they can be features of unordered classification such as "from Beijing", "from Shanghai", or "from Tianjin", "gender is male", "gender is female", etc.

[0080] For example, certain continuous value attribute information in the historical data record can be directly used as the corresponding continuous feature. For example, attribute information such as distance, age, amount, etc. can be directly used as the corresponding continuous features. That is to say, each of the continuous features can be formed by the continuous value attribute information itself among the multiple attribute information. Alternatively, certain attribute information (e.g., continuous value attribute and / or discrete value attribute information) in the historical data record can also be processed to obtain the corresponding continuous feature. For example, the ratio of height to weight can be used as the corresponding continuous feature. In particular, the continuous feature can be formed by performing a continuous transformation on the discrete value attribute information among the multiple attribute information. As an example, the continuous transformation can indicate performing statistics on the values of the discrete value attribute information. For example, the continuous feature can indicate the statistical information of certain discrete value attribute information with respect to the prediction target of the machine learning model. For example, in the example of predicting the purchase probability, the discrete value attribute information of the seller merchant number can be transformed into a probability statistical feature regarding the historical purchase behavior of the corresponding seller merchant code.

[0081] In addition, in addition to the continuous features on which binning operations are performed, the candidate combined feature generation unit 210 can also generate other discrete features. As an alternative, the above features can also be generated by other feature generation devices (not shown). According to an exemplary embodiment of the present invention, any combination can be made among the above features, wherein the continuous features have been converted into binned group features when combined.

[0082] For each continuous feature, the candidate combined feature generation unit 210 can perform at least one binning operation, so as to be able to simultaneously obtain multiple discrete features that characterize certain attributes of the original data record from different perspectives, scales / aspects.

[0083] Here, the binning operation refers to a specific way of discretizing continuous features, that is, dividing the value range of the continuous feature into multiple intervals (i.e., multiple bins), and determining the corresponding binned feature values based on the divided bins. The binning operation can generally be divided into supervised binning and unsupervised binning, and each of these two types includes some specific binning methods. For example, supervised binning includes minimum entropy binning, minimum description length binning, etc., while unsupervised binning includes equal-width binning, equal-depth binning, binning based on k-means clustering, etc. For each binning method, corresponding binning parameters can be set, such as width, depth, etc. It should be noted that according to an exemplary embodiment of the present invention, the binning operation performed by the candidate combined feature generation unit 210 does not limit the type of binning method, nor the parameters of the binning operation, and moreover, the specific representation method of the corresponding generated binned features is also not limited.

[0084] The binning operations performed by the candidate combined feature generation unit 210 may vary in terms of binning methods and / or binning parameters. For example, the at least one binning operation may be binning operations of the same type but with different operation parameters (such as depth, width, etc.), or may be binning operations of different types. Correspondingly, each binning operation can obtain a binning feature, and these binning features together form a binning group feature, which can reflect different binning operations, thereby improving the effectiveness of machine learning materials and providing a better basis for the training / prediction of machine learning models.

[0085] The above shows the process of discretizing continuous features in the first round of iteration. However, it should be understood that according to the exemplary embodiments of the present invention, continuous features can be discretized only for the first round of iteration to obtain discrete features that will always be used for combination subsequently, or discretization can be re-executed for subsequent iterations (for example, for each round of iteration) to obtain discrete features corresponding to the relevant subsequent iterations respectively.

[0086] As an example, the at least one binning operation can be selected from a predetermined number of binning operations for each round of iteration or for all rounds of iteration, where the importance of the binning features corresponding to the selected binning operations is not lower than the importance of the binning features corresponding to the unselected binning operations. Here, the candidate combined feature generation unit 210 can use any means of judging feature importance to measure the importance of each binning feature.

[0087] In subsequent iterations, the candidate combined feature generation unit 210 can generate new candidate combined features according to a search strategy. Here, the search strategy can be aimed at pruning the search tree for combined discrete features to control the number of candidate combined features generated in each round of iteration. For example, the candidate combined feature generation unit 210 can generate new candidate combined features only based on the target combined features selected in the previous round of iteration in each round of iteration.

[0088] The pre-sorting unit 220 is used to perform pre-sorting of the importance of each candidate combined feature in the candidate combined feature set for each iteration. Here, the candidate combined feature set may include candidate combined features generated in one or more iterations. For example, the candidate combined features generated in the current iteration or those generated in the current iteration together with several previous iterations. The pre-sorting unit 220 can use any means of judging the importance of features to measure the importance of each candidate combined feature in the candidate combined feature set. Through pre-sorting, the importance order of the various candidate combined features can be obtained. On this basis, the pre-sorting unit 220 can screen out a part of the candidate combined features from them to form a candidate combined feature pool. Here, the screened candidate combined features can show a certain consistency in terms of the prediction effect, so that only the features with higher importance can be screened out as the target combined features of the machine learning sample.

[0089] Correspondingly, the re-sorting unit 230 is used to perform re-sorting of the importance of each candidate combined feature in the candidate combined feature pool, and select at least one candidate combined feature with higher importance from the candidate combined feature pool as the target combined feature according to the re-sorting result. Here, the re-sorting unit 230 can use any means of judging the importance of features to measure the importance of each candidate combined feature in the candidate combined feature pool. For example, the re-sorting unit 230 can measure the importance of the candidate combined features in the same way as the pre-sorting unit 220, but make a more accurate judgment based on a larger number of data records when judging. The re-sorting unit 230 can select a predetermined number of the most important candidate combined features in the candidate combined feature pool as the target combined features. Here, the target combined features can be directly used as the combined features of the machine learning sample, or the target combined features can be further verified to determine whether to use them as the combined features of the machine learning sample. As an example, if no suitable target combined features are screened out in the current candidate combined feature pool or it is necessary to continue screening for suitable target combined features in the current iteration, a new candidate combined feature pool can be re-determined according to the pre-sorting result. For example, among 100 candidate combined features, the 11th most important to the 20th most important features can be screened out again; or, among 100 candidate combined features, the 2nd most important feature, the 12th most important feature, the 22nd most important feature... the 92nd most important feature can be screened out again. On the other hand, at the end of the screening process in the current iteration, the next iteration can be executed to generate new candidate combined features.

[0090] In Figure 2 In the feature combination device 200 shown, the candidate combined feature generation unit 210, the pre-sorting unit 220, and the re-sorting unit 230 are all involved in determining the importance of features. Correspondingly, as an optional method, the above three units can share some operation parameters or results to save resources.

[0091] Figure 3 A block diagram showing a feature combination device 200 according to another exemplary embodiment of the present invention. In Figure 3 the shown feature combination device 200, in addition to a candidate combination feature generation unit 210, a pre-ranking unit 220, and a re-ranking unit 230, a verification unit 240 is further included, which is used to verify, for each iteration, whether the target combination feature selected by the re-ranking unit 230 is suitable as the combination feature of the machine learning sample.

[0092] Here, the candidate combination feature generation unit 210, the pre-ranking unit 220, and the re-ranking unit 230 can operate in the manner described with reference to Figure 2 and details will not be elaborated here. In addition, the target combination feature selected by the re-ranking unit 230 each time will not be directly used as the combination feature of the machine learning sample, but needs to be verified by the verification unit 240. As an example, the verification unit 240 can verify whether it is suitable as the combination feature of the machine learning sample by integrating the selected target combination feature into the actual machine learning model that will perform prediction for the prediction problem. For example, the verification unit 240 can introduce the target combination feature to be verified into the machine learning model based on the already verified combination features, and verify whether the target combination feature to be verified is suitable as the combination feature of the machine learning sample by measuring the change in the model's effect.

[0093] Figure 1 (Combined with Figure 2 and Figure 3 ) The shown system is designed to generate the combination features of the machine learning sample, and this system can exist independently. Here, it should be noted that the way the system obtains the data record is not restricted. That is, as an example, the data record acquisition device 100 can be a device with the ability to receive and process data records, or can simply be a device that provides the already prepared data records. In addition, the above system can also be integrated into the model training system as a component for completing feature processing.

[0094] Figure 4 A block diagram showing a training system of a machine learning model according to an exemplary embodiment of the present invention. In Figure 4 the shown system, in addition to the above data record acquisition device 100 and feature combination device 200, a machine learning sample generation device 300 and a machine learning model training device 400 are further included.

[0095] Specifically, in Figure 4 the shown system, the data record acquisition device 100 and the feature combination device 200 can operate in the manner described in Figures 1 to 3operate in the manner shown, where, by way of example, the data record acquisition device 100 can acquire historical data records that have been labeled.

[0096] In addition, the machine learning sample generation device 300 is used to generate machine learning samples that at least include a part of the generated combined features. That is to say, in the machine learning samples generated by the machine learning sample generation device 300, it includes a part or all of the combined features generated by the feature combination device 200. In addition, as an alternative, the machine learning samples can also include any other features generated based on the attribute information of the data records, for example, features obtained by performing feature processing on the attribute information of the data records, etc. By way of example, these other features can be generated by the feature combination device 200 or by other devices.

[0097] Specifically, the machine learning sample generation device 300 can generate machine learning training samples. In particular, by way of example, in the case of supervised learning, the machine learning training samples generated by the machine learning sample generation device 300 can include two parts: features and labels.

[0098] The machine learning model training device 400 is used to train a machine learning model based on the machine learning training samples. Here, the machine learning model training device 400 can adopt any appropriate machine learning algorithm (for example, logistic regression) to learn an appropriate machine learning model from the machine learning training samples. By way of example, the machine learning model training device 400 can adopt the same or similar machine learning algorithm as the model used by the combination feature generation device 200 to measure the importance of relevant features.

[0099] In the above example, a relatively stable and better-performing machine learning model can be trained.

[0100] The following Figure 5 is combined with Figure 5 to describe the flowchart of the method for generating combined features of machine learning samples according to an exemplary embodiment of the present invention. Here, by way of example, Figure 1 the method shown can be executed by Figure 5 the system and its devices shown, can also be fully implemented in software by a computer program, or can be executed by a specifically configured computing device Figure 5 the method shown. For ease of description, it is assumed that Figure 1 the method shown is executed by Figure 1 the system shown, and it is assumed that Figure 2 the feature combination device 200 in

[0101] As shown in the figure, in step S100, the historical data record is obtained by the data record acquisition device 100, where the historical data record includes a plurality of attribute information.

[0102] Here, as an example, the data record acquisition device 100 can collect data manually, semi-automatically or fully automatically, or process the collected raw data so that the processed data record has an appropriate format or form. As an example, the data record acquisition device 100 can collect historical data in batches.

[0103] Here, the data record acquisition device 100 can receive the data record manually input by the user through an input device (for example, a workstation). In addition, the data record acquisition device 100 can systematically retrieve the data record from the data source in a fully automatic manner. For example, through a timer mechanism implemented by software, firmware, hardware or a combination thereof, the data source is systematically requested and the requested data is obtained from the response. The data source can include one or more databases or other servers. The fully automatic data acquisition method can be implemented via an internal network and / or an external network, where encrypted data can be transmitted through the Internet. When the server, database, network, etc. are configured to communicate with each other, data collection can be automatically performed without manual intervention. However, it should be noted that there can still be certain user input operations in this way. The semi-automatic method is between the manual method and the fully automatic method. The difference between the semi-automatic method and the fully automatic method is that the trigger mechanism activated by the user replaces, for example, the timer mechanism. In this case, a request to extract data is generated only when a specific user input is received. Preferably, each time data is acquired, the captured data can be stored in a non-volatile memory. As an example, a data warehouse can be used to store the raw data and the processed data collected during the acquisition.

[0104] The above-obtained data records can be from the same or different data sources, that is to say, each data record can also be the splicing result of different data records. For example, in addition to obtaining the information data record filled in by the customer when applying for a credit card from the bank (which includes attribute information fields such as income, education level, job, and asset status), as an example, the data record acquisition device 100 can also obtain other data records of the customer in the bank, such as loan records, daily transaction data, etc. These obtained data records can be spliced into a complete data record. In addition, the data record acquisition device 100 can also obtain data from other private sources or public sources, such as data from data providers, data from the Internet (for example, social websites), data from mobile operators, data from APP operators, data from express delivery companies, data from credit institutions, and so on.

[0105] Optionally, the data record acquisition device 100 may store and / or process the collected data by means of a hardware cluster (such as a Hadoop cluster, a Spark cluster, etc.), for example, storage, classification, and other offline operations. In addition, the data record acquisition device 100 may also perform online stream processing on the collected data.

[0106] As an example, the data record acquisition device 100 may include a data conversion module such as a text analysis module. Correspondingly, in step S100, the data record acquisition device 100 may convert unstructured data such as text into more easily used structured data for further processing or reference in the subsequent process. The text-based data may include emails, documents, web pages, graphics, spreadsheets, call center logs, transaction reports, etc.

[0107] In the steps after obtaining the historical data records, the feature combination device 200 iteratively performs feature combination between at least one discrete feature generated based on the multiple attribute information according to the search strategy to generate candidate combined features, and selects a target combined feature from the generated candidate combined features as the combined feature of the machine learning sample. Among them, for each round of iteration, the feature combination device 200 performs a preliminary sorting of the importance of each candidate combined feature in the candidate combined feature set, filters out a part of the candidate combined features from the candidate combined feature set according to the preliminary sorting result to form a candidate combined feature pool, performs a re-sorting of the importance of each candidate combined feature in the candidate combined feature pool, and selects at least one candidate combined feature with a higher importance from the candidate combined feature pool as the target combined feature according to the re-sorting result.

[0108] The following will detail each step involved in the above processing. First, for the first round of iteration, in step S205, the candidate combined feature generation unit 210 generates at least one discrete feature and / or at least one continuous feature based on the attribute information of the historical data record, and converts the generated continuous feature into a discrete feature.

[0109] Specifically, for at least a part of the attribute information of the historical data record, corresponding continuous features may be generated. According to an exemplary embodiment of the present invention, each continuous feature needs to be converted into a discrete feature when combined with other features. In addition, the candidate combined feature generation unit 210 may also generate discrete features. For example, a certain discrete value attribute information in the historical data record is directly used as a discrete feature, or a discrete feature is obtained by performing feature processing on the attribute information. As an alternative, the above features may also be generated by other feature generation devices (not shown).

[0110] Here, any suitable method can be adopted to discretize continuous features. Preferably, the candidate combined feature generation unit 210 can perform at least one binning operation for each continuous feature to generate a discrete feature composed of at least one binned feature, where each binning operation corresponds to a binned feature. The discrete feature composed of the above binned features can participate in the automatic combination between discrete features instead of the original continuous feature.

[0111] Here, the candidate combined feature generation unit 210 can perform binning operations according to various binning methods and / or binning parameters.

[0112] Taking equal-width binning without supervision as an example, assuming that the value range of the continuous feature is [0, 100] and the corresponding binning parameter (i.e., width) is 50, then 2 bins can be divided. In this case, the continuous feature with a value of 61.5 corresponds to the second bin. If the labels of these two bins are 0 and 1, then the bin label corresponding to the continuous feature is 1. Or, assuming that the binning width is 10, then 10 bins can be divided. In this case, the continuous feature with a value of 61.5 corresponds to the seventh bin. If the labels of these ten bins are from 0 to 9, then the bin label corresponding to the continuous feature is 6. Or, assuming that the binning width is 2, then 50 bins can be divided. In this case, the continuous feature with a value of 61.5 corresponds to the 31st bin. If the labels of these fifty bins are from 0 to 49, then the bin label corresponding to the continuous feature is 30.

[0113] After mapping the continuous feature to multiple bins, the corresponding feature value can be any user-defined value. Here, the binned feature can indicate which bin the continuous feature is assigned to according to the corresponding binning operation. That is, perform binning operations to generate multi-dimensional binned features corresponding to each continuous feature. As an example, each dimension can indicate whether the corresponding continuous feature is assigned to the corresponding bin. For example, use "1" to indicate that the continuous feature is assigned to the corresponding bin, and use "0" to indicate that the continuous feature is not assigned to the corresponding bin. Correspondingly, in the above example, assuming that 10 bins are divided, the binned feature can be a 10-dimensional feature. The binned feature corresponding to the continuous feature with a value of 61.5 can be expressed as [0, 0, 0, 0, 0, 0, 1, 0, 0, 0].

[0114] In addition, as an example, before performing the binning operation, the possible outliers in the data samples can also be removed to reduce the noise in the data records. In this way, the effectiveness of machine learning using binned features can be further improved.

[0115] Specifically, an outlier bin can be additionally set so that continuous features with outliers are assigned to the outlier bin. For example, for a continuous feature with a value range of [0, 1000], a certain number of samples can be selected for pre-binning. For example, equal-width binning can be first performed with a bin width of 10, and then the number of samples in each bin is recorded. For bins with a small number of samples (e.g., less than a threshold), they can be merged into at least one outlier bin. As an example, if the number of samples in the bins at both ends is small, the bins with fewer samples can be merged into an outlier bin, while the remaining bins are retained. Assuming that the number of samples in bins 0-10 is small, bins 0-10 can be merged into an outlier bin, so as to uniformly divide the continuous feature with a value of [0, 100] into the outlier bin.

[0116] According to an exemplary embodiment of the present invention, the at least one binning operation may be binning operations with the same binning method but different binning parameters; or, the at least one binning operation may be binning operations with different binning methods.

[0117] The binning methods here include various binning methods under supervised binning and / or unsupervised binning. For example, supervised binning includes minimum entropy binning, minimum description length binning, etc., while unsupervised binning includes equal-width binning, equal-depth binning, binning based on k-means clustering, etc.

[0118] As an example, at least one binning operation may respectively correspond to equal-width binning operations with different widths. That is to say, the binning method used is the same but the division granularity is different, which enables the generated binned features to better depict the law of the original data record, and thus is more conducive to the training and prediction of the machine learning model. In particular, the different widths used in at least one binning operation may form a geometric sequence numerically. For example, the binning operation can perform equal-width binning according to widths of values 2, 4, 8, 16, etc. Or, the different widths used in at least one binning operation may form an arithmetic sequence numerically. For example, the binning operation can perform equal-width binning according to widths of values 2, 4, 6, 8, etc.

[0119] As another example, at least one binning operation may respectively correspond to equal-depth binning operations with different depths. That is to say, the binning method used in the binning operation is the same but the division granularity is different, which enables the generated binned features to better depict the law of the original data record, and thus is more conducive to the training and prediction of the machine learning model. In particular, the different depths used in the binning operation may form a geometric sequence numerically. For example, the binning operation can perform equal-depth binning according to depths of values 10, 100, 1000, 10000, etc. Or, the different depths used in the binning operation may form an arithmetic sequence numerically. For example, the binning operation can perform equal-depth binning according to depths of values 10, 20, 30, 40, etc.

[0120] For each continuous feature, after obtaining the corresponding at least one binned feature by performing binning operations, a discrete feature corresponding to the continuous feature can be obtained by taking each binned feature as a constituent element, and the discrete feature can be regarded as a set of binned features.

[0121] As described above, according to an exemplary embodiment of the present invention, at least one binning operation needs to be performed on continuous features. Here, the at least one binning operation can be determined by any suitable means. For example, it can be determined with the help of the experience of technicians or business personnel, or can be automatically determined by technical means. As an example, the specific binning operation method can be effectively determined based on the importance of the binned features.

[0122] Correspondingly, the candidate combined feature generation unit 210 can select the at least one binning operation from a predetermined number of binning operations such that the importance of the binned features corresponding to the selected binning operation is not lower than the importance of the binned features corresponding to the unselected binning operations. In this way, the effect of machine learning can be ensured while reducing the size of the combined feature space.

[0123] Specifically, the predetermined number of binning operations can indicate multiple binning operations that differ in binning methods and / or binning parameters. Here, by performing each binning operation, a corresponding binned feature can be obtained. Correspondingly, the candidate combined feature generation unit 210 can determine the importance of these binned features, and then select the binning operation corresponding to the more important binned features as the at least one binning operation to be performed by the candidate combined feature generation unit 210.

[0124] Here, the candidate combined feature generation unit 210 can use any suitable means to automatically determine the importance of the binned features.

[0125] For example, the candidate combined feature generation unit 210 can obtain a single binned feature machine learning model for each binned feature among the binned features corresponding to the predetermined number of binning operations, determine the importance of each binned feature based on the effects of the individual single binned feature machine learning models, and select the at least one binning operation based on the importance of each binned feature, where the single binned feature machine learning model corresponds to each binned feature.

[0126] As an example, assume that for a continuous feature F, there are a predetermined number M (M is an integer greater than 1) of binning operations, corresponding to M binned features f m, where m ∈ [1, M]. Correspondingly, the candidate combined feature generation unit 210 can utilize at least a part of the historical data records to construct M binning single-feature machine learning models (where each binning single-feature machine learning model is based on the corresponding single binning feature f m to make predictions for the machine learning problem), and then measure the performance of these M binning single-feature machine learning models on the same test data set (for example, AUC (Area Under the ROC (Receiver Operating Characteristic) Curve), MAE (Mean Absolute Error), etc.), and determine at least one binning operation to be finally executed based on the ranking of the performance.

[0127] As another example, the candidate combined feature generation unit 210 can obtain a binning overall machine learning model for each binning feature among the binning features corresponding to the predetermined number of binning operations, determine the importance of each binning feature based on the performance of each binning overall machine learning model, and select the at least one binning operation based on the importance of each binning feature, where the binning overall machine learning model corresponds to the binning basic feature subset and each binning feature. As an example, the binning overall machine learning model here can be a Logistic Regression (LR) model; correspondingly, the samples of the binning overall machine learning model are composed of the binning basic feature subset and each binning feature.

[0128] As an example, assume that for the continuous feature F, there are a predetermined number M of binning operations, corresponding to M binning features f m , correspondingly, the candidate combined feature generation unit 210 can utilize at least a part of the historical data records to construct M binning overall machine learning models (where the sample features of each binning overall machine learning model include a fixed binning basic feature subset and the corresponding binning feature f m ), and then measure the performance of these M binning overall machine learning models on the same test data set (for example, AUC, MAE, etc.), and determine at least one binning operation to be finally executed based on the ranking of the performance.

[0129] For another example, the candidate combined feature generation unit 210 may obtain a bin composite machine learning model for each bin feature among the bin features corresponding to the predetermined number of bin operations, determine the importance of each bin feature based on the effects of the respective bin composite machine learning models, and select the at least one bin operation based on the importance of each bin feature. Here, the bin composite machine learning model includes a bin basic sub-model and a bin additional sub-model based on a boosting framework (e.g., a gradient boosting framework). The bin basic sub-model corresponds to a bin basic feature subset, and the bin additional sub-model corresponds to each bin feature.

[0130] As an example, assume that for a continuous feature F, there are a predetermined number M of bin operations, corresponding to M bin features f m , correspondingly, the candidate combined feature generation unit 210 may use at least a part of the historical data records to construct M bin composite machine learning models (where each bin composite machine learning model is based on a fixed bin basic feature subset and the corresponding bin feature f m , and makes predictions for the machine learning problem according to the boosting framework), then measure the effects (e.g., AUC, MAE, etc.) of these M bin composite machine learning models on the same test data set, and determine the at least one bin operation to be finally executed based on the ranking of the effects. Preferably, in order to further improve the operation efficiency and reduce the resource consumption, the candidate combined feature generation unit 210 may construct each bin composite machine learning model by training the bin additional sub-model for each bin feature f m respectively under the condition of a fixed bin basic sub-model.

[0131] According to an exemplary embodiment of the present invention, the bin basic feature subset may be fixedly applied to all relevant bin overall machine learning models or the bin basic sub-models in the bin composite machine learning models. Here, for the first round of iteration, the bin basic feature subset may be empty; or, any feature generated based on the attribute information of the historical data records may be used as the bin basic feature. For example, a part or all of the attribute information of the historical data records may be directly used as the bin basic feature. In addition, as an example, considering the actual machine learning problem, relatively important or basic features may be determined based on estimation or as specified by business personnel as the bin basic features.

[0132] After generating the unit discrete features for generating combined features as described above, in step S210, the candidate combined feature generation unit 210 generates candidate combined features for each iteration according to a search strategy. Here, since the continuous features have been converted into discrete features, any combination can be made among the discrete features as candidate combined features. As an example, the combination between discrete features can be achieved through the Cartesian product. However, it should be noted that the combination method is not limited thereto, and any method capable of combining two or more discrete features can be applied to the exemplary embodiments of the present invention.

[0133] Here, for the first iteration, the candidate combined feature generation unit 210 can directly use each of the discrete features generated in step S205 as candidate combined features.

[0134] Next, in step S220, the pre-ranking unit 220 performs a pre-ranking of the importance of each candidate combined feature in the candidate combined feature set for the first iteration. Here, the candidate combined feature set may include the candidate combined features that need to be pre-ranked in the current iteration, and these candidate combined features may be at least a part of the candidate combined features that have been generated. As an example, the candidate combined feature set in the first iteration may include all the discrete features generated in step S205.

[0135] Here, the pre-ranking unit 220 can use any means for judging the importance of features to measure the importance of each candidate combined feature in the candidate combined feature set.

[0136] For example, the pre-ranking unit 220 can obtain a pre-ranking single-feature machine learning model for each candidate combined feature in the candidate combined feature set, and determine the importance of each candidate combined feature based on the effects of the respective pre-ranking single-feature machine learning models, where the pre-ranking single-feature machine learning model corresponds to each of the candidate combined features.

[0137] As an example, assume that the candidate combined feature set includes N (N is an integer greater than 1) candidate combined features f n , where n ∈ [1, N]. Correspondingly, the pre-ranking unit 220 can use at least a part of the historical data records to construct N pre-ranking single-feature machine learning models (where each pre-ranking single-feature machine learning model makes a prediction for a machine learning problem based on the corresponding single candidate combined feature f n ), then measure the effects (such as AUC, MAE, etc.) of these N pre-ranking single-feature machine learning models on the same test data set, and determine the importance order of each candidate combined feature in the candidate combined feature set based on the ranking of the effects.

[0138] For another example, the pre-ranking unit 220 may obtain a pre-ranking overall machine learning model for each candidate combined feature in the candidate combined feature set, and determine the importance of each candidate combined feature based on the effects of the respective pre-ranking overall machine learning models. Here, the pre-ranking overall machine learning model corresponds to a pre-ranking basic feature subset and each candidate combined feature. As an example, the pre-ranking overall machine learning model here may be an LR model; correspondingly, the samples of the pre-ranking overall machine learning model are composed of a pre-ranking basic feature subset and each candidate combined feature.

[0139] As an example, assume that the candidate combined feature set includes N candidate combined features f n , correspondingly, the pre-ranking unit 220 may use at least a part of the historical data records to construct N pre-ranking overall machine learning models (where the sample features of each pre-ranking overall machine learning model include a fixed pre-ranking basic feature subset and the corresponding candidate combined feature f n ), and then measure the effects of these N pre-ranking overall machine learning models on the same test data set (such as AUC, MAE, etc.), and determine the importance order of each candidate combined feature in the candidate combined feature set based on the ranking of the effects.

[0140] For another example, the pre-ranking unit 220 may obtain a pre-ranking composite machine learning model for each candidate combined feature in the candidate combined feature set, and determine the importance of each candidate combined feature based on the effects of the respective pre-ranking composite machine learning models. Here, the pre-ranking composite machine learning model includes a pre-ranking basic sub-model and a pre-ranking additional sub-model based on a boosting framework (such as a gradient boosting framework), where the pre-ranking basic sub-model corresponds to a pre-ranking basic feature subset, and the pre-ranking additional sub-model corresponds to each candidate combined feature.

[0141] As an example, assume that the candidate combined feature set includes N candidate combined features f n , correspondingly, the pre-ranking unit 220 may use at least a part of the historical data records to construct N pre-ranking composite machine learning models (where each pre-ranking composite machine learning model is based on a fixed pre-ranking basic feature subset and the corresponding candidate combined feature f n , and makes predictions for the machine learning problem according to the boosting framework), and then measure the effects of these N pre-ranking composite machine learning models on the same test data set (such as AUC, MAE, etc.), and determine the importance order of each candidate combined feature in the candidate combined feature set based on the ranking of the effects. Preferably, in order to further improve the operation efficiency and reduce the resource consumption, the pre-ranking unit 220 may, under the condition of a fixed pre-ranking basic sub-model, respectively for each candidate combined feature f nTrain a pre-ranking additional sub-model to construct each pre-ranking composite machine learning model.

[0142] According to an exemplary embodiment of the present invention, the pre-ranking basic feature subset can be fixedly applied to all relevant pre-ranking overall machine learning models or pre-ranking basic sub-models in the pre-ranking composite machine learning model. Here, for the first round of iteration, the pre-ranking basic feature subset can be empty; or, any feature generated based on the attribute information of historical data records can be used as the pre-ranking basic feature. For example, a part or all of the attribute information of the historical data records can be directly used as the pre-ranking basic feature. In addition, as an example, considering the actual machine learning problem, relatively important or basic features can be determined based on estimation or as specified by business personnel as the pre-ranking basic features.

[0143] After determining the importance order of each candidate combination feature in the candidate combination feature set through pre-ranking, the pre-ranking unit 220 can screen out at least a part from the candidate combination features based on the ranking result to form a candidate combination feature pool. As described above, important candidate combination features with consistency in the prediction effect can be preferentially screened to form the candidate combination feature pool, so as to effectively determine the combination features that can ultimately form the machine learning samples. For example, the pre-ranking unit 220 can screen out candidate combination features with higher importance from the candidate combination feature set according to the pre-ranking result to form the candidate combination feature pool.

[0144] Assume that the candidate combination feature set in the first round of iteration includes 1000 discrete features as candidate combination features. The pre-ranking unit 220 can screen out the 10 most important discrete features in the pre-ranking result from them to form the candidate combination feature pool.

[0145] Next, in step S230, the re-ranking unit 230 re-ranks the importance of each candidate combination feature in the candidate combination feature pool, and selects at least one candidate combination feature with higher importance from the candidate combination feature pool as the target combination feature according to the re-ranking result.

[0146] Here, the re-ranking unit 230 can use any means of judging feature importance to measure the importance of each candidate combination feature in the candidate combination feature pool.

[0147] For example, for each candidate combination feature in the candidate combination feature pool, the re-ranking unit 230 can obtain a re-ranking single-feature machine learning model, and determine the importance of each candidate combination feature based on the effect of each re-ranking single-feature machine learning model, where the re-ranking single-feature machine learning model corresponds to each candidate combination feature.

[0148] As an example, assume that the candidate combined feature pool includes 10 candidate combined features. Correspondingly, the re-ranking unit 230 can utilize at least a portion of the historical data records to construct 10 re-ranking single-feature machine learning models (where each re-ranking single-feature machine learning model makes predictions for the machine learning problem based on a corresponding single candidate combined feature), then measure the performance of these 10 re-ranking single-feature machine learning models on the same test data set (e.g., AUC, MAE, etc.), and determine the importance order of each candidate combined feature in the candidate combined feature pool based on the ranking of the performance.

[0149] For another example, the re-ranking unit 230 can obtain a re-ranking overall machine learning model for each candidate combined feature in the candidate combined feature pool, and determine the importance of each candidate combined feature based on the performance of each re-ranking overall machine learning model, where the re-ranking composite machine learning model corresponds to the re-ranking basic feature subset and each of the candidate combined features. As an example, the re-ranking overall machine learning model here can be an LR model; correspondingly, the samples of the re-ranking overall machine learning model are composed of the re-ranking basic feature subset and each of the candidate combined features.

[0150] As an example, assume that the candidate combined feature pool includes 10 candidate combined features. Correspondingly, the re-ranking unit 230 can utilize at least a portion of the historical data records to construct 10 re-ranking overall machine learning models (where the sample features of each re-ranking overall machine learning model include a fixed re-ranking basic feature subset and the corresponding candidate combined feature), then measure the performance of these 10 re-ranking overall machine learning models on the same test data set (e.g., AUC, MAE, etc.), and determine the importance order of each candidate combined feature in the candidate combined feature pool based on the ranking of the performance.

[0151] For another example, the re-ranking unit 230 can obtain a re-ranking composite machine learning model for each candidate combined feature in the candidate combined feature pool, and determine the importance of each candidate combined feature based on the performance of each re-ranking composite machine learning model, where the re-ranking composite machine learning model includes a re-ranking basic sub-model and a re-ranking additional sub-model based on a boosting framework (e.g., gradient boosting framework), where the re-ranking basic sub-model corresponds to the re-ranking basic feature subset, and the re-ranking additional sub-model corresponds to each of the candidate combined features.

[0152] As an example, assume that the candidate combined feature pool includes 10 candidate combined features. Correspondingly, the re-ranking unit 230 can use at least a part of the historical data records to construct 10 re-ranking composite machine learning models (where each re-ranking composite machine learning model is based on a fixed re-ranking basic feature subset and the corresponding candidate combined feature, and makes predictions for the machine learning problem according to the boosting framework), then measure the effects of these 10 re-ranking composite machine learning models on the same test data set (such as AUC, MAE, etc.), and determine the importance order of each candidate combined feature in the candidate combined feature pool based on the ranking of the effects. Preferably, in order to further improve the operation efficiency and reduce the resource consumption, the re-ranking unit 230 can construct each re-ranking composite machine learning model by training a re-ranking additional sub-model for each candidate combined feature respectively under the condition of a fixed re-ranking basic sub-model.

[0153] According to an exemplary embodiment of the present invention, the re-ranking basic feature subset can be fixedly applied to all relevant re-ranking overall machine learning models or the re-ranking basic sub-models in the re-ranking composite machine learning models. Here, for the first round of iteration, the re-ranking basic feature subset can be empty; or, any feature generated based on the attribute information of the historical data records can be used as the re-ranking basic feature. For example, a part or all of the attribute information of the historical data records can be directly used as the re-ranking basic feature. In addition, as an example, considering the actual machine learning problem, relatively important or basic features can be determined based on estimation or as specified by business personnel as the re-ranking basic features.

[0154] After determining the importance order of each candidate combined feature in the candidate combined feature pool through re-ranking, the re-ranking unit 230 can screen out at least one relatively important candidate combined feature from the candidate combined feature pool based on the ranking result as the target combined feature. Assume that the candidate combined feature pool in the first round of iteration includes 10 discrete features as candidate combined features, and the re-ranking unit 230 can screen out the most important 1 discrete feature in the re-ranking result as the target combined feature.

[0155] According to an exemplary embodiment of the present invention, the operation resources can be further effectively controlled by sharing the same model part.

[0156] As an example, when the candidate combination feature generation unit 210, the pre-ranking unit 220, and / or the re-ranking unit 230 respectively perform importance ranking of relevant features based on their respective boosting framework composite machine learning models, for example, a common basic sub-model part can be trained based on a relatively large number of historical data records (e.g., all historical data records), and this part can be used as a fixed model part and respectively used as the binning basic sub-model in the binning composite machine learning model, the pre-ranking basic sub-model in the pre-ranking composite machine learning model, and / or the re-ranking basic sub-model in the re-ranking composite machine learning model. Further, in the case of sharing the basic sub-model, the binning additional sub-model, the pre-ranking additional sub-model, and / or the re-ranking additional sub-model corresponding to each feature whose importance is to be determined can be trained in parallel, so that multiple models can be trained simultaneously with only one read operation of the historical data records.

[0157] In addition, according to an exemplary embodiment of the present invention, the effect of the combined features can be further ensured by controlling the scale of the sample training set, the sample training order, and / or the quality of the sample training set of the relevant model part.

[0158] As an example, the pre-ranking unit 220 can train a pre-ranking single-feature machine learning model based on a relatively small number of historical data records, and the re-ranking unit 230 can train a re-ranking single-feature machine learning model based on a relatively large number of historical data records; or, the pre-ranking unit 220 can train a pre-ranking overall machine learning model based on a relatively small number of historical data records, and the re-ranking unit 230 can train a re-ranking overall machine learning model based on a relatively large number of historical data records; or, the pre-ranking unit 220 can train a pre-ranking additional sub-model based on a relatively small number of historical data records, and the re-ranking unit 230 can train a re-ranking additional sub-model based on a relatively large number of historical data records. Here, the historical data records used by the re-ranking unit 230 can include at least a part of the historical data records used by the pre-ranking unit 220, or the historical data records used by the re-ranking unit 230 may not include any historical data records used by the pre-ranking unit 220. In addition to the difference in the scale of the sample training set, the pre-ranking unit 220 and the re-ranking unit 230 can use the same historical data record set, but only the training order is different. Thus, it can be seen that the feature combination device 200 can perform pre-ranking based on the first quantity of historical data records and perform re-ranking based on the second quantity of historical data records, and the second quantity is not less than the first quantity. In addition, the pre-ranking unit 220 can also use a sample training set with a different quality from that of the re-ranking unit 230. For example, the pre-ranking unit 220 can use a sample training set with a lower quality, and the re-ranking unit 230 can use a sample training set with a higher quality. In this way, even if the re-ranking unit 230 uses a smaller sample training set, the effect of the re-ranking related model can be ensured.

[0159] However, it should be noted that the exemplary embodiments of the present invention are not limited thereto. Instead, any method can be adopted to construct their respective basic sub-models separately, and any appropriate training data set can also be used.

[0160] After screening out the target combined features in the first round of iteration, in step S235, it is determined whether the condition for terminating the iteration is satisfied. Here, any condition for terminating the iteration can be preset in advance, for example, the number of target combined features that have been obtained, the number of iteration rounds that have been executed, etc. When the iteration termination condition is satisfied, the generation process of the combined features can be terminated; otherwise, the method can return to step S205 or step S210 to perform the next round of iteration.

[0161] Specifically, assuming that at least one binning operation performed when converting continuous features into discrete features is selected from a predetermined number of binning operations for each round of iteration, the method needs to return to step S205 so that the candidate combined feature generation unit 210 re-selects at least one binning operation for discretization for the second round of iteration.

[0162] Here, the candidate combined feature generation unit 210 can convert the continuous features into corresponding discrete features again in various ways similar to the first round to perform subsequent feature combination.

[0163] In particular, in the case of selecting binning operations for each round of iteration, assuming that the candidate combined feature generation unit 210 uses a binning overall machine learning model or a binning composite machine learning model to measure the importance of binning features, the target combined features selected in each round of iteration can be added to the binning basic feature subset as new discrete features, that is, the binning basic feature subset can include the target combined features selected before the current round of iteration. Here, the binning basic feature subset on which the binning basic sub-model is based can be updated with the iteration of generating the target combined features. Specifically, the target combined features selected in the first round of iteration can be added to the binning basic feature subset of the first round of iteration to form the binning basic feature subset of the second round of iteration.

[0164] After the candidate combined feature generation unit 210 obtains the discrete features converted from the continuous features again as described above, the method can proceed to step S210; or, in the case where at least one binning operation performed when converting continuous features into discrete features is selected from a predetermined number of binning operations all at once for all rounds of iteration, the method can directly start the next round of iteration from step S210 without executing step S205 again.

[0165] Specifically, in step S210, the candidate combined feature generation unit 210 can generate candidate combined features for the second round of iteration according to a search strategy. As an example, in the first round of iteration, first-order discrete features are selected as the target combined features. Correspondingly, in the second round of iteration, the candidate combined feature generation unit 210 can obtain second-order or higher-order candidate combined features by performing Cartesian product combination on the features.

[0166] For example, the following will describe an example of iteratively generating combined features by the candidate combined feature generation unit 210 in conjunction with Figure 6 the search tree shown. The search tree can be based on a heuristic search strategy such as beam search, where one layer of the search tree can correspond to a feature combination of a specific order.

[0167] Referring to Figure 6 , for ease of description, assume that the unit discrete features that can be combined include feature A, feature B, feature C, feature D, and feature E. As an example, features A, B, and C can be discrete features formed by the discrete value attribute information recorded in historical data itself, while features D and E can be discrete features converted from continuous features through corresponding binning operations in each round of iteration.

[0168] According to the search strategy, the importance of the features under re-ranking can be used as an index to sort each node of the search tree, and then a part of the nodes can be selected to continue expanding in the next layer. For example, assume that in the first round of iteration, finally, two nodes, feature B and feature E, which are first-order features, are selected as the target combined features. Then, in the second round of iteration, the candidate combined feature generation unit 210 can generate feature BA, feature BC, feature BD, feature BE, feature EA, feature EB, feature EC, and feature ED, which are second-order combined features, based on feature B and feature E. Here, as an example, combined features with only sequential changes (e.g., feature BE and feature EB) can be regarded as the same feature, so only one of them is retained after deduplication processing. As described above, the candidate combined feature generation unit 210 can generate candidate combined features for the next round of iteration by combining the target combined features selected in the current round of iteration with at least one discrete feature generated based on multiple attribute information recorded in historical data. Correspondingly, assume that feature BC and feature EA are selected in the second round of iteration, then as Figure 6 shown, continue the iteration in the above manner until a specific cut-off condition is met, such as an order limit, etc. Here, the nodes selected in each layer (shown by solid lines) can be used as the target combined features for subsequent processing, such as being used as the finally adopted sample features or for further verification, while the remaining features (shown by dashed lines) are pruned.

[0169] The above shows an example in which the candidate combination feature generation unit 210 generates candidate combination features step by step. In this example, the candidate combination features generated in the second round of iteration include feature BA, feature BC, feature BD, feature BE, feature EA, feature EB, feature EC, and feature ED.

[0170] Correspondingly, in step S220, the pre-ranking unit 220 performs pre-ranking of the importance of each candidate combination feature in the candidate combination feature set for the second round of iteration. Here, the candidate combination feature set may include candidate combination features that need to be pre-ranked in the current round of iteration. As an example, the candidate combination feature set may include candidate combination features generated in the current round of iteration. For example, the features BA, BC, BD, BE, EA, EB, EC, and ED generated in the second round of iteration; as another example, the candidate combination feature set may not only include candidate combination features generated in the current round of iteration, but may further include candidate combination features generated in the previous round of iteration that have not been selected as the target combination feature. For example, the features BA, BC, BD, BE, EA, EB, EC, and ED generated in the second round of iteration together with the non-target candidate features generated in the first round of iteration, that is, features A, C, and D. In this way, it is possible to more comprehensively measure candidate combination features while ensuring the operation efficiency. It should be noted that according to the exemplary embodiment of the present invention, a part of all candidate combination features generated in the current round and / or previous rounds of iteration may be selected to enter the candidate combination feature set, rather than necessarily using all currently existing candidate combination features.

[0171] Here, the pre-ranking unit 220 may rank the importance of each candidate combination feature in the candidate combination feature set in various ways similar to the first round.

[0172] In particular, in the case where the pre-ranking unit 220 uses a pre-ranking overall machine learning model or a pre-ranking composite machine learning model to measure the importance of candidate combination features, the target combination feature selected in each round of iteration may be added as a new discrete feature to the pre-ranking basic feature subset. That is, the pre-ranking basic feature subset may include the target combination features selected before the current round of iteration. Here, the pre-ranking basic feature subset on which the pre-ranking basic sub-model is based may be updated as the iteration of generating the target combination feature progresses. Specifically, the target combination feature selected in the first round of iteration may be added to the pre-ranking basic feature subset of the first round of iteration to form the pre-ranking basic feature subset of the second round of iteration.

[0173] After the pre-sorting unit 220 obtains a new candidate combined feature pool through pre-sorting processing in step S220, in step S230, the re-sorting unit 230 re-sorts the importance of each candidate combined feature in the candidate combined feature pool, and selects at least one candidate combined feature with higher importance from the candidate combined feature pool as the target combined feature according to the re-sorting result.

[0174] Here, the re-sorting unit 230 can sort the importance of each candidate combined feature in the candidate combined feature pool in various ways similar to the first round.

[0175] In particular, when the re-sorting unit 230 uses a re-sorting overall machine learning model or a re-sorting composite machine learning model to measure the importance of candidate combined features, the target combined features selected in each round of iteration can be added as new discrete features to the re-sorting basic feature subset. That is, the re-sorting basic feature subset can include the target combined features selected before the current round of iteration. Here, the re-sorting basic feature subset on which the re-sorting basic sub-model is based can be updated as the iteration for generating the target combined features progresses. Specifically, the target combined features selected in the first round of iteration can be added to the re-sorting basic feature subset of the first round of iteration to form the re-sorting basic feature subset of the second round of iteration.

[0176] Here, in addition to adopting Figure 6 the method of generating candidate combined features step by step, according to an exemplary embodiment of the present invention, candidate combined features can also be generated more effectively in each round of iteration. Specifically, in step S210, the candidate combined feature generation unit 210 can generate candidate combined features for the next round of iteration by pairwise combining the target combined features selected in the current round of iteration and the previous round of iteration. In this way, valuable combination methods can be more concentratedly mined.

[0177] According to an exemplary embodiment of the present invention, when the iteration termination condition is satisfied, the sum of the target combined features selected in each round of iteration can be used as the combined feature set of the machine learning sample. In particular, when the binning basic feature subset, the pre-sorting basic feature subset, and / or the re-sorting basic feature subset are updated with the target combined features selected in each round, the combined features in the above subsets in the last round of iteration can be used as the combined feature set of the machine learning sample.

[0178] Figure 7 The flowchart showing the training method of the machine learning model according to an exemplary embodiment of the present invention. In Figure 7 the method shown, in addition to the above steps S100, step S205, step S210, step S220, step S230, and step S235, the method further includes step S300 and step S400.

[0179] Specifically, in the Figure 7 method shown, steps S100, S205, S210, S220, S230, and S235 may be similar to the Figure 5 corresponding steps shown, and the details will not be elaborated here.

[0180] In addition, in step S300, a machine learning training sample including at least a part of the generated combined features may be generated by the machine learning sample generation device 300. In the case of supervised learning, the machine learning training sample may include two parts: features and labels.

[0181] In step S400, the machine learning model training device 400 may train a machine learning model based on the machine learning training sample. Here, the machine learning model training device 400 may use an appropriate machine learning algorithm to learn an appropriate machine learning model from the machine learning training sample. As an example, the appropriate machine learning algorithm may be the same as or different from the machine learning algorithms on which the binning single - feature machine learning model, pre - sorting single - feature machine learning model, re - sorting single - feature machine learning model, binning overall machine learning model, pre - sorting overall machine learning model, re - sorting overall machine learning model, binning composite machine learning model (or binning basic sub - model or binning additional sub - model), pre - sorting composite machine learning model (or pre - sorting basic sub - model or pre - sorting additional sub - model), or re - sorting composite machine learning model (or re - sorting basic sub - model or re - sorting additional sub - model) are based.

[0182] After training the machine learning model, the trained machine learning model may be used for prediction.

[0183] According to an exemplary embodiment of the present invention, in order to further enhance the effectiveness of the target combined features, the target combined features may be further verified, and only the target combined features that pass the verification may be used as the combined features of the machine learning sample. Figure 8 The flowchart showing a method for generating combined features of a machine learning sample according to another exemplary embodiment of the present invention is shown. In this example, for each iteration, it may also be checked whether the selected target combined features are suitable as the combined features of the machine learning sample.

[0184] Referring to Figure 8 , steps S100, S205, S210, S220, and S230 are similar to the Figure 5 corresponding steps shown, and the details will not be elaborated here.

[0185] In addition, after obtaining the target combined feature in step S230, the method proceeds to step S240, where it can be verified by the verification unit 240 whether the target combined feature obtained in step S230 is suitable as the combined feature of the machine learning sample.

[0186] As an example, the verification unit 240 can use the change in the effect of the machine learning model after introducing the selected target combined feature for verification, where the sample features of the machine learning model include the target combined features that have passed the verification. Specifically, the verification unit 240 can construct a machine learning model based on the target combined features that have passed the verification. For example, the samples of this machine learning model can at least include those target combined features that have passed the verification before, and can further include other features. Here, the machine learning model can be based on a similar feature subset as the binning basic feature sub-model, the pre-ranking basic feature sub-model, and / or the re-ranking basic feature sub-model, and can be trained based on a relatively large number of historical data records. Optionally, the machine learning model is not based on the boosting framework, so that it can more accurately verify whether the selected target combined feature is truly helpful for performing predictions for machine learning problems.

[0187] Here, the verification unit 240 can determine whether the change in the model effect meets the requirements (for example, the enhancement of the effect meets the expectation or the weakening of the effect is acceptable) after the above-mentioned machine learning model introduces the newly selected target combined feature in this round of iteration. Specifically, the verification unit 240 can determine whether the model effect has been enhanced (for example, whether the enhancement of the model effect reaches the predetermined enhancement degree); or, the verification unit 240 can determine whether the model effect has only decreased slightly (for example, whether the decrease in the model effect is lower than the predetermined decrease degree, in which case the decrease in the model effect can be ignored). When the prediction effect of the model meets the requirements, it can be determined that the selected target combined feature is suitable as the combined feature of the machine learning sample.

[0188] Correspondingly, in the case where the verification result is that the selected target combined feature is suitable as the combined feature of the machine learning sample, the selected target combined feature can be used as the combined feature of the machine learning sample and the next round of iteration can be executed; in the case where the verification result is that the selected target combined feature is not suitable as the combined feature of the machine learning sample, another part of the candidate combined features can be screened out from the candidate combined feature set according to the pre-ranking result to form a new candidate combined feature pool.

[0189] As an example, in step S240, the verification unit 240 can determine whether the machine learning model performs better on the same data test set after introducing the newly selected target combined feature. If it is determined that the newly selected target combined feature brings better prediction effect, it means that the corresponding target combined feature can be used as the combined feature of the machine learning sample, and the method proceeds to step S235 to determine whether the iteration termination condition is satisfied.

[0190] When the iteration termination condition is satisfied, the method ends, and all the target combined features that have passed the verification currently can be used as the combined features finally adopted by the machine learning sample. Otherwise, the method returns to step S205 (or step S210) to perform the next round of iteration.

[0191] If it is determined that the newly selected target combined feature does not bring better prediction effect, the method proceeds to step S245 to determine whether to continue screening other target combined features in this round of iteration. Here, if the screening termination condition is satisfied (for example, a predetermined number of target combined features have been verified in this round of iteration, all the target combined features have been verified in this round of iteration, etc.), the method executes step S235. Otherwise, the method can return to step S220, where the pre-sorting unit 220 reconstructs the candidate combined feature pool so that the re-sorting unit 230 re-screens the target combined features. Here, the pre-sorting unit 220 can reconstruct a new candidate combined feature pool according to the pre-sorting result again. For example, if the first important to the tenth important features have been screened before, the eleventh important to the twentieth important features can be continuously screened, and so on.

[0192] It should be noted that the above verification steps can also be applied to Figure 7 the method shown, and will not be elaborated here.

[0193] Figures 1 to 4 The devices and their units shown can be respectively configured as software, hardware, firmware, or any combination of the above to perform specific functions. For example, these devices or units can correspond to dedicated integrated circuits, can also correspond to pure software code, and can also correspond to modules combined with software and hardware. In addition, one or more functions implemented by these devices or units can also be uniformly executed by components in a physical entity device (such as a processor, a client, or a server, etc.).

[0194] The above references Figures 1 to 8A method and system for generating combined features of machine learning samples according to an exemplary embodiment of the present invention, and a corresponding machine learning model training method and system are described. It should be understood that the above method can be implemented by a program recorded on a computer-readable medium. For example, according to an exemplary embodiment of the present invention, a computer-readable medium for generating combined features of machine learning samples can be provided, wherein a computer program for performing the following method steps is recorded on the computer-readable medium: (A) obtaining historical data records, wherein the historical data records include a plurality of attribute information; and (B) iteratively performing feature combination among at least one discrete feature generated based on the plurality of attribute information according to a search strategy to generate candidate combined features, and selecting target combined features from the generated candidate combined features as the combined features of machine learning samples, wherein for each round of iteration, pre-ranking the importance of each candidate combined feature in the candidate combined feature set; screening out a part of the candidate combined features from the candidate combined feature set according to the pre-ranking result to form a candidate combined feature pool; re-ranking the importance of each candidate combined feature in the candidate combined feature pool; and selecting at least one candidate combined feature with higher importance from the candidate combined feature pool as the target combined feature.

[0195] The computer program in the above computer-readable medium can run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. It should be noted that the computer program can also be used to execute additional steps other than the above steps or perform more specific processing when executing the above steps. The content of these additional steps and further processing has been described with reference to Figures 1 to 8 and will not be repeated here to avoid redundancy.

[0196] It should be noted that the combined feature generation system and the machine learning model training system according to the exemplary embodiment of the present invention can fully rely on the operation of the computer program to implement the corresponding functions, that is, each device corresponds to the respective steps in the functional architecture of the computer program, so that the entire system is called by a dedicated software package (for example, a lib library) to implement the corresponding functions.

[0197] On the other hand, Figures 1 to 4 Each of the devices or units shown can also be implemented by hardware, software, firmware, middleware, microcode, or any combination thereof. When implemented by software, firmware, middleware, or microcode, the program code or code segment for performing the corresponding operations can be stored in a computer-readable medium such as a storage medium, so that the processor can execute the corresponding operations by reading and running the corresponding program code or code segment.

[0198] For example, an exemplary embodiment of the present invention can also be implemented as a computing device, which includes a storage component and a processor. A set of computer-executable instructions is stored in the storage component. When the set of computer-executable instructions is executed by the processor, the combined feature generation method or the machine learning model training method is executed.

[0199] Specifically, the computing device can be deployed in a server or a client, or can be deployed on a node device in a distributed network environment. In addition, the computing device can be a PC computer, a tablet device, a personal digital assistant, a smart phone, a web application, or other devices capable of executing the above instruction set.

[0200] Here, the computing device does not have to be a single computing device, but can also be a collection of devices or circuits that can execute the above instructions (or instruction sets) alone or jointly. The computing device can also be a part of an integrated control system or a system manager, or can be configured as a portable electronic device that can be interconnected locally or remotely (for example, via wireless transmission).

[0201] In the computing device, the processor can include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor can also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0202] Certain operations described in the combined feature generation method and the machine learning model training method according to the exemplary embodiments of the present invention can be implemented in software, certain operations can be implemented in hardware, and in addition, these operations can also be implemented in a combination of software and hardware.

[0203] The processor can run instructions or code stored in one of the storage components, where the storage component can also store data. The instructions and data can also be sent and received via a network interface device through a network, where the network interface device can use any known transmission protocol.

[0204] The storage component can be integrated with the processor. For example, RAM or flash memory is arranged inside an integrated circuit microprocessor, etc. In addition, the storage component can include independent devices, such as an external disk drive, a storage array, or other storage devices that can be used by any database system. The storage component and the processor can be operatively coupled, or can communicate with each other, for example, through an I / O port, a network connection, etc., so that the processor can read files stored in the storage component.

[0205] In addition, the computing device may further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the computing device may be connected to each other via a bus and / or a network.

[0206] The operations involved in the combined feature generation method and the corresponding machine learning model training method according to an exemplary embodiment of the present invention may be described as various interconnected or coupled functional blocks or functional diagrams. However, these functional blocks or functional diagrams may be equally integrated into a single logical device or operate with non-exact boundaries.

[0207] For example, as described above, the computing device for generating combined features of machine learning samples according to an exemplary embodiment of the present invention may include a storage component and a processor. Among them, a set of computer-executable instructions is stored in the storage component. When the set of computer-executable instructions is executed by the processor, the following steps are performed: (A) obtaining historical data records, where the historical data records include a plurality of attribute information; and (B) iteratively performing feature combination among at least one discrete feature generated based on the plurality of attribute information according to a search strategy to generate candidate combined features, and selecting a target combined feature from the generated candidate combined features as the combined feature of the machine learning sample. For each round of iteration, pre-rank the importance of each candidate combined feature in the candidate combined feature set; screen out a part of the candidate combined features from the candidate combined feature set according to the pre-ranking result to form a candidate combined feature pool; re-rank the importance of each candidate combined feature in the candidate combined feature pool; and select at least one candidate combined feature with higher importance from the candidate combined feature pool as the target combined feature.

[0208] The above describes various exemplary embodiments of the present invention. It should be understood that the above description is only exemplary and not exhaustive. The present invention is not limited to the disclosed exemplary embodiments. Many modifications and changes are obvious to those of ordinary skill in the art without departing from the scope and spirit of the present invention. Therefore, the protection scope of the present invention should be subject to the scope of the claims.

Claims

1. A model training method executed by a computing device, comprising: (A) obtaining historical data records input by a user, wherein the historical data records include a plurality of attribute information, the plurality of attribute information relates to the attribute information of one or more of individuals, objects, organizations, enterprises, units, institutions, projects, and events, and the types of the plurality of attribute information include: text data and / or numerical data; (B) iteratively performing feature combination among at least one discrete feature generated based on the plurality of attribute information according to a search strategy to generate candidate combined features, and selecting a target combined feature from the generated candidate combined features as the combined feature of a machine learning sample, wherein, for each round of iteration, pre-ranking the importance of each candidate combined feature in the candidate combined feature set; screening out a part of the candidate combined features from the candidate combined feature set according to the pre-ranking result to form a candidate combined feature pool; re-ranking the importance of each candidate combined feature in the candidate combined feature pool; and selecting at least one candidate combined feature with higher importance from the candidate combined feature pool as the target combined feature according to the re-ranking result; wherein the search strategy is designed to prune a search tree for combined discrete features to control the number of candidate combined features generated in each round of iteration; (C) generating a machine learning sample including at least a part of the combined features; and (D) training the model based on the machine learning sample.

2. The method according to claim 1, wherein Pre-ranking is performed based on a first quantity of historical data records, re-ranking is performed based on a second quantity of historical data records, and the second quantity is not less than the first quantity.

3. The method according to claim 1, wherein Screening out candidate combined features with higher importance from the candidate combined feature set according to the pre-ranking result to form a candidate combined feature pool.

4. The method according to claim 1, wherein The candidate combined feature set includes the candidate combined features generated in the current round of iteration; or, the candidate combined feature set includes the candidate combined features generated in the current round of iteration and the candidate combined features generated in the previous round of iteration that have not been selected as the target combined features.

5. The method according to claim 1, wherein, Generating candidate combined features for the next round of iteration by combining the target combined feature selected in the current round of iteration with the at least one discrete feature; or, generating candidate combined features for the next round of iteration by pairwise combining the target combined features selected in the current round of iteration and the previous round of iteration.

6. The method according to claim 1, wherein The at least one discrete feature includes discrete features converted from continuous features generated based on the plurality of attribute information through the following processing: for each continuous feature, performing at least one binning operation to generate discrete features composed of at least one binned feature, wherein each binning operation corresponds to one binned feature.

7. The method according to claim 6, wherein, The at least one binning operation is selected from a predetermined number of binning operations for each round of iteration or for all rounds of iteration, and the importance of the binned features corresponding to the selected binning operation is not lower than the importance of the binned features corresponding to the unselected binning operations.

8. The method according to claim 7, wherein The at least one binning operation is selected through the following process: for each binning feature among the binning features corresponding to the predetermined number of binning operations, a binning single-feature machine learning model is obtained, the importance of each binning feature is determined based on the effects of the respective binning single-feature machine learning models, and the at least one binning operation is selected based on the importance of each binning feature, where the binning single-feature machine learning model corresponds to each binning feature.

9. The method according to claim 7, wherein, The at least one binning operation is selected through the following process: for each binning feature among the binning features corresponding to the predetermined number of binning operations, a binning overall machine learning model is obtained, the importance of each binning feature is determined based on the effects of the respective binning overall machine learning models, and the at least one binning operation is selected based on the importance of each binning feature, where the binning overall machine learning model corresponds to the binning basic feature subset and each binning feature.

10. The method according to claim 7, wherein, The at least one binning operation is selected through the following process: for each binning feature among the binning features corresponding to the predetermined number of binning operations, a binning composite machine learning model is obtained, the importance of each binning feature is determined based on the effects of the respective binning composite machine learning models, and the at least one binning operation is selected based on the importance of each binning feature, where the binning composite machine learning model includes a binning basic sub-model and a binning additional sub-model based on a boosting framework, where the binning basic sub-model corresponds to the binning basic feature subset and the binning additional sub-model corresponds to each binning feature.

11. The method according to claim 9 or 10, wherein, The binning basic feature subset includes the target combination features selected before the current round of iteration.

12. The method according to claim 1, wherein Pre-sorting is performed through the following process: for each candidate combination feature in the candidate combination feature set, a pre-sorting single-feature machine learning model is obtained, and the importance of each candidate combination feature is determined based on the effects of the respective pre-sorting single-feature machine learning models, where the pre-sorting single-feature machine learning model corresponds to each candidate combination feature.

13. The method according to claim 1, wherein, Pre-sorting is performed through the following process: for each candidate combination feature in the candidate combination feature set, a pre-sorting overall machine learning model is obtained, and the importance of each candidate combination feature is determined based on the effects of the respective pre-sorting overall machine learning models, where the pre-sorting overall machine learning model corresponds to the pre-sorting basic feature subset and each candidate combination feature.

14. The method according to claim 1, wherein, Pre-sorting is performed through the following process: for each candidate combination feature in the candidate combination feature set, a pre-sorting composite machine learning model is obtained, and the importance of each candidate combination feature is determined based on the effects of the respective pre-sorting composite machine learning models, where the pre-sorting composite machine learning model includes a pre-sorting basic sub-model and a pre-sorting additional sub-model based on a boosting framework, where the pre-sorting basic sub-model corresponds to the pre-sorting basic feature subset and the pre-sorting additional sub-model corresponds to each candidate combination feature.

15. The method according to claim 13 or 14, wherein The pre-sorting basic feature subset includes the target combination features selected before the current round of iteration.

16. The method according to claim 1, wherein, Re - sorting is performed through the following process: for each candidate combined feature in the candidate combined feature pool, a re - sorting single - feature machine learning model is obtained, and the importance of each candidate combined feature is determined based on the effects of the respective re - sorting single - feature machine learning models, where the re - sorting single - feature machine learning model corresponds to each of the candidate combined features.

17. The method according to claim 1, wherein, Re - sorting is performed through the following process: for each candidate combined feature in the candidate combined feature pool, a re - sorting overall machine learning model is obtained, and the importance of each candidate combined feature is determined based on the effects of the respective re - sorting overall machine learning models, where the re - sorting composite machine learning model corresponds to the re - sorting basic feature subset and each of the candidate combined features.

18. The method according to claim 1, wherein Re - sorting is performed through the following process: for each candidate combined feature in the candidate combined feature pool, a re - sorting composite machine learning model is obtained, and the importance of each candidate combined feature is determined based on the effects of the respective re - sorting composite machine learning models, where the re - sorting composite machine learning model includes a re - sorting basic sub - model based on a boosting framework and a re - sorting additional sub - model, where the re - sorting basic sub - model corresponds to the re - sorting basic feature subset, and the re - sorting additional sub - model corresponds to each of the candidate combined features.

19. The method according to claim 17 or 18, wherein, The re - sorting basic feature subset includes the target combined features selected before the current round of iteration.

20. The method according to claim 1, wherein Step (B) further includes: for each round of iteration, checking whether the selected target combined features are suitable as the combined features of the machine learning samples.

21. The method according to claim 20, wherein, In step (B), the effect change of the machine learning model based on the target combined features that have passed the check after introducing the selected target combined features is used to check whether the selected target combined features are suitable as the combined features of the machine learning samples.

22. The method according to claim 21, wherein, In the case where the check result is that the selected target combined features are suitable as the combined features of the machine learning samples, the selected target combined features are used as the combined features of the machine learning samples, and the next round of iteration is executed; in the case where the check result is that the selected target combined features are not suitable as the combined features of the machine learning samples, another part of the candidate combined features is screened out from the candidate combined feature set according to the pre - sorting result to form a new candidate combined feature pool.

23. A computer-readable medium for model training, wherein, A computer program for executing the method according to any one of claims 1 to 22 is recorded on the computer - readable medium.

24. A computing device for model training, comprising a storage component and a processor, wherein, A set of computer - executable instructions is stored in the storage component, and when the set of computer - executable instructions is executed by the processor, the method according to any one of claims 1 to 22 is executed.

25. A model training system, comprising: A data record acquisition device for acquiring historical data records input by a user, where the historical data records include a plurality of attribute information, the plurality of attribute information relates to the attribute information of one or more of individuals, objects, organizations, enterprises, units, institutions, projects, and events, and the types of the plurality of attribute information include: text data and / or numerical data; A feature combination device, configured to iteratively perform feature combination among at least one discrete feature generated based on the multiple attribute information according to a search strategy to generate candidate combined features, and select a target combined feature from the generated candidate combined features as the combined feature of a machine learning sample. Wherein, for each iteration, the feature combination device performs a preliminary sorting of the importance of each candidate combined feature in the candidate combined feature set, screens out a part of the candidate combined features from the candidate combined feature set according to the preliminary sorting result to form a candidate combined feature pool, performs a re - sorting of the importance of each candidate combined feature in the candidate combined feature pool, and selects at least one candidate combined feature with a higher importance from the candidate combined feature pool as the target combined feature according to the re - sorting result. Wherein, the search strategy is designed to perform pruning on a search tree regarding combined discrete features to control the number of candidate combined features generated in each iteration. A machine learning sample generation device, configured to generate machine learning samples that at least include a part of the combined features; and A machine learning model training device, configured to train the model based on the machine learning samples.

26. The system according to claim 25, wherein, The feature combination device performs preliminary sorting based on a first quantity of historical data records and performs re - sorting based on a second quantity of historical data records, and the second quantity is not less than the first quantity.

27. The system according to claim 25, wherein, The feature combination device screens out candidate combined features with higher importance from the candidate combined feature set according to the preliminary sorting result to form a candidate combined feature pool.

28. The system according to claim 25, wherein, The candidate combined feature set includes the candidate combined features generated in the current iteration; or, the candidate combined feature set includes the candidate combined features generated in the current iteration and the candidate combined features generated in the previous iteration that have not been selected as the target combined features.

29. The system according to claim 25, wherein the feature The combination device generates candidate combined features for the next iteration by combining the target combined feature selected in the current iteration with the at least one discrete feature; or, the feature combination device generates candidate combined features for the next iteration by pairwise combining the target combined features selected in the current iteration and the previous iteration.

30. The system according to claim 25, wherein, The at least one discrete feature includes discrete features converted from continuous features generated based on the multiple attribute information through the following processing: for each continuous feature, at least one binning operation is performed to generate discrete features composed of at least one binned feature, wherein each binning operation corresponds to a binned feature.

31. The system according to claim 30, wherein, The at least one binning operation is selected from a predetermined number of binning operations for each iteration or for all iterations, wherein the importance of the binned features corresponding to the selected binning operations is not lower than the importance of the binned features corresponding to the unselected binning operations.

32. The system according to claim 31, wherein, The feature combination device selects the at least one binning operation through the following process: for each binning feature among the binning features corresponding to the predetermined number of binning operations, a binning single-feature machine learning model is obtained, the importance of each binning feature is determined based on the effects of the respective binning single-feature machine learning models, and the at least one binning operation is selected based on the importance of each binning feature, where the binning single-feature machine learning model corresponds to each binning feature.

33. The system according to claim 31, wherein, The feature combination device selects the at least one binning operation through the following process: for each binning feature among the binning features corresponding to the predetermined number of binning operations, a binning overall machine learning model is obtained, the importance of each binning feature is determined based on the effects of the respective binning overall machine learning models, and the at least one binning operation is selected based on the importance of each binning feature, where the binning overall machine learning model corresponds to the binning basic feature subset and each binning feature.

34. The system according to claim 31, wherein, The feature combination device selects the at least one binning operation through the following process: for each binning feature among the binning features corresponding to the predetermined number of binning operations, a binning composite machine learning model is obtained, the importance of each binning feature is determined based on the effects of the respective binning composite machine learning models, and the at least one binning operation is selected based on the importance of each binning feature, where the binning composite machine learning model includes a binning basic sub-model and a binning additional sub-model based on a boosting framework, where the binning basic sub-model corresponds to the binning basic feature subset, and the binning additional sub-model corresponds to each binning feature.

35. The system according to claim 33 or 34, wherein The binning basic feature subset includes the target combination features selected before the current round of iteration.

36. The system according to claim 25, wherein, The feature combination device performs pre-ranking through the following process: for each candidate combination feature in the candidate combination feature set, a pre-ranking single-feature machine learning model is obtained, and the importance of each candidate combination feature is determined based on the effects of the respective pre-ranking single-feature machine learning models, where the pre-ranking single-feature machine learning model corresponds to each candidate combination feature.

37. The system according to claim 25, wherein, The feature combination device performs pre-ranking through the following process: for each candidate combination feature in the candidate combination feature set, a pre-ranking overall machine learning model is obtained, and the importance of each candidate combination feature is determined based on the effects of the respective pre-ranking overall machine learning models, where the pre-ranking overall machine learning model corresponds to the pre-ranking basic feature subset and each candidate combination feature.

38. The system according to claim 25, wherein The feature combination device performs pre-ranking through the following process: for each candidate combination feature in the candidate combination feature set, a pre-ranking composite machine learning model is obtained, and the importance of each candidate combination feature is determined based on the effects of the respective pre-ranking composite machine learning models, where the pre-ranking composite machine learning model includes a pre-ranking basic sub-model and a pre-ranking additional sub-model based on a boosting framework, where the pre-ranking basic sub-model corresponds to the pre-ranking basic feature subset, and the pre-ranking additional sub-model corresponds to each candidate combination feature.

39. The system according to claim 37 or 38, wherein, The pre-sorted basic feature subset includes the target combined features selected before the current round of iteration.

40. The system according to claim 25, wherein, The feature combination device performs re-sorting through the following process: for each candidate combined feature in the candidate combined feature pool, a re-sorted single-feature machine learning model is obtained, and the importance of each candidate combined feature is determined based on the effects of the respective re-sorted single-feature machine learning models, where the re-sorted single-feature machine learning model corresponds to each of the candidate combined features.

41. The system according to claim 25, wherein, The feature combination device performs re-sorting through the following process: for each candidate combined feature in the candidate combined feature pool, a re-sorted overall machine learning model is obtained, and the importance of each candidate combined feature is determined based on the effects of the respective re-sorted overall machine learning models, where the re-sorted composite machine learning model corresponds to the re-sorted basic feature subset and each of the candidate combined features.

42. The system according to claim 25, wherein The feature combination device performs re-sorting through the following process: for each candidate combined feature in the candidate combined feature pool, a re-sorted composite machine learning model is obtained, and the importance of each candidate combined feature is determined based on the effects of the respective re-sorted composite machine learning models, where the re-sorted composite machine learning model includes a re-sorted basic sub-model based on a boosting framework and a re-sorted additional sub-model, where the re-sorted basic sub-model corresponds to the re-sorted basic feature subset, and the re-sorted additional sub-model corresponds to each of the candidate combined features.

43. The system according to claim 41 or 42, wherein, The pre-sorted basic feature subset includes the target combined features selected before the current round of iteration.

44. The system according to claim 25, wherein, The feature combination device also checks, for each round of iteration, whether the selected target combined features are suitable as combined features of machine learning samples.

45. The system according to claim 44, wherein, The feature combination device uses the change in the effect of the machine learning model based on the target combined features that have passed the check after introducing the selected target combined features to check whether the selected target combined features are suitable as combined features of machine learning samples.

46. The system according to claim 45, wherein, In the case where the check result is that the selected target combined features are suitable as combined features of machine learning samples, the feature combination device takes the selected target combined features as the combined features of machine learning samples and performs the next round of iteration; in the case where the check result is that the selected target combined features are not suitable as combined features of machine learning samples, the feature combination device filters out another part of the candidate combined features from the candidate combined feature set according to the pre-sorted result to form a new candidate combined feature pool.

Citation Information

Patent Citations

  • Method and system for generating combined features of machine learning sample

    CN107679549A