Model Feature Acquisition Method and Apparatus, Electronic Device, Storage Medium

By performing filtered feature selection, gradual regression traversal and optimal feature traversal on the initial candidate feature set, the optimal feature set matching the target model is selected, which solves the problem of low feature screening efficiency in massive feature data, and achieves efficient feature screening and model effect improvement.

CN114510980BActive Publication Date: 2025-06-10TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011184273.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-29
Publication Date
2025-06-10
Estimated Expiration
2040-10-29

AI Technical Summary

Technical Problem

How to efficiently filter out feature data suitable for the target model from massive feature data to avoid interference with the model effect by redundant and invalid features.

Method used

By performing filtered feature selection processing on the initial candidate feature set, gradually regression traversal and optimal feature traversal, and gradually filter out the optimal feature set that matches the target model.

Benefits of technology

It realizes efficient screening of feature data matching the target model from massive feature data, improves feature screening efficiency, reduces information loss, and improves the model effect of the target model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114510980B_ABST
    Figure CN114510980B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a method and apparatus for obtaining model features. The method includes: performing a filtering feature selection process on an initial candidate feature set to filter out initial candidate features in the initial candidate feature set that do not meet the set threshold conditions, obtaining a candidate feature set; performing a stepwise regression traversal on the candidate features included in the candidate feature set to select the candidate feature with the smallest regression coefficient as the target feature, obtaining a target feature set, where the number of target features included in the target feature set meets the set quantity conditions; performing a traversal of the optimal features on the target feature set to obtain an optimal feature set, where the optimal features included in the optimal feature set make the value of the target feature function optimal, and the target feature function is used to evaluate the model effect of the target model. The optimal feature set screened based on the embodiments of the present application can effectively improve the model effect of the target model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing. Specifically, it relates to a method and device for obtaining model features, an electronic device, and a computer-readable storage medium. Background Art

[0002] With the continuous development of computer processing technology, the form of data is becoming more and more complex, and at the same time, the dimension of data is also getting higher and higher. Therefore, how to screen out the feature data applicable to the target model from a large amount of feature data to avoid the interference of redundant and invalid feature data on the model effect of the target model is a problem that technicians need to continuously explore. Summary of the Invention

[0003] To solve the above technical problems, embodiments of the present application provide a method and device for obtaining model features, an electronic device, and a computer-readable storage medium. Based on the embodiments of the present application, model features can be obtained, and the feature data matching the target model can be efficiently screened out from a large amount of feature data.

[0004] According to one aspect of the embodiments of the present application, a method for obtaining model features is provided, including: performing a filtering-based feature selection process on an initial candidate feature set to filter out the initial candidate features in the initial candidate feature set that do not meet the set threshold conditions, obtaining a candidate feature set; performing a stepwise regression traversal on the candidate features included in the candidate feature set to select the candidate feature with the smallest regression coefficient as the target feature, obtaining a target feature set, where the number of target features included in the target feature set meets the set number condition; performing an optimal feature traversal on the target feature set to obtain an optimal feature set, where the optimal features included in the optimal feature set make the value of the target feature function optimal, and the target feature function is used to evaluate the model effect of the target model.

[0005] According to one aspect of the embodiments of the present application, a device for obtaining model features is provided, including: a feature filtering module configured to perform a filtering-based feature selection process on an initial candidate feature set to filter out the initial candidate features in the initial candidate feature set that do not meet the set threshold conditions, obtaining a candidate feature set; a stepwise regression module configured to perform a stepwise regression traversal on the candidate features included in the candidate feature set to select the candidate feature with the smallest regression coefficient as the target feature, obtaining a target feature set, where the number of target features included in the target feature set meets the set number condition; an optimal traversal module configured to perform an optimal feature traversal on the target feature set to obtain an optimal feature set, where the optimal features included in the optimal feature set make the value of the target feature function optimal, and the target feature function is used to evaluate the model effect of the target model.

[0006] According to one aspect of the embodiments of the present application, an electronic device is provided, including a processor and a memory. A computer-readable instruction is stored on the memory, and when the computer-readable instruction is executed by the processor, the above-described model feature acquisition method is implemented.

[0007] According to one aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer-readable instruction is stored. When the computer-readable instruction is executed by a processor of a computer, the computer is caused to execute the above-described model feature acquisition method.

[0008] According to one aspect of the embodiments of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the model feature acquisition method provided in the above various alternative embodiments.

[0009] In the technical solution provided by the embodiments of the present application, first, through a filtering feature selection process on a large number of initial candidate features, a large amount of feature data useless for the target model can be quickly filtered out. Then, by performing stepwise regression traversal on the candidate features obtained in the previous step, the candidate features are screened in combination with the correlation between the candidate features to obtain a target feature set that is further adapted to the target model. Finally, by traversing the optimal features of the target feature set, further feature screening is performed in combination with the influence of the target features on the model effect of the target model, so as to obtain an optimal feature set with a high matching degree with the target model.

[0010] It can be seen that the embodiments of the present application actually provide a feature screening framework, and through this framework, features with a smaller data volume and a higher matching degree with the target model can be gradually screened out, greatly improving the efficiency of feature screening from a large amount of feature data, and at the same time, reducing information loss during the feature screening process.

[0011] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts. In the drawings:

[0013] Figure 1 It is a schematic diagram of an implementation environment related to the present application;

[0014] Figure 2 It is a flowchart of a model feature acquisition method shown in an exemplary embodiment of the present application;

[0015] Figure 3 It is Figure 2 A flowchart of step S110 in the illustrated embodiment in an exemplary embodiment;

[0016] Figure 4 It is Figure 3 A flowchart of step S113 in the illustrated embodiment in an exemplary embodiment;

[0017] Figure 5 It is Figure 2 A flowchart of step S130 in the illustrated embodiment in an exemplary embodiment;

[0018] Figure 6 It is Figure 2 A flowchart of step S150 in the illustrated embodiment in an exemplary embodiment;

[0019] Figure 7 It is a flowchart of a model feature acquisition method shown in another exemplary embodiment;

[0020] Figure 8 It is a schematic flowchart of a model feature acquisition process related to an embodiment of the present application;

[0021] Figure 9 It is a block diagram of a model feature acquisition device shown in an exemplary embodiment of the present application;

[0022] Figure 10 It shows a schematic diagram of the structure of a computer system of an electronic device suitable for implementing an embodiment of the present application. Detailed implementation manners

[0023] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0024] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0025] The flowcharts shown in the drawings are only exemplary descriptions and do not necessarily include all contents and operations / steps, nor do they have to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.

[0026] In this application, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0027] First of all, it should be noted that artificial intelligence (AI) refers to the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is also the study of the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning and decision-making.

[0028] Machine learning (ML) is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence.

[0029] Selecting appropriate model features can not only reduce the feature data required for training a machine learning model, but also alleviate the problem of feature dimensions, enhance the machine learning model's understanding of features, thereby accelerating the learning process of the machine learning model and improving the model performance of the machine learning model. Therefore, it is very necessary to select model features suitable for the machine learning model from a large amount of feature data for the construction of the machine learning model.

[0030] Please refer to Figure 1 , Figure 1It is a schematic diagram of an implementation environment involved in this application. The implementation environment includes a terminal 100 and a data server 200, and the terminal 100 and the data server 200 communicate through a wired or wireless network.

[0031] The machine learning model is deployed on the data server 200, enabling the data server 200 to provide data services related to the machine learning model to the terminal 100. For example, if the machine learning model is a risk control model, the data server 200 provides prediction and control services for user risks to the terminal 100.

[0032] To obtain a machine learning model with better model effects, it is necessary to select model features suitable for the machine learning model from a large amount of feature data, and construct the machine learning model based on the selected model features.

[0033] Therefore, in some embodiments, the data server 200 performs processing methods such as filtering feature selection processing, stepwise regression traversal, and optimal feature traversal on the initial candidate feature set in sequence, so as to screen out the optimal feature set that matches the machine learning model from the initial candidate feature set. Based on the screened optimal feature set, the data server 200 can perform processing such as training the deployed machine learning model, thereby obtaining a machine learning model with better model effects, and further providing better services for the client 100.

[0034] In other embodiments, it is also possible to screen out the optimal feature set that matches the machine learning model from the initial candidate feature set through a dedicated model construction server ( Figure 1 not shown in the figure), train the machine learning model based on the screened optimal feature set, and deploy the trained machine learning model to the data server 200, thereby reducing the data processing pressure on the data server 200.

[0035] It should be noted that the terminal 100 can be any electronic device capable of running a video playback client, such as a smart phone, a tablet, a laptop computer, a computer, etc.; the data server 200 and the model construction server can be independent physical servers, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. This is not limited here.

[0036] Please refer to Figure 2 , Figure 2It is a flowchart of a model feature acquisition method shown in an exemplary embodiment of the present application.

[0037] This method can be applicable to Figure 1 the shown implementation environment. For example, it can be specifically executed by the data server 200 or the model construction server ( Figure 1 not shown in Figure 1 ) in the shown embodiment environment to screen and obtain an optimal feature set that matches the machine learning model deployed in the data server 200 from the initial candidate feature set. This method can also be applicable to any other application scenarios that require model feature screening, and no limitation is imposed here.

[0038] As Figure 2 shown, this method may include steps S110 to S150, which are introduced in detail as follows:

[0039] Step S110: Perform a filtering-based feature selection process on the initial candidate feature set to filter out the initial candidate features in the initial candidate feature set that do not meet the set threshold conditions, and obtain a candidate feature set.

[0040] First, it should be noted that the initial candidate features, candidate features, target features, and optimal features involved in this embodiment are essentially all feature data, but different names are used to distinguish different stages in the feature data screening process, so as to accurately understand the process of screening model features suitable for the target model from a large amount of feature data in this embodiment.

[0041] It should also be noted that the initial candidate feature set refers to the data source to be subjected to feature screening, which contains a large amount of initial candidate features. For example, these large amounts of initial candidate features can include feature data obtained through various methods, and no limitation is imposed here.

[0042] Not all the initial candidate features contained in the initial candidate feature set are necessarily useful for the target model. Therefore, there must be a large number of redundant features in the initial candidate feature set, which will affect the model effect of the target model. Exemplarily, if the initial candidate feature set is input into the target model, due to the too high complexity of the feature space, problems such as the target model being unable to run, the running speed of the target model being too slow to meet the actual application requirements, and the poor model training effect of the target model are likely to occur. Based on this, it is very necessary to perform feature screening on the initial candidate feature set in this embodiment to improve the model effect of the target model.

[0043] In this embodiment, the filtering-based feature selection process for the initial candidate feature set refers to the process of scoring each initial candidate feature according to the divergence of the feature and setting threshold conditions for feature screening. Based on the filtering-based feature selection process performed in this embodiment, the initial candidate features that do not meet the set threshold conditions in the initial candidate feature set are filtered out.

[0044] Exemplarily, the filtering-based feature selection process for the initial candidate feature set in this embodiment may include performing feature screening processing on the initial candidate feature set based on the Information Value (IV) score, and / or performing feature screening processing based on the Population Stability Index (PSI) score.

[0045] Among them, the Information Value score is used to characterize the discrimination degree of the initial candidate feature for the target model. The higher the Information Value score of the initial candidate feature, the higher the discrimination degree of the initial candidate feature for the target model, and the more conducive it is to improving the model effect of the target model. Therefore, the feature screening process for the initial candidate feature set based on the Information Value score in this embodiment is also a process of filtering out the initial candidate features with an Information Value score less than the first score threshold from the initial candidate feature set.

[0046] The Population Stability score is used to characterize the contribution degree of the initial candidate feature to the stability of the target model. The lower the Population Stability score of the initial candidate feature, the higher the contribution degree of the initial candidate feature to the stability of the target model, and the more conducive it is to improving the model effect of the target model. Thus, the feature screening process for the initial candidate feature set based on the Population Stability score in this embodiment is also a process of filtering out the initial candidate features with a Population Stability score greater than or equal to the second score threshold from the initial candidate feature set.

[0047] It can be seen that the filtering-based feature selection process in this embodiment has strong versatility and low complexity, and is applicable to large-scale feature data sets to quickly filter out a large number of initial candidate features that are not relevant to the target model and obtain a candidate feature set.

[0048] Step S130: Perform stepwise regression traversal on the candidate features included in the candidate feature set to select the candidate feature with the smallest regression coefficient as the target feature, and obtain a target feature set, where the number of target features included in the target feature set meets the set number condition.

[0049] In this embodiment, a stepwise regression traversal is performed on the candidate features included in the candidate feature set. This is a feature screening process for the candidate feature set that takes into account the correlation between candidate features. Specifically, a regression model is constructed based on the candidate features in the candidate feature set. If a significant regression model can be obtained, it indicates that the candidate feature is a useful feature for the target model. Therefore, the candidate feature is included in the target feature set as a target feature. If the constructed regression model is not significant, it indicates that the candidate feature is a useless feature for the target model and needs to be filtered out. By repeatedly executing this process, a target feature set with a feature quantity meeting the set quantity condition and having a high correlation with the target model can be obtained.

[0050] Exemplarily, performing a stepwise regression traversal on the candidate features included in the candidate feature set may include a forward stepwise regression traversal of the candidate features and / or a bidirectional stepwise regression traversal of the candidate features.

[0051] Among them, performing a forward stepwise regression traversal on the candidate features means that in each round of forward stepwise regression process, the candidate feature with the smallest regression coefficient and a regression coefficient less than the first regression threshold is selected as the target feature and included in the target feature set. Thus, a target feature set containing a specified number of target features is obtained. The obtained target feature set has a high correlation with the target model, thereby improving the model effect of the target model.

[0052] Performing a bidirectional stepwise regression traversal on the candidate features means that in each round of bidirectional stepwise regression process, not only the candidate feature with the smallest regression coefficient and a regression coefficient less than the first regression threshold is selected as the target feature and included in the target feature set, but also the target feature in the target feature set with the largest regression coefficient and a regression coefficient greater than the second regression threshold is deleted from the target feature set. Thus, a target feature set containing a specified number of target features is obtained. The obtained target feature set has a higher correlation with the target model to further improve the model effect of the target model.

[0053] Therefore, in this embodiment, a stepwise regression traversal is performed on the candidate features included in the candidate feature set, taking into account the correlation between candidate features. This makes the target feature set obtained by feature screening of the candidate feature set in this embodiment have a very high correlation with the target model. The target model constructed based on the target feature set obtained in this embodiment also has a better model effect compared to the target model constructed based on the candidate set.

[0054] Step S150: Perform a traversal of the optimal features on the target feature set to obtain an optimal feature set. The optimal features included in the optimal feature set make the value of the target feature function optimal. The target feature function is used to evaluate the model effect of the target model.

[0055] In this embodiment, traversing the optimal features of the target feature set is a process of scoring the feature subsets in the target feature set based on the model effect of the target model and selecting the feature subset with the best model effect as the optimal feature set.

[0056] It should be noted that the model effect of the target model can be specifically evaluated according to the target feature function. For example, if the AUC (Area Under Curve, defined as the area enclosed by the ROC (Receiver Operating Characteristic) curve and the coordinate axes) value is used to evaluate the model effect of the target model, the target feature function can be the calculation function corresponding to the AUC value. The optimal target feature function means that the value of the target feature function is the largest.

[0057] If the KS (Kolmogorov-Smirnov Curve, K-S test) value is used to evaluate the model effect of the target model, the target feature function can be the calculation function corresponding to the KS value. Correspondingly, the optimal target feature function means that the value of the target feature function is optimal. Generally speaking, the larger the KS value, the better the target model can distinguish between positive and negative samples. However, in some specific cases, if the KS value is very large, it means that the target model over-distinguishes between positive and negative samples, which instead makes the model effect of the target model poor. Therefore, if the KS value of the target model can reach the optimal score range, for example, the optimal score range is 61% - 75%, it is determined that the value of the target feature function is optimal.

[0058] Exemplarily, traversing the optimal features of the target feature set in this embodiment may include a process of performing a sequential forward selection process on the target feature set and / or a sequential floating forward selection process on the target feature set.

[0059] Among them, performing a sequential forward selection process on the target feature set means that in each round of the sequential forward selection process, a target feature that makes the value of the target feature function optimal is selected and added to the optimal feature set. Any of the obtained optimal features will make the target model have a better model effect. Therefore, the target model constructed based on the optimal feature set will also have a good model effect.

[0060] Performing sequential floating forward selection processing on the target feature set means that in each round of sequential forward selection processing, not only is a target feature that optimizes the value of the target feature function selected and added to the optimal feature set, but also an optimal feature is removed from the optimal feature set. After this optimal feature is removed, the remaining optimal features can still optimize the value of the target feature function. Thus, by performing sequential floating forward selection processing on the target feature set in this embodiment, further feature screening can be performed on the target feature set, and the target model constructed based on the optimal feature set will also have better model performance.

[0061] In this embodiment, by traversing the optimal features of the target feature set, the obtained optimal feature set usually has good classification performance, so it is very suitable for some classification models, such as risk control models. Risk control models usually have strong interpretability, and the correlation between model features also affects the interpretability of model features. The feature screening performed in step S130 of this embodiment, that is, the model feature screening implemented based on the correlation between model features, enables the optimal feature set finally obtained in this embodiment to achieve effects such as reducing the overfitting risk of the risk control model, shortening the model training time, and avoiding the feature dimension disaster.

[0062] Moreover, the model feature acquisition performed in this embodiment can be used as a feature screening framework to gradually screen out features with a smaller data volume and a higher matching degree with the target model through this framework. At the same time, the involved feature data processing volume gradually becomes more complex, greatly improving the efficiency of feature screening from a large amount of feature data and also reducing information loss during the feature screening process.

[0063] Figure 3 is Figure 2 The flowchart of step S110 in the illustrated embodiment in an exemplary embodiment.

[0064] As Figure 3 As shown, performing filter-based feature selection processing on the initial candidate feature set to filter out the initial candidate features in the initial candidate feature set that do not meet the set threshold conditions may include steps S111 to S113, which are introduced in detail as follows:

[0065] Step S111, calculate the information value scores of the initial candidate features contained in the initial candidate feature set.

[0066] First, it should be noted that the initial candidate features contained in the initial candidate set can usually be classified into two categories: continuous initial candidate features and nominal initial candidate features. Among them, the values of continuous initial candidate features are continuous numerical values, and the values of nominal initial candidate features are discrete nominal variables.

[0067] For the continuous initial candidate features in the initial candidate set, it is impossible to directly evaluate the discrimination of the continuous initial candidate features for the target model. It is necessary to discretize the continuous initial candidate features and then evaluate the information value score of the continuous initial candidate features based on the results obtained from the discretization process.

[0068] Exemplarily, first, it is necessary to perform feature binning on the continuous initial candidate features to obtain multiple feature bins corresponding to the continuous initial candidate features. Exemplarily, feature binning can be performed on the continuous initial candidate features in ways such as equal-frequency binning, equal-distance binning, decision tree binning, etc., and this is not restricted herein.

[0069] For the nominal initial candidate features in the initial candidate feature set, feature binning can be performed according to the actual value-taking situations. For example, if the number of discrete values contained in the nominal initial candidate feature is small, each discrete value can be used as a feature bin; if the number of discrete values contained in the nominal initial candidate feature is large, the nominal initial candidate feature is first converted into a one-hot (unique hot code) feature, and then feature binning is performed on the obtained one-hot feature. Thus, the sum of the information value scores of each feature bin can be used as the information value score of the nominal initial candidate feature.

[0070] After obtaining multiple feature bins corresponding to the initial candidate features, calculate the information value scores of each feature bin according to the number of positive samples and negative samples contained in each feature bin, and the sum of the information value scores of each feature bin can be used as the information value score of the continuous initial candidate feature. Among them, each initial candidate feature usually carries multiple user labels corresponding to different users, and the user labels are usually "label 0" or "label 1". Therefore, the positive samples and negative samples contained in each feature bin can be determined based on the type to which the user label belongs.

[0071] For easy understanding, for example, the initial candidate feature simultaneously contains feature information corresponding to multiple users, and the user label corresponding to each user is "label 0" or "label 1". After performing feature binning on the initial candidate feature, the discrete data corresponding to different users are contained in each obtained feature bin. Therefore, the discrete data corresponding to "label 0" are used as negative samples, and the discrete data corresponding to "label 1" are used as positive samples.

[0072] The information value scores of each feature bin can be calculated through the following formula:

[0073]

[0074] If the initial candidate features are subjected to feature binning to obtain n feature bins, IV i represents the information value score within the i-th feature bin, represents the ratio of the number of bad samples in the i-th feature bin to the total number of bad samples, represents the ratio of the number of good samples in the i-th feature bin to the total number of good samples. Among them, the total number of bad samples is the total amount of bad samples contained in the n feature bins, and the total number of good samples is the total amount of bad samples contained in the n feature bins.

[0075] After obtaining the information value scores of each feature bin, the information value score IV of the initial candidate features can be calculated based on the following formula:

[0076]

[0077] Step S113, filter out the initial candidate features with information value scores less than the first score threshold from the initial candidate feature set to obtain the candidate feature set.

[0078] As described in the foregoing embodiments, the information value score is used to characterize the discrimination degree of the initial candidate features for the target model. The higher the information value score of the initial candidate features, the higher the discrimination degree of the initial candidate features for the target model, and the more conducive to improving the model effect of the target model.

[0079] For example, usually, the initial candidate features with information value scores less than or equal to 0.02 are determined to have no discrimination degree for the target model; the initial candidate features with information value scores greater than 0.02 and less than or equal to 0.1 are determined to have weak discrimination degree for the target model; the initial candidate features with information value scores greater than 0.1 and less than or equal to 0.3 are determined to have moderate discrimination degree for the target model; the initial candidate features with information value scores greater than 0.3 are determined to have strong discrimination degree for the target model.

[0080] However, since the information value score is related to the number of feature bins, it is not entirely the case that the larger the information value score, the higher the discrimination degree of the initial candidate features for the target model. Usually, when the information value score is greater than 0.5, it may be questioned that the discrimination degree of the initial candidate features for the target model is too high and not realistic enough. Therefore, in the process of feature screening of the initial candidate feature set based on the information value score of the initial candidate features in this embodiment, the first score threshold needs to be set to a smaller value, such as 0.01, or it can be adjusted according to the actual situation to appropriately relax the set threshold conditions, so that too much feature data will not be lost in this feature screening stage.

[0081] It can be seen from this that in this embodiment, the initial candidate features with information value scores less than the first score threshold are filtered out from the initial candidate feature set, which can not only quickly screen out a candidate feature set with high discrimination for the target model, but also will not cause the problem of a large amount of loss of feature data, which is beneficial to the subsequent stage of feature screening processing.

[0082] Figure 4 Yes Figure 3 Step S113 in the illustrated embodiment is a flowchart of an exemplary embodiment.

[0083] As Figure 4 As shown, filtering out the initial candidate features with information value scores less than the first score threshold from the initial candidate feature set to obtain a candidate feature set may include steps S210 to S230, which are introduced in detail as follows:

[0084] Step S210, after filtering out the initial candidate features with information value scores less than the first score threshold from the initial candidate feature set, calculate the population stability scores of each initial candidate feature in the filtered initial candidate feature set.

[0085] First of all, it should be noted that in this embodiment, based on filtering out a large number of initial candidate features irrelevant to the target model from the initial candidate feature set according to the information value score, the remaining initial candidate features are further screened according to the population stability score to further filter out the initial candidate features with poor stability.

[0086] The population stability scores of each initial candidate feature in the filtered initial candidate feature set can be calculated through the following process:

[0087] First, the filtered initial candidate feature set needs to be divided into a training sample set and a test sample set. Exemplarily, each user corresponds to multiple initial candidate features, so the training sample set and the test sample set can be divided based on different users. For example, it can be randomly determined that each user corresponds to the training sample set or the test sample set, and then the multiple initial candidate features corresponding to the user are included in the corresponding training sample set or test sample set.

[0088] Therefore, each user in this embodiment will correspond to a data set label, and this data set label corresponds to the training sample set or the test sample set. And since each initial candidate feature corresponds to multiple users at the same time, each initial candidate feature will carry multiple data set labels.

[0089] Then, binning processing is performed on each initial candidate feature in the filtered initial candidate feature set to obtain multiple feature bins corresponding to each initial candidate feature. In this embodiment, the chi-square binning method can be used to perform feature binning processing on continuous initial candidate features, or other feature binning processing methods can be used. The discrete initial candidate features can be processed using the method of the foregoing embodiment, and this is not limited herein.

[0090] Finally, according to the distribution ratio of the training samples and test samples contained in each feature bin, the population stability score of each feature bin is calculated, and thus the sum of the population stability scores of each feature bin is used as the population stability score of the corresponding initial candidate feature.

[0091] It should be noted that the training samples and test samples contained in each feature bin can be determined based on the dataset labels. Still taking an example for illustration, the initial candidate feature contains feature information corresponding to multiple users, and the dataset label corresponding to each user is "training sample set" or "test sample set". After performing feature binning processing on the initial candidate feature, the discrete data corresponding to different users are contained in each obtained feature bin. Therefore, the discrete data corresponding to the "training sample set" are used as training samples, and the discrete data corresponding to the "test sample set" are used as test samples.

[0092] Exemplarily, the population stability score of each feature bin can be calculated through the following formula:

[0093]

[0094] If n feature bins are obtained after performing feature binning processing on the initial candidate feature, PSI i represents the population stability score within the i-th feature bin, represents the ratio of the number of test samples in the i-th feature bin to the total number of test samples, represents the ratio of the number of training samples in the i-th feature bin to the total number of training samples.

[0095] After obtaining the population stability scores of each feature bin, the population stability score PSI of the initial candidate feature can be calculated based on the following formula:

[0096]

[0097] Step S230, select the initial candidate features with population stability scores less than the second score threshold as candidate features to obtain a candidate feature set.

[0098] As described in the foregoing embodiments, the group stability score is used to characterize the contribution degree of the initial candidate features to the stability of the target model. The lower the group stability score of the initial candidate features, the higher the contribution degree of the initial candidate features to the stability of the target model, and the more conducive it is to improving the model effect of the target model.

[0099] For example, usually, the initial candidate features with a group stability score less than 10% are determined to be able to make the target model stable; the initial candidate features with a group stability score greater than or equal to 10% and less than or equal to 20% are determined to be able to make the target model generally stable; the initial candidate features with a group stability score greater than 20% are determined to have poor stability of the target model. In actual situations, the threshold of the group stability score for different stability degrees can be adaptively adjusted.

[0100] In the embodiment, the second score threshold can be set to 10% so that the initial candidate features with a group stability score less than the second score threshold are selected as candidate features in this embodiment, and the obtained candidate feature set can make the target model have better stability, thereby being able to improve the model effect of the target model.

[0101] It should be noted that the process of feature screening according to the group stability score in this embodiment is after the process of feature screening according to the information value score, so as to quickly filter out a large number of features irrelevant to the target model based on the information value score first, and then filter out some features with poor stability based on the information value score. The two are combined with each other to ensure that the selected candidate feature set is more suitable for the target model and makes the target model have a higher model effect.

[0102] Figure 5 Yes Figure 2 The flowchart of step S130 in the illustrated embodiment in an exemplary embodiment.

[0103] As Figure 5 As shown, performing stepwise regression traversal on the candidate features included in the candidate feature set may include steps S131 to S133, which are introduced in detail as follows:

[0104] Step S131, performing forward stepwise regression traversal and two-way stepwise regression traversal on the candidate features included in the candidate feature set respectively to obtain a specified number of target features respectively.

[0105] First of all, it should be noted that the forward stepwise regression traversal of the candidate features included in the candidate feature set in this embodiment may include the following process:

[0106] In each round of forward stepwise regression process, candidate features are selected without replacement from the candidate feature set. A first regression model is constructed based on the selected candidate features and the selected candidate feature subset, and the first regression model is verified based on the candidate feature set to obtain the first regression coefficients of each candidate feature contained in the candidate feature set.

[0107] If the smallest first regression coefficient obtained in each round of forward stepwise regression process is less than the first regression threshold, the candidate feature corresponding to the smallest first regression coefficient is merged into the selected candidate feature subset, and the candidate feature corresponding to the smallest first regression coefficient is deleted from the candidate feature set, so as to update the selected candidate feature subset and the candidate feature set in the next round based on the updated selected candidate feature subset and the candidate feature set.

[0108] If it is determined that the number of candidate features contained in the currently updated selected candidate feature subset reaches the specified number, the candidate features contained in the currently updated selected candidate feature subset are used as target features.

[0109] It should be understood that in each round of forward stepwise regression process, the selected candidate feature subset refers to the feature set composed of the candidate features selected in the previous forward stepwise regression process. Therefore, during the forward stepwise regression traversal, the number of features contained in the selected candidate feature subset will continuously increase.

[0110] To facilitate the understanding of the above forward stepwise regression traversal process, the following will describe the forward stepwise regression traversal process through an exemplary implementation manner:

[0111] First, in any round of forward stepwise regression process, a candidate feature x (x ∈ M) is randomly selected without replacement from the candidate feature set M, and the candidate feature x is merged with the selected candidate feature subset N to construct a first regression model. The first regression model is verified according to the candidate feature set M to obtain the first regression coefficients ρ of each candidate feature contained in the candidate feature set M. Then, the candidate feature k (k ∈ M) with the smallest ρ value is found. If the smallest ρ value is less than the preset first regression threshold e, it means that the candidate feature k is a useful feature for the target model. Therefore, the candidate feature k is merged into the selected candidate feature subset N, and the candidate feature k is deleted from the candidate feature set M, thereby realizing the update of the selected candidate feature subset N and the candidate feature set M. Based on the updated N and M, the next round of forward stepwise regression process can be continued.

[0112] Repeat the above process until the number of features in N reaches the preset specified number, that is, the forward stepwise regression traversal of the candidate features contained in the candidate feature set is completed to obtain the specified number of target features.

[0113] It should also be noted that the process of performing two-way stepwise regression traversal on the candidate features contained in the candidate feature set in this embodiment is similar to, but not exactly the same as, the process of performing forward stepwise regression traversal. The detailed process is as follows:

[0114] In any round of two-way stepwise regression process, candidate features are still selected from the candidate feature set without replacement. A first regression model is constructed based on the selected candidate features and the selected candidate feature subset, and the first regression model is verified based on the candidate feature set to obtain the first regression coefficients of each candidate feature contained in the candidate feature set. If the smallest first regression coefficient obtained is less than the first regression threshold, the candidate feature corresponding to the smallest first regression coefficient is merged into the selected candidate feature subset, and the candidate feature corresponding to the smallest first regression coefficient is deleted from the candidate feature set. This process is the same as the forward stepwise regression process.

[0115] However, in each round of two-way stepwise regression process, a second regression model is also constructed based on the updated selected candidate feature subset, and the second regression model is verified based on the updated selected candidate feature subset. If the largest second regression coefficient obtained in each round is greater than the second regression threshold, the candidate feature corresponding to the largest second regression coefficient is deleted from the updated selected candidate feature subset to further update the selected candidate feature subset. Then, based on the candidate feature set updated in each round and the further updated selected candidate feature subset, the next round of updating of the selected candidate feature subset and the candidate feature set is performed until the number of candidate features contained in the updated selected candidate feature subset reaches the specified number, and thus the specified number of candidate features are used as target features.

[0116] That is to say, after each round of two-way stepwise regression process finishes the process of each round of forward stepwise regression, that is, after merging candidate feature k into the selected candidate feature subset N and deleting candidate feature subset k from the candidate feature set M, a second regression model is also constructed for the updated selected candidate feature subset N and statistical verification is performed to find the candidate feature y (y ∈ N) with the largest second regression coefficient. If the second regression coefficient of candidate feature y is greater than the preset second regression threshold f, it means that candidate feature y is a useless feature for the target model, so the candidate feature needs to be deleted from the selected candidate feature subset N. Based on the updated N and M, the next round of two-way stepwise regression process can continue.

[0117] Repeat the above process until the number of features in N reaches the preset specified number, that is, the two-way stepwise regression traversal of the candidate features contained in the candidate feature set is completed to obtain the specified number of target features.

[0118] Step S133: Perform deduplication on all the obtained target features to obtain a target feature set with duplicate target features filtered out.

[0119] It can be seen that during each round of forward stepwise regression traversal, the features in the selected candidate feature subset only enter and do not exit, so the traversal speed is relatively fast, but redundant features may be filtered out; while during each round of bidirectional stepwise regression traversal, new features will be added to the selected candidate feature subset and features will also be deleted, so the traversal speed is relatively slow, but the features filtered out can better match the target model.

[0120] In this embodiment, forward stepwise regression traversal and bidirectional stepwise regression traversal are respectively performed on the candidate features included in the candidate feature set to respectively obtain a specified number of target features, and then deduplication is performed on all the obtained target features to obtain a target feature set with duplicate target features filtered out. The target feature set obtained in this embodiment combines the advantages of forward stepwise regression traversal and bidirectional stepwise regression traversal, so it is more suitable for the target model, thereby further improving the model effect of the target model.

[0121] Figure 6 Yes Figure 2 The flowchart of step S150 in the illustrated embodiment is in a flowchart of an exemplary embodiment.

[0122] Such as Figure 6 As shown, traversing the optimal features of the target feature set to obtain the optimal feature set may include steps S151 to S153, which are introduced in detail as follows:

[0123] Step S151: Perform sequential forward selection processing and sequential floating forward selection processing on the target feature set respectively to select the optimal features from the target feature set.

[0124] First, the process of performing sequential forward selection processing on the target feature set includes the following:

[0125] During each round of sequential forward selection processing, select target features from the target feature set. If the combination of the selected target features and the selected target feature subset makes the value of the target feature function optimal, then add the selected target features as the optimal features to the selected target feature subset until all the target features in the target feature set are selected, and use the finally updated selected target feature subset as the optimal feature set.

[0126] It should be understood that during each round of sequential forward selection processing, the selected target feature subset refers to the feature set composed of the target features selected during the historical sequential forward selection processing.

[0127] For example, in each round of sequential forward selection processing, a target feature a is selected from the target feature set and added to the selected target feature subset B to optimize the target feature function J(B + a). Simply put, in each round, a target feature that optimizes the value of the target feature function is selected from the target feature set and added to the optimal feature set until all the target features in the target feature set are selected. Among them, the optimization of the target feature function J(B + a) is used to characterize that the selected target feature a enables the target model to have the optimal model effect.

[0128] Performing sequential floating forward selection processing on the target feature set includes the following process:

[0129] In each round of sequential forward selection processing, a target feature is selected from the target feature set. If the combination of the selected target feature and the selected target feature subset optimizes the value of the target feature function, the selected target feature is added to the selected target feature subset;

[0130] A target feature is selected again from the updated selected target feature subset in each round. If the other target features after filtering out the re-selected target feature optimize the value of the target feature function, the re-selected target feature is deleted from the updated selected target feature subset to further update the selected target feature subset;

[0131] Based on the further updated selected target feature subset in each round, the next round of updating of the selected target feature subset is performed until all the target features in the target feature set are selected, and the finally updated selected target feature subset is used as the optimal feature set.

[0132] Specifically, after each round of sequential forward selection processing is completed in each round of sequential floating forward selection processing, that is, after a target feature a is selected from the target feature set and added to the selected target feature subset B in each round to optimize the target feature function J(B + a), a target feature c is also selected from the currently updated selected target feature subset B to optimize the target feature function J(B - c), and then the target feature c is deleted from the currently updated selected target feature subset B.

[0133] It should be understood that the optimization of the target feature function J(B - c) is used to characterize that even if the target feature c is deleted from the selected target feature subset B, the model effect of the target model is still optimized. Therefore, the target feature c is a useless feature for the target model. Therefore, deleting the target feature c can increase the improvement of the model effect of the obtained optimal feature set on the target model.

[0134] Step S153: Remove duplicates from all the selected optimal features to obtain an optimal feature set that filters out duplicate optimal features.

[0135] In each round of the sequential forward selection process, the features in the selected target feature subset only enter and do not exit, so the screening speed is relatively fast, but redundant features may be screened out; while in each round of the sequential floating forward selection process, new features are added to the selected target feature subset and some features are also deleted, so the screening speed is slower, but the screened features can better match the target model.

[0136] Therefore, the optimal feature set obtained in this embodiment combines the advantages of both the sequential forward selection process and the sequential floating forward selection process, so it is more suitable for the target model and can further improve the model effect of the target model.

[0137] Figure 7 It is a flowchart of a model feature acquisition method shown according to another exemplary embodiment.

[0138] As Figure 7 shown, on the basis of the embodiment shown in Figure 2 , this model feature method further includes steps S310 to S330, which are introduced in detail as follows:

[0139] Step S310: Evaluate the model effect of the target model according to the optimal feature set to obtain the evaluation index corresponding to the target model.

[0140] First of all, it should be noted that in this embodiment, the model effect of the target model is evaluated according to the optimal feature set, which is a process of running the target model according to the optimal feature set and calculating the evaluation index corresponding to the target model based on the running result of the target model to evaluate the model effect based on the evaluation index. As mentioned above, the evaluation index can be the AUC value or the KS value.

[0141] Taking the evaluation of the model effect of the target model based on the KS value as an example, the KS value is calculated based on the KS curve, where the horizontal axis of the KS curve is different classification thresholds and the vertical axis is the true positive rate TP r and the false positive rate FP r change curve, and the calculation formula of the KS value is as follows:

[0142] KS = max|TP r - FP r |

[0143] It can be seen that the KS value is specifically the maximum value of the difference between the true positive rate TP r and the false positive rate FP r represented by the vertical axis of the KS curve. If the binary classification scenario is used as an example, the true positive rate TPr and false positive rate FP r They can be calculated using the following formulas:

[0144]

[0145] Among them, TP represents the number of positive samples correctly predicted as positive samples by the target model, FN represents the number of positive samples incorrectly predicted as negative samples by the target model, FP represents the number of negative samples incorrectly predicted as positive samples by the target model, and TN represents the number of negative samples correctly predicted as negative samples by the target model.

[0146] Step S330: If the evaluation index does not meet the set index threshold range, the set threshold condition is adjusted to reselect the optimal feature set from the initial candidate feature set based on the adjusted set threshold condition.

[0147] Generally speaking, the larger the KS value, the higher the target model's ability to distinguish between positive and negative samples, and the better the classification effect of the target model. However, not all cases indicate that the higher the KS value, the better the classification effect of the target model. For example, in a credit reporting model, if the credit reporting model completely misclassifies positive and negative samples, the resulting KS value is still high, and the credit reporting model expects the resulting credit score distribution to be normally distributed. However, if the KS value is too large, such as 90%, it means that the credit reporting model distinguishes too much between positive and negative samples, and a normally distributed credit score distribution cannot be obtained, and therefore a better model effect cannot be obtained.

[0148] For example, the classification effect of the target model corresponding to the KS value can be shown in the following Table 1:

[0149]

[0150]

[0151] Table 1

[0152] Based on this, the indicator threshold range can be set according to the expected model effect of the target model. For example, in scenarios with high requirements for classification effect, the indicator threshold range can be set to 61%-75%, and in some scenarios with lower requirements for classification effect, the indicator threshold range can be set to 51%-60%. This is not restricted here.

[0153] If the evaluation metrics obtained based on step S310 do not meet the set metric threshold range, it indicates that the optimal feature set obtained by screening cannot enable the target model to have good model performance. Therefore, it is necessary to adjust the screening parameters involved in the feature screening process, such as the set first score threshold, second score threshold, first regression threshold, and second regression threshold, so as to reselect the optimal feature set from the initial candidate feature set based on the adjusted set threshold conditions until the optimal feature set finally obtained enables the target model to have better model performance.

[0154] Therefore, in this embodiment, the model performance of the target model is evaluated according to the optimal feature set obtained by screening, so as to further judge the feature performance of the optimal feature set obtained by screening based on the obtained evaluation metrics, thereby ensuring that the finally obtained optimal feature set has a very high matching degree with the target model, and the target model will also obtain very good model performance by running the finally obtained optimal feature set.

[0155] Figure 8 It is a schematic flowchart of a model feature acquisition process involved in an embodiment of the present application.

[0156] Such as Figure 8 shown, this exemplary model feature acquisition process includes steps S10 to S60, which are introduced in detail as follows:

[0157] First, feature screening is performed on the initial candidate feature set according to the feature IV (i.e., information value score) value, specifically selecting the initial candidate features with an IV value greater than or equal to 0.01. Then, feature screening is performed on the initial candidate feature set that has filtered out the initial candidate features with an IV value less than 0.01 according to the PSI (i.e., population stability score) value, specifically selecting the initial candidate features with a PSI value less than 10%, so as to obtain a candidate feature set composed of the selected initial candidate features.

[0158] Then, forward stepwise regression traversal and bidirectional stepwise regression traversal are respectively performed on the candidate feature set to respectively obtain a specified number of target features, and duplicate removal processing is performed on all the obtained target features to obtain a target feature set that has filtered out the duplicate target features.

[0159] Finally, sequential forward selection processing and sequential floating forward selection processing are respectively performed on the target feature set to select the optimal features from the target feature set, and duplicate removal processing is performed on all the selected optimal features to obtain an optimal feature set that has filtered out the duplicate optimal features.

[0160] Based on this process, the optimal feature set that matches the target model can be screened from the initial candidate feature set. Running the target model based on the obtained optimal feature combination can make the target model have better model performance.

[0161] In the above embodiments of the present application, the feature screening process can be specifically implemented by programming in Python (an interpreted computer coding language), or can also be implemented by programming in other types of computer languages, and this is not limited herein.

[0162] In the application scenario where the target model is a risk control model, for example, the risk control model can be a credit investigation model, etc. Traditional risk control models are all implemented based on SAS (Statistical Analysis System) language, and only recently have they gradually been converted to be implemented based on Python language. Therefore, the present application's implementation of model feature screening based on Python language can achieve the conversion of risk control model feature screening from SAS language to Python language.

[0163] Moreover, the present application can form a fixed and automated model feature screening framework, which is beneficial for screening a large number of features, and the stability of the feature screening effect is good. It also greatly improves the efficiency of screening features from a large number of features, and at the same time reduces the information loss during the feature screening process.

[0164] Figure 9 It is a block diagram of a model feature acquisition device shown in an exemplary embodiment of the present application.

[0165] As Figure 9 shown, the model feature acquisition device includes:

[0166] A feature filtering module 410 configured to perform a filtering-based feature selection process on the initial candidate feature set to filter out the initial candidate features in the initial candidate feature set that do not meet the set threshold conditions, and obtain a candidate feature set; a stepwise regression module 430 configured to perform a stepwise regression traversal on the candidate features included in the candidate feature set to select the candidate feature with the smallest regression coefficient as the target feature, and obtain a target feature set, where the number of target features included in the target feature set meets the set number condition; an optimal traversal module 450 configured to perform an optimal feature traversal on the target feature set to obtain an optimal feature set, where the optimal features included in the optimal feature set make the value of the target feature function optimal, and the target feature function is used to evaluate the model performance of the target model.

[0167] In another exemplary embodiment, the feature filtering module 410 includes:

[0168] An information value score calculation unit, configured to calculate the information value scores of the initial candidate features included in the initial candidate feature set; an information value score screening unit, configured to filter out the initial candidate features with information value scores less than a first score threshold from the initial candidate feature set to obtain a candidate feature set.

[0169] In another exemplary embodiment, the information value score calculation unit includes:

[0170] A first feature binning unit, configured to perform feature binning on the initial candidate features included in the initial candidate set to obtain a plurality of feature bins corresponding to the initial candidate features; a first bin score calculation subunit, configured to calculate the information value scores of the respective feature bins according to the numbers of positive samples and negative samples included in the respective feature bins, and use the sum of the information value scores of the respective feature bins as the information value score of the initial candidate feature.

[0171] In another exemplary embodiment, the information value score screening unit includes:

[0172] A population stability calculation subunit, configured to calculate the population stability scores of the respective initial candidate features in the filtered initial candidate feature set after filtering out the initial candidate features with information value scores less than a first score threshold from the initial candidate feature set; a population stability screening subunit, configured to select the initial candidate features with population stability scores less than a second score threshold as candidate features to obtain a candidate feature set.

[0173] In another exemplary embodiment, the population stability calculation subunit includes:

[0174] A sample set partitioning subunit, configured to partition the filtered initial candidate feature set into a training sample set and a test sample set; a second feature binning unit, configured to perform binning on the respective initial candidate features in the filtered initial candidate feature set to obtain a plurality of feature bins corresponding to the respective initial candidate features; a second bin score calculation subunit, configured to calculate the population stability scores of the respective feature bins according to the distribution ratios of the training samples and test samples included in the respective feature bins, and use the sum of the population stability scores of the respective feature bins as the population stability score of the corresponding initial candidate feature.

[0175] In another exemplary embodiment, the stepwise regression module 430 includes:

[0176] A regression traversal unit, configured to perform forward stepwise regression traversal and bidirectional stepwise regression traversal on the candidate features included in the candidate feature set respectively to obtain a specified number of target features; a target feature deduplication unit, configured to perform deduplication processing on all the obtained target features to obtain a target feature set after filtering out the duplicate target features.

[0177] In another exemplary embodiment, the regression traversal unit includes:

[0178] A first regression verification subunit, configured to, in each round of forward stepwise regression process, select candidate features from the candidate feature set without replacement, construct a first regression model according to the selected candidate features and the selected candidate feature subset, and verify the first regression model based on the candidate feature set to obtain the first regression coefficients of the respective candidate features included in the candidate feature set; a first screening and updating subunit, configured to, if the smallest first regression coefficient obtained in each round of forward stepwise regression process is less than the first regression threshold, merge the candidate feature corresponding to the smallest first regression coefficient into the selected candidate feature subset, and delete the candidate feature corresponding to the smallest first regression coefficient from the candidate feature set, so as to update the selected candidate feature subset and the candidate feature set in the next round based on the updated selected candidate feature subset and the candidate feature set; a first target feature obtaining subunit, configured to, if it is determined that the number of candidate features included in the currently updated selected candidate feature subset reaches the specified number, use the candidate features included in the currently updated selected candidate feature subset as target features.

[0179] In another exemplary embodiment, the regression traversal unit includes:

[0180] A second regression verification subunit, configured to, in each round of bidirectional stepwise regression process, select candidate features from the candidate feature set without replacement, construct a first regression model according to the selected candidate features and the selected candidate feature subset, and verify the first regression model based on the candidate feature set to obtain the first regression coefficients of the respective candidate features included in the candidate feature set; a second screening and updating subunit, configured to, if the smallest first regression coefficient obtained in each round of bidirectional stepwise regression process is less than the first regression threshold, merge the candidate feature corresponding to the smallest first regression coefficient into the selected candidate feature subset, and delete the candidate feature corresponding to the smallest first regression coefficient from the candidate feature set; a re-regression verification subunit, configured to construct a second regression model according to the updated selected candidate feature subset, and verify the second regression model based on the updated selected candidate feature subset, and if the largest second regression coefficient obtained in each round is greater than the second regression threshold, delete the candidate feature corresponding to the largest second regression coefficient from the updated selected candidate feature subset, so as to further update the selected candidate feature subset; a second target feature obtaining subunit, configured to, based on the candidate feature set updated in each round and the further updated selected candidate feature subset, update the selected candidate feature subset and the candidate feature set in the next round until the number of candidate features included in the updated selected candidate feature subset reaches the specified number, and use the specified number of candidate features as target features.

[0181] In another exemplary embodiment, the optimal traversal module 450 includes:

[0182] A forward selection processing unit configured to perform sequential forward selection processing and sequential floating forward selection processing on the target feature set respectively to select optimal features from the target feature set; an optimal feature deduplication unit configured to perform deduplication processing on all the selected optimal features to obtain an optimal feature set with duplicate optimal features filtered out.

[0183] In another exemplary embodiment, the forward selection processing unit includes:

[0184] A first target feature selection subunit configured to select a target feature from the target feature set in each round of sequential forward selection processing. If the combination of the selected target feature and the selected target feature subset makes the value of the target feature function optimal, the selected target feature is added as an optimal feature to the selected target feature subset until all the target features in the target feature set are selected; a first optimal feature set acquisition subunit configured to use the finally updated selected target feature subset as the optimal feature set.

[0185] In another exemplary embodiment, the forward selection processing unit includes:

[0186] A second target feature selection subunit configured to select a target feature from the target feature set in each round of sequential forward selection processing. If the combination of the selected target feature and the selected target feature subset makes the value of the target feature function optimal, the selected target feature is added to the selected target feature subset; a target feature re-selection subunit configured to re-select a target feature from the updated selected target feature subset in each round. If the other target features after filtering out the re-selected target feature make the value of the target feature function optimal, the re-selected target feature is deleted from the updated selected target feature subset to further update the selected target feature subset; a second optimal feature acquisition subunit configured to update the next-round selected target feature subset based on the further updated selected target feature subset in each round until all the target features in the target feature set are selected, and use the finally updated selected target feature subset as the optimal feature set.

[0187] In another exemplary embodiment, the model feature acquisition device further includes:

[0188] A model evaluation module configured to evaluate the model effect of the target model according to the optimal feature set to obtain an evaluation index corresponding to the target model; a threshold adjustment module configured to adjust the set threshold condition if the evaluation index does not meet the set index threshold interval, so as to re-select the optimal feature set from the initial candidate feature set based on the adjusted set threshold condition.

[0189] It should be noted that the device provided in the above embodiments and the method provided in the above embodiments belong to the same concept. The specific manners in which each module and unit perform operations have been described in detail in the method embodiments and will not be elaborated here.

[0190] An embodiment of the present application further provides an electronic device, including a processor and a memory. Among them, computer-readable instructions are stored on the memory, and when the computer-readable instructions are executed by the processor, the method for obtaining model features as described above is implemented.

[0191] Figure 10 The structural schematic diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown.

[0192] It should be noted that Figure 10 The computer system 1600 of the electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0193] As Figure 10 shown, the computer system 1600 includes a central processing unit (CPU) 1601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1602 or the program loaded from the storage section 1608 into the random access memory (RAM) 1603, such as executing the method described in the above embodiments. In the RAM 1603, various programs and data required for system operations are also stored. The CPU 1601, the ROM 1602, and the RAM 1603 are connected to each other through a bus 1604. The input / output (I / O) interface 1605 is also connected to the bus 1604.

[0194] The following components are connected to the I / O interface 1605: an input section 1606 including a keyboard, a mouse, etc.; an output section 1607 including such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1608 including a hard disk, etc.; and a communication section 1609 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1609 performs communication processing via a network such as the Internet. The drive 1610 is also connected to the I / O interface 1605 as required. A removable medium 1611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1610 as required so that a computer program read from it can be installed into the storage section 1608 as required.

[0195] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 1609, and / or installed from the removable medium 1611. When the computer program is executed by the central processing unit (CPU) 1601, various functions defined in the system of the present application are executed.

[0196] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0197] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0198] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the unit itself.

[0199] Another aspect of this application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for obtaining model features as described above is implemented. The computer-readable storage medium can be included in the electronic device described in the above embodiments, or can exist alone without being assembled into the electronic device.

[0200] Another aspect of this application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for obtaining model features provided in the above various embodiments.

[0201] The above content is only a preferred exemplary embodiment of this application and is not used to limit the implementation of this application. Those of ordinary skill in the art can easily make corresponding adaptations or modifications according to the main idea and spirit of this application. Therefore, the protection scope of this application should be subject to the protection scope required by the claims.

Claims

1. A method for obtaining model features, characterized in that, it includes: Performing a filtering feature selection process on the initial candidate feature set to filter out the initial candidate features in the initial candidate feature set that do not meet the set threshold conditions, and obtaining a candidate feature set; Performing stepwise regression traversal on the candidate features included in the candidate feature set to select the candidate feature with the smallest regression coefficient as the target feature, and obtaining a target feature set, where the number of target features included in the target feature set meets the set quantity condition; Performing traversal of the optimal features on the target feature set to obtain an optimal feature set, where the optimal features included in the optimal feature set make the value of the target feature function optimal, and the target feature function is used to evaluate the model effect of the target model; wherein, performing a filtering feature selection process on the initial candidate feature set to filter out the initial candidate features in the initial candidate feature set that do not meet the set threshold conditions, and obtaining a candidate feature set, includes: Calculating the information value score of the initial candidate features included in the initial candidate feature set; After filtering out the initial candidate features with information value scores less than the first score threshold from the initial candidate feature set, calculating the population stability scores of the individual initial candidate features in the filtered initial candidate feature set; Selecting the initial candidate features with population stability scores less than the second score threshold as candidate features to obtain the candidate feature set.

2. The method according to claim 1, characterized in that, calculating the information value score of the initial candidate features included in the initial candidate feature set, includes: Performing feature binning on the initial candidate features included in the initial candidate feature set to obtain multiple feature bins corresponding to the initial candidate features; Calculating the information value scores of the respective feature bins according to the number of positive samples and negative samples included in each feature bin, and taking the sum of the information value scores of the respective feature bins as the information value score of the initial candidate feature.

3. The method according to claim 1, characterized in that, after filtering out the initial candidate features with information value scores less than the first score threshold from the initial candidate feature set, calculating the population stability scores of the individual initial candidate features in the filtered initial candidate feature set, includes: Dividing the filtered initial candidate feature set into a training sample set and a test sample set; Performing binning on the individual initial candidate features in the filtered initial candidate feature set to obtain multiple feature bins corresponding to the individual initial candidate features; Calculating the population stability scores of the respective feature bins according to the distribution ratio of the training samples and test samples included in each feature bin, and taking the sum of the population stability scores of the respective feature bins as the population stability score of the corresponding initial candidate feature.

4. The method according to claim 1, characterized in that, performing stepwise regression traversal on the candidate features included in the candidate feature set, includes: Perform forward stepwise regression traversal and bidirectional stepwise regression traversal on the candidate features included in the candidate feature set respectively to obtain a specified number of target features respectively. Perform duplicate removal processing on all the obtained target features to obtain a target feature set in which duplicate target features are filtered out.

5. The method according to claim 4, wherein, Performing forward stepwise regression traversal on the candidate features included in the candidate feature set includes: In each round of forward stepwise regression process, select candidate features from the candidate feature set without replacement, construct a first regression model according to the selected candidate features and the selected candidate feature subset, and verify the first regression model based on the candidate feature set to obtain the first regression coefficients of each candidate feature included in the candidate feature set. If the minimum first regression coefficient obtained in each round of forward stepwise regression process is less than the first regression threshold, merge the candidate feature corresponding to the minimum first regression coefficient into the selected candidate feature subset, and delete the candidate feature corresponding to the minimum first regression coefficient from the candidate feature set, so as to update the selected candidate feature subset and the candidate feature set in the next round based on the updated selected candidate feature subset and candidate feature set. If it is determined that the number of candidate features included in the currently updated selected candidate feature subset reaches the specified number, use the candidate features included in the currently updated selected candidate feature subset as target features.

6. The method according to claim 4, wherein, Performing bidirectional stepwise regression traversal on the candidate features included in the candidate feature set includes: In each round of bidirectional stepwise regression process, select candidate features from the candidate feature set without replacement, construct a first regression model according to the selected candidate features and the selected candidate feature subset, and verify the first regression model based on the candidate feature set to obtain the first regression coefficients of each candidate feature included in the candidate feature set. If the minimum first regression coefficient obtained in each round of bidirectional stepwise regression process is less than the first regression threshold, merge the candidate feature corresponding to the minimum first regression coefficient into the selected candidate feature subset, and delete the candidate feature corresponding to the minimum first regression coefficient from the candidate feature set. Construct a second regression model according to the updated selected candidate feature subset, and verify the second regression model based on the updated selected candidate feature subset. If the maximum second regression coefficient obtained in each round is greater than the second regression threshold, delete the candidate feature corresponding to the maximum second regression coefficient from the updated selected candidate feature subset to further update the selected candidate feature subset. Based on the candidate feature set updated in each round and the further updated selected candidate feature subset, update the selected candidate feature subset and the candidate feature set in the next round until the number of candidate features included in the updated selected candidate feature subset reaches the specified number, and use the specified number of candidate features as target features.

7. The method according to claim 1, wherein, Traverse the optimal features of the target feature set to obtain an optimal feature set, including: Perform sequential forward selection processing and sequential floating forward selection processing on the target feature set respectively to select optimal features from the target feature set; Deduplicate all the selected optimal features to obtain an optimal feature set with duplicate optimal features filtered out.

8. The method according to claim 7, wherein, Performing sequential forward selection processing on the target feature set includes: In each round of sequential forward selection processing, select a target feature from the target feature set. If the combination of the selected target feature and the selected target feature subset makes the value of the target feature function optimal, then add the selected target feature as an optimal feature to the selected target feature subset until all the target features in the target feature set are selected; Take the finally updated selected target feature subset as the optimal feature set.

9. The method according to claim 7, wherein, Performing sequential floating forward selection processing on the target feature set includes: In each round of sequential forward selection processing, select a target feature from the target feature set. If the combination of the selected target feature and the selected target feature subset makes the value of the target feature function optimal, then add the selected target feature to the selected target feature subset; Select a target feature again from the updated selected target feature subset. If the other target features after filtering out the reselected target feature make the value of the target feature function optimal, then delete the reselected target feature from the updated selected target feature subset to further update the selected target feature subset; Based on the further updated selected target feature subset in each round, update the selected target feature subset in the next round until all the target features in the target feature set are selected, and take the finally updated selected target feature subset as the optimal feature set.

10. The method according to claim 1, wherein, The method further includes: Evaluate the model effect of the target model according to the optimal feature set to obtain an evaluation index corresponding to the target model; If the evaluation index does not meet the set index threshold range, adjust the set threshold condition to reselect the optimal feature set from the initial candidate feature set based on the adjusted set threshold condition.

11. A model feature acquisition device, wherein, It includes: A feature filtering module configured to perform filtering-based feature selection processing on the initial candidate feature set to filter out the initial candidate features in the initial candidate feature set that do not meet the set threshold condition, and obtain a candidate feature set; A stepwise regression module configured to perform stepwise regression traversal on the candidate features included in the candidate feature set to select the candidate feature with the smallest regression coefficient as the target feature, and obtain a target feature set, where the number of target features included in the target feature set meets the set quantity condition; An optimal traversal module, configured to traverse the target feature set for optimal features to obtain an optimal feature set, where the optimal features included in the optimal feature set enable the value of the target feature function to be optimal, and the target feature function is used to evaluate the model effect of the target model; Among them, the feature filtering module is further configured to perform the following steps: Calculate the information value scores of the initial candidate features included in the initial candidate feature set; After filtering out the initial candidate features with information value scores less than the first score threshold from the initial candidate feature set, calculate the population stability scores of the respective initial candidate features in the filtered initial candidate feature set; Select the initial candidate features with population stability scores less than the second score threshold as candidate features to obtain the candidate feature set.

12. An electronic device, characterized in that, it includes: a memory storing computer-readable instructions; a processor that reads the computer-readable instructions stored in the memory to execute the method according to any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, computer-readable instructions are stored thereon, and when the computer-readable instructions are executed by a processor of a computer, the computer is caused to execute the method according to any one of claims 1-10.

14. A computer program product, characterized in that, it includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Filtering-package combination flow feature selection method based on support vector machine

    CN108319987A

  • Multi-round cyclic feature selection method and device during model training

    CN110298389A