Model training method and device based on dominant characteristic distillation and electronic equipment
By using advantageous feature distillation technology in credit model training, the knowledge of the teacher model is transferred to the student model, which solves the problem of unstable external data quality of third-party, improves the generalization performance and robustness of the model, and achieves a more efficient and reliable credit model score.
Patent Information
- Application Number
- CN202311602360.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-28
- Publication Date
- 2025-05-30
AI Technical Summary
During credit model training, the quality of external data of third parties is unstable, which affects the accuracy and reliability of model training and online services.
The model training method based on advantageous feature distillation is adopted, and the key features and knowledge learned by the large teacher model are transferred to the small student model through transfer learning technology of the teacher model and the student model, thereby reducing the impact of data source changes on model scores and improving the generalization performance and robustness of the model.
It improves the generalization performance and robustness of the model, reduces the impact of data source changes, reduces the number of model parameters, calculation complexity and resource consumption, and improves the accuracy of the model.
Smart Images

Figure CN120068995A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of machine learning technologies, and more particularly to credit risk control technologies, and specifically to a model training method, apparatus, and electronic device based on dominant feature distillation. Background Art
[0002] In the field of Internet finance, credit risk control models have become a very important part and have developed rapidly. Traditional financial institutions need to conduct loans through methods such as offline customer signing and asset guarantee. The entire process is relatively complex and requires a large amount of human resources. At the same time, it is difficult to quickly judge the risk credit status of customers. Internet finance mainly provides online services. Credit models based on technologies such as big data, machine learning, and artificial intelligence have been able to evaluate the credit of individuals or enterprises more accurately and efficiently. Furthermore, it can review and disburse loans more quickly, benefiting many customers with loan needs. At the same time, it can help Internet financial institutions understand the needs, risks, and credit status of customers.
[0003] The inventors found that with the continuous expansion of financial business, on the premise of safety and compliance, more and more enterprises access and use third-party data to fully explore customer information and assist in business decision-making. However, the problem of unstable quality of third-party external data is a difficult problem that cannot be ignored in the current credit model training and online services.
[0004] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept. Therefore, it may include information that does not form the prior art known to ordinary skilled artisans in this country. Summary of the Invention
[0005] This summary of the disclosure is used to introduce concepts in a concise form, which will be described in detail in the subsequent detailed implementation section. This summary of the disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.
[0006] Some embodiments of the present disclosure propose a model training method based on dominant feature distillation, a model training apparatus based on dominant feature distillation, an electronic device, a computer-readable medium, and a computer program product to solve one or more of the technical problems mentioned in the above background art section.
[0007] In a first aspect, some embodiments of the present disclosure provide a model training method based on dominant feature distillation, including: inputting sample data into a pre-trained teacher model to obtain the output result of the teacher model, where the teacher model is trained using the sample data; training a plurality of student models according to the sample conventional feature data, corresponding sample labels, and the output result of the teacher model in the sample data, where at least one of the following is different among the plurality of student models: the input target sample data, the output prediction target, and the prediction accuracy; determining the fusion method of the models based on the target prediction data, and fusing the plurality of trained student models to obtain a target model for analyzing the target prediction data.
[0008] In some embodiments, training a plurality of student models according to the sample conventional feature data, corresponding sample labels, and the output result of the teacher model in the sample data includes: for each student model among the plurality of student models, inputting the sample conventional feature data in the target sample data into the student model to obtain the student data output by the hidden layer of the student model; performing a first-stage training on the student model based on the comparison result between the student data and the teacher data to adjust the model parameters in the hidden layer of the student model; where the teacher data is the data output by the hidden layer of the teacher model obtained by inputting the target sample data into the teacher model.
[0009] In some embodiments, performing a first-stage training on the student model based on the comparison result between the student data and the teacher data to adjust the model parameters in the hidden layer of the student model includes: inputting the student data into a fully connected layer to obtain convolutional student data to match the data dimension of the teacher data; using the mean square error between the convolutional student data and the teacher data as the target loss function for the first-stage training, and adjusting the model parameters in the hidden layer of the student model according to the comparison result between the value of the target loss function and a first threshold.
[0010] In some embodiments, training a plurality of student models according to the sample conventional feature data, corresponding sample labels, and the output result of the teacher model in the sample data further includes: for each student model among the plurality of student models, inputting the sample conventional feature data in the target sample data into the student model after the first-stage training and outputting the student prediction data; performing a second-stage training on the student model based on the student prediction data, the teacher prediction data, and the corresponding sample labels to adjust the model parameters of the student model; where the teacher prediction data is the prediction data output by inputting the target sample data into the teacher model.
[0011] In some embodiments, based on the student prediction data, the teacher prediction data, and the corresponding sample labels, the student model is trained in a second stage to adjust the model parameters of the student model, including: determining a first loss value between the student prediction data and the teacher prediction data, and determining a second loss value between the student prediction data and the corresponding sample labels; taking the weighted sum of the first loss value and the second loss value as the target loss function for the second-stage training, and further adjusting the model parameters of the student model according to the comparison result between the value of the target loss function and a second threshold.
[0012] In some embodiments, the method further includes: extracting original feature data from the multi-dimensional data obtained offline, and performing feature engineering processing on the original feature data; dividing the processed feature data into dominant feature data and regular feature data, and labeling each piece of feature data to generate sample data for model training, where the dominant feature data is feature data that cannot be obtained online.
[0013] In some embodiments, the method further includes: in response to receiving a query request representing user credit, obtaining regular feature data related to the query request online; inputting the regular feature data obtained online into the target model, fusing the credit prediction data of multiple student models in a determined fusion manner, and outputting the fusion result as the credit prediction result of the target model.
[0014] In a second aspect, some embodiments of the present disclosure provide a model training apparatus based on dominant feature distillation, including: a teacher model input unit configured to input sample data into a pre-trained teacher model to obtain the output result of the teacher model, where the teacher model is trained using the sample data; a student model training unit configured to train multiple student models according to the sample regular feature data in the sample data, the corresponding sample labels, and the output result of the teacher model, where there is at least one difference among the multiple student models in the following: the input target sample data, the output prediction target, and the prediction accuracy; a model fusion unit configured to determine the fusion manner of the model based on the target prediction data, and fuse the multiple trained student models to obtain a target model for analyzing the target prediction data.
[0015] In some embodiments, the student model training unit includes a first training subunit configured to, for each student model among the multiple student models, input the sample regular feature data in the target sample data into the student model to obtain student data output by the hidden layer of the student model; performing a first-stage training on the student model based on the comparison result between the student data and the teacher data to adjust the model parameters in the hidden layer of the student model; where the teacher data is data output by the hidden layer of the teacher model obtained by inputting the target sample data into the teacher model.
[0016] In some embodiments, the first training subunit is further configured to input the student data into the fully connected layer to obtain the convolutional student data to match the data dimension of the teacher data; use the mean square error between the convolutional student data and the teacher data as the target loss function for the first-stage training, and adjust the model parameters in the hidden layer of the student model according to the comparison result between the value of the target loss function and the first threshold.
[0017] In some embodiments, the student model training unit further includes a second training subunit, which is configured to, for each student model among multiple student models, input the sample regular feature data in the target sample data into the student model completed in the first-stage training, and output the student prediction data; perform second-stage training on the student model based on the student prediction data, the teacher prediction data, and the corresponding sample label to adjust the model parameters of the student model; wherein, the teacher prediction data is the prediction data obtained by inputting the target sample data into the teacher model.
[0018] In some embodiments, the second training subunit is further configured to determine a first loss value between the student prediction data and the teacher prediction data, and determine a second loss value between the student prediction data and the corresponding sample label; use the weighted sum of the first loss value and the second loss value as the target loss function for the second-stage training, and further adjust the model parameters of the student model according to the comparison result between the value of the target loss function and the second threshold.
[0019] In some embodiments, the device further includes a feature processing unit, which is configured to extract the original feature data from the multi-dimensional data obtained offline, perform feature engineering processing on the original feature data; divide the processed feature data into dominant feature data and regular feature data, and label each feature data to generate the sample data for model training, wherein the dominant feature data is the feature data that cannot be obtained online.
[0020] In some embodiments, the device further includes an online prediction unit, which is configured to, in response to receiving a query request characterizing the user's credit, obtain the regular feature data related to the query request online; input the regularly obtained feature data online into the target model, fuse the credit prediction data of multiple student models in a determined fusion manner, and output the fusion result as the credit prediction result of the target model.
[0021] In a third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device having one or more programs stored thereon, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the model training method based on dominant feature distillation described in any implementation manner of the first aspect above.
[0022] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the model training method based on dominant feature distillation described in any implementation manner of the first aspect above is implemented.
[0023] In a fifth aspect, some embodiments of the present disclosure provide a computer program product including a computer program, where the computer program implements the model training method based on dominant feature distillation described in any implementation manner of the first aspect above when executed by a processor.
[0024] The above various embodiments of the present disclosure have the following beneficial effects: The model training method based on dominant feature distillation in some embodiments of the present disclosure can improve the generalization performance and robustness of the model while ensuring the prediction accuracy of the model, and reduce the impact of data source changes. Specifically, in order to alleviate the problems proposed in the background art, the related art gradually reduces the dependence on third-party data and improves the stability and quality of data. In addition to strict external data review, some common technical solutions in the current credit scenario are as follows: 1) Backup data source: During the model training process, in addition to using a third-party data source, a backup data source can also be prepared. If a data source goes offline, the backup data source can be started in a timely manner to ensure that the model can continue to work; 2) Data filling: Machine learning algorithms (KNN, random forest, Markov chain Monte Carlo algorithm) or other data mining deep networks are used to predict or fill missing data; 3) Model monitoring: A model monitoring and early warning trigger mechanism is established to timely detect problems such as data fluctuations or offline.
[0025] However, these technologies have the following problems: 1) The backup data source will bring pressure on data costs and data storage; switching from an important data source to a backup data source requires a large amount of analysis, data support, A / B testing (a new page optimization method), and even after evaluation, the model needs to be retrained and iterated from scratch, which is a huge project; 2) Data filling, based on statistical rules and filling methods that have not been effectively and sufficiently trained, may result in different risk meanings before and after filling and cannot effectively distinguish missing data in different situations, reducing the variable interpretability and model controllability, and causing a significant attenuation of the model effect; the method of learning and fitting the data sample rules based on machine learning algorithms (such as decision trees, random forests, etc.) requires a large number of samples to be trained in advance, consuming a large amount of computing resources and time, and is prone to model bias and overfitting; 3) Model monitoring and early warning, the early warning mechanism cannot fundamentally solve the problem of missing important data sources and model failure.
[0026] Based on this, the model training method based on dominant feature distillation according to some embodiments of the present disclosure proposes an integrated model modeling and deployment solution based on a teacher-student model and dominant feature distillation, especially for the credit risk control scenario. The method in this embodiment can transfer the key features and knowledge learned by a large teacher model to a small student model through dominant feature distillation and transfer learning techniques, so as to reduce the impact and damage of data source changes on model scoring, and reduce the number of model parameters, computational complexity and resource consumption while improving the accuracy. And since multiple independent student models are trained simultaneously, the optimal way of fusing student models can be flexibly selected based on the latest sample distribution, and then the final model score can be obtained by integrating the voting results of multiple models. This can make the model score have better robustness and generalization ability to achieve the automatic adaptation of the model to data source changes. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more obvious. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.
[0028] Figure 1 is a flowchart of some embodiments of the model training method based on dominant feature distillation of the present disclosure;
[0029] Figure 2 is a flowchart of some other embodiments of the model training method based on dominant feature distillation of the present disclosure;
[0030] Figure 3 is a flowchart of some embodiments of the training of the student model in the model training method of the present disclosure;
[0031] Figure 4 is a schematic diagram of some application scenarios of the model training method of the present disclosure;
[0032] Figure 5 is a schematic structural diagram of some embodiments of the model training device based on dominant feature distillation of the present disclosure;
[0033] Figure 6 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0035] In addition, it should be noted that for the sake of convenience of description, only the parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0036] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.
[0037] In addition, the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless clearly specified otherwise in the context, it should be understood as "one or more".
[0038] Figure 1 Flow 100 showing some embodiments of a model training method based on privileged features distillation according to the present disclosure is shown. The model training method may include the following steps:
[0039] Step 101, inputting sample data into a pre-trained teacher model to obtain the output result of the teacher model.
[0040] It should be noted that privileged features distillation (PFD) is generally a transfer learning technique and a model compression method. This technique helps a simple model (Student Model) to be trained by leveraging the powerful representation learning ability of a complex model (Teacher Model). It can distill and transfer the important features and knowledge in a large and complex neural network to a small network, reducing the number of model parameters and computational resource consumption, and improving the generalization performance and robustness of the model. In addition, the teacher model is usually pre-trained before training the student model.
[0041] In some embodiments, the execution entity (such as a model training server) of the model training method based on dominant feature distillation can communicate and connect with other electronic devices (such as a third-party server or a database) through a wired connection or a wireless connection. Here, the execution entity can perform offline training on the teacher model using sample data. That is, the sample feature data in the sample data can be input into the teacher model to obtain the prediction data of the teacher model. Then, the loss function value can be determined based on the prediction data and the sample label data. If the loss function value meets the expected requirements, such as being less than the expected threshold, it can indicate that the training of the teacher model is completed.
[0042] In addition, when the training of the teacher model is completed, the execution entity can use the teacher model to train the student model. That is, the sample data can be input into the trained teacher model to obtain the output result of the teacher model for training the student model.
[0043] It can be understood that the network structure of the teacher model here can be set according to actual needs. Usually, the dominant feature distillation adopts the Teacher-Student mode. Generally, a complex and large model is used as the Teacher, that is, the teacher role. As an example, the teacher model can adopt LSTM (Long Short Term Memory), Bert model structure, etc. Among them, Bert is the abbreviation of Bidirectional Encoder Representation from Transformers, that is, the Encoder of the bidirectional Transformer. This can make the learning ability of the teacher model stronger and can fully mine the dominant features in the data, etc.
[0044] In some embodiments, the execution entity can obtain sample data from historical data. Specifically, first, the original feature data can be extracted from the multi-dimensional data obtained offline, and the original feature data can be processed by feature engineering. Then, the processed feature data can be divided into dominant feature data and conventional feature data. After that, each feature data can be labeled to generate the sample data for model training. Among them, the dominant feature data is usually the feature data that cannot be obtained online, that is, it can often only be obtained through offline analysis.
[0045] It can be understood that feature engineering, feature extraction, or feature discovery is generally a process of extracting features (characteristics, attributes, properties) from the original data using domain knowledge. Its motivation is to use these additional features to improve the quality of the results of the machine learning process, rather than just providing the original data to the machine learning process. That is to say, feature engineering is usually the basis of model training, especially for model training based on dominant feature distillation.
[0046] As an example, for the training of a credit scenario risk control model based on privileged feature distillation, such as Figure 2 the feature ETL module shown in. Here, ETL is the abbreviation of Extract-Transform-Load, which is usually used to describe the process of extracting, transforming, and loading data from the source end to the destination end. For the feature ETL module here, first, it is necessary to obtain data sources in various dimensions such as consumption ability, asset status, credit transactions, basic information, social network, etc. from the data stream (such as third-party shopping platforms, savings platforms, etc.), and extract the original features from the data sources. Then, the features can be refined into categories and feature domains according to business understanding, expert rules, correlation, etc. For example, for shopping platform data, it can be divided into favorite feature data, browsing feature data, add-to-cart feature data, and order feature data according to scenarios. Another example is that the asset status can be divided into stored-value assets, wealth management assets, and assets to be repaid. And the data in each feature domain can be cleaned, processed, and constructed. Common feature engineering can include: outlier and missing value handling, data bucketing, standardization processing, feature derivation, feature screening, and dimensionality reduction, etc.
[0047] After that, the large pool of cleaned features can be divided into a privileged feature library (Privileged Features) and a regular feature library (Regular Features). Here, features with high discrimination, high-dimensional sparsity, but only available offline can be classified into Privileged Features. Among them, the classification method can adopt importance ranking such as single-variable analysis shap (SHapley Additive exPlanations, quantifying feature contributions), such as high PSI (Population Stability Index) stability, IV (a commonly used indicator in credit models, generally used for variable screening), and KS (an indicator used to measure the effect of classification models), and the feature importance in XGB (extreme gradient boosting), such as the total_gain gain size ranking, etc. For example, for the credit risk control scenario, the user's historical repayment time, the historical probability of on-time repayment, etc. are crucial for the user's risk rating, so these data can be used as privileged features.
[0048] Step 102: Train multiple student models according to the sample regular feature data, corresponding sample labels, and the output results of the teacher model in the sample data.
[0049] In some embodiments, the execution entity may train multiple student models based on the sample general feature data, the corresponding sample labels, and the output results of the teacher model in the sample data. Here, the student models may adopt models with relatively simple network structures but deeper and narrower, such as classification models. Among them, at least one of the following is different among the multiple student models: the input target sample data, the predicted target output, and the prediction accuracy.
[0050] Here, we upgrade the student model in the advantage feature distillation to multiple (at least two) student models. The number of student models can be customized and comprehensively determined according to multiple factors such as the application scenario and the gain of the evaluation effect. Increasing the number of student models can, on the one hand, introduce more student network models to further improve their integration ability and accuracy, with higher fault tolerance; on the other hand, it can create customized learning models for different scenarios and training samples, making the application scenarios more diverse and more compatible. For example, some student models can analyze user data with rising risks, and some student models can analyze user data with decreasing risks.
[0051] In some embodiments, as Figure 3 shown, the execution entity may train the student models in two stages. The first stage is PDF feature distillation to learn the hidden (including) layer. At this time, for each student model among the multiple student models, the execution entity may input the sample general feature data in the target sample data into the student model to obtain the student data output by the hidden layer of the student model. And the target sample data may be input into the teacher model to obtain the data output by the hidden layer of the teacher model. In this way, based on the comparison result between the student data and the teacher data, the first-stage training of the student model can be performed to adjust the model parameters in the hidden layer of the student model.
[0052] As Figure 3 shown in the left training process in, the Hint layer hidden layer output of the Teacher guides the training of the Student network. Here, the first h layer parameters of the Teacher network are used as W Hint and the first g layer parameters of the Student network are used as W Guided . Among them, the Hint is generally defined as the hidden layer output of the teacher, which is used to guide the learning process of the student. Similarly, a hidden layer is selected from the student and called the guided layer. At this time, the error between the two outputs can be used as the optimization objective function, and then the parameters in the hidden layer of the student model can be adjusted.
[0053] Optionally, to adapt to the Teacher network structure, as Figure 3As shown, a convolutional layer, i.e., a fully connected layer, can be added to the outermost layer of the Student network. At this time, student data can be input into the fully connected layer to obtain convolutional student data to match the data dimension of the teacher data. That is to say, the network parameter size of the hidden layer of the student model is adapted to that of the teacher model. Then, the mean square error between the convolutional student data and the teacher data can be used as the target loss function for the first-stage training. Furthermore, the model parameters in the hidden layer of the student model can be adjusted according to the comparison result between the value of the target loss function and the first threshold.
[0054] As an example, the optimization objective function for this stage of learning can be: the mean square error of the data outputs of the Teacher and Student networks, i.e.,
[0055]
[0056] where W r is the mapping function of the fully connected layer; u h (x; W Hint ) represents the output of the hidden layer of the teacher model; v g (x; W Guided ) represents the output of the hidden layer of the student model.
[0057] Furthermore, after the first-stage training is completed, the execution entity can perform the second-stage training on the student model. As shown in the right training process in Figure 3 , the second stage is target knowledge distillation. At this time, for each student model among multiple student models, the execution entity can input the sample regular feature data in the target sample data into the student model completed in the first-stage training to output student prediction data. And input the target sample data into the teacher model to output the prediction data. Based on the student prediction data, the teacher prediction data, and the corresponding sample labels, the second-stage training of the student model is performed to adjust the model parameters of the student model.
[0058] In some embodiments, the execution entity can determine the first loss value between the student prediction data and the teacher prediction data, and can determine the second loss value between the student prediction data and the corresponding sample labels. Then, the weighted sum of the first loss value and the second loss value can be used as the target loss function for the second-stage training. And according to the comparison result between the value of the target loss function and the second threshold, the model parameters of the student model are further adjusted.
[0059] That is to say, the W obtained from the first-stage training of the Student network GuidedThe parameters can be used as the initial values of the parameters of the student model in the second stage. In the second stage, only Soft-target and Hard-target can be used to train the Student network. That is, for the trained Teacher network, use the high-temperature distilled T high The class probabilities of the Softmax (activation function) layer generated are used as Soft-target, and combined with the original label Hard-target to jointly train the Student network. The optimization objective function for its training can include two parts, the cross-entropy loss of the softmax outputs of the Student and the Teacher under the same temperature condition, and the cross-entropy loss function of the Softmax output of the Student and the original label. That is
[0060]
[0061] where α and β represent adjustable weight parameters; represents the first loss value above; H(y, P S ) represents the second loss value above.
[0062] In some alternative implementation manners, to simplify the training process, the execution subject can also only adopt the above-mentioned second-stage training to train multiple student models, which will not be elaborated here.
[0063] It can be understood that in this embodiment, the teacher model is used to assist the student model in two-stage training. The teacher model fully excavates the dominant features and has strong learning ability. It can transfer the knowledge (intermediate layer features, hidden layer data) it has learned to the student model with relatively weak learning ability, so as to enhance the generalization ability of the student model. And in the first stage, the hidden layer parameters of the student model are trained and adjusted specifically, which is not only beneficial to improving the learning ability of the student model, but also can improve the accuracy of its prediction results in the second stage, thus improving the training efficiency of the second stage.
[0064] Step 103, determine the fusion method of the model based on the target prediction data, and fuse the trained multiple student models to obtain a target model for analyzing the target prediction data.
[0065] In some embodiments, the execution entity may determine the fusion method of the model based on the target prediction data. And it can fuse multiple trained student models to obtain a target model for analyzing the target prediction data. The fusion method here can be set according to the actual situation, such as weighted summation or binary tree. For example, when determining the user credit risk level, the prediction result of the student model for predicting low risk can be set with a higher weight, while the prediction result of the student model for predicting high risk can be set with a lower weight or a negative weight.
[0066] That is to say, in the offline environment, this embodiment trains the student model and the teacher model simultaneously: including multiple student models and one teacher model. The teacher model uses a deeper network structure to guide the student models using shallow networks to learn. The teacher model additionally utilizes the dominant features, and its accuracy is thus higher. By transferring the knowledge distilled from the teacher model to the student models, a relatively high accuracy can be maintained even when the number of model parameters and the computational complexity are small.
[0067] Currently, most of the credit models adopted by Internet financial institutions are constructed based on technologies such as big data and machine learning. These models use a large amount of historical data such as their own data assets or third-party data to build models, analyze and predict, and automatically process a large amount of data to predict the default probability of loans. With more and more enterprises accessing and using third-party data, how to cope with the computational pressure, unstable interface performance, sudden offline of data sources, etc. brought by the massive big data, resulting in situations such as online score deviation, model failure, and decision-making errors, is an urgent problem to be solved currently.
[0068] From the above description, it can be seen that the model training method based on dominant feature distillation in some embodiments of the present disclosure proposes an integrated model modeling and deployment solution based on the teacher-student model and dominant feature distillation, especially for the credit risk control scenario. The method in this embodiment can transfer the key features and knowledge learned by the large teacher model to the small student models through dominant feature distillation and transfer learning techniques, so as to achieve the purpose of reducing the impact and damage of data source changes on model scoring, reducing the number of model parameters, computational complexity, and resource consumption, while improving the accuracy. And because multiple independent student models are trained simultaneously, the optimal student model fusion method can be flexibly selected based on the latest sample distribution situation, and then the final model score can be obtained by integrating the voting results of multiple models. This can make the model score have better robustness and generalization ability.
[0069] That is to say, the method of this embodiment can apply the distillation of advantageous features to the development of credit risk control model scoring, and the main technical problems that can be solved are: 1) the problem that the core features of batch offline training cannot be obtained or are seriously missing after the model is launched; 2) data samples are difficult to obtain and the cost is high; 3) model complexity, the dimension of the input features is too high, and the model online reasoning speed is slow.
[0070] It is understandable that the specific processing and parameter adjustment methods in the above model training process can be determined according to the specific problem and data situation. For example, the depth selection of the student model can be determined based on the sample size, data quality, training algorithm, etc. of the data stream in the application scenario. The selection of temperature T can be determined based on the analysis of the scale of the student model, the degree of attention to negative labels during training, etc.
[0071] In addition, the two-stage training can adopt serial training or joint training mode. That is, during the model training process, the teacher network can be trained in advance, and then the student network can be trained in serial training; or sampling joint training can be adopted, that is, the teacher network and the student network are jointly trained. This will speed up the training, but it will require higher computing resources such as GPU (Graphics Processing Unit).
[0072] In some embodiments, we can deploy the trained target model online. It should be noted that when serving online, only conventional features and student models are usually extracted for deployment. That is, a simple model is applied to online deployment production. Since the input does not depend on the dominant features, while ensuring the consistency of features used for offline training and online deployment, the computational complexity can be greatly reduced, and the efficiency of online reasoning can be improved.
[0073] like Figure 4 As shown, in response to receiving a query request characterizing user credit, conventional feature data related to the query request can be obtained online. Afterwards, the conventional feature data obtained online is input into the target model, so that the credit prediction data of multiple student models are fused according to a determined fusion method, and the fusion result is output as the credit prediction result of the target model. In other words, when a user initiates a query request, after verification, the above-mentioned conventional feature library that is ready is obtained online, and the matched feature data is used as the input of N student models that have been trained offline. After the inference operation, the output probabilities of the N student models are obtained. Finally, a comprehensive model scoring result can be obtained through an integrated algorithm such as Bagging (Bootstrap aggregating).
[0074] It can be understood that the execution entity for online analysis (such as a credit platform server) can be the same as or different from the execution entity in Figure 1 the embodiment.
[0075] With further reference to Figure 5 , as an implementation of the method shown above Figures 1 to 3 , the present disclosure provides some embodiments of a model training apparatus based on dominant feature distillation. These apparatus embodiments correspond to those method embodiments shown in Figures 1 to 3 . The apparatus can be specifically applied to various electronic devices.
[0076] As shown in Figure 5 , a model training apparatus 500 based on dominant feature distillation in some embodiments may include: a teacher model input unit 501 configured to input sample data into a pre-trained teacher model to obtain an output result of the teacher model, where the teacher model is trained using the sample data; a student model training unit 502 configured to train a plurality of student models according to sample conventional feature data in the sample data, corresponding sample labels, and the output result of the teacher model, where at least one of the following is different among the plurality of student models: the input target sample data, the output prediction target, and the prediction accuracy; a model fusion unit 503 configured to determine a fusion method of the models based on target prediction data and fuse the plurality of trained student models to obtain a target model for analyzing the target prediction data.
[0077] In some embodiments, the student model training unit 502 may include a first training subunit (not shown in the figure) configured to, for each student model among the plurality of student models, input the sample conventional feature data in the target sample data into the student model to obtain student data output by the hidden layer of the student model; perform first-stage training on the student model based on the comparison result between the student data and the teacher data to adjust the model parameters in the hidden layer of the student model; where the teacher data is data obtained by inputting the target sample data into the teacher model and output by the hidden layer of the teacher model.
[0078] In some embodiments, the first training subunit may be further configured to input the student data into a fully connected layer to obtain convolutional student data to match the data dimension of the teacher data; use the mean square error between the convolutional student data and the teacher data as the target loss function for the first-stage training, and adjust the model parameters in the hidden layer of the student model according to the comparison result between the value of the target loss function and a first threshold.
[0079] In some embodiments, the student model training unit 502 may further include a second training subunit (not shown in the figure), which is configured to, for each of the multiple student models, input the sample general feature data in the target sample data into the student model that has completed the first-stage training, and output student prediction data; based on the student prediction data, the teacher prediction data, and the corresponding sample labels, perform second-stage training on the student model to adjust the model parameters of the student model; wherein the teacher prediction data is the prediction data obtained by inputting the target sample data into the teacher model.
[0080] In some embodiments, the second training subunit may be further configured to determine a first loss value between the student prediction data and the teacher prediction data, and determine a second loss value between the student prediction data and the corresponding sample labels; take the weighted sum of the first loss value and the second loss value as the target loss function for the second-stage training, and further adjust the model parameters of the student model according to the comparison result between the value of the target loss function and a second threshold.
[0081] In some embodiments, the device 500 may further include a feature processing unit (not shown in the figure), which is configured to extract original feature data from the multi-dimensional data obtained offline, perform feature engineering processing on the original feature data; divide the processed feature data into dominant feature data and general feature data, and label each feature data to generate sample data for model training, where the dominant feature data is feature data that cannot be obtained online.
[0082] In some embodiments, the device 500 may further include an online prediction unit (not shown in the figure), which is configured to, in response to receiving a query request characterizing the user's credit, obtain online the general feature data related to the query request; input the online obtained general feature data into the target model, fuse the credit prediction data of the multiple student models according to the determined fusion method, and output the fusion result as the credit prediction result of the target model.
[0083] It can be understood that the various units described in the model training device 500 based on dominant feature distillation correspond to the respective steps in the method described in the reference. Figures 1 to 3 Therefore, the operations, features, and beneficial effects described above for the method also apply to the model training device 500 based on dominant feature distillation and the units included therein, and will not be elaborated herein.
[0084] Next, refer to Figure 6 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing some embodiments of the present disclosure. Figure 6 The electronic device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.
[0085] As shown Figure 6 in FIG. 4, the electronic device 600 may include a processing device 601 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in the read-only memory (ROM) 602 or a program loaded from the storage device 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0086] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic disk, a memory card, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 FIG. 4 shows an electronic device 600 having various devices, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included. Figure 6 Each block shown in FIG. 4 may represent a device or, as needed, multiple devices.
[0087] Specifically, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such some embodiments, the computer program may be downloaded and installed from a network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the methods of some embodiments of the present disclosure are performed.
[0088] It should be noted that the computer-readable media described in some embodiments of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0089] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed network.
[0090] The above computer-readable medium may be included in the above electronic device; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to: input sample data into a pre-trained teacher model to obtain an output result of the teacher model, where the teacher model is trained using the sample data; train a plurality of student models according to the sample regular feature data, the corresponding sample labels, and the output result of the teacher model in the sample data, where at least one of the following is different among the plurality of student models: the input target sample data, the output prediction target, and the prediction accuracy; determine a model fusion method based on the target prediction data, and fuse the plurality of trained student models to obtain a target model for analyzing the target prediction data.
[0091] In addition, computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0092] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0093] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes a teacher model input unit, a student model training unit, and a model fusion unit. Among them, the names of these units do not constitute a limitation to the unit itself in some cases. For example, the teacher model input unit can also be described as "a unit that inputs sample data into a pre-trained teacher model to obtain the output result of the teacher model".
[0094] The functions described above in this article can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0095] Some embodiments of the present disclosure also provide a computer program product, including a computer program which, when executed by a processor, implements any one of the above-mentioned model training methods based on dominant feature distillation.
[0096] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A model training method based on dominant feature distillation, including: Inputting sample data into a pre-trained teacher model to obtain the output result of the teacher model, where the teacher model is trained using the sample data; Training multiple student models according to the sample conventional feature data, corresponding sample labels in the sample data, and the output result of the teacher model, where at least one of the following is different among the multiple student models: the input target sample data, the output prediction target, and the prediction accuracy; Determining the fusion method of the model based on the target prediction data, and fusing the trained multiple student models to obtain a target model for analyzing the target prediction data.
2. The model training method according to claim 1, wherein, the training of the multiple student models according to the sample conventional feature data, corresponding sample labels in the sample data, and the output result of the teacher model includes: For each student model among the multiple student models, inputting the sample conventional feature data in the target sample data into the student model to obtain the student data output by the hidden layer of the student model; Based on the comparison result between the student data and the teacher data, performing the first-stage training on the student model to adjust the model parameters in the hidden layer of the student model; wherein the teacher data is the data output by the hidden layer of the teacher model obtained by inputting the target sample data into the teacher model.
3. The model training method according to claim 2, wherein, the performing the first-stage training on the student model to adjust the model parameters in the hidden layer of the student model based on the comparison result between the student data and the teacher data includes: Inputting the student data into a fully connected layer to obtain convolutional student data to match the data dimension of the teacher data; Taking the mean square error between the convolutional student data and the teacher data as the target loss function for the first-stage training, and adjusting the model parameters in the hidden layer of the student model according to the comparison result between the value of the target loss function and the first threshold.
4. The model training method according to claim 2, wherein, the training of the multiple student models according to the sample conventional feature data, corresponding sample labels in the sample data, and the output result of the teacher model further includes: For each student model among the multiple student models, inputting the sample conventional feature data in the target sample data into the student model after the first-stage training, and outputting student prediction data; Based on the student prediction data, teacher prediction data, and corresponding sample labels, performing the second-stage training on the student model to adjust the model parameters of the student model; wherein the teacher prediction data is the prediction data output by inputting the target sample data into the teacher model.
5. The model training method according to claim 4, wherein, Performing a second-stage training on the student model based on the student prediction data, the teacher prediction data, and the corresponding sample labels to adjust the model parameters of the student model, including: Determining a first loss value between the student prediction data and the teacher prediction data, and determining a second loss value between the student prediction data and the corresponding sample labels; Performing a weighted summation of the first loss value and the second loss value as the target loss function for the second-stage training, and further adjusting the model parameters of the student model according to the comparison result between the value of the target loss function and a second threshold.
6. The model training method according to claim 1, wherein, the method further includes: Extracting original feature data from the multi-dimensional data obtained offline, and performing feature engineering processing on the original feature data; Dividing the processed feature data into dominant feature data and conventional feature data, and tagging each feature data to generate the sample data for model training, wherein the dominant feature data is feature data that cannot be obtained online.
7. The model training method according to any one of claims 1-6, wherein, the method further includes: In response to receiving a query request characterizing user credit, obtaining conventional feature data related to the query request online; Inputting the conventional feature data obtained online into the target model, fusing the credit prediction data of the multiple student models according to the determined fusion method, and outputting the fusion result as the credit prediction result of the target model.
8. A model training device based on dominant feature distillation, including: A teacher model input unit configured to input sample data into a pre-trained teacher model to obtain the output result of the teacher model, wherein the teacher model is trained using the sample data; A student model training unit configured to train multiple student models according to the sample conventional feature data in the sample data, the corresponding sample labels, and the output result of the teacher model, wherein at least one of the following is different among the multiple student models: the input target sample data, the output prediction target, and the prediction accuracy; A model fusion unit configured to determine the fusion method of the model based on the target prediction data, and fuse the multiple trained student models to obtain a target model for analyzing the target prediction data.
9. An electronic device, including: One or more processors; A storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the model training method based on dominant feature distillation according to any one of claims 1-7.
10. A computer-readable medium having a computer program stored thereon, wherein, when the computer program is executed by a processor, implementing the model training method based on dominant feature distillation according to any one of claims 1-7.
11. A computer program product comprising a computer program which, when executed by a processor, implements the method for training a model based on dominant feature distillation according to any one of claims 1-7.