Data Determination Method, Apparatus, Computer-Readable Storage Medium, and Electronic Device

By predicting and attributing the credit feature set, the abnormal feature categories and information in the credit feature set are accurately determined, and the problem of inaccurate explanation of the credit default risk prediction results in the existing technology is solved, and high-accurate risk warning and risk control effects are achieved.

CN114693428BActive Publication Date: 2025-06-03INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210266082.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-17
Publication Date
2025-06-03
Estimated Expiration
2042-03-17

AI Technical Summary

Technical Problem

The reasons for the prediction results of credit default risk in the prior art are inaccurate and cannot meet the needs of the risk control field.

Method used

By obtaining the credit feature set of the target object, making predictions and obtaining the prediction results and the first result, then when the prediction results represent that the target object has a default risk, each feature information in the credit feature set is attributively calculated to obtain the second result, and finally determining the abnormal feature category and abnormal feature information in the credit feature set based on the first result and the second result.

Benefits of technology

It realizes an accurate explanation of the prediction results of credit default risk, improves the accuracy of judgment of risk warning causes, meets the needs of the risk control field, and improves risk control efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114693428B_ABST
    Figure CN114693428B_ABST
Patent Text Reader

Abstract

The present invention discloses a data determination method, apparatus, computer-readable storage medium and electronic device. It relates to the field of fintech or other fields. The method includes: obtaining a credit feature set of a target object, where the credit feature set includes a plurality of feature information belonging to at least one feature category; performing a prediction on the credit feature set to obtain a prediction result and a first result; in the case where the prediction result indicates that the target object has a default risk, performing attribution calculation on each feature information in the credit feature set of the target object to obtain a second result; determining an abnormal feature category and abnormal feature information in the credit feature set based on the first result and the second result, where the abnormal feature category is the feature category with the highest contribution degree to the prediction result, and the abnormal feature information is the feature information with the highest contribution degree to the prediction result. The present invention solves the technical problem that the existing technology has inaccurate judgment on the cause of the credit default risk prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of fintech, and in particular, to a data determination method, apparatus, computer-readable storage medium, and electronic device. Background Art

[0002] In the field of risk control, in the credit scenario, the post-loan monitoring of customers is an indispensable environment. It is particularly important to accurately identify the customer groups that may have default risks, and then strengthen supervision or have account managers conduct manual collection to reduce the proportion of bad debts.

[0003] Among them, the field of risk control has high requirements for the interpretability of prediction results. That is, for customer groups that may have default risks, it is necessary to specifically analyze the reasons for their risk warnings. However, in the prior art, the judgment of the reasons for the occurrence of credit default risk prediction results is inaccurate, thus unable to meet the risk control requirements.

[0004] For the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] Embodiments of the present invention provide a data determination method, apparatus, computer-readable storage medium, and electronic device to at least solve the technical problem that the prior art inaccurately judges the reasons for the occurrence of credit default risk prediction results.

[0006] According to one aspect of the embodiments of the present invention, a data determination method is provided, including: obtaining a credit feature set of a target object, where the credit feature set includes a plurality of feature information belonging to at least one feature category; performing a prediction on the credit feature set to obtain a prediction result and a first result, where the prediction result represents whether the target object has a default risk, and the first result includes a first feature importance ratio of each feature category to the prediction result, and a second feature importance ratio of each feature information to the prediction result; in the case where the prediction result represents that the target object has a default risk, performing an attribution calculation on each feature information in the credit feature set of the target object to obtain a second result; determining an abnormal feature category and abnormal feature information in the credit feature set based on the first result and the second result, where the abnormal feature category is the feature category with the highest contribution degree to the prediction result, and the abnormal feature information is the feature information with the highest contribution degree to the prediction result.

[0007] Further, the second result includes an attribution value corresponding to each feature category and an attribution value corresponding to each feature information.

[0008] Further, the data determination method further includes: predicting the credit feature set based on a preset model to obtain a prediction result; determining a second feature importance ratio corresponding to each feature information based on the feature importance of each feature information, where the feature importance represents the influence degree of the feature information on the prediction result; adding up the second feature importance ratios of the feature information corresponding to each feature category to obtain the category feature importance ratio of each feature category; adding up the second feature importance ratios of all the feature information in the credit feature set to obtain the total feature importance ratio; determining a first feature importance ratio based on the category feature importance ratio and the total feature importance ratio.

[0009] Further, the data determination method further includes: determining the feature importance of each feature information; adding up the feature importance of all the feature information in the credit feature set to obtain the total feature importance; determining a second feature importance ratio based on the feature importance and the total feature importance.

[0010] Further, the data determination method further includes: determining an attribution value of each feature information based on the credit feature set; adding up the attribution values of the feature information belonging to the same feature category to obtain the attribution value corresponding to each feature category.

[0011] Further, the data determination method further includes: obtaining an explanation result based on the first result and the second result, where the explanation result represents the contribution degree of each feature category to the prediction result and the contribution degree of each feature information to the prediction result; determining an abnormal feature category and abnormal feature information in the credit feature set based on the explanation result.

[0012] Further, the data determination method further includes: multiplying the attribution value of each feature category by the corresponding first feature importance ratio to obtain a first value corresponding to each feature category; determining the contribution degree of each feature category to the prediction result based on the first value corresponding to each feature category; multiplying the attribution value of each feature information by the corresponding second feature importance ratio to obtain a second value corresponding to each feature information; determining the contribution degree of each feature information to the prediction result based on the second value corresponding to each feature information.

[0013] Further, the data determination method further includes: before predicting the credit feature set to obtain a prediction result and a first result, obtaining an initial credit feature set of at least one historical credit object, where the initial credit feature set includes a plurality of initial feature information belonging to different initial feature categories; processing the initial feature information corresponding to each initial feature category based on at least one trained first prediction model to obtain the initial feature importance corresponding to each initial feature information, where each initial feature category corresponds to a first prediction model; screening out target feature information from the initial feature information corresponding to each initial feature category based on the initial feature importance, where the target feature information corresponds to the feature information; training a second prediction model based on the target feature information to obtain a target prediction model, where the target prediction model is used to predict each credit feature set and obtain a prediction result and a first result.

[0014] According to another aspect of the embodiments of the present invention, there is also provided a data determination device, including: an acquisition module, configured to acquire a credit feature set of a target object, where the credit feature set includes a plurality of feature information belonging to at least one feature category; a prediction module, configured to predict the credit feature set to obtain a prediction result and a first result, where the prediction result represents whether the target object has a default risk, and the first result includes a first feature importance ratio of each feature category to the prediction result and a second feature importance ratio of each feature information to the prediction result; a calculation module, configured to perform attribution calculation on each feature information in the credit feature set of the target object when the prediction result represents that the target object has a default risk to obtain a second result; a determination module, configured to determine an abnormal feature category and abnormal feature information in the credit feature set based on the first result and the second result, where the abnormal feature category is the feature category with the highest contribution degree to the prediction result, and the abnormal feature information is the feature information with the highest contribution degree to the prediction result.

[0015] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium, including: a computer program is stored in the computer-readable storage medium, where the computer program is configured to execute the above data determination method when running.

[0016] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, where the electronic device includes one or more processors; a memory, configured to store one or more programs, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement running the program, where the program is configured to execute the above data determination method when running.

[0017] According to another aspect of the embodiments of the present invention, there is also provided a computer program product, including a computer program / instructions, which when executed by a processor, implement the above data determination method.

[0018] In the embodiments of the present invention, a method is adopted to judge the cause of the credit default risk prediction result based on the feature importance ratio and the attribution value. By obtaining the credit feature set of the target object and predicting the credit feature set, a prediction result and a first result are obtained. Then, when the prediction result indicates that the target object has a default risk, attribution calculation is performed on each feature information in the credit feature set of the target object to obtain a second result. Thus, based on the first result and the second result, the abnormal feature category and abnormal feature information in the credit feature set are determined. Among them, the credit feature set includes multiple feature information belonging to at least one feature category, the prediction result indicates whether the target object has a default risk, the first result includes the first feature importance ratio of each feature category to the prediction result, and the second feature importance ratio of each feature information to the prediction result. The abnormal feature category is the feature category with the highest contribution degree to the prediction result, and the abnormal feature information is the feature information with the highest contribution degree to the prediction result.

[0019] It is easy to notice that in the above process, based on the first feature importance ratio and the second feature importance ratio in the first result, the feature category and feature information with relatively high influence on the prediction result can be determined. Performing attribution calculation on each feature information in the credit feature set of the target object to obtain the second result can be used to determine the initial contribution degree of each feature information and each feature category to the prediction result. Further, by combining the first result and the second result, the final contribution degree of each feature information and each feature category to the prediction result can be determined, so that the abnormal feature category and abnormal feature information in the credit feature set can be accurately determined, realizing the judgment of the cause of the prediction result based on multiple factors, thereby improving the accuracy of the judgment of the risk warning cause, meeting the risk control requirements and improving the risk control efficiency.

[0020] It can be seen that the solution provided by the present application achieves the purpose of judging the cause of the credit default risk prediction result based on the feature importance ratio and the attribution value, thereby realizing the technical effect of improving the accuracy of the judgment of the risk warning cause, and further solving the technical problem that the cause of the credit default risk prediction result in the prior art is judged inaccurately. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0022] Figure 1 It is a schematic diagram of an optional data determination method according to an embodiment of the present invention;

[0023] Figure 2 It is a schematic diagram of an optional data determination method according to an embodiment of the present invention;

[0024] Figure 3 It is a schematic diagram of an optional training model according to an embodiment of the present invention;

[0025] Figure 4 It is a schematic diagram of an optional data determination device according to an embodiment of the present invention;

[0026] Figure 5 It is a schematic diagram of an optional electronic device according to an embodiment of the present invention. Detailed implementation manners

[0027] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these process, method, product or device.

[0029] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties.

[0030] Embodiment 1

[0031] According to an embodiment of the present invention, an embodiment of a data determination method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0032] Figure 1 is a schematic diagram of an alternative data determination method according to an embodiment of the present invention, as Figure 1 shown, the method includes the following steps:

[0033] Step S101, obtain a credit feature set of a target object, where the credit feature set includes a plurality of feature information belonging to at least one feature category.

[0034] In step S101, the credit feature set of the target object can be obtained through devices such as electronic devices, application systems, servers, etc. In this embodiment, the credit feature set of the target object is obtained through a risk warning system. Among them, the risk warning system can obtain the credit feature set based on at least one of a preset database, the Internet, a cloud server, or other devices for information reading. The credit feature set includes a plurality of feature information, each feature information corresponds to a feature category, and the feature categories corresponding to different feature information can be the same or different. Optionally, each feature information is at least composed of a feature identifier and a feature value, and the feature information is a feature that may affect the customer's credit, such as customer basic information, customer credit investigation, customer assets, customer turnover, customer tax, etc.

[0035] Optionally, in this embodiment, the target object is a customer who has already obtained credit. Preferably, the target customer is a customer who has already obtained credit but has not defaulted. Some of the feature categories and feature information in the credit feature set are shown in the following table:

[0036] Table 1 - Examples of Feature Categories and Feature Information

[0037]

[0038]

[0039] In Table 1, taking the feature categories and feature information corresponding to serial numbers 1-3 as an example, among them, "corporate turnover" is the feature category, and "average monthly transaction balance of corporate turnover in the last N 1 months", "average monthly transaction balance of corporate turnover in the last N 2 months", "average monthly transaction balance of corporate turnover in the last N 3 months" are the feature information corresponding to "corporate turnover". The corresponding relationships of other feature categories and feature information are the same as the foregoing content, so they will not be elaborated here.

[0040] It should be noted that by obtaining the credit feature set of the target object, it is convenient to subsequently determine whether the target object has default risk based on the feature credit set, and the reasons when there is default risk.

[0041] Step S102: Predict the credit feature set to obtain a prediction result and a first result, where the prediction result indicates whether the target object has default risk, and the first result includes the first feature importance ratio of each feature category to the prediction result and the second feature importance ratio of each feature information to the prediction result.

[0042] In step S102, the risk warning system can predict the credit feature set based on the trained prediction model to obtain a prediction result and a first result. In this embodiment, the aforementioned trained prediction model can be an Xgboost model, and the parameters of this model can be set as follows: the learning rate is set to 0.1, the number of trees is 241, the maximum depth of the tree is 5, the optimization function is: binary: logistic, and the random seed is 27.

[0043] Optionally, during the prediction process, the prediction model can not only predict whether the target object has default risk, but also give a preliminary explanation for the reason of the prediction result. Specifically, the prediction model can calculate the feature importance degree (i.e., the influence degree) of each feature information on the prediction result, determine the feature importance degree of each feature information, and then determine the feature importance ratio (i.e., the second feature importance ratio) of each feature information based on the feature importance degree of each feature information. And determine the feature importance ratio of each feature category (i.e., the first feature importance ratio) based on the feature importance ratio of each feature information.

[0044] It should be noted that by predicting the credit feature set to obtain a prediction result and a first result, an accurate prediction of whether the target object has default risk is achieved. At the same time, based on the first feature importance ratio and the second feature importance ratio in the first result, the feature categories and feature information with relatively high influence on the prediction result can be determined, so as to facilitate the subsequent determination of the reasons when the target object has default risk.

[0045] Step S103: When the prediction result indicates that the target object has default risk, perform attribution calculation on each feature information in the credit feature set of the target object to obtain a second result.

[0046] In step S103, when the prediction result indicates that the target object has a default risk, the risk warning system can perform attribution calculation on each feature information in the credit feature set of the target object based on the interpretation model. Among them, in this embodiment, the interpretation model is the Tree SHAP model, and the second result includes but is not limited to the attribution value corresponding to each feature information. The attribution value corresponding to each feature information can represent the initial contribution degree of this feature information to the prediction result, and then the initial contribution degree of each feature category to the prediction result can be determined. It should be emphasized that when the prediction result indicates that the target object does not have a default risk, the risk warning system can also perform attribution calculation on each feature information in the credit feature set of the target object based on the interpretation model.

[0047] It should be noted that by performing attribution calculation on each feature information in the credit feature set of the target object, the attribution value of each feature information can be obtained, thereby determining the initial contribution degree of each feature information and each feature category to the prediction result, which is convenient for subsequent determination of the reason when the target object has a default risk.

[0048] Step S104, determining the abnormal feature category and abnormal feature information in the credit feature set based on the first result and the second result, where the abnormal feature category is the feature category with the highest contribution degree to the prediction result, and the abnormal feature information is the feature information with the highest contribution degree to the prediction result.

[0049] In step S104, the risk warning system can multiply the feature importance ratio (i.e., the second feature importance ratio) corresponding to each feature information in the first result by the attribution value of each feature information in the second result to determine the final contribution degree of each feature information to the prediction result; it can also add the feature importance ratio (i.e., the second feature importance ratio) corresponding to each feature information in the first result to the attribution value of each feature information in the second result to determine the final contribution degree of each feature information to the prediction result; it can also multiply the feature importance ratio (i.e., the second feature importance ratio) corresponding to each feature information in the first result and the attribution value of each feature information in the second result by the corresponding weight coefficients respectively and then sum them to determine the final contribution degree of each feature information to the prediction result.

[0050] Optionally, the risk warning system can determine the attribution value of each feature category based on the attribution value of each feature information in the second result, and can determine the final contribution degree of each feature category to the prediction result from the foregoing methods by a method that is the same as or different from the method for determining the final contribution degree of each feature information to the prediction result.

[0051] Further, after the risk warning system determines the final contribution degree of each feature category or each feature information to the prediction result, the risk warning system can select the feature category with the highest contribution degree to the prediction result as the abnormal feature category, and select the feature information with the highest contribution degree to the prediction result as the abnormal feature information. Among them, the abnormal feature category represents the most important category reason for the target customer to have default risk, and the abnormal feature information represents the most important feature reason for the target customer to have default risk. The feature categories corresponding to the abnormal feature category and the abnormal feature information may be the same or different.

[0052] It should be noted that by combining the first result and the second result to determine the abnormal feature category and the abnormal feature information in the credit feature set, and judging the cause of the prediction result based on various factors, the accuracy of the judgment is improved. That is, under the condition of meeting indicators such as precision and recall rate, the interpretability of the prediction result is improved, so as to meet the risk control requirements and improve the risk control efficiency.

[0053] Currently, for the post-loan warning service in the risk control field, the existing technical solutions usually focus on rule control, extract rules from information such as the customer's credit information, personal bank statements, corporate bank statements, and customer assets, and screen customers who may have the risk of overdue default through rules. In addition, machine learning linear models and non-linear modeling can also be used to score each borrower to determine customers who may have the risk of overdue default. However, in the above two solutions, when predicting and interpreting in the way of logistic regression and scoring card model modeling or rule control, although it has strong interpretability and is easy to find the specific reasons for the customer's risk, the rule effect is poor, and there is a certain gap between the indicators of the linear model and the non-linear model; when predicting and interpreting in the way of modeling with integrated models such as Xgboost and Gradient Boosting Decision Tree (GBDT), although its indicator effect can meet the requirements, the interpretability of the non-linear model for the prediction result is low, and the specific reasons for this result cannot be inferred from the predicted result. Therefore, the solutions in the existing technology often cannot fully meet both the model indicator effect and the interpretability of the model.

[0054] Based on the solution defined in the above steps S101 to S104, it can be known that in the embodiment of the present invention, a method of judging the cause of the credit default risk prediction result based on the feature importance ratio and the attribution value is adopted. By obtaining the credit feature set of the target object and predicting the credit feature set, a prediction result and a first result are obtained. Then, when the prediction result indicates that the target object has a default risk, attribution calculation is performed on each feature information in the credit feature set of the target object to obtain a second result. Thus, based on the first result and the second result, the abnormal feature category and the abnormal feature information in the credit feature set are determined. Among them, the credit feature set includes multiple feature information belonging to at least one feature category, the prediction result indicates whether the target object has a default risk, the first result includes the first feature importance ratio of each feature category to the prediction result, and the second feature importance ratio of each feature information to the prediction result. The abnormal feature category is the feature category with the highest contribution degree to the prediction result, and the abnormal feature information is the feature information with the highest contribution degree to the prediction result.

[0055] It is easy to notice that in the above process, based on the first feature importance ratio and the second feature importance ratio in the first result, the feature category and the feature information with relatively high influence degree on the prediction result can be determined. Performing attribution calculation on each feature information in the credit feature set of the target object to obtain the second result can be used to determine the initial contribution degree of each feature information and each feature category to the prediction result. Further, by combining the first result and the second result, the final contribution degree of each feature information and each feature category to the prediction result can be determined, so that the abnormal feature category and the abnormal feature information in the credit feature set can be accurately determined, realizing the judgment of the cause of the prediction result based on multiple factors, and further improving the accuracy of the judgment of the risk warning cause, meeting the risk control requirements and improving the risk control efficiency.

[0056] Thus, it can be seen that the solution provided by the present application achieves the purpose of judging the cause of the credit default risk prediction result based on the feature importance ratio and the attribution value, thereby realizing the technical effect of improving the accuracy of the judgment of the risk warning cause, and further solving the technical problem that the cause of the credit default risk prediction result in the prior art is judged inaccurately.

[0057] In an alternative embodiment, the second result includes the attribution value corresponding to each feature category and the attribution value corresponding to each feature information.

[0058] Optionally, in this embodiment, after the interpretation model calculates the attribution value for each feature information in the credit feature set, the interpretation model can also determine the attribution value of each feature category based on the attribution value corresponding to each feature information, so as to use the attribution value corresponding to each feature category and the attribution value corresponding to each feature information as the second result.

[0059] It should be noted that by determining the attribution value corresponding to each feature category, it is used to determine the abnormal feature category.

[0060] In an alternative embodiment, in the process of predicting the credit feature set to obtain the prediction result and the first result, the risk warning system can predict the credit feature set based on a preset model to obtain the prediction result, then determine the second feature importance ratio corresponding to each feature information based on the feature importance of each feature information, then add up the second feature importance ratios of the feature information corresponding to each feature category to obtain the category feature importance ratio of each feature category, and then add up the second feature importance ratios of all feature information in the credit feature set to obtain the total feature importance ratio, so as to determine the first feature importance ratio based on the category feature importance ratio and the total feature importance ratio. Among them, the feature importance represents the degree of influence of the feature information on the prediction result.

[0061] Among them, in the process of determining the second feature importance ratio corresponding to each feature information based on the feature importance of each feature information, the risk warning system can determine the feature importance of each feature information based on a preset model, then add up the feature importance of all feature information in the credit feature set to obtain the total feature importance, so as to determine the second feature importance ratio based on the feature importance and the total feature importance.

[0062] Optionally, the preset model is an Xgboost model. During the growth of each tree in Xgboost, the feature to be split currently is selected by calculating the Gain coefficient of each feature. Specifically, the calculation process is as follows:

[0063]

[0064] Among them, is the score of the left subtree, is the score of the right subtree, is the score without splitting, and γ is the complexity cost after adding a new leaf node. By calculating the Gani of each feature information, the Gani value of each feature in each tree of the Xgboost model can be obtained, and taking the average value represents the feature importance degree of each feature information. The formula is as follows:

[0065]

[0066] Among them, Fea Gani represents the feature importance of the feature information, and Gain i represents the Gani value of the feature in the i-th tree, and n represents the total number of trees in the Xgboost model.

[0067] Furthermore, after determining the importance of each feature information, the risk warning system can calculate the second feature importance ratio of the feature information based on a preset model, and the formula is as follows:

[0068]

[0069] Among them, Fea_rate represents the second feature importance ratio, and Fea i represents the feature importance of the i-th feature information, and m represents the total number of feature information in the credit feature set. Optionally, the sum of the second feature importance ratios of all feature information is 1, and the larger the second feature importance ratio of a certain feature information, the more important the feature information is and the greater the impact on the model. In this embodiment, the second feature importance ratios of some feature information are as follows:

[0070] Table 2 Feature Information - Second Feature Importance Ratio

[0071] Index Information Second Feature Importance Ratio Average Transaction Balance of Corporate Bank Flows in the Past N Months 0.03 Maximum Transaction Balance of Corporate Bank Flows in the Past N Months 0.013 Standard Deviation of Transaction Balance of Corporate Bank Flows in the Past N Months 0.063 ... ... Total 1

[0072] It can be seen from Table 2 that the second feature importance ratio of "the average balance of the legal person's transaction in the recent N months" is relatively higher, indicating that its impact on the model is relatively greater and this feature information is relatively more important.

[0073] It should be noted that by determining the second feature importance ratio based on the feature importance of each feature information, the accurate determination of the importance degree of each feature information among all feature information is realized, so as to facilitate obtaining a more accurate first result.

[0074] Furthermore, after determining the second feature importance ratio, the risk warning system can calculate the first feature importance ratio of the feature category based on a preset model, and the formula is as follows:

[0075]

[0076] Among them, Fea_cls_rate(c) represents the first feature importance ratio corresponding to the feature category c, and Fea_rate(c) i represents the second feature importance ratio of the i-th feature information among all feature information corresponding to the feature category c, m represents the total number of feature information corresponding to the feature category c, and Fea_rate jThe second feature importance ratio represents the j-th feature information, and n represents the total number of feature information in the combination of credit features. Among them, the sum of the first feature importance ratios of all feature categories is 1. The larger the first feature importance ratio of a certain feature category, the greater the influence of this feature category on the model.

[0077] In this embodiment, the first feature importance ratios of some feature categories are shown as follows:

[0078] Table 3 Feature Category - First Feature Importance Ratio

[0079] Index Sub - category Example Feature Importance Corporate Bank Flows 0.231 Personal Bank Flows 0.349 Personal Assets 0.153 Corporate Assets 0.323 ... ... Total 1

[0080] It can be seen from Table 3 that the first feature importance ratio of "corporate cash flow" is relatively higher, indicating that its influence on the model is relatively greater, and this feature category is relatively more important.

[0081] It should be noted that by determining the first feature importance ratio based on the second feature importance ratio of each feature information, the accurate determination of the importance degree of each feature category among all feature categories is realized, thus facilitating the acquisition of a more accurate first result.

[0082] In an alternative embodiment, in the process of performing attribution calculation on each feature information in the credit feature set of the target object to obtain the second result, the risk warning system can determine the attribution value of each feature information based on the credit feature set through the interpretation model, and perform an addition process on the attribution values of the feature information belonging to the same feature category to obtain the attribution value corresponding to each feature category.

[0083] Optionally, in this embodiment, after determining the first result, the risk warning system can input each feature information, the first result, and the structure file of the aforementioned preset model into the interpretation model to determine the second result based on the interpretation model.

[0084] Specifically, the interpretation model is the Tree SHAP model. This model explains the original complex model through a simple interpretation model, which is defined as an approximation of any interpretation of the original model. It represents the Shapley value interpretation as an additive feature attribution method, and interprets the predicted value of the model as the sum of the attribution values of each feature information. The expression formula is as follows:

[0085]

[0086] Among them, g is the interpretation model, and φ 0 is the constant of the interpretation model, and φ j is the Shapley value (i.e., the attribution value) of each feature information. And the attribution value of each feature can be calculated based on the following formula:

[0087]

[0088] Among them, φ j represents the Shapley value of the j-th feature information, {x1, …, xp} represents the set of all feature information (i.e., the credit feature set), p represents the total number of feature information in the credit feature set, and there are p! combination cases under any permutation and combination. {x1, …, xp}\{x j} represents the possible set of all feature information that does not include {x j}, and fx(S) represents the prediction of the feature subset S, represents the proportion of the combination of feature information of the subset S, and the sum of the proportions of the combinations of feature information of all possible subsets S is equal to 1. Thus, the determination of the Shapley value of each feature information is realized.

[0089] Furthermore, after determining the attribution value of each feature information, the attribution values of the feature information corresponding to the same feature category are added together to obtain the attribution value of the corresponding feature category.

[0090] It should be noted that based on the credit feature set, the attribution value of each feature information is determined, which realizes the accurate determination of the attribution value of each feature information, thus facilitating the obtaining of a more accurate second result.

[0091] In an alternative embodiment, in the process of determining the abnormal feature category and abnormal feature information in the credit feature set based on the first result and the second result, the risk warning system can obtain an explanation result through the explanation model based on the first result and the second result, and thus determine the abnormal feature category and abnormal feature information in the credit feature set based on the explanation result. Among them, the explanation result represents the contribution degree of each feature category to the prediction result, and the contribution degree of each feature information to the prediction result.

[0092] Optionally, the risk warning system can calculate the second feature importance ratio corresponding to each feature information in the first result and the attribution value of each feature information in the second result to determine the contribution degree of each feature information to the prediction result, and can also calculate the first feature importance ratio corresponding to each feature category in the first result and the attribution value of each feature category in the second result to determine the contribution degree of each feature category to the prediction result, so as to select the feature information with the highest contribution degree to the prediction result from all feature information as the abnormal feature information, and select the feature category with the highest contribution degree to the prediction result from all feature categories as the abnormal feature category.

[0093] It should be noted that by determining the contribution degree of each feature category to the prediction result and the contribution degree of each feature information to the prediction result, the abnormal feature information and abnormal feature category can be screened out more accurately.

[0094] In an optional embodiment, in the process of obtaining the explanation result based on the first result and the second result, the explanation model can multiply the attribution value of each feature category with the corresponding first feature importance ratio to obtain the first numerical value corresponding to each feature category, thereby determining the contribution of each feature category to the prediction result based on the first numerical value corresponding to each feature category; and multiply the attribution value of each feature information with the corresponding second feature importance ratio to obtain the second numerical value corresponding to each feature information, thereby determining the contribution of each feature information to the prediction result based on the second numerical value corresponding to each feature information.

[0095] Specifically, since the second feature importance ratio of each feature information calculated by the Xgboost model can characterize the importance of each feature information to the model training and prediction process, and the sum of all second feature importance ratios is 1, it can be used as an algorithm in the weight optimization interpretation model to obtain a more accurate interpretation result. The calculation formula is as follows:

[0096]

[0097] Among them, g(z') represents the explanation model corresponding to the z-th target object, φ 0 is the constant of the explanation model, n represents the total number of feature information, φ j Indicates the Shapley value of the i-th feature information, Fea_rate i The second feature importance ratio of the i-th feature information is represented and used as the weight of the attribution value of the i-th feature information. Thus, the contribution of each feature information to the prediction result can be determined.

[0098] Furthermore, since the amount of feature information may be large, each feature information may be combined according to the feature category to which it belongs, and the contribution of the feature category to the prediction result may be calculated. The formula is as follows:

[0099]

[0100] Among them, g(z',c) represents the explanation model corresponding to the z-th target object, n represents the number of feature information corresponding to feature category c, and φ i (c) represents the attribution value of the i-th feature information among all feature information corresponding to feature category c, Fea_cls_rate(c) represents the first feature importance ratio of feature category c, and it is used as the weight of the attribution value of feature category c, and j represents the total number of feature categories.

[0101] It should be noted that by using the first result as the weight to calculate the interpretation result of the second result, the contribution degree of the feature category and feature information to the prediction result can be calculated more accurately, which is convenient for more accurately screening out abnormal feature information and abnormal feature categories.

[0102] In an alternative embodiment, optionally, the working process of an alternative risk warning system in the present application is described. As Figure 2 shown, the credit feature set of the target object is input into a preset model to obtain a prediction result and a first result, and then the credit feature set, the first result, and the preset model structure are input into an interpretation model to obtain an interpretation result. Finally, the risk warning system outputs abnormal feature information and abnormal feature categories based on the interpretation result, and at the same time outputs the prediction result and the first result. The information output by the risk warning system is as follows:

[0103] Table 4 Information Output by the Risk Warning System (Feature Information)

[0104]

[0105]

[0106] In Table 4, the target object ID is used to determine the identity of the target object. The value corresponding to the first result can represent the feature importance or the second feature importance ratio corresponding to the abnormal feature information. The value corresponding to the prediction result is used to represent whether the target object has a default risk. For example, when the value is 0, it represents that the target object has a default risk, and when the value is 1, it represents that the target object does not have a default risk. The feature information corresponding to the abnormal feature information (1) represents the feature information with the highest contribution degree to the prediction result, and the feature information corresponding to the abnormal feature information (2) represents the feature information with the second highest contribution degree to the prediction result.

[0107] Table 5 Information Output by the Risk Warning System (Feature Category)

[0108] Target Object ID Second Result Prediction Result Abnormal Feature Category (1) Abnormal Feature Category (2) 0101*******1277 0.003459497 0 Personal Bank Flows Personal Credit 0101*******6518 0.006582637 0 Personal Credit Report Enterprise Basic Information 0101*******1428 0.039281327 0 Personal Bank Flows Personal Assets

[0109] In Table 5, the target object ID is used to determine the identity of the target object. The value corresponding to the second result can represent the sum of the feature importance or the first feature importance ratio of all the feature information corresponding to the abnormal feature category. The value corresponding to the prediction result is used to represent whether the target object has a default risk. For example, when the value is 0, it represents that the target object has a default risk, and when the value is 1, it represents that the target object does not have a default risk. The feature category corresponding to the abnormal feature category (1) represents the feature category with the highest contribution degree to the prediction result, and the feature category corresponding to the abnormal feature category (2) represents the feature category with the second highest contribution degree to the prediction result.

[0110] In an alternative embodiment, before predicting the credit feature set to obtain a prediction result and a first result, the risk warning system may obtain the initial credit feature sets of at least one historical credit object, and then process the initial feature information corresponding to each initial feature category based on at least one trained first prediction model to obtain the initial feature importance corresponding to each initial feature information. Then, based on the initial feature importance, target feature information is screened out from the initial feature information corresponding to each initial feature category, so as to train the second prediction model based on the target feature information to obtain a target prediction model. Each initial feature category corresponds to a first prediction model. The initial credit feature set includes multiple initial feature information belonging to different initial feature categories. The target feature information corresponds to the feature information. The target prediction model is used to predict each credit feature set and obtain a prediction result and a first result.

[0111] Optionally, the training process of the foregoing preset model is described. Specifically, in this embodiment, the risk warning system may obtain the relevant field contents of the overdue date and items from the data table containing the historical customer loan overdue information, and then explore useful fields (i.e., unprocessed initial feature information) from the business data tables including customer basic information, customer credit investigation, customer assets, customer turnover, industrial and commercial information, customer taxes, etc., and explore whether the time span of the data meets the sample construction requirements and system requirements. After that, referring to the expert business experience, the RFM method can be used to perform feature derivation on the unprocessed initial feature information to enrich the feature variables in each dimension, and the unprocessed initial feature information after feature derivation is screened based on relevant methods such as the correlation, IV value, and PSI of the features to obtain the initial feature information, so as to depict a more comprehensive and accurate customer risk profile.

[0112] Furthermore, after determining the initial feature information to be counted, the risk warning system may be based on the operating credit customers from X month in 2018 to Y month in 2020, and use the time span from X month in 2018 to M month in 2020 as the training sample and the time span from N month in 2020 to Y month in 2020 as the verification sample. And the good and bad customers are marked by the condition of whether there is an overdue of T days or more, where including the occurrence of bad and overdue of corporate loans; the occurrence of bad and overdue of corporate credit investigation loans and letters of credit; the existence of records of being dishonestly executed in corporate credit investigation; the occurrence of bad of corporate operating express loans; the occurrence of bad of personal loans; the occurrence of bad of personal credit investigation loans or credit cards; the existence of bad debts, asset disposals, surety compensations or records of being dishonestly executed in personal credit investigation; the occurrence of bad of personal operating express loans and other situations and exceeding the preset number of days T are regarded as bad customers, and the non-occurrence of the foregoing situations is regarded as good customers. Thus, the sample construction is realized.

[0113] Even further, as Figure 3As shown, the risk warning system can split the sample data according to feature categories, and input the sample data corresponding to different feature categories into different trained first prediction models. For example, in Figure 3 , the sample data corresponding to the four feature categories of "corporate transaction records", "personal transaction records", "corporate assets", and "personal assets" are respectively input into four first prediction models. Then, based on at least one trained first prediction model, the initial feature information corresponding to each initial feature category is processed to obtain the initial feature importance corresponding to each initial feature information. Among them, the first prediction model is preferably an Xgboost model.

[0114] Among them, the risk warning system can pre-train the first prediction model. During the training process of the first prediction model, the risk warning system can split the sample data according to time span and feature categories, successively train and verify the sample data corresponding to each feature category, and perform model tuning under different data sets to obtain a trained first prediction model.

[0115] After determining the initial feature importance corresponding to each initial feature information in each first prediction model, as Figure 3 shown, the risk warning system can select the top 10 initial feature information with the highest feature importance in each initial feature category as the target feature information, and this target feature information is used to determine the credit feature set of the target object that needs to be obtained when the preset model is actually applied.

[0116] Furthermore, as Figure 3 shown, the foregoing target feature information is input into the second prediction model for training to obtain a target prediction model, that is, the preset model. Among them, the second prediction model is preferably an Xgboost model. In this embodiment, the learning rate of the obtained target prediction model is set to 0.1, the number of trees is 241, the maximum depth of the trees is 5, the optimization function is: binary: logistic, and the random seed is 27.

[0117] It can be seen that the solution provided by this application achieves the purpose of judging the cause of the credit default risk prediction result based on the feature importance ratio and attribution value, improves the business effect while ensuring the business understandability, greatly reduces the labor cost of business personnel for customer tracking and monitoring, realizes the technical effect of improving the accuracy of judging the cause of risk warning, and thus solves the technical problem that the existing technology cannot accurately judge the cause of the credit default risk prediction result. It should be emphasized that this application can be applied to scenarios of predicting credit default risks and causes of credit users in the field of fintech, can also be applied to other scenarios in the field of fintech, and can also be applied to other fields.

[0118] Example 2

[0119] According to an embodiment of the present invention, an embodiment of a data determination device is provided, wherein, Figure 4 is a schematic diagram of an optional data determination device according to an embodiment of the present invention, as Figure 4 shown, the device includes:

[0120] An acquisition module 401, configured to acquire a credit feature set of a target object, wherein the credit feature set includes a plurality of feature information belonging to at least one feature category;

[0121] A prediction module 402, configured to predict the credit feature set to obtain a prediction result and a first result, wherein the prediction result represents whether the target object has a default risk, and the first result includes a first feature importance ratio of each feature category to the prediction result, and a second feature importance ratio of each feature information to the prediction result;

[0122] A calculation module 403, configured to perform attribution calculation on each feature information in the credit feature set of the target object when the prediction result represents that the target object has a default risk, to obtain a second result;

[0123] A determination module 404, configured to determine an abnormal feature category and abnormal feature information in the credit feature set based on the first result and the second result, wherein the abnormal feature category is the feature category with the highest contribution degree to the prediction result, and the abnormal feature information is the feature information with the highest contribution degree to the prediction result.

[0124] It should be noted that the above acquisition module 401, prediction module 402, calculation module 403, and determination module 404 correspond to steps S101 to S104 in the above embodiment. The examples and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment 1.

[0125] Optionally, the second result includes an attribution value corresponding to each feature category and an attribution value corresponding to each feature information.

[0126] Optionally, the prediction module further includes: a sub-prediction module for predicting the credit feature set based on a preset model to obtain a prediction result; a first sub-determination module for determining a second feature importance ratio corresponding to each feature information based on the feature importance of each feature information, where the feature importance characterizes the influence degree of the feature information on the prediction result; a first processing module for adding up the second feature importance ratios of the feature information corresponding to each feature category to obtain a category feature importance ratio for each feature category; a second processing module for adding up the second feature importance ratios of all the feature information in the credit feature set to obtain a total feature importance ratio; a second sub-determination module for determining a first feature importance ratio based on the category feature importance ratio and the total feature importance ratio.

[0127] Optionally, the first sub-determination module further includes: a third sub-determination module for determining the feature importance of each feature information; a third processing module for adding up the feature importance of all the feature information in the credit feature set to obtain a total feature importance; a fourth sub-determination module for determining the second feature importance ratio based on the feature importance and the total feature importance.

[0128] Optionally, the calculation module further includes: a fifth sub-determination module for determining an attribution value for each feature information based on the credit feature set; a fourth processing module for adding up the attribution values of the feature information belonging to the same feature category to obtain an attribution value corresponding to each feature category.

[0129] Optionally, the determination module further includes: a fifth processing module for obtaining an explanation result based on the first result and the second result, where the explanation result characterizes the contribution degree of each feature category to the prediction result and the contribution degree of each feature information to the prediction result; a sixth processing module for determining an abnormal feature category and abnormal feature information in the credit feature set based on the explanation result.

[0130] Optionally, the fifth processing module further includes: a seventh processing module for multiplying the attribution value of each feature category by the corresponding first feature importance ratio to obtain a first value corresponding to each feature category; a sixth determination module for determining the contribution degree of each feature category to the prediction result based on the first value corresponding to each feature category; an eighth processing module for multiplying the attribution value of each feature information by the corresponding second feature importance ratio to obtain a second value corresponding to each feature information; a seventh sub-determination module for determining the contribution degree of each feature information to the prediction result based on the second value corresponding to each feature information.

[0131] Optionally, the data determination module further includes: a sub-acquisition module configured to acquire an initial credit feature set of at least one historical credit object, where the initial credit feature set includes a plurality of initial feature information belonging to different initial feature categories; a ninth processing module configured to process the initial feature information corresponding to each initial feature category based on at least one trained first prediction model to obtain the initial feature importance corresponding to each initial feature information, where each initial feature category corresponds to a first prediction model; a screening module configured to screen out target feature information from the initial feature information corresponding to each initial feature category based on the initial feature importance, where the target feature information corresponds to the feature information; a training module configured to train a second prediction model based on the target feature information to obtain a target prediction model, where the target prediction model is configured to predict each credit feature set and obtain a prediction result and a first result.

[0132] Embodiment 3

[0133] On the other hand, according to an embodiment of the present invention, there is also provided a computer-readable storage medium storing a computer program, where the computer program is configured to execute the above data determination method when running.

[0134] Embodiment 4

[0135] On the other hand, according to an embodiment of the present invention, there is also provided an electronic device, where Figure 5 is a schematic diagram of an optional electronic device according to an embodiment of the present invention, as Figure 5 shown, the electronic device includes one or more processors; a memory configured to store one or more programs, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement a program for running, where the program is configured to execute the above data determination method when running.

[0136] Embodiment 5

[0137] On the other hand, according to an embodiment of the present invention, there is also provided a computer program product including computer program / instructions, and the computer program / instructions implement the above data determination method when executed by a processor.

[0138] The above serial numbers of the embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.

[0139] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0140] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0141] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0142] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0143] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present invention. The aforementioned storage medium includes: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs, etc., which can store program codes.

[0144] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A data determination method, characterized in that, it includes: Obtain a credit feature set of a target object, where the credit feature set includes multiple feature information belonging to at least one feature category; Predict the credit feature set to obtain a prediction result and a first result, where the prediction result represents whether the target object has a default risk, and the first result includes a first feature importance ratio of each feature category to the prediction result, and a second feature importance ratio of each feature information to the prediction result; When the prediction result represents that the target object has a default risk, perform attribution calculation on each feature information in the credit feature set of the target object to obtain a second result; Determine an abnormal feature category and abnormal feature information in the credit feature set based on the first result and the second result, where the abnormal feature category is the feature category with the highest contribution degree to the prediction result, and the abnormal feature information is the feature information with the highest contribution degree to the prediction result; Predict the credit feature set to obtain a prediction result and a first result, including: predicting the credit feature set based on a preset model to obtain the prediction result; determining the second feature importance ratio corresponding to each feature information based on the feature importance degree of each feature information, where the feature importance degree represents the influence degree of the feature information on the prediction result; adding the second feature importance ratios of the feature information corresponding to each feature category to obtain the category feature importance ratio of each feature category; adding the second feature importance ratios of all feature information in the credit feature set to obtain the total feature importance ratio; determining the first feature importance ratio based on the category feature importance ratio and the total feature importance ratio; Determine an abnormal feature category and abnormal feature information in the credit feature set based on the first result and the second result, including: obtaining an explanation result based on the first result and the second result, where the explanation result represents the contribution degree of each feature category to the prediction result, and the contribution degree of each feature information to the prediction result; determining the abnormal feature category and abnormal feature information in the credit feature set based on the explanation result; Obtain an explanation result based on the first result and the second result, including: multiplying the attribution value of each feature category by the corresponding first feature importance ratio to obtain a first value corresponding to each feature category; determining the contribution degree of each feature category to the prediction result based on the first value corresponding to each feature category; multiplying the attribution value of each feature information by the corresponding second feature importance ratio to obtain a second value corresponding to each feature information; determining the contribution degree of each feature information to the prediction result based on the second value corresponding to each feature information.

2. The method according to claim 1, characterized in that, The second result includes the attribution value corresponding to each feature category and the attribution value corresponding to each feature information.

3. The method according to claim 1, wherein, determining the second feature importance ratio corresponding to each feature information based on the feature importance of each feature information includes: determining the feature importance of each feature information; performing an addition process on the feature importances of all feature information in the credit feature set to obtain a total feature importance; determining the second feature importance ratio based on the feature importance and the total feature importance.

4. The method according to claim 2, wherein, performing an attribution calculation on each feature information in the credit feature set of the target object to obtain a second result, including: determining the attribution value of each feature information based on the credit feature set; performing an addition process on the attribution values of the feature information belonging to the same feature category to obtain the attribution value corresponding to each feature category.

5. The method according to claim 1, wherein, before predicting the credit feature set to obtain a prediction result and a first result, the method further includes: acquiring an initial credit feature set of at least one historical credit object, wherein the initial credit feature set includes a plurality of initial feature information belonging to different initial feature categories; processing the initial feature information corresponding to each initial feature category based on at least one trained first prediction model to obtain the initial feature importance corresponding to each initial feature information, wherein each initial feature category corresponds to a first prediction model; screening out target feature information from the initial feature information corresponding to each initial feature category based on the initial feature importance, wherein the target feature information corresponds to the feature information; training a second prediction model based on the target feature information to obtain a target prediction model, wherein the target prediction model is used to predict each credit feature set and obtain a prediction result and a first result.

6. A data determination device, wherein, comprising: an acquisition module, configured to acquire a credit feature set of a target object, wherein the credit feature set includes a plurality of feature information belonging to at least one feature category; a prediction module, configured to predict the credit feature set to obtain a prediction result and a first result, wherein the prediction result represents whether the target object has a default risk, and the first result includes a first feature importance ratio of each feature category to the prediction result and a second feature importance ratio of each feature information to the prediction result; a calculation module, configured to perform an attribution calculation on each feature information in the credit feature set of the target object to obtain a second result when the prediction result represents that the target object has a default risk. A determination module, configured to determine an abnormal feature category and abnormal feature information in the credit feature set based on the first result and the second result, where the abnormal feature category is the feature category with the highest contribution degree to the prediction result, and the abnormal feature information is the feature information with the highest contribution degree to the prediction result; Wherein, the prediction module further includes: a sub-prediction module, configured to predict the credit feature set based on a preset model to obtain the prediction result; a first sub-determination module, configured to determine a second feature importance ratio corresponding to each feature information based on the feature importance degree of each feature information, where the feature importance degree characterizes the influence degree of the feature information on the prediction result; a first processing module, configured to perform an addition process on the second feature importance ratios of the feature information corresponding to each feature category to obtain a category feature importance ratio of each feature category; a second processing module, configured to perform an addition process on the second feature importance ratios of all feature information in the credit feature set to obtain a total feature importance ratio; a second sub-determination module, configured to determine the first feature importance ratio based on the category feature importance ratio and the total feature importance ratio; Wherein, the determination module further includes: a fifth processing module, configured to obtain an explanation result based on the first result and the second result, where the explanation result characterizes the contribution degree of each feature category to the prediction result and the contribution degree of each feature information to the prediction result; a sixth processing module, configured to determine an abnormal feature category and abnormal feature information in the credit feature set based on the explanation result; The fifth processing module further includes: a seventh processing module, configured to multiply the attribution value of each feature category by the corresponding first feature importance ratio to obtain a first value corresponding to each feature category; a sixth determination module, configured to determine the contribution degree of each feature category to the prediction result based on the first value corresponding to each feature category; an eighth processing module, configured to multiply the attribution value of each feature information by the corresponding second feature importance ratio to obtain a second value corresponding to each feature information; a seventh sub-determination module, configured to determine the contribution degree of each feature information to the prediction result based on the second value corresponding to each feature information.

7. A computer-readable storage medium, characterized in that, a computer program is stored in the computer-readable storage medium, where the computer program is configured to execute the data determination method described in any one of claims 1 to 5 when running.

8. An electronic device, characterized in that, the electronic device includes one or more processors; a memory, configured to store one or more programs, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement a program for running, where the program is configured to execute the data determination method described in any one of claims 1 to 5 when running.

9. A computer program product, including a computer program / instructions, It is characterized in that when the computer program / instructions are executed by a processor, the data determination method described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Data prediction method and computer readable storage medium

    CN110349012A

  • Method and device for dynamically generating early warning rules

    CN111369344A