Model interpretation information generation method and device, equipment, medium and program product

By standardizing and clustering the user sample set and combining the regression model construction method, the problem of the simple machine learning model poor fitting ability to unstructured data and the black box model cannot explain group samples is solved, achieving high-accurate feature importance interpretation.

CN120105358APending Publication Date: 2025-06-06JINGDONG TECH HLDG CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311648655.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Simple machine learning models have poor fitting ability to unstructured data, resulting in poor accuracy in interpreting the importance of sample features, and methods such as SHAP values ​​and LIME of complex black box models cannot explain the importance of feature for population samples.

Method used

By standardizing the user sample set and clustering based on the standardized user sample set and output score set, a set of user sample sets is obtained. Then, a regression model is constructed for user samples in each user sample group that meet the preset score conditions, and a model interpretation information corresponding to the black box model is generated.

Benefits of technology

The accuracy of the interpretation of the importance of sample characteristics is improved, and the interpretation of the importance of characteristic of population samples is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105358A_ABST
    Figure CN120105358A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a model interpretation information generation method and device, equipment, a medium and a program product. A specific embodiment of the method comprises the steps of performing standardization processing on each target user feature of each user sample; standardizing an output score set of the user sample set corresponding to the black box model; clustering the user sample set; for each user sample group, executing the following steps: selecting each user sample meeting a preset first score condition from the user sample group as a first user sample set; generating a first regression model according to the first user sample set and each output score corresponding to the first user sample set in the output score set; and generating model interpretation information corresponding to the first user sample set and the black box model according to the model parameter information of the first regression model. The implementation mode is related to credible artificial intelligence, the accuracy of importance interpretation of sample features is improved, and feature importance interpretation can be carried out for group samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular to a method, apparatus, device, medium, and program product for generating model interpretation information. Background Art

[0002] In order to effectively describe the model's scoring logic for users and enhance the interpretability of the scoring, modelers usually use simple machine learning models such as LR and XGB. Since such models can output the importance of features, they can effectively describe the model's scoring logic in the overall scoring customer group. In addition, for some complex black box models, methods such as SHAP value and LIME are usually used to explain the importance of features for a single sample.

[0003] However, the inventors found that when the above method is used to interpret the scores output by the model, the following technical problems often occur: simple machine learning models have poor fitting capabilities for unstructured data, resulting in poor accuracy in interpreting the importance of sample features. Methods such as SHAP values ​​and LIME used in complex black box models cannot interpret the importance of features for group samples.

[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the invention

[0005] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.

[0006] Some embodiments of the present disclosure propose a method, apparatus, electronic device, computer-readable medium, and computer program product for generating model explanation information to solve one or more of the technical problems mentioned in the above background technology section.

[0007] In a first aspect, some embodiments of the present disclosure provide a method for generating model explanation information, the method comprising: standardizing each target user feature of each user sample in a user sample set to obtain a standardized user sample set; standardizing the output score set of a black box model corresponding to the above user sample set to obtain a standardized output score set; clustering the above user sample set based on the above standardized user sample set and the above standardized output score set to obtain a user sample group set; for each user sample group in the above user sample group set, performing the following steps: selecting each user sample that meets a preset first score condition from the above user sample group as a first user sample set; generating a first regression model based on the above first user sample set and each output score corresponding to the above first user sample set in the above output score set; generating model explanation information corresponding to the above first user sample set and the above black box model based on model parameter information of the above first regression model.

[0008] Optionally, after selecting each user sample that meets a preset first score condition from the user sample group as the first user sample set, the method further includes: selecting each user sample that meets a preset second score condition from the user sample group as a second user sample set; generating a second regression model based on the second user sample set and each output score corresponding to the second user sample set in the output score set; and generating model explanation information corresponding to the second user sample set and the black box model based on model parameter information of the second regression model.

[0009] Optionally, the above-mentioned generating model interpretation information corresponding to the above-mentioned first user sample set and the above-mentioned black box model based on the model parameter information of the above-mentioned first regression model includes: arranging the various weights included in the model parameter information of the above-mentioned first regression model in descending order to obtain a weight sequence as a first weight sequence; selecting a preset number of first weights from the above-mentioned first weight sequence as a target first weight sequence; and generating model interpretation information corresponding to the above-mentioned first user sample set and the above-mentioned black box model according to the feature name of each target user feature corresponding to the above-mentioned target first weight sequence.

[0010] Optionally, the above-mentioned generating model explanation information corresponding to the first user sample set and the black box model according to the model parameter information of the first regression model includes: for each weight included in the model parameter information of the first regression model, performing the following steps: determining the feature name of the target user feature corresponding to the weight as the first feature name; determining the mean of each target user feature corresponding to the first feature name in the first user sample set as the first feature mean; determining the mode of each target user feature corresponding to the first feature name in the first user sample set as the first feature mode; combining the first feature name, the weight, the first feature mean and the first feature mode into first feature explanation information; and determining each combined first feature explanation information as the model explanation information corresponding to the first user sample set and the black box model.

[0011] Optionally, the above-mentioned generating model interpretation information corresponding to the above-mentioned second user sample set and the above-mentioned black box model based on the model parameter information of the above-mentioned second regression model includes: arranging the various weights included in the model parameter information of the above-mentioned second regression model in descending order to obtain a weight sequence as a second weight sequence; selecting a preset number of second weights from the above-mentioned second weight sequence as a target second weight sequence; and generating model interpretation information corresponding to the above-mentioned second user sample set and the above-mentioned black box model according to the feature name of each target user feature corresponding to the above-mentioned target second weight sequence.

[0012] Optionally, the above-mentioned generating model explanation information corresponding to the second user sample set and the black box model according to the model parameter information of the second regression model includes: for each weight included in the model parameter information of the second regression model, performing the following steps: determining the feature name of the target user feature corresponding to the weight as the second feature name; determining the mean of each target user feature corresponding to the second feature name in the second user sample set as the second feature mean; determining the mode of each target user feature corresponding to the second feature name in the second user sample set as the second feature mode; combining the second feature name, the weight, the second feature mean and the second feature mode into second feature explanation information; determining each combined second feature explanation information as the model explanation information corresponding to the second user sample set and the black box model.

[0013] Optionally, after generating model explanation information corresponding to the second user sample set and the black box model based on the model parameter information of the second regression model, the method further includes: displaying the generated model explanation information corresponding to the first user sample set and the model explanation information corresponding to the second user sample set in a tabular form; in response to detecting a confirmation operation of the model explanation information corresponding to the first user sample set and the model explanation information corresponding to the second user sample set, sending each output score corresponding to the first user sample set and the second user sample set in the output score set to a downstream task end.

[0014] In a second aspect, some embodiments of the present disclosure provide a model explanation information generating device, the device comprising: a first standardization processing unit, configured to perform standardization processing on each target user feature of each user sample in a user sample set to obtain a standardized user sample set; a second standardization processing unit, configured to perform standardization processing on the output score set of a black box model corresponding to the above user sample set to obtain a standardized output score set; a clustering unit, configured to cluster the above user sample set based on the above standardized user sample set and the above standardized output score set to obtain a user sample group set; an execution unit, configured to perform the following steps for each user sample group in the above user sample group set: select each user sample that meets a preset first score condition from the above user sample group as a first user sample set; generate a first regression model according to the above first user sample set and each output score corresponding to the above first user sample set in the above output score set; and generate model explanation information corresponding to the above first user sample set and the above black box model according to the model parameter information of the above first regression model.

[0015] Optionally, after selecting each user sample that meets the preset first score condition from the user sample group as the first user sample set, the model explanation information generating device further includes: a selection unit, a first generation unit, and a second generation unit. The selection unit is configured to select each user sample that meets the preset second score condition from the user sample group as the second user sample set. The first generation unit is configured to generate a second regression model based on the second user sample set and each output score corresponding to the second user sample set in the output score set. The second generation unit is configured to generate model explanation information corresponding to the second user sample set and the black box model based on the model parameter information of the second regression model.

[0016] Optionally, the execution unit is further configured to: arrange the weights included in the model parameter information of the above-mentioned first regression model in descending order to obtain a weight sequence as a first weight sequence; select a preset number of first weights from the above-mentioned first weight sequence as a target first weight sequence; and generate model interpretation information corresponding to the above-mentioned first user sample set and the above-mentioned black box model according to the feature names of each target user feature corresponding to the above-mentioned target first weight sequence.

[0017] Optionally, the execution unit is further configured to: for each weight included in the model parameter information of the above-mentioned first regression model, perform the following steps: determine the feature name of the target user feature corresponding to the above-mentioned weight as the first feature name; determine the mean of each target user feature corresponding to the above-mentioned first feature name in the above-mentioned first user sample set as the first feature mean; determine the mode of each target user feature corresponding to the above-mentioned first feature name in the above-mentioned first user sample set as the first feature mode; combine the above-mentioned first feature name, the above-mentioned weight, the above-mentioned first feature mean and the above-mentioned first feature mode into the first feature explanation information; determine the combined each first feature explanation information as the model explanation information corresponding to the above-mentioned first user sample set and the above-mentioned black box model.

[0018] Optionally, the second generation unit is further configured to: arrange the weights included in the model parameter information of the above-mentioned second regression model in descending order to obtain a weight sequence as a second weight sequence; select a preset number of second weights from the above-mentioned second weight sequence as a target second weight sequence; and generate model interpretation information corresponding to the above-mentioned second user sample set and the above-mentioned black box model according to the feature names of each target user feature corresponding to the above-mentioned target second weight sequence.

[0019] Optionally, the second generation unit is further configured to: for each weight included in the model parameter information of the above-mentioned second regression model, perform the following steps: determine the feature name of the target user feature corresponding to the above-mentioned weight as the second feature name; determine the mean of each target user feature corresponding to the above-mentioned second feature name in the above-mentioned second user sample set as the second feature mean; determine the mode of each target user feature corresponding to the above-mentioned second feature name in the above-mentioned second user sample set as the second feature mode; combine the above-mentioned second feature name, the above-mentioned weight, the above-mentioned second feature mean and the above-mentioned second feature mode into second feature explanation information; determine the combined each second feature explanation information as the model explanation information corresponding to the above-mentioned second user sample set and the above-mentioned black box model.

[0020] Optionally, after generating the model interpretation information corresponding to the second user sample set and the black box model according to the model parameter information of the second regression model, the model interpretation information generating device further includes: a display unit and a sending unit. The display unit is configured to display the generated model interpretation information corresponding to the first user sample set and the model interpretation information corresponding to the second user sample set in a table. The sending unit is configured to send each output score corresponding to the first user sample set and the second user sample set in the output score set to the downstream task end in response to detecting the confirmation operation of the model interpretation information corresponding to the first user sample set and the model interpretation information corresponding to the second user sample set.

[0021] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the above-mentioned first aspect.

[0022] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the above-mentioned first aspect is implemented.

[0023] In a fifth aspect, some embodiments of the present disclosure provide a computer program product, including a computer program, which implements the method described in any implementation manner of the above-mentioned first aspect when executed by a processor.

[0024] The above-mentioned embodiments of the present disclosure have the following beneficial effects: through the model explanation information generation method of some embodiments of the present disclosure, the accuracy of the interpretation of the importance of sample features is improved, and the feature importance can be explained for group samples. Specifically, the reason for the poor accuracy of the interpretation of the importance of sample features and the inability to interpret the feature importance for group samples is that the simple machine learning model has poor fitting ability for unstructured data, resulting in poor accuracy of the interpretation of the importance of sample features. For the SHAP value, LIME and other methods used in the complex black box model, it is impossible to interpret the feature importance for group samples. Based on this, the model explanation information generation method of some embodiments of the present disclosure firstly performs standardization processing on each target user feature of each user sample in the user sample set to obtain a standardized user sample set. Thus, each target user feature of the user sample can be unified to the same dimension. Then, the output score set of the black box model corresponding to the above user sample set is standardized to obtain a standardized output score set. Thus, the output score output by the black box model can be mapped to the same score space. Afterwards, based on the above standardized user sample set and the above standardized output score set, the above user sample set is clustered to obtain a user sample group set. Thus, user samples can be clustered into various classes with user features and scores. Finally, for each user sample group in the above user sample group set, the following steps are performed: Step 1, select each user sample that meets the preset first score condition from the above user sample group as the first user sample set. Thus, each user sample whose score size meets the preset first score condition can be screened out from each user sample group. Step 2, generate a first regression model according to the above first user sample set and each output score corresponding to the above first user sample set in the above output score set. Thus, a regression model can be constructed using each user sample whose score size meets the preset first score condition and each original output score. Step 3, generate model explanation information corresponding to the above first user sample set and the above black box model according to the model parameter information of the above first regression model. Thus, the model explanation information corresponding to the first user sample set and the above black box model can be determined using the model parameter information of the regression model. Also because the output score of the user sample is determined by the black box model, the black box model has a strong fitting ability for unstructured data, thereby improving the accuracy of the interpretation of the importance of sample features. Also, because the model explanation information is for the first user sample set, and the first user sample set is selected from a class of user sample groups and the scores meet the preset first score condition, the first user sample set can represent a class of user group samples, thereby realizing the interpretation of feature importance for group samples. Therefore, the accuracy of the interpretation of the importance of sample features is improved, and the interpretation of feature importance can be performed for group samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0026] Figure 1 is a schematic diagram of an application scenario of a method for generating model explanation information according to some embodiments of the present disclosure;

[0027] Figure 2 is a flowchart of some embodiments of the method for generating model explanation information according to the present disclosure;

[0028] Figure 3 are flow charts of other embodiments of the method for generating model explanation information according to the present disclosure;

[0029] Figure 4 is a schematic diagram of the structure of some embodiments of the model interpretation information generating device according to the present disclosure;

[0030] Figure 5 It is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0031] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0032] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.

[0033] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0034] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0035] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0036] With regard to the collection, storage, and use of user personal information (such as user samples) involved in this disclosure, before performing the corresponding operations, the relevant organizations or individuals shall fulfill their obligations, including conducting personal information security impact assessments, fulfilling the obligation to inform the personal information subject, and obtaining the authorization and consent of the personal information subject in advance.

[0037] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0038] Figure 1 It is a schematic diagram of an application scenario of the method for generating model explanation information according to some embodiments of the present disclosure.

[0039] exist Figure 1 In the application scenario of, first, the computing device 101 can standardize each target user feature of each user sample in the user sample set 102 to obtain a standardized user sample set 103. Then, the computing device 101 can standardize the output score set 105 of the black box model 104 corresponding to the above user sample set 102 to obtain a standardized output score set 106. After that, the computing device 101 can cluster the above user sample set based on the above standardized user sample set 103 and the above standardized output score set 106 to obtain a user sample group set 107. Finally, for each user sample group (for example, user sample group 1071) in the above user sample group set 107, the computing device 101 can perform the following steps: First, select each user sample that meets the preset first score condition from the above user sample group 1071 as the first user sample set 1073. The above user sample group set 107 may include user sample group 1071 and user sample group 1072. The second step is to generate a first regression model 1075 based on the first user sample set 1073 and each output score 1074 corresponding to the first user sample set 1073 in the output score set 105. The third step is to generate model explanation information 1076 corresponding to the first user sample set 1073 based on the model parameter information of the first regression model 1075.

[0040] It should be noted that the computing device 101 can be hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or as a single server or a single terminal device. When the computing device is embodied as software, it can be installed in the hardware devices listed above. It can be implemented as multiple software or software modules for providing distributed services, or as a single software or software module. No specific limitation is made here.

[0041] It should be understood that Figure 1 The number of computing devices in the embodiment is only illustrative. Any number of computing devices may be provided according to implementation requirements.

[0042] Continue to refer Figure 2 , shows a process 200 of some embodiments of the method for generating model explanation information according to the present disclosure. The method for generating model explanation information includes the following steps:

[0043] Step 201 , standardize each target user feature of each user sample in the user sample set to obtain a standardized user sample set.

[0044] In some embodiments, the execution subject (eg Figure 1 The computing device shown in the figure) can perform standardization processing on each target user feature of each user sample in the user sample set to obtain a standardized user sample set. Among them, the above-mentioned user sample set can be attribute-related information of each user including each user feature. Here, the user feature can be continuous data. For example, the user features of the user sample can include but are not limited to at least one of the following: the number of items circulated within a preset time period, the total value of items circulated, the repayment rate on time, and the number of overdue payments. The above-mentioned target user features can be user features with strong pre-screening effect and strong explanatory power. The information value of the above-mentioned target user features can be greater than the first preset threshold, and the statistics of the above-mentioned target user features can be greater than the preset threshold. The above-mentioned target user features can be structured data with clear business descriptions, and the bad debt rate change trend of each segment is consistent with the prior knowledge, and there is a monotonic relationship with the risk label, which can facilitate the subsequent characterization of population attributes.

[0045] In practice, the above-mentioned execution subject can perform z-score normalization processing on each target user feature of each user sample in the user sample set to obtain a standardized user sample set. The standardized user sample includes the feature values ​​of each target user feature after the normalization processing. Optionally, the normalization processing method can also be but not limited to: entropy normalization, maximum and minimum normalization, centralization normalization, and regularization normalization. In this way, it can be ensured that the dimensions of different features can be aligned to facilitate subsequent grouping and clustering.

[0046] Step 202 , normalize the output score set of the black box model corresponding to the user sample set to obtain a standardized output score set.

[0047] In some embodiments, the execution entity may standardize the output score set of the black box model corresponding to the user sample set to obtain a standardized output score set. The black box model may be a neural network model for generating output scores corresponding to each user based on the characteristics of each target user in the user sample set. For example, the black box model may be a DNN. The characteristics of each target user may be used as input features of the black box model. The output score may represent the score of the user. The output score may be a probability value in the range of [0,1]. For example, the output score may represent the score of the user's credit. In practice, for each output score in the output score set, the execution entity may standardize the output score using the following formula to obtain a standardized output score:

[0048]

[0049] Among them, Y s Represents the standardized output score. Base represents a constant, which can be used as the base score after standardization. For example, Base can be 600. Y represents the output score of the black box model. B is a constant, which means that every time the ratio of the black box model's predicted positive probability Y to the negative probability 1-Y doubles, B is added to Base. For example, B can be a value between [10,50].

[0050] Step 203: clustering the user sample set based on the standardized user sample set and the standardized output score set to obtain a user sample group set.

[0051] In some embodiments, the execution subject may cluster the user sample set based on the standardized user sample set and the standardized output score set to obtain a user sample group set. In practice, the execution subject may cluster the user sample set using the standardized user sample and the standardized output score as features for determining the distance between two user samples to obtain a user sample group set. Specifically, the execution subject may use the kmeans algorithm for clustering. The distance between two user samples may be a Euclidean distance.

[0052] Step 204: for each user sample group in the user sample group set, perform the following steps:

[0053] Step 2041: Select each user sample that meets a preset first score condition from the user sample group as a first user sample set.

[0054] In some embodiments, the execution subject may select each user sample that satisfies the preset first score condition from the user sample group as the first user sample set. The preset first score condition may be that the output score corresponding to the user sample is greater than the first preset score. Here, the specific setting of the first preset score is not limited. Thus, each user sample with a higher output score may be screened out from the user sample group, and each user corresponding to the first user sample set may be a high-score group.

[0055] Step 2042: Generate a first regression model according to the first user sample set and each output score in the output score set corresponding to the first user sample set.

[0056] In some embodiments, the execution subject may generate a first regression model based on the first user sample set and the output scores corresponding to the first user sample set in the output score set. In practice, the execution subject may use the output score as the dependent variable, the unstandardized target user features as the independent variable, and the first user sample set and the output scores corresponding to the first user sample set in the output score set to construct a linear regression model as the first regression model. For example, the execution subject may construct a linear regression model using a packaged linear regression model construction method. The loss function of the linear regression model may be:

[0057]

[0058] Wherein, n represents the number of first user samples included in the first user sample set. i Represents the original output score of the black box model corresponding to the i-th first user sample. Represents the output score of the i-th first user sample output by the constructed linear regression model.

[0059] Step 2043: Generate model explanation information corresponding to the first user sample set and the black box model according to the model parameter information of the first regression model.

[0060] In some embodiments, the execution entity may generate model explanation information corresponding to the first user sample set and the black box model based on the model parameter information of the first regression model. The model parameter information may include information of each independent variable. The independent variable information may include weights and feature names. In practice, the execution entity may sort the independent variable information in descending order of weight to obtain an independent variable information sequence as the model explanation information corresponding to the first user sample set and the black box model. The model explanation information corresponding to the first user sample set and the black box model may be understood as a feature that has a high influence on the output score of the black box model corresponding to the first user sample in the first user sample set.

[0061] In some optional implementations of some embodiments, the execution subject may generate model explanation information corresponding to the first user sample set and the black box model according to the model parameter information of the first regression model through the following steps:

[0062] In the first step, the weights included in the model parameter information of the first regression model are arranged in descending order to obtain a weight sequence as a first weight sequence.

[0063] The second step is to select a preset number of first weights from the above-mentioned first weight sequence as the target first weight sequence. Here, the specific setting of the preset number is not limited. In practice, the above-mentioned execution entity can select a first preset number of first weights from the front end of the above-mentioned first weight sequence and a second preset number of first weights from the back end of the above-mentioned first weight sequence to form the target first weight sequence in response to determining that there are negative numbers in the above-mentioned first weight sequence. The sum of the first preset number and the second preset number is the above-mentioned preset number. In response to determining that there are no negative numbers in the above-mentioned first weight sequence, the above-mentioned execution entity can select a preset number of first weights from the front end of the above-mentioned first weight sequence to form the target first weight sequence.

[0064] The third step is to generate model explanation information corresponding to the first user sample set and the black box model according to the feature names of each target user feature corresponding to the first target weight sequence. In practice, for each target first weight in the first target weight sequence, the execution subject can combine the target first weight and the feature name of the target user feature corresponding to the first target weight into independent variable information. Then, the combined independent variable information can be determined as the model explanation information corresponding to the first user sample set and the black box model. In this way, features with strong interpretability can be screened out for model interpretation of the first user sample set.

[0065] In some optional implementations of some embodiments, the execution subject may generate model explanation information corresponding to the first user sample set and the black box model according to the model parameter information of the first regression model through the following steps:

[0066] In the first step, for each weight included in the model parameter information of the first regression model, the following steps are performed:

[0067] In the first sub-step, the feature name of the target user feature corresponding to the weight is determined as the first feature name.

[0068] In the second sub-step, the mean of each target user feature corresponding to the first feature name in the first user sample set is determined as the first feature mean.

[0069] In a third sub-step, the mode of each target user feature corresponding to the first feature name in the first user sample set is determined as the first feature mode.

[0070] The fourth sub-step is to combine the first feature name, the weight, the first feature mean and the first feature mode into first feature explanation information.

[0071] In the second step, the combined first feature explanation information is determined as the model explanation information corresponding to the first user sample set and the black box model. Thus, the model explanation information can be combined from the dimensions of weight, mean and mode, so that the user can understand the model explanation information corresponding to the second user sample set and the black box model from multiple dimensions.

[0072] The above-mentioned embodiments of the present disclosure have the following beneficial effects: through the model explanation information generation method of some embodiments of the present disclosure, the accuracy of the interpretation of the importance of sample features is improved, and the feature importance can be explained for group samples. Specifically, the reason for the poor accuracy of the interpretation of the importance of sample features and the inability to interpret the feature importance for group samples is that the simple machine learning model has poor fitting ability for unstructured data, resulting in poor accuracy of the interpretation of the importance of sample features. For the SHAP value, LIME and other methods used in the complex black box model, it is impossible to interpret the feature importance for group samples. Based on this, the model explanation information generation method of some embodiments of the present disclosure firstly performs standardization processing on each target user feature of each user sample in the user sample set to obtain a standardized user sample set. Thus, each target user feature of the user sample can be unified to the same dimension. Then, the output score set of the black box model corresponding to the above user sample set is standardized to obtain a standardized output score set. Thus, the output score output by the black box model can be mapped to the same score space. Afterwards, based on the above standardized user sample set and the above standardized output score set, the above user sample set is clustered to obtain a user sample group set. Thus, user samples can be clustered into various classes with user features and scores. Finally, for each user sample group in the above user sample group set, the following steps are performed: Step 1, select each user sample that meets the preset first score condition from the above user sample group as the first user sample set. Thus, each user sample whose score size meets the preset first score condition can be screened out from each user sample group. Step 2, generate a first regression model according to the above first user sample set and each output score corresponding to the above first user sample set in the above output score set. Thus, a regression model can be constructed using each user sample whose score size meets the preset first score condition and each original output score. Step 3, generate model explanation information corresponding to the above first user sample set and the above black box model according to the model parameter information of the above first regression model. Thus, the model explanation information corresponding to the first user sample set and the above black box model can be determined using the model parameter information of the regression model. Also because the output score of the user sample is determined by the black box model, the black box model has a strong fitting ability for unstructured data, thereby improving the accuracy of the interpretation of the importance of sample features. Also, because the model explanation information is for the first user sample set, and the first user sample set is selected from a class of user sample groups and the scores meet the preset first score condition, the first user sample set can represent a class of user group samples, thereby realizing the interpretation of feature importance for group samples. Therefore, the accuracy of the interpretation of the importance of sample features is improved, and the interpretation of feature importance can be performed for group samples.

[0073] Further references Figure 3 , which shows a process 300 of another embodiment of the method for generating model explanation information. The process 300 of the method for generating model explanation information includes the following steps:

[0074] Step 301 , standardize each target user feature of each user sample in the user sample set to obtain a standardized user sample set.

[0075] Step 302 , normalize the output score set of the black box model corresponding to the user sample set to obtain a standardized output score set.

[0076] Step 303: clustering the user sample set based on the standardized user sample set and the standardized output score set to obtain a user sample group set.

[0077] Step 304: for each user sample group in the user sample group set, perform the following steps:

[0078] Step 3041: Select each user sample that meets a preset first score condition from the user sample group as a first user sample set.

[0079] Step 3042: Generate a first regression model according to the first user sample set and each output score in the output score set corresponding to the first user sample set.

[0080] Step 3043: Generate model explanation information corresponding to the first user sample set and the black box model according to the model parameter information of the first regression model.

[0081] In some embodiments, the specific implementation of steps 301-3043 and the technical effects brought about can be referred to Figure 2 The corresponding steps 201-2043 in the embodiments are not repeated here.

[0082] Step 3044: Select each user sample that meets the preset second score condition from the user sample group as the second user sample set.

[0083] In some embodiments, the execution subject (eg Figure 1 The computing device shown in the figure) can select each user sample that meets the preset second score condition from the above user sample group as the second user sample set. The above preset second score condition can be that the output score corresponding to the user sample is less than the second preset score. The second preset score is less than the first preset score. Here, the specific setting of the second preset score is not limited. Thus, each user sample with a low output score can be screened out from the user sample group, and each user corresponding to the second user sample set can be used as a low-score group.

[0084] Step 3045: Generate a second regression model according to the second user sample set and each output score in the output score set corresponding to the second user sample set.

[0085] In some embodiments, the execution subject may generate a second regression model based on the second user sample set and the output scores corresponding to the second user sample set in the output score set. In practice, the execution subject may use the output score as the dependent variable, each unstandardized target user feature as the independent variable, and use the second user sample set and the output scores corresponding to the second user sample set in the output score set to construct a linear regression model as the second regression model.

[0086] Step 3046: Generate model explanation information corresponding to the second user sample set and the black box model based on the model parameter information of the second regression model.

[0087] In some embodiments, the execution subject may generate model explanation information corresponding to the second user sample set and the black box model according to the model parameter information of the second regression model. The model parameter information of the second regression model may include information of each independent variable. The independent variable information may include weights and user feature names. In practice, the execution subject may sort the independent variable information included in the model parameter information of the second regression model in descending order of weight, and obtain an independent variable information sequence as the model explanation information corresponding to the second user sample set and the black box model.

[0088] In some optional implementations of some embodiments, the execution subject may generate model explanation information corresponding to the second user sample set and the black box model according to the model parameter information of the second regression model through the following steps:

[0089] In the first step, the weights included in the model parameter information of the second regression model are arranged in descending order to obtain a weight sequence as a second weight sequence.

[0090] The second step is to select a preset number of second weights from the second weight sequence as the target second weight sequence. In practice, the execution subject may, in response to determining that there are negative numbers in the second weight sequence, select a first preset number of second weights from the front end of the second weight sequence and a second preset number of second weights from the back end of the second weight sequence to form the target second weight sequence. The sum of the first preset number and the second preset number is the preset number. In response to determining that there are no negative numbers in the second weight sequence, the execution subject may select a preset number of second weights from the front end of the second weight sequence to form the target second weight sequence.

[0091] The third step is to generate model explanation information corresponding to the second user sample set and the black box model according to the feature names of each target user feature corresponding to the target second weight sequence. In practice, for each target second weight in the target second weight sequence, the execution subject can combine the target second weight and the feature name of the target user feature corresponding to the target second weight into independent variable information. Then, the combined independent variable information can be determined as the model explanation information corresponding to the second user sample set and the black box model. In this way, features with strong interpretability can be screened out for model interpretation of the second user sample set.

[0092] In some optional implementations of some embodiments, the execution subject may generate model explanation information corresponding to the second user sample set and the black box model according to the model parameter information of the second regression model through the following steps, including:

[0093] In the first step, for each weight included in the model parameter information of the second regression model, the following steps are performed:

[0094] In the first sub-step, the feature name of the target user feature corresponding to the weight is determined as the second feature name.

[0095] In the second sub-step, the mean of each target user feature corresponding to the second feature name in the second user sample set is determined as the second feature mean.

[0096] In a third sub-step, the mode of each target user feature corresponding to the second feature name in the second user sample set is determined as the second feature mode.

[0097] The fourth sub-step is to combine the second feature name, the weight, the second feature mean and the second feature mode into second feature explanation information.

[0098] In the second step, the combined second feature explanation information is determined as the model explanation information corresponding to the second user sample set and the black box model. Thus, the model explanation information can be combined from the dimensions of weight, mean and mode, so that the user can understand the model explanation information corresponding to the second user sample set and the black box model from multiple dimensions.

[0099] Optionally, after step 3046, the execution entity may also display the generated model explanation information corresponding to the first user sample set and the generated model explanation information corresponding to the second user sample set in the form of a table. In practice, the execution entity may display the generated model explanation information corresponding to the first user sample set in one table. And the generated model explanation information corresponding to the second user sample set may be displayed in another table. Thus, the user can intuitively view the model explanation information corresponding to the high-scoring population and the low-scoring population in a user sample group.

[0100] Then, in response to the confirmation operation of detecting the model interpretation information corresponding to the first user sample set and the model interpretation information corresponding to the second user sample set, the output scores corresponding to the first user sample set and the second user sample set in the output score set can be sent to the downstream task end. Among them, the confirmation operation can be an operation of confirming and approving the model interpretation information corresponding to the first user sample set and the model interpretation information corresponding to the second user sample set. For example, the confirmation operation can be a selection operation acting on the confirmation control. The downstream task end can be a computing device that needs to receive the output scores corresponding to each user to perform downstream tasks according to the output scores. Here, the downstream task can include but is not limited to at least one of the following: information push, user authority adjustment. In practice, the execution subject can send the output scores corresponding to the first user sample set and the second user sample set in the output score set to the downstream task end through a wired connection or a wireless connection. Thus, after the user confirms the model interpretation information corresponding to the first user sample set and the model interpretation information corresponding to the second user sample set, the output scores corresponding to the first user sample set and the second user sample set can be automatically sent to the downstream task end, so that the downstream task end automatically performs the downstream task.

[0101] It should be noted that the above-mentioned wireless connection methods may include but are not limited to 3G / 4G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other wireless connection methods currently known or to be developed in the future.

[0102] from Figure 3 It can be seen that Figure 2 Compared with the description of some corresponding embodiments, Figure 3 The process 300 of the model explanation information generation method in some corresponding embodiments embodies the step of expanding the second user sample set. Therefore, the schemes described in these embodiments can interpret the importance of user features for high-scoring groups and low-scoring groups in each group of user samples.

[0103] Further references Figure 4 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a device for generating model explanation information. These device embodiments are similar to Figure 2 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0104] like Figure 4 As shown, the model explanation information generating device 400 of some embodiments includes: a first standardization processing unit 401, a second standardization processing unit 402, a clustering unit 403 and an execution unit 404. The first standardization processing unit 401 is configured to perform standardization processing on each target user feature of each user sample in the user sample set to obtain a standardized user sample set; the second standardization processing unit 402 is configured to perform standardization processing on the output score set of the black box model corresponding to the above user sample set to obtain a standardized output score set; the clustering unit 403 is configured to cluster the above user sample set based on the above standardized user sample set and the above standardized output score set to obtain a user sample group set; the execution unit 404 is configured to perform the following steps for each user sample group in the above user sample group set: select each user sample that meets the preset first score condition from the above user sample group as the first user sample set; generate a first regression model according to the above first user sample set and each output score corresponding to the above first user sample set in the above output score set; generate model explanation information corresponding to the above first user sample set and the above black box model according to the model parameter information of the above first regression model.

[0105] Optionally, after selecting each user sample that meets the preset first score condition from the user sample group as the first user sample set, the model explanation information generating device 400 may further include: a selection unit, a first generation unit, and a second generation unit (not shown in the figure). The selection unit is configured to select each user sample that meets the preset second score condition from the user sample group as the second user sample set. The first generation unit is configured to generate a second regression model based on the second user sample set and each output score corresponding to the second user sample set in the output score set. The second generation unit is configured to generate model explanation information corresponding to the second user sample set and the black box model based on the model parameter information of the second regression model.

[0106] Optionally, the execution unit 404 can be further configured to: arrange the weights included in the model parameter information of the above-mentioned first regression model in descending order to obtain a weight sequence as a first weight sequence; select a preset number of first weights from the above-mentioned first weight sequence as a target first weight sequence; and generate model interpretation information corresponding to the above-mentioned first user sample set and the above-mentioned black box model according to the feature names of each target user feature corresponding to the above-mentioned target first weight sequence.

[0107] Optionally, the execution unit 404 can be further configured to: for each weight included in the model parameter information of the above-mentioned first regression model, perform the following steps: determine the feature name of the target user feature corresponding to the above-mentioned weight as the first feature name; determine the mean of each target user feature corresponding to the above-mentioned first feature name in the above-mentioned first user sample set as the first feature mean; determine the mode of each target user feature corresponding to the above-mentioned first feature name in the above-mentioned first user sample set as the first feature mode; combine the above-mentioned first feature name, the above-mentioned weight, the above-mentioned first feature mean and the above-mentioned first feature mode into the first feature explanation information; determine the combined each first feature explanation information as the model explanation information corresponding to the above-mentioned first user sample set and the above-mentioned black box model.

[0108] Optionally, the second generation unit can be further configured to: arrange the weights included in the model parameter information of the above-mentioned second regression model in descending order to obtain a weight sequence as a second weight sequence; select a preset number of second weights from the above-mentioned second weight sequence as a target second weight sequence; and generate model interpretation information corresponding to the above-mentioned second user sample set and the above-mentioned black box model according to the feature names of each target user feature corresponding to the above-mentioned target second weight sequence.

[0109] Optionally, the second generation unit can be further configured to: for each weight included in the model parameter information of the above-mentioned second regression model, perform the following steps: determine the feature name of the target user feature corresponding to the above-mentioned weight as the second feature name; determine the mean of each target user feature corresponding to the above-mentioned second feature name in the above-mentioned second user sample set as the second feature mean; determine the mode of each target user feature corresponding to the above-mentioned second feature name in the above-mentioned second user sample set as the second feature mode; combine the above-mentioned second feature name, the above-mentioned weight, the above-mentioned second feature mean and the above-mentioned second feature mode into second feature explanation information; determine the combined each second feature explanation information as the model explanation information corresponding to the above-mentioned second user sample set and the above-mentioned black box model.

[0110] Optionally, after generating the model interpretation information corresponding to the second user sample set and the black box model according to the model parameter information of the second regression model, the model interpretation information generating device 400 may further include: a display unit and a sending unit (not shown in the figure). The display unit is configured to display the generated model interpretation information corresponding to the first user sample set and the model interpretation information corresponding to the second user sample set in the form of a table. The sending unit is configured to send each output score corresponding to the first user sample set and the second user sample set in the output score set to the downstream task end in response to detecting the confirmation operation of the model interpretation information corresponding to the first user sample set and the model interpretation information corresponding to the second user sample set.

[0111] It is understood that the units described in the device 400 are similar to those described in the reference Figure 2 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the device 400 and the units included therein, and will not be described in detail here.

[0112] Reference below Figure 5 , which shows an electronic device 500 (eg, Figure 1 Schematic diagram of the structure of the computing device in FIG. Figure 5 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0113] like Figure 5 As shown, the electronic device 500 may include a processing device 501 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0114] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 5The electronic device 500 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 5 Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0115] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.

[0116] It should be noted that the computer-readable medium recorded in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0117] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0118] The computer-readable medium may be included in the electronic device; or it may exist independently without being assembled into the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: standardizes each target user feature of each user sample in the user sample set to obtain a standardized user sample set; standardizes the output score set of the black box model corresponding to the user sample set to obtain a standardized output score set; clusters the user sample set based on the standardized user sample set and the standardized output score set to obtain a user sample group set; for each user sample group in the user sample group set, performs the following steps: selects each user sample that meets the preset first score condition from the user sample group as the first user sample set; generates a first regression model according to the first user sample set and each output score corresponding to the first user sample set in the output score set; generates model explanation information corresponding to the first user sample set and the black box model according to the model parameter information of the first regression model.

[0119] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0120] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0121] The units described in some embodiments of the present disclosure may be implemented by software or by hardware. The described units may also be provided in a processor, for example, may be described as: a processor including a first standardization processing unit, a second standardization processing unit, a clustering unit, and an execution unit. The names of these units do not constitute limitations on the units themselves in certain circumstances, for example, the first standardization processing unit may also be described as "a unit that performs standardization processing on each target user feature of each user sample in a user sample set to obtain a standardized user sample set".

[0122] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0123] Some embodiments of the present disclosure also provide a computer program product, including a computer program, which implements any of the above-mentioned model interpretation information generation methods when executed by a processor.

[0124] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) and the technical solutions formed.

Claims

1. A method for generating model explanation information, include: Standardizing each target user feature of each user sample in the user sample set to obtain a standardized user sample set; Normalizing the output score set of the black box model corresponding to the user sample set to obtain a standardized output score set; Based on the standardized user sample set and the standardized output score set, clustering the user sample set to obtain a user sample group set; For each user sample group in the user sample group set, perform the following steps: Selecting each user sample that meets a preset first score condition from the user sample group as a first user sample set; generating a first regression model according to the first user sample set and each output score in the output score set corresponding to the first user sample set; Model explanation information corresponding to the first user sample set and the black box model is generated according to the model parameter information of the first regression model.

2. The method according to claim 1, in, After selecting each user sample satisfying a preset first score condition from the user sample group as a first user sample set, the method further includes: Selecting each user sample that meets a preset second score condition from the user sample group as a second user sample set; generating a second regression model according to the second user sample set and each output score in the output score set corresponding to the second user sample set; Model explanation information corresponding to the second user sample set and the black box model is generated according to the model parameter information of the second regression model.

3. The method according to claim 1, in, The generating, according to the model parameter information of the first regression model, model explanation information corresponding to the first user sample set and the black box model comprises: Arrange the weights included in the model parameter information of the first regression model in descending order to obtain a weight sequence as a first weight sequence; Selecting a preset number of first weights from the first weight sequence as a target first weight sequence; Model explanation information corresponding to the first user sample set and the black box model is generated according to the feature name of each target user feature corresponding to the target first weight sequence.

4. The method according to claim 1, in, The generating, according to the model parameter information of the first regression model, model explanation information corresponding to the first user sample set and the black box model comprises: For each weight included in the model parameter information of the first regression model, the following steps are performed: Determine the feature name of the target user feature corresponding to the weight as the first feature name; Determine the mean of each target user feature corresponding to the first feature name in the first user sample set as a first feature mean; Determine the mode of each target user feature corresponding to the first feature name in the first user sample set as a first feature mode; Combining the first feature name, the weight, the first feature mean, and the first feature mode into first feature explanation information; The combined first feature explanation information is determined as model explanation information corresponding to the first user sample set and the black box model.

5. The method according to claim 2, in, The generating, according to the model parameter information of the second regression model, model explanation information corresponding to the second user sample set and the black box model comprises: Arrange the weights included in the model parameter information of the second regression model in descending order to obtain a weight sequence as a second weight sequence; Selecting a preset number of second weights from the second weight sequence as a target second weight sequence; Model explanation information corresponding to the second user sample set and the black box model is generated according to the feature name of each target user feature corresponding to the target second weight sequence.

6. The method according to claim 2, in, The generating, according to the model parameter information of the second regression model, model explanation information corresponding to the second user sample set and the black box model comprises: For each weight included in the model parameter information of the second regression model, the following steps are performed: Determine the feature name of the target user feature corresponding to the weight as the second feature name; Determine the mean of each target user feature corresponding to the second feature name in the second user sample set as a second feature mean; Determine the mode of each target user feature corresponding to the second feature name in the second user sample set as a second feature mode; combining the second feature name, the weight, the second feature mean and the second feature mode into second feature explanation information; The combined second feature explanation information is determined as model explanation information corresponding to the second user sample set and the black box model.

7. The method according to claim 2, in, After generating model explanation information corresponding to the second user sample set and the black box model according to the model parameter information of the second regression model, the method further includes: Displaying the generated model explanation information corresponding to the first user sample set and the generated model explanation information corresponding to the second user sample set in a table format; In response to detecting a confirmation operation of the model explanation information corresponding to the first user sample set and the model explanation information corresponding to the second user sample set, each output score corresponding to the first user sample set and the second user sample set in the output score set is sent to a downstream task end.

8. A device for generating model explanation information, include: A first standardization processing unit is configured to perform standardization processing on each target user feature of each user sample in the user sample set to obtain a standardized user sample set; A second standardization processing unit is configured to perform standardization processing on the output score set of the black box model corresponding to the user sample set to obtain a standardized output score set; A clustering unit, configured to cluster the user sample set based on the standardized user sample set and the standardized output score set to obtain a user sample group set; The execution unit is configured to perform the following steps for each user sample group in the user sample group set: select each user sample that meets a preset first score condition from the user sample group as a first user sample set; generate a first regression model according to the first user sample set and each output score corresponding to the first user sample set in the output score set; and generate model explanation information corresponding to the first user sample set and the black box model according to model parameter information of the first regression model.

9. An electronic device, include: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

10. A computer readable medium having a computer program stored thereon, in, When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

11. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.