A method and device for explaining a trusted machine learning model based on a formal method

CN118966376BActive Publication Date: 2026-08-28NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411003495.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-08-28
Estimated Expiration
2044-07-24

AI Technical Summary

Technical Problem

但是,现有技术对模型进行可解释性分析时,解释结果的可靠性无法保证

Benefits of technology

[0049]本发明提供了一种基于形式化方法的可信机器学习模型解释方法及装置,其中方法包括:获取目标机器学习模型以及目标机器学习模型的输入特征值;将输入特征值连续变量进行的特征值离散化,得到离散化后的输入特征值;将离散化后的输入特征值采用多值逻辑进行形式化编译,得到形式化编译后的输入特征值变量;根据目标机器学习模型的参数,采用形式化方法进行编译,得到基于阈值的形式化逻辑表达式,并简化至质蕴含项或主蕴含项,得到目标机器学习模型的决策的规则根因;目标机器学习模型的决策规则是根据(树形模型)模型本身的决策树或者(神经网络模型)基于阈值条件生成的博弈决策树,从根节点到叶子节点的一条路径作为一条规则,将决策树从根节点到叶子节点的所有路径归类做析取得到的;根据规则根因,采用自然语言表示,得到目标机器学习模型的解释结果。本发明通过对输入特征值进行形式化编译和离散化处理,结合多值逻辑变量进行形式化解释,有效解决了机器学习模型解释性不足的问题,并且提高了模型解释的可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118966376B_ABST
    Figure CN118966376B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on formalization method's trusted machine learning model explanation method and device, it is related to artificial intelligence technical field, the method includes obtaining target machine learning model and its input characteristic value;Input characteristic value is compiled using formalization method, to obtain continuous variable;Continuous variable is discretized, to obtain the input characteristic value after discretization;The logic variable compilation of characteristic value after discretization is carried out, to obtain the variable after formalization compilation;According to the decision rule of model, the logic expression after formalization compilation is simplified, to obtain the rule reason of the expression of implication, that is, decision;Rule reason is obtained using natural language representation, to obtain the explanation result of model.The application uses formalization method to explain trusted machine learning model, compared with traditional explanation method, can more accurately reveal the internal logic of model decision, to improve the reliability of model explanation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a reliable machine learning model interpretation method and apparatus based on formal methods. Background Technology

[0002] With the rapid development of artificial intelligence, machine learning models such as deep learning are ubiquitous in real life. Decisions made by machine learning models affect all aspects of human life and work. However, machine learning models are black boxes, with the common problems of being unexplainable, unquestionable, and inexplicable. Existing interpretable machine learning technologies are often based on approximation, fitting, and approximation strategies to provide rough and vague explanations of the models. However, in critical fields involving security, such as cybersecurity, which relies on the analysis of binary strings, even a deviation in the interpretation of a single character can cause significant semantic misunderstanding. Therefore, it is urgent to provide logically rigorous, absolutely accurate, and semantically clear and concise explanations for these models.

[0003] Currently, to achieve reliable machine learning, approximate or simulated interpretable machine learning methods are insufficient. It is necessary to formally open the black box of the machine learning model, obtain its formal logical expression, and extract its simplest form—the root cause of the model's decisions. Only after expert review and testing can reliable machine learning be achieved. However, existing technologies cannot guarantee the reliability of interpretation results when performing interpretability analysis on models. Summary of the Invention

[0004] The purpose of this invention is to provide a reliable machine learning model interpretation method and apparatus based on formal methods, which can achieve highly reliable machine learning model interpretation.

[0005] To achieve the above objectives, the present invention provides the following solution:

[0006] In a first aspect, the present invention provides a reliable machine learning model interpretation method based on formal methods, comprising:

[0007] Obtain the target machine learning model and the input feature values ​​of the target machine learning model;

[0008] Based on the input feature values ​​of the target machine learning model, a formal method is used to compile the model to obtain continuous variables of input feature values;

[0009] The continuous input feature value variable is discretized to obtain the discretized input feature value; the discretized input feature value is distributed in multiple discrete feature value partitions;

[0010] The discretized input feature values ​​are formally compiled using Boolean logic or multi-valued logic variables to obtain formally compiled input feature value variables;

[0011] According to the decision rules of the target machine learning model, the formally compiled input feature value variables are represented as a formal logical expression of the target machine learning model, and the formal logical expression is simplified to obtain the rule root causes of the decision of the target machine learning model; the rule root causes are the quality implication term expressions of the target machine learning model.

[0012] Based on the root causes of the rules, the interpretation results of the target machine learning model are obtained using natural language representation.

[0013] Optionally, the target machine learning model includes neural network-based models and tree-based models constructed from decision trees.

[0014] Optionally, when the target machine learning model is a tree-based model constructed from decision trees, the formally compiled input feature value variables are simplified according to the decision rules of the target machine learning model to obtain the root causes of the decision rules of the target machine learning model, specifically including:

[0015] Based on the discrete partitioning of the feature values, a path from the root node to the leaf node in the target machine learning model is taken as a rule to determine the set of disjunctions between all paths, thus obtaining the decision rule of the tree model.

[0016] Based on the decision rules of the tree model, the tree model is simplified to obtain a knowledge compilation logic expression; the knowledge compilation logic expression is a logic expression represented by CNF, DNF, NNF or a subset thereof;

[0017] The knowledge compilation logic expression is simplified to a quality implication term or a principal implication term expression to obtain the rule root cause of the decision of the tree model.

[0018] Optionally, when the target machine learning model is a neural network-based model, the formally compiled input feature value variables are simplified according to the decision rules of the target machine learning model to obtain the root causes of the decision rules of the target machine learning model, specifically including:

[0019] Obtain the weights, biases, and activation functions of the neural network model;

[0020] Based on the weights, biases, activation functions, and continuous input feature values ​​of the neural network model, determine the threshold-based linear function of the neural network model;

[0021] The threshold conditions of the threshold-based linear function are regularized, and the decision tree for each threshold condition in the regularized linear function is calculated.

[0022] Based on the decision tree, the results are divided into decision intervals;

[0023] The decision rules of the neural network model are obtained based on the boundaries of the decision interval of each decision tree.

[0024] Based on the decision rules of the neural network model, the neural network model is simplified to obtain a knowledge compilation logic expression; the knowledge compilation logic expression is a logic expression represented by CNF, DNF, NNF or a subset thereof;

[0025] The knowledge compilation logic expression is simplified to a quality implication term or a principal implication term expression to obtain the rule root causes of the neural network model's decision.

[0026] Optionally, the specific formula for the threshold-based linear function is as follows:

[0027]

[0028] Where T1 and T2 are constants representing the threshold; f(x) is a linear function based on the threshold, f(x)∈{y1,y2,…,y} n}, n∈N, the domain of f(x) is X, and X=∑ i ω i x i ∈X; ω i It is a constant, which is x i The weights;

[0029] Optionally, when the target machine learning model is a model built on a neural network, the formally compiled input feature value variables are simplified according to the decision rules of the target machine learning model to obtain the root causes of the decision rules of the target machine learning model, specifically including:

[0030] Obtain the weights, biases, and activation functions of the neural network model;

[0031] Based on the weights, biases, activation functions, and continuous input feature values ​​of the neural network model, determine the threshold-based linear function of the neural network model;

[0032] The threshold conditions of the threshold-based linear function are regularized, and the decision boundaries between each true and false domains under each threshold condition of the threshold-based linear function are calculated.

[0033] The decision rules of the neural network model are obtained based on the decision boundaries between each tautology domain and each false tautology domain under each threshold condition.

[0034] Based on the decision rules of the neural network model, the neural network model is simplified to obtain a knowledge compilation logic expression; the knowledge compilation logic expression is a logic expression represented by CNF, DNF, NNF or a subset thereof;

[0035] The knowledge compilation logic expression is simplified to a quality implication term or a principal implication term expression to obtain the rule root causes of the neural network model's decision.

[0036] Optionally, the multi-valued logic variables include Boolean variables and multi-valued discrete logic variables.

[0037] Optionally, the continuous input feature values ​​are discretized to obtain discretized input feature values, specifically including:

[0038] The continuous input feature values ​​are discretized using the K-means discretization method, the chi-square Chi1 discretization method, the Chi2 discretization method, the minimum description length principal discretization method, or the optimal binning for scoring modeling discretization method to obtain the discretized input feature values.

[0039] Secondly, the present invention provides a reliable machine learning model interpretation device based on formal methods, comprising:

[0040] The acquisition module is used to acquire the target machine learning model and the input feature values ​​of the target machine learning model;

[0041] The first compilation module is used to compile the input feature values ​​of the target machine learning model using a formal method to obtain continuous variables of input feature values;

[0042] The discrete module is used to discretize the feature values ​​of the continuous input feature value variable to obtain the discretized input feature value; the discretized input feature value is distributed in multiple feature value discrete partitions;

[0043] The second compilation module is used to formally compile the discretized input feature values ​​using Boolean logic or multi-valued logic variables to obtain formally compiled input feature value variables.

[0044] A simplification module is used to represent the formally compiled input feature value variables as a formal logical expression of the target machine learning model according to the decision rules of the target machine learning model, and to simplify the formal logical expression to obtain the rule root causes of the decision of the target machine learning model; the rule root causes are the quality implication term expressions of the target machine learning model.

[0045] The output module is used to obtain the interpretation results of the target machine learning model based on the root causes of the rules, using natural language representation.

[0046] Optionally, the discrete module specifically includes:

[0047] The discrete submodule is used to discretize the feature values ​​of the continuous input feature value variable using the K-means discretization method, the chi-square Chi1 discretization method, the Chi2 discretization method, the minimum description length principal discretization method, or the optimal binning for scoring modeling discretization method, to obtain the discretized input feature values.

[0048] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0049] This invention provides a reliable machine learning model interpretation method and apparatus based on formal methods. The method includes: acquiring a target machine learning model and its input feature values; discretizing the continuous input feature values ​​to obtain discretized input feature values; formally compiling the discretized input feature values ​​using multi-valued logic to obtain formally compiled input feature value variables; compiling the target machine learning model's parameters using formal methods to obtain a threshold-based formal logic expression, and simplifying it to prime or principal implications to obtain the root causes of the target machine learning model's decisions; the decision rules of the target machine learning model are derived from the decision tree of the model itself (tree model) or the game decision tree generated based on threshold conditions (neural network model), with each path from the root node to a leaf node as a rule, and all paths from the root node to the leaf node of the decision tree are classified and extracted; and representing the root causes of the rules using natural language to obtain the interpretation result of the target machine learning model. This invention effectively solves the problem of insufficient interpretability of machine learning models and improves the reliability of model interpretation by formally compiling and discretizing the input feature values ​​and combining them with formal interpretation of multi-valued logical variables. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 This is a schematic diagram of a reliable machine learning model interpretation method based on formal methods provided in Embodiment 1 of the present invention;

[0052] Figure 2 This is a main process framework diagram provided for Embodiment 1 of the present invention;

[0053] Figure 3 This is a flowchart of the continuous variable discretization and Booleanization method provided in Embodiment 1 of the present invention;

[0054] Figure 4 This is a flowchart illustrating the knowledge compilation process, i.e., the formal interpretation process, of the tree-based machine learning model provided in Embodiment 1 of the present invention.

[0055] Figure 5 The flowchart provided in Embodiment 1 of the present invention describes the knowledge compilation and extraction of formal interpretation of machine learning models such as neural networks;

[0056] Figure 6 This is a schematic diagram of a reliable machine learning model interpretation device based on a formal method, provided in Embodiment 2 of the present invention. Detailed Implementation

[0057] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0058] The purpose of this invention is to provide a reliable machine learning model interpretation method and apparatus based on formal methods, which can achieve highly reliable machine learning model interpretation.

[0059] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0060] Example 1

[0061] like Figure 1 As shown, this embodiment provides a reliable machine learning model interpretation method based on formal methods, including:

[0062] Step 101: Obtain the target machine learning model and the input feature values ​​of the target machine learning model;

[0063] Step 102: Based on the input feature values ​​of the target machine learning model, compile them using a formal method to obtain continuous variables of input feature values.

[0064] Step 103: Discretize the continuous input feature value to obtain the discretized input feature value; the discretized input feature value is distributed in multiple discrete feature value partitions.

[0065] Step 104: The discretized input feature values ​​are formally compiled using Boolean logic or multi-valued logic variables to obtain formally compiled input feature value variables.

[0066] Step 105: Based on the decision rules of the target machine learning model, the formally compiled input feature value variables are represented as a formal logical expression of the target machine learning model, and the formal logical expression is simplified to obtain the rule root causes of the decision of the target machine learning model; the rule root causes are the quality implication term expressions of the target machine learning model.

[0067] Step 106: Based on the root causes of the rules, use natural language representation to obtain the explanation results of the target machine learning model.

[0068] Specifically, when performing step 101, the target machine learning model includes a neural network-based model and a tree-based model constructed from decision trees.

[0069] Specifically, to obtain a machine learning model that makes decisions in real life, with functions including but not limited to a classifier, the model includes neural networks and neural network-based models such as deep learning, RNN, transformer, etc., and extracts the parameters of the model (such as the weights, bias, and activation function of the neural network, and the node thresholds of the decision tree model, etc.).

[0070] Specifically, when performing steps 102-104, the following can be done:

[0071] like Figure 3 As shown, compiling the input feature values ​​of the obtained model into logical variables using formal methods includes the following steps:

[0072] 1) Discretize the eigenvalues ​​of the continuous variable. Discretization methods include, but are not limited to, K-means, Chi1, Chi2, minimum description length principal (MDLP), and optimal binning for scoring modeling (Smbinning). After discretizing the eigenvalues, the partitioning points of the continuous variable eigenvalues ​​(F) are obtained, dividing the variable into discrete variables (DF).

[0073] 2) The discretized input feature values ​​are represented as multi-valued logic variables. Specifically, based on the partition points of the continuous variable feature values ​​(F), they are mapped to a set of multiple Boolean variables (BF). For example:

[0074] In some embodiments of this example, when performing step 105, the specific steps may be as follows:

[0075] like Figure 4 As shown, when the target machine learning model is a tree-based model constructed from decision trees, the formally compiled input feature value variables are simplified according to the decision rules of the target machine learning model to obtain the root causes of the decision rules of the target machine learning model, specifically including:

[0076] Step 501: Based on the partitioning of the leaf nodes of the tree model, take a path from the root node to a leaf node in the target machine learning model as a rule, determine the set of disjunctions between all paths, and obtain the decision rule of the tree model.

[0077] Step 502: According to the decision rules of the tree model, the disjunctive set is simplified to obtain the knowledge compilation logic expression; the knowledge compilation logic expression is a logic expression represented by CNF, DNF, NNF or a subset thereof.

[0078] Step 503: Simplify the knowledge compilation logic expression to the quality implication term or principal implication term expression to obtain the rule root cause of the decision of the tree model.

[0079] The knowledge compilation languages ​​used include, but are not limited to, Conjunctive Normal Form (CNF), Disjunctive Normal Form (DNF), Nested Normal Form (NNF), and their subsets (such as Binary Decision Graph OBDD, Simplified Binary Decision Graph ROBDD, Multivariate Decision Graph MDD, DNNF, d-DNNF), and any knowledge compilation language based on directed acyclic graphs.

[0080] Specifically, when executing step 501, the steps can be as follows:

[0081] Based on the obtained discrete eigenvalue expression, the decision tree model is used as a rule (p) based on a path from the root node to the leaf node. i This classifies all paths in a decision tree from the root node to the leaf node. For example, for a classifier with decision values ​​of True and False, all paths in the decision tree can be divided into two categories: paths with a decision value of True. "and "the set of paths that end in False" ".make It can be known For the rules where all decisions in this model are True, This refers to the rule that all decisions in this model are false.

[0082] Specifically, step 502 can be performed as follows:

[0083] Decision trees and decision tree-based models, such as random forests and xboostingtrees, can be represented as CNF or DNF expressions, set union normal form NNF, or any subset thereof (such as ROBDD, multivariate decision graph MDD, DNNF, d-DNNF, etc.).

[0084] In some embodiments of this example, when performing step 105, the specific steps may be as follows:

[0085] like Figure 5 As shown, when the target machine learning model is a neural network-based model, the formally compiled logical expression is simplified according to the decision rules of the target machine learning model to obtain the root causes of the decision rules of the target machine learning model, specifically including:

[0086] Step 511: Obtain the weights, biases, and activation functions of the neural network model.

[0087] Step 512: Determine the threshold-based linear function of the neural network model based on the weights, biases, activation functions, and continuous input feature values ​​of the neural network model.

[0088] Step 513: Regularize the threshold conditions of the threshold-based linear function and calculate the decision tree for each threshold condition in the regularized linear function.

[0089] Step 514: Based on the decision tree, divide the results into decision intervals.

[0090] Step 515: Obtain the decision rules of the neural network model based on the boundaries of the decision interval of each decision tree.

[0091] Step 516: Based on the decision rules of the neural network model, simplify the neural network model to obtain a knowledge compilation logic expression; the knowledge compilation logic expression is a logic expression represented by CNF, DNF, NNF or a subset thereof.

[0092] Step 517: Simplify the knowledge compilation logic expression to the quality implication term or principal implication term expression to obtain the rule root cause of the neural network model's decision.

[0093] Specifically, when performing steps 512-517, the details can be as follows:

[0094] Multiply the weight matrices of each neuron in the neural network model to obtain the final weight values; multiply the bias matrices of each neuron in the neural network model to obtain the final bias values; then, based on the input-output mapping relationship of the activation function of each layer, obtain the threshold-based linear function (L) corresponding to the neural network model.

[0095] Specifically, the formula for the threshold-based linear function is as follows:

[0096]

[0097] Where T1 and T2 are constants representing thresholds; f(x) is a linear function based on the threshold, f(x)∈{y1,y2,…,y} n}, n∈N, the domain of f(x) is X, and X=∑ i ω i x i ∈X; ω i It is a constant, which is x i weight These are called threshold conditions.

[0098] Then, taking the threshold condition (Con) of the threshold-based linear function (L) obtained in the previous step, and traversing all mappings from input to output, a decision tree (DT) is constructed. i Each path in this decision tree from the root node to a leaf node represents an interpretation of a decision, and because the path construction is logically expressed, it is a formal interpretation. Traversing the decision tree with threshold conditions yields all formal interpretations of the model; in other words, the problem of formal machine learning interpretation is transformed into a decision tree traversal problem.

[0099] Take all the paths (np) from the root node to the leaf node obtained in the previous step, and classify them according to the values ​​of their leaf nodes (i.e., the final decision values). For example, for a classifier with decision values ​​of True and False, all paths in the decision tree can be divided into two categories: paths with a result of True. "" and "the set of paths that end in False" ".make It can be known This represents the input interval where all decisions of the model are True. This represents the input interval where all decisions of the model are False.

[0100] Based on the path classification obtained in the previous step, merge similar terms within each category and simplify them to prime or principal implication terms. For example, in the above case, ... Simplify to the extreme Will Simplify to the extreme This refers to the explanation of the quality or principal implications of the model, i.e., the root causes of decisions.

[0101] The root causes of the decisions obtained from the previous simplification step are represented in natural language and returned to the user.

[0102] In some embodiments of this example, when performing step 105, the specific steps may be as follows:

[0103] When the target machine learning model is a tree-based model constructed from decision trees, the formally compiled input feature value variables are simplified according to the decision rules of the target machine learning model to obtain the root causes of the decision rules of the target machine learning model, specifically including:

[0104] Step 521: Obtain the weights, biases, and activation functions of the neural network model.

[0105] Step 522: Determine the threshold-based linear function of the neural network model based on the weights, biases, activation functions, and continuous input feature values ​​of the neural network model.

[0106] Step 523: Regularize the threshold conditions of the threshold-based linear function and calculate the decision boundary between each true and false domain under each threshold condition of the threshold-based linear function.

[0107] Step 524: Based on the decision boundaries between each true and false domain under each threshold condition, obtain the decision rules of the neural network model.

[0108] Step 525: Based on the decision rules of the neural network model, simplify the neural network model to obtain a knowledge compilation logic expression; the knowledge compilation logic expression is a logic expression represented by CNF, DNF, NNF or a subset thereof.

[0109] Step 526: Simplify the knowledge compilation logic expression to the quality implication term or principal implication term expression to obtain the rule root cause of the neural network model's decision.

[0110] Specifically, this method is a fast way to interpret neural network models, and step 523 can be performed as follows:

[0111] Definition 2. (Truthless Domain / False Domain): For a function Where ω i It is the weight, a positive constant.

[0112] set up Denotes a domain interval of Con(x), where They represent x respectively i The upper and lower limits of the value, Let Con(x) be the input vector space, if for any If Con(x) always holds true, then it is called... for The "true domain"; similarly, if for any If Con(x) is always false, then it is called... for The "permanent domain".

[0113] During the traversal, by observing the precision of the values ​​taken on both sides of the inequality sign, the calculation of the decision interval can be further simplified.

[0114] The specific method for dividing the input space into decision intervals is as follows:

[0115] Step a) Regularize all threshold conditions (Con) of the threshold-based linear function, i.e., place those threshold conditions with the same sign on the same side of the inequality until there are no negative terms on either side of the inequality. For simplicity, the two sides of the inequality are referred to as the red side and the blue side, respectively. The eigenvalues ​​contained on both sides are sorted from largest to smallest as follows: and The weights corresponding to each feature value are as follows: and The precision of the values ​​on both sides is the minimum absolute value of the weights of the eigenvalues ​​on both sides, and the maximum values ​​of each eigenvalue are respectively... and The value ranges for both sides are [L] p H p ] and [Ln H n ].

[0116] Step b) can be explained starting from either side of the threshold condition inequality. For simplicity, we'll start with the red side. Beginning with the minimum range of red's values, and using the blue side's precision as the minimum interval, we iterate through all possible red values. For each value (Target), we divide it sequentially by the blue side's weight value. Round the quotient up to obtain the sequence This allows us to obtain a formal interpretation with only one value.

[0117] Step c) For each value obtained in the previous step like calculate As the new target value.

[0118] Step d) Take the new target value obtained in the previous step and remove each weight value of the blue side starting from the second weight value in turn. Round the quotient up to the nearest integer. That is, a formal interpretation using two variables

[0119] By analogy, we can obtain a formal interpretation of the model using three, four, or any finite number of positive integer variables (limited to the total number of red eigenvalues ​​(M)).

[0120] Application examples in real-world network DoS intrusion detection scenarios:

[0121] The trained model is a three-layer neural network, with 10, 100, and 1 neurons in each layer, respectively. The activation function for the intermediate layers is Identity, and the activation function for the output layer is sigmoid. The input-output mapping is logically the same as the Identity function, so no distinction is made here. According to step 4.1, the threshold condition for the model in this paper is:

[0122] -1.07A-0.19B-0.09C+0.4D+0.299E-0.07F+0.31G+0.42H-1.21I-0.05≥-4.74.

[0123] After normalization in step a), the threshold condition of the model is transformed into:

[0124] 0.42H+0.4D+0.31G+0.299E+4.74≥1.21I+1.07A+0.19B+0.09C+0.07F+0.05J.

[0125] According to step b), the red side inputs are 0.42H, 0.4D, 0.31G, 0.299E and threshold T = 4.74, respectively, and the blue side inputs are 1.21I, 1.07A, 0.19B, 0.097C, 0.07F and 0.05J.

[0126] The value ranges for both sides defined in step a) are [L p H p ] = [6.17, 61805], [L n H n = [2.687, 33.422].

[0127] Based on steps c) and d), the formal interpretation using one variable is calculated as follows:

[0128] Rule 1: If an input instance has negative variables (I, A, B, C, F, J) where B≤19, A≤4, or I≤3, and other negative variables are equal to 1, then regardless of the value of the positive variables, the input instance will be identified as a DoS intrusion by the model.

[0129] Formal interpretation using two variables:

[0130] Rule 2: If an input instance has negative variables where F < 40 and J takes any value, or F < 41 and J < 15, or F < 42 and J < 14, or F < 43 and J < 12, or F < 44 and J < 11, or F < 45 and J < 10, or F is any value and J < 8, and other negative variables are equal to 1, then regardless of the value of the positive variables, the input instance will definitely be identified as a DoS intrusion by the model.

[0131] In some embodiments of this example, the multi-valued logic variables include Boolean variables and multi-valued discrete logic variables.

[0132] In some embodiments, the local interpretation method of the neural network model can be as follows:

[0133] Local explanations, which explain why the model makes decisions for a specific instance, are the same as global explanations, which refer to the model's decision rules regardless of a specific instance.

[0134] 1) Given a specific instance, the feature values ​​of that instance are... in The eigenvalues ​​with positive weights are referred to as the red side for simplicity. The eigenvalues ​​with negative weights are called blue squares.

[0135] After being discretized and compiled, it is transformed into a formal logical expression (which may be a Boolean logical expression, a multi-valued logical expression, or other logical expressions, depending on the logic used).

[0136] Perform steps 521-522 to obtain the threshold-based linear function (L) of the model.

[0137] Choose any side to begin the explanation. First, substitute the feature values ​​of that instance into the side corresponding to the threshold condition of the function. For simplicity, we will start with the red side. Starting from the minimum value range of the red side, and using the precision of the blue side as the minimum interval unit, we will traverse all possible values ​​of the red side. Following steps a)-d), we will calculate the decision boundary of the instance, simplify and translate it into natural language, and return it to the user.

[0138] Example 2.

[0139] To execute the method corresponding to Embodiment 1 above and achieve the corresponding functions and technical effects, a reliable machine learning model interpretation device based on formal methods is provided below, such as... Figure 6 As shown, it includes the following modules:

[0140] The acquisition module 601 is used to acquire the target machine learning model and the input feature values ​​of the target machine learning model.

[0141] The first compilation module 602 is used to compile the input feature values ​​of the target machine learning model using a formal method to obtain continuous variables of input feature values.

[0142] Discretization module 603 is used to discretize the feature values ​​of the continuous input feature value variable to obtain the discretized input feature value; the discretized input feature value is distributed in multiple feature value discrete partitions.

[0143] The second compilation module 604 is used to formally compile the discretized input feature values ​​using Boolean logic or multi-valued logic variables to obtain formally compiled input feature value variables.

[0144] The simplification module 605 is used to represent the formally compiled input feature value variables as a formal logical expression of the target machine learning model according to the decision rules of the target machine learning model, and to simplify the formal logical expression to obtain the rule root causes of the decision of the target machine learning model; the rule root causes are the quality implication term expressions of the target machine learning model.

[0145] The output module 606 is used to obtain the interpretation result of the target machine learning model by using natural language representation based on the root causes of the rules.

[0146] Specifically, the discrete module includes:

[0147] The discrete submodule is used to discretize the feature values ​​of the continuous input feature value variable using the K-means discretization method, the chi-square Chi1 discretization method, the Chi2 discretization method, the minimum description length principal discretization method, or the optimal binning for scoring modeling discretization method, to obtain the discretized input feature values.

[0148] In summary, the present invention has the following beneficial effects:

[0149] 1) The explanation of the machine learning model provided by the method of this invention is rigorous, logically sufficient and complete, and has spatiotemporal stability. In contrast, most existing interpretable machine learning models are based on empiricism and use approximation methods to make speculative explanations of the model. They do not have spatiotemporal stability, that is, for the same input, different runs may give different explanations, and they do not have logical completeness and sufficiency.

[0150] The reason is that this invention uses formal methods to model, simplify, and extract rules from the model. Specifically, this includes: the compilation language of first-order logic is a formal method; multi-valued logic reasoning is also a formal method; and the use of compilation languages ​​based on CNF, etc., to describe the knowledge of the model is a formal method based on first-order logic reasoning.

[0151] 2) Compared to other formally interpretable machine learning methods, this invention has the advantages of low time complexity and low space consumption. This is because existing formally interpreted machine learning methods, due to their use of first-order logic-based reasoning, encounter the combinatorial explosion problem, resulting in NP-hard time complexity. This invention uses the decision tree interval boundary value calculation method based on game trees in step four, which can directly calculate the boundary value without traversing all possibilities, saving a significant amount of computation. Therefore, rule extraction can be completed in an average time of less than 0.001 seconds (this time was obtained through testing in a real-world scenario; the test environment was: running on...). A 64-bit Ubuntu 20.04 system running on a Xeon(R) CPU E5-2678 v3 processor (2.50GHz x 24, 31.2Gbit cache, 500GB hard drive).

[0152] The method proposed in this invention can interpret not only decision trees and tree-based models, but also neural networks, neural network-based models, and other machine learning models. This is because the underlying idea—"treating a machine learning model as a chip, using formal knowledge to compile and describe the model's internal structure as a logic circuit, and then using a method based on formal verification of the logic circuit to extract the model's internal logic, thereby obtaining a formal interpretation"—is applicable to all types of models.

[0153] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0154] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A formal method-based trusted machine learning model explanation method, characterized in that, include: Obtain the target machine learning model and the input feature values ​​of the target machine learning model; Based on the input feature values ​​of the target machine learning model, a formal method is used to compile the model to obtain continuous variables of input feature values; Discretize the continuous input feature values ​​to obtain the discretized input feature values; The discretized input feature values ​​are distributed across multiple discrete feature value partitions; The discretized input feature values ​​are formally compiled using Boolean logic or multi-valued logic variables to obtain formally compiled input feature value variables; According to the decision rules of the target machine learning model, the formally compiled input feature value variables are represented as a formal logical expression of the target machine learning model, and the formal logical expression is simplified to obtain the rule root causes of the decision of the target machine learning model; the rule root causes are the quality implication term expressions of the target machine learning model. The target machine learning model includes neural network-based models and tree-based models constructed from decision trees; When the target machine learning model is a neural network-based model, the formally compiled input feature value variables are simplified according to the decision rules of the target machine learning model to obtain the root causes of the decision rules of the target machine learning model, specifically including: Obtain the weights, biases, and activation functions of the neural network model; Based on the weights, biases, activation functions, and continuous input feature values ​​of the neural network model, determine the threshold-based linear function of the neural network model; The threshold conditions of the threshold-based linear function are regularized, and the decision tree for each threshold condition in the regularized linear function is calculated. Based on the decision tree, the results are divided into decision intervals; The decision rules of the neural network model are obtained based on the boundaries of the decision interval of each decision tree. Based on the decision rules of the neural network model, the neural network model is simplified to obtain a knowledge compilation logic expression; the knowledge compilation logic expression is a logic expression represented by CNF, DNF, NNF or a subset thereof; The knowledge compilation logic expression is simplified to the expression of quality implication term or principal implication term to obtain the root cause of the decision of the neural network model; The specific formula for the threshold-based linear function is as follows: ; in, T 1 ,T 2 The constant represents the threshold. It is a threshold-based linear function. , The domain is And there are ; It is a constant. The weights; , ; Based on the root causes of the rules, the interpretation results of the target machine learning model are obtained using natural language representation.

2. The method for interpreting a reliable machine learning model based on a formal method according to claim 1, wherein when the target machine learning model is a tree-based model constructed based on a decision tree, it is characterized in that, Based on the decision rules of the target machine learning model, the formally compiled input feature value variables are simplified to obtain the root causes of the decision rules of the target machine learning model, specifically including: Based on the discrete partitioning of the feature values, a path from the root node to the leaf node in the target machine learning model is taken as a rule to determine the set of disjunctions between all paths, thus obtaining the decision rule of the tree model. Based on the decision rules of the tree model, the tree model is simplified to obtain a knowledge compilation logic expression; the knowledge compilation logic expression is a logic expression represented by CNF, DNF, NNF or a subset thereof; The knowledge compilation logic expression is simplified to a quality implication term or a principal implication term expression to obtain the rule root cause of the decision of the tree model.

3. The reliable machine learning model interpretation method based on formal methods according to claim 1, wherein when the target machine learning model is a model constructed based on a neural network, it is characterized in that, Based on the decision rules of the target machine learning model, the formally compiled input feature value variables are simplified to obtain the root causes of the decision rules of the target machine learning model, specifically including: Obtain the weights, biases, and activation functions of the neural network model; Based on the weights, biases, activation functions, and continuous input feature values ​​of the neural network model, determine the threshold-based linear function of the neural network model; The threshold conditions of the threshold-based linear function are regularized, and the decision boundaries between each true and false domains under each threshold condition of the threshold-based linear function are calculated. The decision rules of the neural network model are obtained based on the decision boundaries between each tautology domain and each false tautology domain under each threshold condition. Based on the decision rules of the neural network model, the neural network model is simplified to obtain a knowledge compilation logic expression; the knowledge compilation logic expression is a logic expression represented by CNF, DNF, NNF or a subset thereof; The knowledge compilation logic expression is simplified to a quality implication term or a principal implication term expression to obtain the rule root causes of the neural network model's decision.

4. The reliable machine learning model interpretation method based on formal methods according to claim 1, characterized in that, The multi-valued logic variables include Boolean variables and multi-valued discrete logic variables.

5. The reliable machine learning model interpretation method based on formal methods according to claim 1, characterized in that, Discretizing the continuous input feature values ​​to obtain discretized input feature values ​​specifically includes: The continuous input feature values ​​are discretized using the K-means discretization method, the chi-square Chi1 discretization method, the Chi2 discretization method, the minimum description length principal discretization method, or the optimal binning for scoring modeling discretization method to obtain the discretized input feature values.

6. A reliable machine learning model interpretation device based on formal methods, used to implement the reliable machine learning model interpretation method based on formal methods as described in claim 1, characterized in that, include: The acquisition module is used to acquire the target machine learning model and the input feature values ​​of the target machine learning model; The first compilation module is used to compile the input feature values ​​of the target machine learning model using a formal method to obtain continuous variables of input feature values; The discrete module is used to discretize the feature values ​​of the continuous input feature value variable to obtain the discretized input feature value; The discretized input feature values ​​are distributed across multiple discrete feature value partitions; The second compilation module is used to formally compile the discretized input feature values ​​using Boolean logic or multi-valued logic variables to obtain formally compiled input feature value variables. A simplification module is used to represent the formally compiled input feature value variables as a formal logical expression of the target machine learning model according to the decision rules of the target machine learning model, and to simplify the formal logical expression to obtain the rule root causes of the decision of the target machine learning model; the rule root causes are the quality implication term expressions of the target machine learning model. The output module is used to obtain the interpretation results of the target machine learning model based on the root causes of the rules, using natural language representation.

7. The reliable machine learning model interpretation device based on formal methods according to claim 6, characterized in that, The discrete module specifically includes: The discrete submodule is used to discretize the feature values ​​of the continuous input feature value variable using the K-means discretization method, the chi-square Chi1 discretization method, the Chi2 discretization method, the minimum description length principal discretization method, or the optimal binning for scoring modeling discretization method, to obtain the discretized input feature values.