Financial system drill scene key fault risk analysis method based on tree model
Through the tree model-based method, classify the fault data of the financial system, analyze the characteristic importance and predict the risk, the problem of low accuracy of traditional analysis methods is solved, and more efficient identification and optimization of fault risk in the financial system is achieved.
Patent Information
- Application Number
- CN202510622727.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-15
AI Technical Summary
When traditional financial system failure risk analysis methods deal with technical failure risks caused by system code or infrastructure in financial service platforms, there are problems such as data silos, insufficient data governance capabilities, and insufficient processing of technical failure complexity and uncertainty of risk prediction models, resulting in low analysis accuracy.
Using a tree model-based method, we collect log files of the financial system in the chaos engineering fault drill, use Gaussian hybrid model and K-means clustering analysis to classify fault data, combine the random forest decision tree and XGBoost model to perform feature importance analysis and risk prediction on the fault data, and finally determine the key fault through comprehensive analysis and output the critical fault path.
It improves the accuracy and efficiency of financial system failure risk analysis, can more accurately identify key failures, optimize system stability, and reduce service interruptions and economic losses.
Smart Images

Figure CN120125346A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of computer information processing, and particularly to a method for analyzing key fault risks in a financial system drill scenario based on a tree model. Background Art
[0002] With the rapid development of fintech, financial service platforms play an increasingly important role in providing efficient and convenient financial services. These platforms usually rely on complex financial system codes to support core functions such as data management and service operation. However, the stability and reliability of financial systems face many challenges, such as service outages, network delays, data loss, etc. These faults may lead to service interruptions, degraded user experience, and even economic losses, seriously affecting the normal operation of the platform. Therefore, financial institutions need to use advanced technical means to predict and analyze potential faults in the operation of financial systems, in order to identify key risk points and optimize system stability.
[0003] Traditional methods for analyzing financial system fault risks lack targeted analysis tools and methods for technical fault risks caused by system codes or infrastructure in financial service platforms. When dealing with such problems, the existing technologies often face the following limitations: First, the problem of data islands, which makes it difficult to integrate and analyze fault data; second, insufficient data governance capabilities, making it difficult to extract valuable features from complex system logs; third, the existing risk prediction models are insufficient in dealing with the complexity and uncertainty of technical faults, and it is difficult to accurately identify key faults.
[0004] Therefore, the accuracy of current methods for analyzing financial system fault risks is relatively low. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a method, device, computer device, and computer-readable storage medium for analyzing key fault risks in a financial system drill scenario based on a tree model, which can improve the accuracy of the financial system fault risk analysis method.
[0006] A method for analyzing key fault risks in a financial system drill scenario based on a tree model, the method comprising: Collect the log files of the module under test in the financial system during the chaos engineering fault drill, analyze the log files, extract the fault data in the log files, and obtain a fault data set; Classify the fault data set using the Gaussian mixture model and the K-means clustering analysis method to determine the fault types to which the respective fault data in the fault data set belong; Use a decision tree constructed based on random forest to analyze the respective fault data in the fault data set to determine the feature importance of the respective fault data; Input the feature importance of each fault data into the XGBoost model for predictive analysis to obtain the risk prediction values of each fault data; Conduct comprehensive analysis based on the fault types and risk prediction values to which each fault data belongs, determine the key faults, and output the key fault path containing the key faults.
[0007] In one embodiment, classifying the fault data set using the Gaussian mixture model and the K-means clustering analysis method to determine the fault types to which each fault data in the fault data set belongs, including: Fit each fault data in the fault data set through the Gaussian mixture model to obtain the load vector of each fault data; Perform K-means clustering by calculating the distance between the load vector of each fault data and the cluster center to obtain the fault types to which each fault data in the fault data set belongs.
[0008] In one embodiment, fitting each fault data in the fault data set through the Gaussian mixture model to obtain the load vector of each fault data, including: The total number of fault data in the fault data set is n. For any fault data i, i ∈ 1, 2,... n, denote its number of observations as , and the set of observation records is , where represents the th observation value of fault data i, represents the real number space of the observation value or distribution parameter, and d is the dimension of the observation value; Assumption 1: There are K potential true distributions , …, , …, . The observation of any fault data is sampled from one of the potential true distributions, where is the kth potential true distribution, which is used to describe the distribution characteristics of the observation data of the fault data; let represent the relationship between fault data i and the potential true distribution. If the potential true distribution corresponding to fault data i is , then the observation of this fault data i satisfies , …, ~ , denoted as = k; for any k and k', k' ∈ 1, 2,... K, k ≠ k', and are different potential true distributions, where is another potential true distribution; Hypothesis 2: There are G Gaussian components, and the mean and variance of each Gaussian component g are , , where g ∈ 1, 2, …, G; for any two Gaussian components g and g', g' ∈ 1, 2, …, G, g ≠ g', there is ≠ , ≠ ; Hypothesis 3: The latent true distribution , …, , …, is in the space spanned by G Gaussian components; for any latent true distribution , the distribution density is expressed as , where = ([[]] , …, ) is the weight vector of the latent true distribution with respect to G Gaussian components, and satisfies , represents the Gaussian distribution density function with , as parameters; Considering the sample joint distribution based on Hypotheses 1 - 3, denote the relevant parameters of the Gaussian components as , for each fault data i (1 ≤ i ≤ n), denote the vector with respect to G Gaussian components as = ([[]] ), which satisfies , the distribution of each fault data i is represented by a mixture distribution, and the density function of the mixture distribution is , based on the probability density function of the mixture distribution, the observation corresponding probability of the fault data i is: ; where represents the probability density function, represents the Gaussian distribution density function; Denote the set of all fault data coefficient vectors as , then the observation information of n fault data is decomposed into two parts. The first part is the common factor represented in the form of a distribution, and the first part is determined by the parameter Θ; the second part is the load vector represented in the form of a vector, and the second part is described by Φ; Based on Hypotheses 1 - 3, parametric modeling is performed on the generation mechanism of the fault data, and the maximum likelihood estimation method is used to estimate the overall load vector Φ; the log - likelihood function of all fault data observations is: ; Introduce latent variables to describe the correspondence between each observation record at the finest granularity and Gaussian components. For the j-th observation of any fault data i , if comes from Gaussian component g, then = 0. Denote the set of latent variables for all observations as: ; After introducing the set of latent variables , the log-likelihood function of all fault data observations is expressed as: ; where the superscript represents transpose; Use the EM algorithm to perform iterative optimization on . Denote the number of iterations of the EM algorithm as T, and denote the parameter estimates at the t-th iteration as and , t = 0, 1, 2,..., T; the initial values of the parameter estimates are and ; In each iteration of the EM algorithm, two steps need to be performed. The first step is to calculate the conditional expectation of the log-likelihood function under the current parameter estimates, and this step is called the E-step; the second step is to find the parameter values that maximize the conditional expectation as the updated parameter estimates, and this step is called the M-step; at the t-th iteration, update the parameter estimates and to and . The specific processes of the E-step and M-step that need to be executed in sequence are as follows: E-step: Given and , find the expectation of the log-likelihood function l(Θ, Φ, Ψ) with respect to the set of latent variables Ψ as the objective function, and this objective function is: ; where represents expectation; Denote the conditional expectation of as , and obtain it through conditional probability calculation: ; where represents the weight of fault data i at the t-th iteration on the g-th Gaussian component, represents the mean on the g-th Gaussian component at the t-th iteration, Denote the variance on the \(g\)-th Gaussian component at the \(t\)-th iteration; M-step: Find the parameters that maximize the objective function as the updated parameter estimates and ; The specific update formulas are obtained by solving: , , ; where, is the new mean of the fault data \(i\) on the \(g\)-th Gaussian component, is the new weight of the fault data \(i\) on the \(g\)-th Gaussian component, and the superscript represents the transpose; Iteratively repeat the E-step and M-step until the estimation converges or reaches the set termination condition, and output the final estimated values , where, is the estimated value of the mean on the \(g\)-th Gaussian component, is the estimated value of the variance on the \(g\)-th Gaussian component, is the relative importance of the fault data \(i\) on the \(g\)-th Gaussian component, and the load vector of each fault data \(i\) is .
[0009] In one embodiment, the number of Gaussian components \(G\) is determined using the Bayesian information criterion, where the Bayesian information criterion has the expression: : where, represents the value of the log-likelihood function obtained by substituting the parameter estimates obtained by iterative optimization given the number of Gaussian components as \(G\), and \(d\) is the parameter dimension of each Gaussian component.
[0010] In one embodiment, using the decision tree constructed based on the random forest to analyze each fault data in the fault dataset to determine the feature importance of each fault data, including: Suppose there are \(n\) fault data, \(I\) decision trees, \(C\) categories, and the Gini index of the node \(q\) of the \(\alpha\)-th decision tree is calculated as: ; where, represents the proportion of the category in the node \(q\) of the \(\alpha\)-th decision tree, represents the proportion of the category in the node \(q\) of the \(\alpha\)-th decision tree; The importance of the failure data \(i\) at the node \(q\) of the \(\alpha\)-th decision tree, that is, the change in the Gini index of the failure data \(i\) before and after branching at the node \(q\) of the \(\alpha\)-th decision tree is: ; Wherein, and respectively represent the Gini indices of the two new nodes after branching; If the nodes where the failure data \(i\) appears in the \(\alpha\)-th decision tree are the set \(Q\), then the importance of the failure data \(i\) in the \(\alpha\)-th decision tree is: ; Assume that there are \(I\) decision trees in the random forest, and the Gini index of the failure data \(i\) in the random forest is: ; Wherein, is the Gini index of the failure data \(i\) in the random forest; Normalize the Gini index of the failure data \(i\) in the random forest to obtain the feature importance of the failure data \(i\): ; Wherein, is the feature importance of the failure data \(i\).
[0011] In one embodiment, the XGBoost model expression is: ; Wherein, \(H\) represents the total number of trees generated by iterative prediction, represents a function in the function space \(F\); is the risk prediction value of the failure data \(i\); The objective function of the XGBoost model is: ; Wherein, is the loss function, is the penalty term, and \(n\) represents the total number of failure data; the form is: ; Wherein, \(M\) is the number of leaf nodes of the tree, is the weight of the leaf node \(m\), \(\gamma\) and \(\tau\) are pre-designed hyperparameters, is the mathematical model of the \(t\)-th iteration; The process of iteratively solving the XGBoost model is: ; Wherein, is the initial risk prediction value of the XGBoost model for fault data i, is the risk prediction value of the XGBoost model for fault data i after the first iteration, is the risk prediction value of the XGBoost model for fault data i after the second iteration, is the risk prediction value of the XGBoost model for fault data i after the t-th iteration, is the mathematical model at the first iteration, is the mathematical model at the second iteration, is the mathematical model at the v-th iteration, is the risk prediction value of the XGBoost model for fault data i after the (t - 1)-th iteration; The objective function of the XGBoost model for the t-th iteration is: ; where, represents the sum of the risk prediction values of the previous t - 1 iterations for fault data i, represents the base learning rate of the XGBoost model; The objective function of the XGBoost model for the t-th iteration is approximated using the Taylor formula as: , ; where, is the first-order derivative of the feature importance of fault data i; represents the second-order derivative of the feature importance of fault data i.
[0012] In one embodiment, the comprehensive analysis based on the fault type and risk prediction value of each fault data to determine the key fault and output the key fault path including the key fault includes: According to the fault type and risk prediction value of each fault data, a risk value weighting formula is used for comprehensive analysis to obtain the final risk value of each fault data; According to the final risk value of each fault data, the fault data with the final risk value exceeding the threshold is determined, and the fault represented by the fault data is determined as the key fault, and the key fault path including the key fault is output.
[0013] In one embodiment, the risk value weighting formula is: , ; where, is the final risk value of fault data i, is the risk coefficient of fault data i, is the number of fault data contained in the cluster to which fault data i belongs, is the distance between the load vector of fault data i and the cluster center.
[0014] A key fault risk analysis device for a financial system drill scenario based on a tree model, comprising: A data collection module, configured to collect log files of the module under test in the financial system during a chaos engineering fault drill, analyze the log files, extract fault data from the log files, and obtain a fault data set; A clustering module, configured to classify the fault data set using a Gaussian mixture model and K-means clustering analysis method to determine the fault types to which the respective fault data in the fault data set belong; A fault data analysis module, configured to analyze each fault data in the fault data set using a decision tree constructed based on a random forest to determine the feature importance of each fault data; A risk prediction module, configured to input the feature importance of each fault data into an XGBoost model for prediction analysis to obtain a risk prediction value for each fault data; A key fault analysis module, configured to perform comprehensive analysis based on the fault types and risk prediction values to which the respective fault data belong, determine key faults, and output a key fault path including the key faults.
[0015] A computer device, comprising a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the key fault risk analysis method for a financial system drill scenario based on a tree model are implemented.
[0016] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the key fault risk analysis method for a financial system drill scenario based on a tree model are implemented.
[0017] The above-mentioned key fault risk analysis method, device, computer equipment and computer-readable storage medium for the financial system drill scenario based on the tree model collect the log files of the module under test in the financial system during the chaos engineering fault drill, analyze the log files, extract the fault data from the log files to obtain a fault data set, classify the fault data set using the Gaussian mixture model and the K-means clustering analysis method, determine the fault types to which each fault data in the fault data set belongs, use a decision tree constructed based on a random forest to analyze each fault data in the fault data set, determine the feature importance of each fault data, input the feature importance of each fault data into the XGBoost model for predictive analysis, obtain the risk prediction values of each fault data, and perform comprehensive analysis based on the fault types and risk prediction values to which each fault data belongs to determine the key faults and output the key fault paths containing the key faults. Thereby, the accuracy and efficiency of the fault risk analysis of the financial system are improved. Brief Description of the Drawings
[0018] Figure 1 It is a schematic flowchart of a key fault risk analysis method for a financial system drill scenario based on a tree model in an embodiment. Detailed Description of the Embodiment
[0019] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0020] In one embodiment, as Figure 1 shown, a key fault risk analysis method for a financial system drill scenario based on a tree model is provided, including the following steps: Step S220: Collect the log files of the module under test in the financial system during the chaos engineering fault drill, analyze the log files, and extract the fault data from the log files to obtain a fault data set.
[0021] Among them, the fault data is represented in a structured manner.
[0022] In one example, during the chaos engineering automated fault drill, system data and monitoring information during the drill are captured, and the collected fault data is structured to ensure clear insight and maximum value from the fault drill.
[0023] Step S240: Classify the fault data set using the Gaussian mixture model and the K-means clustering analysis method to determine the fault types to which each fault data in the fault data set belongs.
[0024] Among them, each piece of fault data in the fault dataset is clustered by K-means into similar groups or clusters, that is, each group or cluster represents a fault type.
[0025] Among them, the fault types can include service downtime, network latency, data loss, and so on.
[0026] In one embodiment, the Gaussian mixture model and K-means clustering analysis method are used to classify the fault dataset to determine the fault type to which each piece of fault data in the fault dataset belongs, including: Each piece of fault data in the fault dataset is fitted through the Gaussian mixture model to obtain the load vector of each piece of fault data; by performing K-means clustering on the distance between the load vector of each piece of fault data and the cluster center, the fault type to which each piece of fault data in the fault dataset belongs is obtained.
[0027] Among them, for the clustering objects represented by the distribution function, in order to achieve the clustering division of complex data, parametric data modeling can be relied on the Gaussian mixture model. Fault data often has a complex distribution, and a feasible way to model the data distribution is to adopt a finite mixture model.
[0028] Among them, the GMM (Gaussian Mixture Module) not only realizes the modeling of complex distributions, but also provides a vector expression based on Gaussian components for all sample observations. This expression can be further applied to clustering and classification algorithms. Consider modeling a set of independent and identically distributed samples with a GMM containing G Gaussian components, where the parameters of each Gaussian component g (1≤g≤G) are: mean , variance , and weight .
[0029] Among them, the fault data has a hierarchical nested structure and does not satisfy the independence assumption. Therefore, on the basis of the GMM, it is necessary to further design a clustering algorithm for the distribution function. First, all data observations are fitted through the GMM at the observation level to obtain the vector representation of each observation. For each object, the vector representation of the clustering object is indirectly obtained through the average value of the corresponding observation vectors, so as to apply common clustering algorithms such as K-means to divide the clustering objects represented by the vectors. This method is denoted as GMM+K-means.
[0030] Among them, the process of using this clustering analysis method to divide the structured fault data into similar groups or clusters is as follows: 1. Set the fault factor model; 2. Solve the fault factor model; 3. After obtaining the load vector of each fault data according to the iterative estimation, use the K-means clustering method to obtain the clustering division result of all fault data.
[0031] Among them, after obtaining the load vector of each fault data according to the iterative estimation, the clustering division result of all fault data can be obtained through the K-means clustering method. For the k-th cluster, examine the load vectors of all samples divided into this cluster in this round of iteration, and use the mean value of their estimated values as the vector representation of the cluster center. The cluster center obtained in this way has good representativeness for the vectors within the current cluster. After determining the cluster center in each round of iteration, the cluster center with the smallest Euclidean distance from its load vector can be found for each fault data, and the fault data can be divided into the corresponding cluster. Repeat the iteration until convergence or the termination condition is reached, and the clustering division of all fault data can be obtained.
[0032] In one embodiment, each fault data in the fault dataset is fitted through a Gaussian mixture model to obtain the load vector of each fault data, including: The total number of fault data in the fault dataset is n. For any one fault data i, i ∈ 1, 2, …, n, denote its number of observations as , and the set of observation records is , where represents the -th observation value of fault data i, represents the real number space of observation values or distribution parameters, and d is the dimension of the observation values; Hypothesis 1: There are K potential true distributions , …, , …, , and the observation of any one fault data is sampled from one of the potential true distributions, where is the k-th potential true distribution, which is used to describe the distribution characteristics of the observed data of the fault data; let represent the relationship between fault data i and the potential true distribution. If the potential true distribution corresponding to fault data i is , then the observation of this fault data i satisfies , …, ~ , denoted as = k; for any k and k', k' ∈ 1, 2, …, K, k ≠ k', and are different potential true distributions, where is another potential true distribution; Hypothesis 2: There are G Gaussian components, and the mean and variance of each Gaussian component g are , , g ∈ 1, 2, …, G; for any two Gaussian components g and g', g' ∈ 1, 2, …, G, g ≠ g', there is ≠ , ≠ ; Hypothesis 3: Latent true distribution 、 …、 、…、 In the space spanned by G Gaussian components; for any latent true distribution , the distribution density is expressed as , where = ([[]] ,…, ) is the weight vector of the latent true distribution with respect to G Gaussian components, and satisfies , denotes the Gaussian distribution density function with , as parameters; Considering the sample joint distribution based on Hypotheses 1 - 3, denote the parameters related to the Gaussian components as , for each fault data i (1 ≤ i ≤ n), denote the vector with respect to G Gaussian components as = ([[]] ), which satisfies , the distribution of each fault data i is represented by a mixture distribution, and the density function of the mixture distribution is , based on the probability density function of the mixture distribution, the observation corresponding probability of the fault data i is:[[]] ; where denotes the probability density function, denotes the Gaussian distribution density function; Denote the set of all fault data coefficient vectors as , then the observation information of n fault data is decomposed into two parts. The first part is the common factor represented in the form of a distribution, and the first part is determined by the parameter Θ; the second part is the load vector represented in the form of a vector, and the second part is described by Φ.[[]]
[0033] Among them, since this application uses parametric statistics to handle the clustering problem of the distribution function, it is necessary to make assumptions about the data structure, etc., and construct a suitable clustering method based on the assumptions.[[]]
[0034] It should be understood that if the hierarchical structure between the object and the observation is not considered, Hypotheses 1 and 2 are basically the same as the general GMM hypotheses. Combining Hypotheses 1 and 3, it is required that the latent true distribution function has identifiability. Specifically, for any 1 ≤ k ≠ k' ≤ K, the corresponding weight vectors and They are not exactly the same. The model under this assumption is called the fault factor model. In this fault factor model, the observation of each object (i.e., fault data) is generated by the combined action of a common factor represented by Gaussian components and a loading vector. Therefore, estimating this fault factor model can decompose the fault data information into a common factor and a loading vector, where the loading vector contains the personalized characteristics of each fault and is suitable for clustering.
[0035] Among them, based on Assumption 1, for any fault data i, we can obtain . When Φ is known, for any 1 ≤ i ≠ j ≤ n, if and only if . Therefore, within the framework of the fault factor model, the core goal is to use the observation information to estimate the loading vector Φ of all clustering objects. After obtaining the estimate of the loading vector, objects with similar vector representations are assigned to the same cluster, and objects with dissimilar vectors are assigned to different clusters. Since the dimension of any vector in Φ is fixed at G, after obtaining the estimate of Φ, traditional clustering methods such as K-means can be used for rapid clustering.
[0036] Based on Assumptions 1 - 3, parameterize the generation mechanism of fault data for modeling, and use the maximum likelihood estimation method to estimate the overall loading vector Φ; the log-likelihood function of all fault data observations is: ; Introduce the latent variable to describe the correspondence between each observation record at the finest granularity and the Gaussian components. For the j-th observation of any fault data i, if comes from the Gaussian component g, then = 0. Denote the set of latent variables of all observations as: ; After introducing the set of latent variables , the log-likelihood function of all fault data observations is expressed as: ; Among them, the superscript represents the transpose; Use the EM algorithm to iteratively optimize . Denote the number of iterations of the EM algorithm as T, and denote the parameter estimates at the t-th iteration as and , t = 0, 1, 2…, T; the initial values of the parameter estimates are and ; In each iteration of the EM algorithm, two steps need to be executed. The first step is to calculate the conditional expectation of the log-likelihood function under the current parameter estimate, and this step is called the E-step. The second step is to find the parameter values that maximize the conditional expectation as the updated parameter estimate, and this step is called the M-step. In the t-th iteration, the parameter estimates and are updated to and . The specific processes of the E-step and M-step that need to be executed in sequence are as follows: E-step: Given and , find the expectation of the log-likelihood function l(Θ, Φ, Ψ) with respect to the set of latent variables Ψ as the objective function, and this objective function is: ; Among them, represents the expectation; The conditional expectation of is abbreviated as , and is calculated through the conditional probability: ; Among them, represents the weight of the failure data i in the g-th Gaussian component in the t-th iteration, represents the mean on the g-th Gaussian component in the t-th iteration, represents the variance on the g-th Gaussian component in the t-th iteration; M-step: Find the parameters that maximize the objective function as the updated parameter estimates and ; The specific update formulas obtained by solving are: , , ; Among them, is the new mean of the failure data i on the g-th Gaussian component, is the new weight of the failure data i on the g-th Gaussian component, and the superscript represents the transpose; Iteratively repeat the E-step and M-step until the estimation converges or reaches the set termination condition, and output the final estimated values , among which, is the estimated value of the mean on the g-th Gaussian component, is the estimated value of the variance on the g-th Gaussian component, is the relative importance of the fault data i on the g-th Gaussian component, and the load vector of each fault data i .
[0037] Among them, for in the above formula, taking the derivative with respect to each parameter, an explicit solution of the parameter estimation cannot be obtained. Therefore, the EM algorithm is considered, and the parameter estimation is obtained through iterative optimization. To simplify the calculation of the maximum likelihood estimation, the latent variable is considered to be introduced to describe the correspondence between each observation record at the finest granularity and the Gaussian component.
[0038] Among them, regarding the selection of the number of Gaussian components G, when a larger value is assigned to G, the data can be more fully fitted, but it may also bring more redundant parameters. The selection of G actually involves the problem of the selection of the fault factor model. To balance the complexity of the fault factor model and the data fitting effect, the Bayesian Information Criterion (BIC) can be considered to help determine the appropriate number of components.
[0039] In one embodiment, the number of Gaussian components G is determined using the Bayesian Information Criterion, where the Bayesian Information Criterion has the following expression: ; Among them, represents the value of the log-likelihood function obtained by substituting the parameter estimation obtained through iterative optimization given that the number of Gaussian components is G, and d is the parameter dimension of each Gaussian component.
[0040] Among them, the smaller the value of BIC, the more appropriate the corresponding G is, and a better data fitting effect can be obtained with fewer model parameters.
[0041] Step S260, use the decision tree constructed based on the random forest to analyze each fault data in the fault dataset, and determine the feature importance of each fault data.
[0042] Among them, the random forest is a classifier that uses multiple decision trees to train and predict samples.
[0043] Among them, to solve the problem that the traditional single key fault path search algorithm based on clustering analysis is difficult to comprehensively consider in the case of multiple parameters and has poor accuracy and interpretability in complex environments, a decision tree model is also used to assist in analyzing the fault data.
[0044] Among them, the method of using the Gini index to evaluate is used to obtain the feature importance of each fault data. In the random forest, whenever a node splits, the algorithm calculates the reduction in Gini impurity through this split. The average value of the reduction in Gini impurity for each feature split in all decision trees is used to obtain the Gini importance of the feature.
[0045] In one embodiment, decision trees constructed based on random forests are used to analyze each piece of fault data in the fault dataset to determine the feature importance of each piece of fault data, including: Suppose there are n pieces of fault data, I decision trees, and C categories. The Gini index of node q in the α-th decision tree is calculated as follows: ; where, represents the proportion of category in node q of the α-th decision tree, represents the proportion of category in node q of the α-th decision tree; The importance of fault data i in node q of the α-th decision tree, that is, the change in the Gini index of fault data i before and after branching in node q of the α-th decision tree is: ; where, and represent the Gini indices of the two new nodes after branching respectively; If the nodes where fault data i appears in the α-th decision tree are the set Q, then the importance of fault data i in the α-th decision tree is: ; Assume that there are I decision trees in the random forest. The Gini index of fault data i in the random forest is: ; where, is the Gini index of fault data i in the random forest; Normalize the Gini index of fault data i in the random forest to obtain the feature importance of fault data i: ; where, is the feature importance of fault data i.
[0046] Step S280: Input the feature importance of each piece of fault data into the XGBoost model for predictive analysis to obtain the risk prediction values of each piece of fault data.
[0047] Among them, the XGBoost algorithm of the XGBoost model is the GBDT algorithm mode, belonging to the category of Boosting iterative and tree methods. The idea of this algorithm is to grow a tree through feature analysis and continuously add a tree. For each added tree, some leaf nodes in each leaf will be reached, and each leaf node corresponds to a score. Therefore, the final result is simply to add up the scores corresponding to each leaf, which is the risk prediction value of the fault data.
[0048] Among them, the XGBoost algorithm adopts a step-by-step forward addition mode, and after generating a weak learner in one iteration, it no longer requires recalculating a certain relationship.
[0049] In one embodiment, the XGBoost model expression is: ; Among them, H represents the total number of trees generated by iterative prediction, represents a function in the function space F; is the risk prediction value of the fault data i; The objective function of the XGBoost model is: ; Among them, is the loss function, is the penalty term, and n represents the total number of fault data; the form is: ; Among them, M is the number of leaf nodes of the tree, is the weight of the leaf node m, γ and τ are pre-designed hyperparameters, is the mathematical model of the t-th iteration.
[0050] Among them, when introducing a regularization term, the calculation will select a simplified and excellent performance mode. The rightmost regularization term in the loss function is used to control the overfitting ability of the weakest learner in each iteration, without involving the set of the last module.
[0051] The process of iteratively solving the XGBoost model is: ; Among them, is the initial risk prediction value of the XGBoost model for the fault data i, is the risk prediction value of the XGBoost model for the fault data i after the first iteration, is the risk prediction value of the XGBoost model for the fault data i after the second iteration, is the risk prediction value of fault data i after the t-th iteration of the XGBoost model, is the mathematical model at the 1st iteration, is the mathematical model at the 2nd iteration, is the mathematical model at the v-th iteration, is the risk prediction value of fault data i after the (t - 1)-th iteration of the XGBoost model.
[0052] Among them, one iteration of the XGBoost model will form a new decision tree. The decision tree is established according to the residual from the actual value, and the established decision tree will predict the result for the first time. Thus, the prediction result of the t-th tree can be obtained, which is numerically equal to the prediction results of the previous t - 1 trees plus the performance of the t-th tree.
[0053] The objective function of the XGBoost model for the t-th iteration is: ; where, represents the sum of the risk prediction values of fault data i for the previous t - 1 iterations, represents the base learning rate of the XGBoost model; Approximating the objective function of the t-th iteration of the XGBoost model using the Taylor formula gives: , ; where, is the first-order derivative of the feature importance of fault data i; represents the second-order derivative of the feature importance of fault data i.
[0054] Among them, when calculating the t-th tree of the model, the results and model structures of the previous t - 1 trees are fixed. Therefore, it can be seen that the objective function is transformed into a function about and . During the operation of the results, the accuracy of the results is determined by the error function of the samples. During the process of obtaining the final results, the regression processes are similar and end when reaching the selected values of the model structure parameters.
[0055] Step S300, based on the comprehensive analysis of the fault types and risk prediction values of each fault data, determine the key faults and output the key fault paths containing the key faults.
[0056] Among them, the key faults can be determined as the key monitoring objects under the tested module of the financial system, improving the stability of the financial system.
[0057] In one embodiment, comprehensive analysis is performed based on the fault type and risk prediction value to which each fault data belongs, to determine critical faults, and output a critical fault path including the critical faults, including: According to the fault type and risk prediction value to which each fault data belongs, comprehensive analysis is performed using a risk value weighting formula to obtain the final risk value of each fault data; According to the final risk value of each fault data, determine the fault data whose final risk value exceeds the threshold, and determine the fault represented by the fault data as a critical fault, and output a critical fault path including the critical fault.
[0058] Among them, the critical fault is a certain fault in the directly locked log file; the critical fault path is a complete fault chain with the critical fault, which may include the log data before the occurrence of the critical fault and the subsequent chained log data.
[0059] In one embodiment, the risk value weighting formula is: , ; Among them, is the final risk value of fault data i, is the risk coefficient of fault data i, is the number of fault data included in the cluster to which fault data i belongs, is the distance between the load vector of fault data i and the cluster center.
[0060] The above-mentioned key fault risk analysis method for the financial system drill scenario based on the tree model collects the log files of the modules under test in the financial system during the chaos engineering fault drill, analyzes the log files, extracts the fault data in the log files to obtain a fault data set, classifies the fault data set using the Gaussian mixture model and the K-means clustering analysis method to determine the fault type to which each fault data in the fault data set belongs, uses the decision tree constructed based on the random forest to analyze each fault data in the fault data set to determine the feature importance of each fault data, inputs the feature importance of each fault data into the XGBoost model for prediction analysis to obtain the risk prediction value of each fault data, and performs comprehensive analysis based on the fault type and risk prediction value to which each fault data belongs to determine critical faults, and outputs a critical fault path including the critical faults. Thereby, the accuracy and efficiency of the fault risk analysis of the financial system are improved.
[0061] It should be understood that although Figure 1The steps in the flowchart are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 at least a part of the steps in Figure 1 may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in rotation with at least a part of other steps or sub-steps or stages of other steps.
[0062] In one embodiment, a key fault risk analysis device for a financial system drill scenario based on a tree model is provided, including: a data collection module, a clustering module, a fault data analysis module, a risk prediction module, and a key fault analysis module.
[0063] The data collection module is configured to collect log files of the module under test in the financial system during the chaos engineering fault drill, analyze the log files, extract fault data from the log files, and obtain a fault data set.
[0064] The clustering module is configured to classify the fault data set using a Gaussian mixture model and K-means clustering analysis method to determine the fault types to which the respective fault data in the fault data set belong.
[0065] The fault data analysis module is configured to analyze the respective fault data in the fault data set using a decision tree constructed based on a random forest to determine the feature importance of each fault data.
[0066] The risk prediction module is configured to input the feature importance of each fault data into an XGBoost model for predictive analysis to obtain a risk prediction value for each fault data.
[0067] The key fault analysis module is configured to perform comprehensive analysis based on the fault types to which the respective fault data belong and the risk prediction values to determine key faults, and output a key fault path including the key faults.
[0068] The above key fault risk analysis device for the financial system drill scenario based on the tree model collects the log files of the module under test in the financial system during the chaos engineering fault drill, analyzes the log files, extracts the fault data in the log files to obtain a fault data set, classifies the fault data set using the Gaussian mixture model and the K-means clustering analysis method to determine the fault types to which each fault data in the fault data set belongs, analyzes each fault data in the fault data set using a decision tree constructed based on a random forest to determine the feature importance of each fault data, inputs the feature importance of each fault data into the XGBoost model for predictive analysis to obtain the risk prediction values of each fault data, and performs a comprehensive analysis based on the fault types to which each fault data belongs and the risk prediction values to determine the key faults and outputs the key fault paths including the key faults. Thereby, the accuracy and efficiency of the fault risk analysis of the financial system are improved.
[0069] For the specific limitations of the key fault risk analysis device for the financial system drill scenario based on the tree model, reference can be made to the limitations of the key fault risk analysis method for the financial system drill scenario based on the tree model in the above text, which will not be elaborated here. Each module in the above key fault risk analysis device for the financial system drill scenario based on the tree model can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form so that the processor can call and execute the operations corresponding to the above modules.
[0070] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above key fault risk analysis method for the financial system drill scenario based on the tree model are implemented.
[0071] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above key fault risk analysis method for the financial system drill scenario based on the tree model are implemented.
[0072] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0073] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0074] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation of the scope. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for analyzing key failure risks in financial system drill scenarios based on a tree model, characterized in that: The tree model-based financial system exercise scenario critical failure risk analysis method includes: Collect log files of the tested modules of the financial system during the chaos engineering fault drill, analyze the log files, extract fault data from the log files, and obtain a fault data set; The fault data set is classified using a Gaussian mixture model and a K-means clustering analysis method to determine the fault type to which each fault data in the fault data set belongs; Using a decision tree constructed based on random forest to analyze each fault data in the fault data set, and determine the feature importance of each fault data; The feature importance of each fault data is input into the XGBoost model for prediction analysis to obtain the risk prediction value of each fault data; A comprehensive analysis is performed based on the fault type and risk prediction value of each fault data, critical faults are determined, and a critical fault path including the critical faults is output.
2. The method for analyzing critical failure risks in financial system drill scenarios based on a tree model according to claim 1 is characterized in that: The method of classifying the fault data set using a Gaussian mixture model and a K-means clustering analysis method to determine the fault type to which each fault data in the fault data set belongs includes: Fitting each fault data in the fault data set through a Gaussian mixture model to obtain a load vector for each fault data; By performing K-means clustering on the distance between the load vector of each fault data and the cluster center, the fault type to which each fault data in the fault data set belongs is obtained.
3. The method for analyzing critical failure risks in financial system drill scenarios based on a tree model according to claim 2 is characterized in that: The fitting of each fault data in the fault data set by a Gaussian mixture model to obtain a load vector for each fault data includes: The total number of fault data in the fault data set is n. For any fault data i, i∈1, 2, ... n, the number of observations is , the set of observation records is ,in, Indicates the fault data i Observations, The real number space representing the observation value or distribution parameter, d is the dimension of the observation value; Assumption 1: There are K potential true distributions , …、 , …, , any observation of fault data is sampled from one of the potential true distributions, where is the kth potential true distribution, which is used to describe the distribution characteristics of the observed data of the fault data; let Represents the relationship between fault data i and the potential true distribution. If the fault data i corresponds to the potential true distribution , then the observation of the fault data i satisfies ,…, ~ , denoted as = k; for any k and k', k'∈1, 2, ...K, k≠k', and are different potential true distributions, where is another potential true distribution; Assumption 2: There are G Gaussian components, and the mean and variance of each Gaussian component g are , , g∈1,2,…G; for any two Gaussian components g and g', g'∈1,2,…G, g≠g', ≠ , ≠ ; Assumption 3: Underlying True Distribution , …、 , …, In the space spanned by G Gaussian components; for any potential true distribution , The distribution density of ,in, =( ,…, ) is the underlying true distribution Regarding the weight vector of G Gaussian components, and satisfying , Indicates , is the Gaussian distribution density function of the parameter; Based on assumptions 1 to 3, consider the joint distribution of samples and record the Gaussian component related parameters as , for each fault data i (1≤i≤n), the vector of G Gaussian components is recorded as =( ),satisfy , the distribution of each fault data i is represented by a mixed distribution, and the density function of the mixed distribution is , based on the probability density function of the mixed distribution, the observed corresponding probability of fault data i is for: ; in, represents the probability density function, represents the Gaussian distribution density function; The set of all fault data coefficient vectors is , then the observation information of n fault data is decomposed into two parts, the first part is the common factor expressed in the form of distribution, the first part is determined by the parameter Θ; the second part is the load vector expressed in the form of vector, the second part is described by Φ; Based on assumptions 1 to 3, the generation mechanism of fault data is modeled parametrically, and the maximum likelihood estimation method is used to estimate the overall load vector Φ; the log-likelihood function of all fault data observations is for: ; Introducing hidden variables To describe the correspondence between each observation record and Gaussian component at the finest granularity, for any fault data i, the jth observation ,like From the Gaussian component g, then = 0, and the set of hidden variables for all observations is recorded as: ; Introducing a set of hidden variables After that, the log-likelihood function of all failure data observations is It is expressed as: ; Among them, the superscript represents transpose; Using EM algorithm Perform iterative optimization, record the number of EM algorithm iterations as T, and record the parameter estimate of the tth iteration as and , t=0, 1, 2…, T; the initial value of the parameter estimate is and Each iteration of the EM algorithm requires two steps. The first step is to calculate the conditional expectation of the log-likelihood function under the current parameter estimate. This step is called the E-step. The second step is to find the parameter value that maximizes the conditional expectation as the updated parameter estimate. This step is called the M-step. At the tth iteration, the parameter estimate is and Updated to and The specific processes of E-step and M-step that need to be executed in sequence are as follows: E-step: Given and When , we find the expectation of the log-likelihood function l(Θ,Φ,Ψ) with respect to the latent variable set Ψ as the objective function. for: ; in, Expressed as expectation; Will The conditional expectation of , calculated by conditional probability: ; in, represents the weight of the fault data i of the t-th iteration on the g-th Gaussian component, represents the mean of the g-th Gaussian component at the t-th iteration, represents the variance of the g-th Gaussian component at the t-th iteration; M-step: Find the objective function The maximized parameter is used as the updated parameter estimate and ; The specific update formula obtained by solving is: , , ; in, is the new mean of the fault data i on the g-th Gaussian component, is the new weight of fault data i on the gth Gaussian component, with superscript represents transpose; Iterate and repeat the E-step and M-step until the estimate converges or reaches the set termination condition, and output the final estimate ,in, is the estimated value of the mean on the g-th Gaussian component, is the estimated value of the variance on the g-th Gaussian component, is the relative importance of fault data i on the g-th Gaussian component, and the load vector of each fault data i is .
4. The method for analyzing critical failure risks in financial system drill scenarios based on a tree model according to claim 1 is characterized in that: The use of a decision tree constructed based on a random forest to analyze each fault data in the fault data set to determine the feature importance of each fault data includes: Suppose there are n fault data, I decision trees, C categories, and the Gini index of node q of the αth decision tree is The calculation formula is: ; in, Represents the category in node q of the αth decision tree The proportion of Represents the category in node q of the αth decision tree The proportion of The importance of fault data i at node q of the αth decision tree, that is, the change in the Gini index of fault data i before and after the branching of node q of the αth decision tree is: ; in, and Respectively represent the Gini index of the two new nodes after branching; The nodes where the fault data i appears in the αth decision tree are set Q, so the importance of the fault data i in the αth decision tree is for: ; Assuming that there are I decision trees in the random forest, the Gini index of fault data i in the random forest is: ; in, is the Gini index of fault data i in random forest; Normalize the Gini index of fault data i in the random forest to obtain the feature importance of fault data i: ; in, is the feature importance of fault data i.
5. The method for analyzing critical failure risks in financial system drill scenarios based on a tree model according to claim 4 is characterized in that: The XGBoost model expression is: ; Among them, H represents the total number of trees generated by iterative prediction, represents a function in the function space F; is the risk prediction value of fault data i; The objective function of the XGBoost model for: ; in, is the loss function, is a penalty term, n represents the total number of fault data; the form is: ; Among them, M is the number of leaf nodes of the tree, is the weight of leaf node m, γ and τ are pre-designed hyperparameters, is the mathematical model of the tth iteration; The process of iteratively solving the XGBoost model is: ; in, is the initial risk prediction value of the XGBoost model for fault data i, is the risk prediction value of fault data i after the first iteration of the XGBoost model, is the risk prediction value of fault data i after the second iteration of the XGBoost model, is the risk prediction value of fault data i after the tth iteration of the XGBoost model, is the mathematical model for the first iteration, is the mathematical model for the second iteration, is the mathematical model at the vth iteration, is the risk prediction value of fault data i after the t-1th iteration of the XGBoost model; The objective function of the XGBoost model for the tth iteration is for: ; in, represents the sum of the risk prediction values of the fault data i in the first t−1 iterations, Represents the base learning rate of the XGBoost model; The objective function of the t-th iteration of the XGBoost model approximated by Taylor's formula is: , ; in, is the first-order derivative of the feature importance of fault data i; The second-order derivative representing the feature importance of fault data i.
6. The method for analyzing critical failure risks in financial system drill scenarios based on a tree model according to claim 1 is characterized in that: The method of performing a comprehensive analysis based on the fault type and risk prediction value of each fault data to determine the critical fault and outputting the critical fault path including the critical fault includes: According to the fault type and risk prediction value of each fault data, a comprehensive analysis is performed using the risk value weighting formula to obtain the final risk value of each fault data; According to the final risk value of each fault data, the fault data whose final risk value exceeds the threshold is determined, the fault represented by the fault data is determined as a critical fault, and a critical fault path including the critical fault is output.
7. The method for analyzing critical failure risks in financial system drill scenarios based on a tree model according to claim 6 is characterized in that: The risk value weighting formula is: , ; in, is the final risk value of failure data i, is the risk factor of failure data i, is the number of faulty data contained in the cluster to which faulty data i belongs, is the distance between the load vector of fault data i and the cluster center.
8. A device for analyzing key failure risks in financial system drill scenarios based on a tree model, characterized in that: include: A data collection module is used to collect log files of the tested modules of the financial system during the chaos engineering fault drill, analyze the log files, extract fault data from the log files, and obtain a fault data set; A clustering module, used to classify the fault data set using a Gaussian mixture model and a K-means clustering analysis method, and determine the fault type to which each fault data in the fault data set belongs; A fault data analysis module, used to analyze each fault data in the fault data set using a decision tree constructed based on random forest, and determine the feature importance of each fault data; The risk prediction module is used to input the feature importance of each fault data into the XGBoost model for prediction analysis to obtain the risk prediction value of each fault data; The critical fault analysis module is used to perform a comprehensive analysis based on the fault type and risk prediction value of each fault data, determine the critical fault, and output the critical fault path containing the critical fault.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the tree model-based financial system exercise scenario critical failure risk analysis method described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the tree model-based financial system exercise scenario critical failure risk analysis method described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Fan key part fault diagnosis method
CN111444940A
Power transmission system fault classification method based on GA and XGBoost-RF stacking algorithm
CN119089322A