Defective asset case retrieval analysis method and system
By cleaning and feature extraction of corporate financial data, combining dynamic weight logistic regression rules and multi-layer knowledge graph technology, high-risk paths and fine-screen high-risk cases are solved, and the problem that traditional rule engines are difficult to penetrate multi-layer relationships is improved, and asset recovery rate and identification accuracy are improved.
Patent Information
- Application Number
- CN202510211589.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional rule engines are difficult to penetrate multi-layer relationships, which leads to financial institutions being unable to effectively recover hidden assets when pursuing non-performing assets, resulting in a decrease in asset recovery rate and increasing losses.
By cleaning, standardizing and feature extraction of corporate financial data, market information and historical cases, structured data is generated; based on dynamic weight logistic regression rules, a multi-layer knowledge graph is constructed, and a space-time graph attention network is used to calculate risk propagation weights, identify high-risk paths, and through a graph-enhanced hybrid gradient enhancement model, combining structured features and graph embedding features, fine-screen high-risk cases.
It improves the accuracy of the initial screening stage, reduces misjudgment, enhances the ability to identify high-risk cases, improves asset recovery rate, and reduces potential losses of financial institutions.
Smart Images

Figure CN120147016A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of case retrieval and analysis, and more specifically, to a method and system for retrieving and analyzing non-performing asset cases. Background Art
[0002] Non-performing asset cases refer to assets formed by financial institutions or enterprises during their operation processes that cannot recover the principal and interest on time, or assets with high default risks; through the retrieval and analysis of non-performing asset cases, these high-risk assets are identified, evaluated, and managed to reduce the potential losses of financial institutions or enterprises.
[0003] Then, enterprises conceal assets through complex equity structures or related-party transactions, and traditional rule engines are difficult to penetrate multiple layers of relationships, resulting in the inability of financial institutions to effectively recover the concealed assets when pursuing non-performing assets, leading to a decrease in the asset recovery rate and an increase in losses. Therefore, a method and system for retrieving and analyzing non-performing asset cases are provided. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and system for retrieving and analyzing non-performing asset cases to solve the problem in the above background art that traditional rule engines are difficult to penetrate multiple layers of relationships, resulting in the inability of financial institutions to effectively recover the concealed assets when pursuing non-performing assets, leading to a decrease in the asset recovery rate and an increase in losses.
[0005] To achieve the above purpose, the present invention provides a method for retrieving and analyzing non-performing asset cases, including the following steps:
[0006] S1. Clean, standardize, and extract features from the collected enterprise financial data, market information, and historical cases to generate structured data;
[0007] S2. Based on the dynamic weight logistic regression rule, perform a preliminary screening on the cleaned structured data to obtain a high-risk candidate case set, and extract features from the high-risk candidate case set to obtain the structured features of the high-risk candidate case set
[0008] S3. Construct a multi-layer knowledge graph based on SPO triples, and introduce time and space dimensions for the edge attributes of the multi-layer knowledge graph to dynamically update the edge attributes;
[0009] S4. Calculate the risk propagation weight s of the multi-layer knowledge graph with dynamically updated edge attributes based on the spatio-temporal graph attention network (ST-GAT) t , identify high-risk paths, output high-risk subgraphs, and locate hidden assets through a subgraph search algorithm to update the node attributes of the multi-layer knowledge graph;
[0010] S5. The GraphSAGE algorithm generates the graph embedding vectors of the enterprise association network based on the updated graph spectrum, and dynamically splices the graph embedding vectors with the structured features of the high-risk candidate case set to obtain the hybrid features The spliced hybrid features are used as the input of the graph-enhanced hybrid gradient boosting model to obtain the highly screened high-risk cases;
[0011] S6. The graph-enhanced hybrid gradient boosting model is trained using a progressive training strategy to balance the contributions of the structured features and the graph features to the prediction results.
[0012] As a further improvement of this technical solution, the dynamic weight logistic regression rule extracts the initial rule set R from the structured data D based on the pre-generated rule, and uses the DWLR algorithm to perform weight assignment and threshold adjustment on the initial rule set R, and combines the Monte Carlo rule search to iteratively optimize the rule combination to obtain the high-risk candidate case set;
[0013] The cleaned structured data where y i ∈{0,1}, and the initial rule set is generated by the DWLR algorithm as where the mathematical expression of the DWLR algorithm is:
[0014]
[0015] In the formula, represents the feature vector of the i-th sample, with dimension d; y i ∈{0,1} represents the label of the i-th sample, 0 represents low risk, and 1 represents high risk; m represents the total number of samples; r j represents the j-th rule; n represents the total number of rules; represents the initial rule space generated by the pre-generated rule; β represents the polarization degree of controlling the weight distribution; θ represents the comprehensive scoring threshold, the condition for determining high-risk cases; w j represents the dynamic weight of the rule r j ; S(x) represents the comprehensive scoring function. When S(x)>θ, it is judged as a high-risk candidate case, x represents the feature vector of the input sample; Responee Time is the response time; h represents the preset value of the response time; φ j (x) represents the indicator function, which judges whether x triggers the rule r j ; j represents the rule index variable; w j represents the dynamic weight of the j-th rule; n represents the total number of rules; F1-score(R) represents the overall performance evaluation index of the rule set R, which comprehensively combines the harmonic mean of the precision rate and the recall rate;
[0016] Among them, the weight calculation expression is:
[0017]
[0018] In the formula, P(r j ) represents the precision of rule r j ; P(r k ) represents the precision of rule r k ; k represents the rule index variable; exp(*) represents the natural exponential function.
[0019] Specifically, the specific steps for the pre-generation rule to extract the candidate rule set from the structured data D are as follows:
[0020] Specifically, perform equal-frequency binning (quantile cut) on continuous features (such as asset-liability ratio, current ratio) to discretize continuous values and convert them into interval conditions that can be processed by logical rules, generating interval-type rules:
[0021]
[0022] Among them, q p is the p-th quantile (p = 0.2, 0.4, 0.6, 0.8);
[0023] Generate enumeration-type rules for categorical features (such as collateral type, industry classification):
[0024]
[0025] Feature interaction rules:
[0026] Automatically generate second-order cross rules (example: product rule of debt ratio and number of lawsuits):
[0027]
[0028] Among them, γ interaction is determined by grid search (such as taking the 75th percentile of the distribution);
[0029] Use decision trees to assist in generating complex rules:
[0030]
[0031] Furthermore, for each rule r j , calculate its precision P(r j ):
[0032]
[0033] In the formula, TP(r j) The number of true high-risk cases correctly judged by the rule as high-risk; FP(r j ) The number of low-risk cases wrongly judged by the rule as high-risk;
[0034] Take the cleaned structured data as input data, where y i ∈ {0, 1}; Retain P(r j ) > γ precision rules to generate an initial rule set where γ precision is the threshold for judging whether feature interaction triggers a rule; Each rule is in the form of:
[0035]
[0036] where is an indicator function, which takes the value of 1 when the condition is satisfied, otherwise 0, f k is the k-th feature involved in the rule; OP is a logical operator, OP ∈ {>, <, =, ≠}, v k is the threshold or class value of feature f k (interval endpoints of binning, enumerated values of categorical features, etc.); K represents the number of conditions included in rule r j (K = 3 means the rule is composed of 3 conditions combined); X represents the feature vector of the input sample (i.e., x i ).
[0037] As a further improvement of this technical solution, the specific steps for the Monte Carlo rule search to iteratively optimize the rule combination to obtain a high-risk candidate case set are as follows:
[0038] S1.1. Set Monte Carlo parameters: number of iterations T, temperature decay coefficient α, initial temperature τ 0 , F1-score threshold γ F1 , response time upper limit h;
[0039] S1.2. Randomly select K rules from the initial rule set R to form an initial planning combination R current , calculate its comprehensive scoring function S(x), and evaluate the F1-score and response time of R current . If F1-score(R current ) ≥ γ F1 and Response Time < h, then save it as a candidate rule;
[0040] S1.3. Randomly perform one of the following operations on the current rule combination R current :
[0041] Select one from R\Rcurrent Randomly select a rule from it and add it;
[0042] Randomly remove a rule from R; current Randomly select a rule from R;
[0043] and replace it with another unused rule in the initial rule set R; current Randomly select a rule r;
[0044] and slightly adjust its threshold v; j ; k ;
[0045] S1.4. Recalculate the weight w according to the DWLR algorithm for the new rule combination R; new ; j ;
[0046] S1.5. Calculate the F1-score and response time for the new rule combination R; new ;
[0047] S1.6. If R satisfies F1-score(R)≥γ and Response Time < h, new then calculate the energy difference ΔE = F1-score(R)-F1-score(R); current ; F1 and accept the new rule combination R as the new current solution with probability;
[0048] where τ represents the current temperature parameter in the simulated annealing algorithm, controlling the sensitivity of the acceptance probability; current ; new ;
[0049] If the F1-score of the new rule combination R is higher and the response time is compliant, directly accept it; ; new as the new current solution;
[0050] In the formula, τ represents the current temperature parameter in the simulated annealing algorithm, controlling the sensitivity of the acceptance probability;
[0051] If the F1-score of the new rule combination R is higher and the response time is compliant, directly accept it; new ;
[0052] S1.7. Update the temperature τ = α·τ (simulated annealing mechanism), where α represents the temperature decay coefficient; 0 ;
[0053] Repeat steps S1.1 - S1.5 until the maximum number of iterations is reached;
[0054] S1.8. Traverse all candidate rule combinations, retain the combinations that satisfy F1-score > γ, and sort them in descending order of F1-score; F1 ;
[0055] S1.9. Select the Top-K combinations, merge their rules and remove duplicates to form the optimized high-risk candidate case set R final 。
[0056] As a further improvement of this technical solution, in S3, construct a multi-layer knowledge graph, introduce time and space dimensions for the edge attributes of the multi-layer knowledge graph, and the specific steps involved are as follows:
[0057] S2.1. Use the BiLSTM-CRF model to model the text sequence and define the entity recognition probability function:
[0058]
[0059] where h t is the hidden state output by BiLSTM, A is the CRF transition matrix, and y is the label sequence (such as BIOES annotation);
[0060] Vectorize the attribute values in the structured data:
[0061] v attr = MLP([v num ⊕ v cat )
[0062] v num is the normalized value of the numerical attribute, v cat is the embedding vector of the categorical attribute, and ⊕ represents the concatenation operation;
[0063] S2.2. Use the Span-based joint extraction model to define the entity-relationship joint probability for generating entity-relationship triples (SPO) as the edges of the knowledge graph:
[0064] P(s, o, r / C) = σ(W r · [e s ⊕ e o ⊕ e s,o )
[0065] where e s , e o are the entity span embeddings, e s,o is the context interaction feature, and W r is the relationship classification weight matrix;
[0066] S2.3. Based on the top-down ontology definition, construct a concept hierarchy tree:
[0067]
[0068] where c i is the entity type (such as enterprise, collateral), parent(c i) is the parent class concept;
[0069] Fuse the Jaccard coefficient and the embedded cosine similarity by weighting in the entity alignment similarity calculation:
[0070]
[0071] A e is the set of entity attributes, v e is the entity embedding vector, and α is the weight coefficient;
[0072] Adopt weighted voting for multi-source conflicting attribute values:
[0073]
[0074] In the formula, w i is the data source confidence weight;
[0075] S2.4. Calculate the node risk propagation probability based on the random walk algorithm:
[0076]
[0077] In the formula, d ij represents the shortest path length from the current node v i to the target node v j , and γ is the attenuation coefficient; represents the set of neighbor nodes of node v i ; exp(-γ·d ij ) represents the unnormalized propagation probability weight of node v j ; exp(-γ·d ik ) represents the unnormalized propagation probability weight of node v k ; represents that node k belongs to the neighbor set of v i ;
[0078] S2.5. In the entity-relationship triples generated in step S2.2, attach spatio-temporal attribute fields to each edge and adjust the node risk propagation probability:
[0079] If the edge in the knowledge graph is e = (s, r, o), then after attaching the time and space dimension attributes, it is expressed as:
[0080] e ST = <s, r, o; A time , A space >;
[0081] In the formula, s and o represent the start and end entities of the edge, is the relationship type; A timeRepresents the time dimension attribute; A space Represents the space dimension attribute; e ST Represents an edge containing spatio-temporal dynamic attributes;
[0082] The adjusted node risk propagation probability is:
[0083]
[0084] In the formula, w edge Represents the spatio-temporal weight, w edge =γ·Σ t τ(t)·f decay (t)+(1 - γ)·g(x, y)·PolicyRisk, where γ∈[0,1] represents the spatio-temporal weight balance coefficient, (x, y) represents the target position coordinates, and γ represents the space dimension weight coefficient.
[0085] As a further improvement of this technical solution, in S4, the risk propagation weight s of the multi-layer knowledge graph after dynamically updating the edge attributes is calculated through the spatio-temporal graph attention network t and the high-risk subgraph is output. The specific steps involved are:
[0086] S3.1. Perform spatio-temporal encoding on node features and edge attributes:
[0087]
[0088] Among them, Represents the result obtained after performing spatio-temporal encoding on node features and edge attributes; Is the structured feature of the high-risk candidate case set output by S2, Is the time decay function value, Is the spatial attribute interpolation result; W t Represents a learnable weight matrix for linearly transforming the input concatenated features; ReLU represents an activation function for introducing non-linearity, and its definition is ReLU(x)=max(0, x);
[0089] S3.2. Calculate the dynamic attention coefficient a ij :
[0090]
[0091] Among them, Represents the spatio-temporal weight of the edge between node i and node j; γ represents a hyperparameter; f decay (t ij ) represents the time decay function for measuring the influence of time on the edge weight; PolicyRisk(g ij) represents the policy risk function, which is used to evaluate the policy risk between the nodes connected by the edge; LeakyReLU(*) represents the activation function, where, a represents a small positive number (usually 0.01), which is used to avoid the problem of gradient disappearance, x represents the input value of the activation function; W represents the learnable weight matrix; represents the spatio-temporal weight of the edge between node i and node k; h i represents the feature representation obtained by node i after spatio-temporal encoding; h j represents the feature representation obtained by node j after spatio-temporal encoding; a T represents the learnable attention vector; t ij represents the time factor related to the edge between node i and node j;
[0092] S3.3. Calculate the risk propagation weight s ij according to the dynamic attention coefficient a t :
[0093]
[0094] In the formula, represents the set of neighbor nodes of node i; a ij represents the dynamic attention coefficient;
[0095] S3.4. Based on the risk propagation weight s t screen the high-risk edges and construct the candidate subgraph G risk :
[0096] G risk = {(i, j)|s t (i, j)>m}
[0097] In the formula, m is the preset threshold; both i and j represent the nodes in the graph; s t (i, j) represents the risk conduction intensity from node i to node j;
[0098] S3.5. For the path P=(v risk , v 1 ,..., v 1 ,... v n ) in the candidate subgraph G
[0099]
[0100] In the formula, Decay(Δt)=e -λΔt represents the time decay function, Δt represents the time interval of the edge, λ represents the control decay rate; v k represents the k-th node on the path P; s t (vk , v k+1 ) represents the risk conduction intensity from node k to node k + 1 in path P;
[0101] S3.6. Use the greedy algorithm to select the top K paths with the highest scores, and merge nodes and edges to generate a high-risk subgraph G hight-risk .
[0102] As a further improvement of this technical solution, the specific steps involved in locating hidden assets and updating the graph node attributes through the subgraph search algorithm are as follows:
[0103] S4.1. Based on the edges in the high-risk subgraph G hight-risk and their risk propagation weights s t , construct a weighted adjacency matrix A, where A ij = s t (i, j). Optimize the modularity Q through the Louvain algorithm, and the objective function is:
[0104]
[0105] where m is the total edge weight of the subgraph; k i is the weighted degree of node i; k j is the weighted degree of node j; δ determines the community membership; δ(c i , c j ) represents the exponential function; A ij represents the weight of the edge between nodes i and j in the adjacency matrix;
[0106] S4.2. Set a path score threshold θ', and filter high-risk paths;
[0107] For each node v in each community C i , perform a depth-first search to explore paths with a length ≤ L. Among them, the path score is:
[0108]
[0109] In the formula, Decay(Δt uv ) represents the time decay function, Decay(Δt) = e -λΔt , where Δt is the time interval of the edge, and λ represents the decay rate parameter; Score(P) represents the risk score of the path); (u, v) represents an edge in path P, connecting nodes u and v; s t (u, v) represents the risk propagation weight from node u to node v; Δt uv represents the time interval of edge (u, v);;
[0110] Retain the paths with Score(P) ≥ θ' and store them in the candidate path set Pcandidate ;
[0111] S4.3. Update the attribute vector of each hidden asset node v:
[0112]
[0113] where f(*) represents the mapping function; represents the structured feature vector of node v; ⊕ represents the vector concatenation operation; IsHiddenAsset represents whether node v is a hidden asset; HiddenValue represents the value attribute of the hidden asset;
[0114] If the hidden asset node v is marked as a hidden asset, update the attributes of the upstream nodes along its associated edges in the reverse direction:
[0115] where,
[0116] In the formula, represents the structured feature vector of the upstream node u; RiskExposure represents the risk exposure attribute of the node; += represents the accumulation operation; represents the set of neighbor nodes of node v; u represents an upstream node of node v; α represents the risk exposure coefficient; s t (u, v) represents the risk propagation weight from node u to node v;
[0117] S4.4. When a new asset node is added, only recalculate the modularity Q for the affected community C i to avoid global reconstruction.
[0118] As a further improvement of this technical solution, use GraphSAGE to generate the graph embedding vector of the enterprise association network:
[0119]
[0120] In the formula, represents the embedding of node v at the k-th layer; W (k) represents the trainable weight matrix at the k-th layer; the final output
[0121] Feature dynamic concatenation:
[0122] Perform feature dynamic concatenation on the graph embedding vector and the structured features of the high-risk candidate case set to obtain the hybrid feature z i :
[0123]
[0124] The concatenated hybrid feature z iInput the tree model and the graph adapter simultaneously;
[0125] The concatenated hybrid features serve as the input carrier for the dual-path model in GE-HGBM, separating and processing the two types of features; among them, the tree model only receives the structured feature part, and the GNN only receives the graph embedding part;
[0126] Concatenation operation --- Force alignment during the data preprocessing stage to prevent feature misalignment caused by index errors (such as the structured features of sample A being mismatched with the graph embedding of sample B);
[0127] The Graph-Enhanced Hybrid Gradient Boosting Model (GE-HGBM model) is an improved model based on the fusion of graph features, and its mathematical expression is:
[0128]
[0129] In the formula, represents the predicted value of the i-th sample; represents the structured feature vector of the i-th sample; represents the final graph embedding vector of node v; λ represents the graph feature fusion coefficient; f k represents the gradient boosting tree base model; g represents the graph feature adapter; K represents the number of base models of the gradient boosting tree, and k represents the index variable; f k (*) represents the prediction of the gradient boosting tree base model on the input feature *; g(*) represents the transformation of the graph feature adapter on the input graph embedding vector *;
[0130] Integrate the spatio-temporal dynamic features and the time decay mechanism into the Graph-Enhanced Hybrid Gradient Boosting Model, then the optimized Graph-Enhanced Hybrid Gradient Boosting Model is:
[0131]
[0132] In the formula, s t is the spatio-temporal risk weight output by ST-GAT; w t is the time decay coefficient; T represents the total number of time steps; t represents the index variable of the time step;
[0133] Composite loss function:
[0134]
[0135] In the formula, represents the composite loss function; Ω(f k ) represents the complexity regularization term (number of leaf nodes / weights) of the k-th tree; W (k)denotes the weight matrix of the k-th layer of GraphSAGE; α denotes the regularization strength hyperparameter of the gradient boosting tree; β denotes the regularization strength hyperparameter of the graph feature encoding; γ denotes the regularization strength hyperparameter of the spatio-temporal dynamic features; denotes the cross-entropy loss term; denotes the gradient boosting tree regularization term; denotes the graph feature encoding regularization term; denotes the spatio-temporal dynamic regularization term; denotes the predicted value of the i-th sample; y i denotes the true label of the i-th sample; n denotes the total number of samples, and i denotes the index variable; U (t) denotes the weight matrix of the t-th layer of ST-GAT.
[0136] As a further improvement of this technical solution, the progressive training strategy trains the graph-enhanced hybrid gradient boosting model, and the progressive training strategy is as follows:
[0137] Phase 1:
[0138] Only use to train the model and fix these parameters {f k};
[0139] Phase 2:
[0140] Input to train the parameters of the model g(*) and GraphSAGE
[0141] Phase 3:
[0142] Unfreeze all parameters and g(*), and perform end-to-end optimization to fine-tune the overall model.
[0143] On the other hand, the present invention provides a non-performing asset case retrieval and analysis system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program to implement the steps of the non-performing asset case retrieval and analysis method described in any one of the above.
[0144] Compared with the prior art, the beneficial effects of the present invention are:
[0145] 1. In this non-performing asset case retrieval and analysis method and system, through dynamic weight logistic regression (DWLR) combined with Monte Carlo rule search, the rule weights and combinations are dynamically optimized to improve the accuracy in the preliminary screening stage, reduce misjudgments (such as low-risk cases being mislabeled as high-risk), and at the same time, the F1-score constraint is used to balance precision and recall to ensure the comprehensiveness and reliability of the high-risk candidate set.
[0146] 2. In the bad asset case retrieval and analysis method and system, a multi-layer knowledge graph introducing spatio-temporal attributes is constructed, and the spatio-temporal graph attention network (ST-GAT) is combined to dynamically calculate the risk propagation weight, and the risk conduction probability between nodes is quantified by the random walk algorithm, which is used to capture hidden assets in complex association relationships such as guarantee chains and equity substitution, and reflect the risk fluctuations over time and region in real time (such as seasonal changes in judicial enforcement efficiency and regional economic policy risk conduction), so as to improve the dynamics and accuracy of risk path identification.
[0147] 3. In the bad asset case retrieval and analysis method and system, the structured features (such as financial data) and graph embedding features (enterprise association network vectors generated by GraphSAGE) are dynamically spliced, and the dual-channel feature fusion is realized through the graph-enhanced hybrid gradient boosting model (GE-HGBM), taking into account the statistical laws of numerical features and the topological characteristics of the graph structure, and solving the problem of missed detection caused by traditional models ignoring association relationships (such as hidden risks of isolated nodes), so as to improve the robustness of high-risk case prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0148] Figure 1 It is the overall method flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0149] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0150] Embodiment 1:
[0151] Please refer to Figure 1 As shown, this embodiment provides a bad asset case retrieval and analysis method, including the following steps:
[0152] S1. Clean, standardize and extract features from the collected enterprise financial data, market information and historical cases to generate structured data;
[0153] The structured data mainly comes from original data such as enterprise financial data, market information, and judicial litigation records; these data are usually in tabular form (such as CSV, database tables), and contain clear fields (such as enterprise name, debt amount, collateral valuation, etc.);
[0154] Feature extraction: Extract key features from the original data (such as the credit rating of the debtor, the market value of the collateral, industry risks, etc.);
[0155] The key features include basic features, risk features, and association features;
[0156] Among them, the basic features include asset - liability ratio, current ratio, and cash - flow coverage ratio;
[0157] The risk features include industry risk index (based on policy sensitivity / cycle volatility), collateral value decay rate, and debtor's historical default record;
[0158] The association features include supply - chain dependence degree, number of equity penetration levels, and hidden association paths in the knowledge graph (such as mining through the guarantee chain and implicit share - holding relationships);
[0159] S2. Perform preliminary screening on the cleaned structured data based on the dynamic - weight logistic regression rule to obtain a set of high - risk candidate cases. Feature extraction is performed on the set of high - risk candidate cases (the set of high - risk candidate cases preferentially uses filtering methods (chi - square test / MI algorithm) for fast noise reduction, and then combines embedded methods (XGBoost feature importance) for fine - tuning and selection. Then, feature transformation is carried out (numerical features eliminate the dimension difference through the Z - score standardization algorithm, while categorical features select the one - hot / target encoding algorithm according to the cardinality to convert discrete features into a numerical form that can be understood by the model), and the structured features of the set of high - risk candidate cases are obtained
[0160] In this embodiment, the dynamic - weight logistic regression rule extracts an initial rule set R from the structured data D based on the pre - generated rule, and uses the DWLR algorithm to perform weight assignment and threshold adjustment on the initial rule set R. The rule combination is iteratively optimized by combining the Monte Carlo rule search, setting the threshold of F1 - score, and eliminating the rules with F1 - score ≤ threshold to obtain a set of high - risk candidate cases;
[0161] The cleaned structured data where y i ∈{0,1}, and the initial rule set generated by the DWLR algorithm is Among them, the mathematical expression of the DWLR algorithm is:
[0162]
[0163] In the formula, represents the feature vector of the i - th sample, with dimension d; y i ∈{0,1} represents the label of the i - th sample, 0 represents low - risk, and 1 represents high - risk; m represents the total number of samples; r j represents the j - th rule; n represents the total number of rules; denotes the initial rule space generated by the pre-generated rules; β represents the polarization degree of controlling the weight distribution; θ represents the comprehensive scoring threshold, the condition for determining high-risk cases; w j denotes the dynamic weight of rule r j ; S(x) represents the comprehensive scoring function. When S(x)>θ, it is judged as a high-risk candidate case, where x represents the feature vector of the input sample; Response Time is the response time; h represents the preset value of the response time; φ j (x) represents the indicator function to judge whether x triggers rule r j ; j represents the rule index variable; w j denotes the dynamic weight of the j-th rule; n represents the total number of rules; F1-score(R) represents the overall performance evaluation index of the rule set R, which synthesizes the harmonic mean of precision and recall;
[0164] Among them, the weight calculation expression is:
[0165]
[0166] Specifically, when β↑ (for example, β = 5), the weight distribution tends to be polarized, and high-precision rules dominate; when β↓ (for example, β = 1), the weight distribution tends to be uniform;
[0167] In the formula, P(r j ) represents the precision of rule r j ; P(r k ) represents the precision of rule r k ; k represents the rule index variable; exp(*) represents the natural exponential function.
[0168] Specifically, the specific steps for the pre-generated rules to extract the candidate rule set from the structured data D are as follows:
[0169] Specifically, for continuous features (such as asset-liability ratio, current ratio), equal-frequency binning (quantile cut) is performed to discretize the continuous values and convert them into interval conditions that can be processed by logical rules, generating interval-type rules:
[0170]
[0171] Among them, q p is the p-th quantile (p = 0.2, 0.4, 0.6, 0.8);
[0172] Generate enumeration-type rules for categorical features (such as collateral type, industry classification):
[0173]
[0174] Feature interaction rules:
[0175] Automatically generate second-order cross rules (e.g., the product rule of debt ratio and the number of litigation cases):
[0176]
[0177] where γ interaction is determined by grid search (e.g., taking the 75th percentile of the distribution);
[0178] Use decision trees to assist in generating complex rules:
[0179]
[0180] Furthermore, for each rule r j , calculate its precision P(r j ):
[0181]
[0182] In the formula, TP(r j ) represents the number of true high-risk cases correctly judged as high-risk by the rule; FP(r j ) represents the number of low-risk cases wrongly judged as high-risk by the rule;
[0183] Take the cleaned structured data as the input data, where
[0184] where γ precision is the threshold for judging whether feature interaction triggers the rule; the form of each rule is:
[0185]
[0186] where is the indicator function, taking the value of 1 when the condition is satisfied and 0 otherwise, f k is the k-th feature involved in the rule; OP is the logical operator, OP ∈ {>, <, =, ≠}, v k is the threshold or category value of feature f k (the interval endpoint of binning, the enumerated value of categorical features, etc.); K represents the number of conditions included in rule r j (K = 3 means the rule is composed of 3 conditions combined); X represents the feature vector of the input sample (i.e., x i );
[0187] Furthermore, the specific steps for iterative optimization of the rule combination by Monte Carlo rule search to obtain the set of candidate high-risk cases are:
[0188] S1.1. Set Monte Carlo parameters: number of iterations T, temperature decay coefficient α, initial temperature τ 0 , F1-score threshold γ F1 , upper limit of response time h;
[0189] S1.2. Randomly select K rules from the initial rule set R to form the initial planning combination R current , calculate its comprehensive scoring function S(x), and evaluate R current 's F1-score and response time. If F1-score(R current ) ≥ γ F1 and Response Time < h, then save it as a candidate rule;
[0190] S1.3. Randomly perform one of the following operations on the current rule combination R current :
[0191] Randomly select a rule from R\R current and add it;
[0192] Randomly remove a rule from R current ;
[0193] Randomly select a rule in R current and replace it with another unused rule in the initial rule set R;
[0194] Randomly select a rule r j , and fine-tune its threshold v k (such as adjusting the bin interval endpoint q p to an adjacent quantile);
[0195] S1.4. According to the new rule combination R new , recalculate the weight w j according to the DWLR algorithm;
[0196] S1.5. Calculate the F1-score and response time of the new rule combination R new ;
[0197] S1.6. If R new satisfies F1-score(R current ) ≥ γ F1 and Response Time < h, then calculate the energy difference ΔE = F1-score(R current ) - F1-score(R new );
[0198] Accept the new rule combination R with probability new as the new current solution;
[0199] In the formula, τ represents the current temperature parameter in the simulated annealing algorithm, which controls the sensitivity of the acceptance probability;
[0200] If the new rule combination R new has a higher F1-score and the response time is compliant, directly accept it;
[0201] S1.7. Update the temperature τ = α·τ 0 (simulated annealing mechanism), where α represents the temperature decay coefficient;
[0202] Repeat steps S1.1 - S1.5 until the maximum number of iterations is reached;
[0203] S1.8. For all candidate rule combinations traversed, retain the combinations that satisfy F1-score > γ F1 and sort them in descending order of F1-score;
[0204] S1.9. Select the Top-K combinations (among all candidate rule combinations, sort them from high to low according to F1-score, and select the top K rule combinations), merge their rules and remove duplicates to form the optimized high-risk candidate case set R final .
[0205] S3. Construct a multi-layer knowledge graph based on SPO triples, and introduce time and space dimensions for the edge attributes of the multi-layer knowledge graph to dynamically update the edge attributes;
[0206] Among them, for the time dimension: the trading frequency of collateral (reflecting the change of asset liquidity over time), the seasonal fluctuation of judicial execution efficiency;
[0207] For the space dimension: the regional economic risk index (the difference in judicial execution efficiency in different regions), the geographical distribution of industry policy sensitivity;
[0208] In this embodiment, the specific steps involved in constructing a multi-layer knowledge graph and introducing time and space dimensions for the edge attributes of the multi-layer knowledge graph are as follows:
[0209] S2.1. Entity boundary detection: Use the BiLSTM-CRF model to model the text sequence, and define the entity recognition probability function:
[0210]
[0211] Among them, h t is the hidden state output by BiLSTM, A is the CRF transition matrix, and y is the label sequence (such as BIOES annotation);
[0212] Attribute embedding:
[0213] Vectorize the attribute values in structured data (such as registered capital, industry classification):
[0214] v attr = MLP([v num ⊕ v cat )
[0215] v num is the standardized value of numerical attributes, v cat is the embedding vector of categorical attributes, and ⊕ represents the concatenation operation;
[0216] Entity boundary detection identifies entities in the text, and attribute embedding vectorizes the attributes of entities, providing clear entities and computable attribute representations for relation extraction and graph construction; specifically, entity boundary detection provides entity candidates for subsequent relation extraction (S2.2), and attribute embedding provides a computable similarity basis for entity alignment (S2.3);
[0217] S2.2. Adopt a Span-based joint extraction model to define the entity-relation joint probability for generating entity-relation triples (SPO) as the edges of the knowledge graph:
[0218] P(s, o, r / C) = σ(W r · [e s ⊕ e o ⊕ e s,o )
[0219] where e s , e o are entity span embeddings, e s,o is the context interaction feature, and W r is the relation classification weight matrix;
[0220] S2.3. Based on the top-down ontology definition, construct a concept hierarchy tree:
[0221]
[0222] where c i is the entity type (such as enterprise, collateral), and parent(c i ) is the parent concept;
[0223] Specifically, the hierarchical structure definition of the knowledge graph:
[0224] Concept layer
[0225] The concept layer tree is formally represented as a directed acyclic graph (DAG):
[0226]
[0227] ε sub ={ ( c i , subClassOf, c j ) | c j ∈parent ( c j )} ;
[0228] In the formula, represents the concept set (such as "enterprise", "collateral"); ε sub represents the inheritance relationship edge;
[0229] Entity layer
[0230] Graph structure of entity and relationship:
[0231]
[0232] represents the set of entity nodes; ε represents the set of relationship edges; A time represents the time dimension attribute; A space represents the space dimension attribute;
[0233] Fuse the Jaccard coefficient (used for explicit discrete attribute matching (such as industry classification, registered capital binning interval) to ensure strict alignment of key business attributes (such as enterprise names in judicial litigation records)) and the embedding cosine similarity (used for numerical attribute continuity analysis (such as the actual value of registered capital) and cross-modal semantic association (such as the semantic similarity between "manufacturing industry" and "equipment manufacturing industry")) in entity alignment similarity calculation, and avoid missing high-risk entities due to name differences (such as "XX Group" vs "XX Holdings"):
[0234]
[0235] A e is the set of entity attributes, v e is the entity embedding vector, and α is the weight coefficient (usually taken as 0.6);
[0236] Conflict resolution decision function:
[0237] Adopt weighted voting for multi-source conflict attribute values (such as collateral valuation):
[0238]
[0239] In the formula, w iis the confidence weight of the data source (for example, the weight of the judicial evaluation report W = 0.8, and the enterprise disclosure w = 0.5);
[0240] S2.4. Calculate the node risk propagation probability based on the random walk algorithm:
[0241]
[0242] In the formula, d ij represents the shortest path length from the current node v i to the target node v j , and γ is the attenuation coefficient; represents the set of neighbor nodes of node v i ; exp(-γ·d ij ) represents the unnormalized propagation probability weight of node v j ; exp(-γ·d ik ) represents the unnormalized propagation probability weight of node v k ; represents that node k belongs to the neighbor set of v i ;
[0243] S2.5. In the entity-relationship triple (SPO) generated in step S2.2, attach spatio-temporal attribute fields to each edge and adjust the node risk propagation probability:
[0244] If the edge in the knowledge graph is e = (s, r, o), then after attaching the time and space dimension attributes, it is represented as:
[0245] e ST =<s, r, o; A time , A space >;
[0246] In the formula, s and o represent the start and end entities of the edge, is the relationship type; A time represents the time dimension attribute; A space represents the space dimension attribute; e ST represents an edge containing spatio-temporal dynamic attributes;
[0247] A time ={τ(t), f decay (t)};
[0248]
[0249] In the formula, f decay (t) represents the timeliness attenuation function (such as the seasonal attenuation of judicial execution efficiency); τ(t) represents the time stamp; β represents the initial efficiency value; λ represents the attenuation rate; t 0Indicates the reference time; t represents the current time;
[0250] Spatial dimension attribute A space Used to define the spatial attribute as a geocoding or regional label:
[0251] A space = {RegionID, g(latitude and longitude), PolicyRisk};
[0252] In the formula, RegionID represents the administrative region code; g(*) represents the spatial interpolation function; PolicyRisk represents the policy sensitivity,
[0253] Endow the edge attributes of the knowledge graph with dynamic characteristics of time and space dimensions, used to capture the timeliness and geographical differences of the relationships between entities, and support dynamic risk analysis (such as seasonal fluctuations in judicial enforcement efficiency, regional economic risk transmission);
[0254] The adjusted node risk propagation probability is:
[0255]
[0256] In the formula, w edge represents the spatio-temporal weight, w edge = γ·∑ t τ(t)·f decay (t)+(1 - γ)·g(x, y)·PolicyRisk, where γ∈[0,1] represents the spatio-temporal weight balance coefficient, (x, y) represents the target position coordinates, and γ represents the spatial dimension weight coefficient.
[0257] S4. Calculate the risk propagation weight s t (the intensity of risk conduction between nodes) of the multi-layer knowledge graph after dynamically updating the edge attributes based on the spatio-temporal graph attention network (ST-GAT), identify high-risk paths, output high-risk subgraphs, and locate hidden assets (such as hidden assets associated through guarantee chains and equity substitution) through subgraph search algorithms (such as community detection, path tracing), and update the node attributes of the multi-layer knowledge graph (mark hidden assets);
[0258] In this embodiment, calculating the risk propagation weight s t and outputting high-risk subgraphs of the multi-layer knowledge graph after dynamically updating the edge attributes through the spatio-temporal graph attention network involve the following specific steps:
[0259] S3.1. Perform spatio-temporal encoding on node features and edge attributes:
[0260]
[0261] Among them, It represents the result obtained after spatio-temporal encoding of node features and edge attributes; It is the structured feature of the set of high-risk candidate cases output by S2, It is the value of the time decay function, It is the result of spatial attribute interpolation; W t It represents a learnable weight matrix used to perform a linear transformation on the input concatenated features; ReLU represents an activation function used to introduce non-linearity, and its definition is ReLU(x) = max(0, x);
[0262] S3.2. Calculate the dynamic attention coefficient a ij :
[0263]
[0264] Among them, It represents the spatio-temporal weight of the edge between node i and node j; γ represents a hyperparameter; f decay (t ij ) represents the time decay function used to measure the impact of time on the edge weight; PolicyRisk(g ij ) represents the policy risk function used to evaluate the policy risk between the nodes connected by the edge; LeakyReLU(*) represents the activation function, where, a represents a small positive number (usually 0.01) used to avoid the problem of gradient disappearance, χ represents the input value of this activation function; W represents a learnable weight matrix; It represents the spatio-temporal weight of the edge between node i and node k; h i It represents the feature representation obtained after spatio-temporal encoding of node i; h j It represents the feature representation obtained after spatio-temporal encoding of node j; a T It represents a learnable attention vector; t ij It represents the time factor related to the edge between node i and node j;
[0265] S3.3. Calculate the risk propagation weight s according to the dynamic attention coefficient a ij Calculate the risk propagation weight s t :
[0266]
[0267] In the formula, the risk propagation weight s t represents the risk conduction intensity from node i to node j; represents the set of neighbor nodes of node i; a ij represents the dynamic attention coefficient;
[0268] S3.4. Based on the risk propagation weight s t Filter high-risk edges and construct a candidate subgraph G risk :
[0269] G risk ={(i, j)|s t (i, j)>m}
[0270] where m is a preset threshold (such as the Top 10% quantile of the risk propagation weight), used to eliminate low-risk edges; G risk represents the candidate subgraph, which is a graph structure obtained by filtering high-risk edges; both i and j represent nodes in the graph, here used to refer to node pairs; s t (i, j) represents the risk conduction intensity from node i to node j, that is, the risk propagation weight;
[0271] S3.5. For the path P=(v risk , v 1 ,..., v 1 ) in the candidate subgraph G n , calculate the path risk score Score(P):
[0272]
[0273] where Decay(Δt)=e -λΔt represents the time decay function, Δt represents the time interval of the edge, λ represents the control decay rate; v k represents the k-th node on the path P; s t (v k , v k+1 ) represents the risk conduction intensity from node k to node k+1 on the path P, that is, the risk propagation weight;
[0274] S3.6. Adopt the greedy algorithm to select the top K paths with the highest scores, and merge nodes and edges to generate a high-risk subgraph G hight-risk .
[0275] Furthermore, the specific steps involved in locating hidden assets and updating graph node attributes through the subgraph search algorithm are as follows:
[0276] S4.1. Based on the edges in the high-risk subgraph G hight-risk and their risk propagation weight s t , construct a weighted adjacency matrix A, where A ij =s t (i,j), optimize the modularity Q through the Louvain algorithm, and the objective function is:
[0277]
[0278] Among them, m is the total edge weight of the sub - graph; k i is the weighted degree of node i; k j is the weighted degree of node j; δ is used to judge community membership; δ(c i , c j ) represents the exponential function; A ij represents the weight of the edge between node i and node j in the adjacency matrix;
[0279] S4.2. Set the path score threshold θ', and filter out high - risk paths;
[0280] For each node ν in each community C i , perform a depth - first search (DFS) to explore paths with length ≤ L. Among them, the path score is:
[0281]
[0282] In the formula, Decay(Δt uv ) represents the time decay function, Decay(Δt)=e -λΔt , where Δt is the time interval of the edge, and λ represents the decay rate parameter; Score(P) represents the risk score of path P; (u, v) represents an edge in path P, connecting node u and node v; s t (u, v) represents the risk propagation weight from node u to node v; Δt uv represents the time interval of edge (u, v);;
[0283] Retain the paths with Score(R)≥θ' and store them in the candidate path set P candidate ;
[0284] S4.3. For each hidden asset node v, update its attribute vector:
[0285]
[0286] In the formula, f(*) represents the mapping function (using the linear normalization method to map Score(P) to a specific range, such as [0, 1]); represents the structural feature vector of node v; ⊕ represents the vector concatenation operation; IsHiddenAsset represents whether node v is a hidden asset; HiddenValue represents the value attribute of the hidden asset;
[0287] If the hidden asset node v is marked as a hidden asset, then update the attributes of the upstream nodes along its associated edges in the reverse direction:
[0288] Among them,
[0289] In the formula, represents the structured feature vector of the upstream node u; RiskExposure represents the risk exposure attribute of the node; += represents the accumulation operation; represents the set of neighbor nodes of node v; u represents an upstream node of node v; α represents the risk exposure coefficient; s t (u, v) represents the risk propagation weight from node u to node v;
[0290] S4.4. When a new asset node is added, only the affected community C i recalculates the modularity Q to avoid global reconstruction.
[0291] In this embodiment, a multi-layer knowledge graph incorporating spatio-temporal attributes is constructed. Combining the spatio-temporal graph attention network (ST-GAT), the risk propagation weight is dynamically calculated, and the risk conduction probability between nodes is quantified through the random walk algorithm, which is used to capture hidden assets in complex association relationships such as guarantee chains and equity trusteeships, and to reflect the fluctuations of risks over time and regions in real time (such as seasonal changes in judicial enforcement efficiency and regional economic policy risk conduction), improving the dynamics and accuracy of risk path identification.
[0292] S5. The GraphSAGE algorithm generates the graph embedding vector of the enterprise association network based on the updated graph (the input of GraphSAGE is the structure of the enterprise association network (the knowledge graph constructed by SPO triples), and the initial attributes of its nodes and edges come from structured data (such as enterprise financial data, judicial records)), and dynamically splices the graph embedding vector with the structured features of the high-risk candidate case set to obtain the mixed feature z through feature dynamic splicing i , and the spliced mixed feature z i is used as the input of the graph-enhanced hybrid gradient boosting model to obtain the highly screened high-risk cases (judging the risk level of the cases);
[0293] In this embodiment, the use of GraphSAGE to generate the graph embedding vector of the enterprise association network:
[0294]
[0295] In the formula, represents the embedding of node v at the k-th layer; W (k) represents the trainable weight matrix at the k-th layer; the final output
[0296] Feature dynamic splicing:
[0297] Dynamically splice the graph embedding vector with the structured features of the high-risk candidate case set to obtain the mixed feature z i :
[0298]
[0299] The concatenated hybrid feature z i Input the tree model and the graph adapter simultaneously;
[0300] The concatenated hybrid feature serves as the input carrier of the dual-path model in GE-HGBM, separating and processing the two types of features; among them, the tree model only receives the structured feature part, and the GNN only receives the graph embedding part;
[0301] Concatenation operation --- Force alignment in the data preprocessing stage to prevent feature misalignment caused by index errors (such as the structured features of sample A being misaligned with the graph embedding of sample B);
[0302] The Graph-Enhanced Hybrid Gradient Boosting Model (GE-HGBM model) is an improved model based on the fusion of graph features, and its mathematical expression is:
[0303]
[0304] In the formula, represents the predicted value of the i-th sample; represents the structured feature vector of the i-th sample; represents the final graph embedding vector of node v; λ represents the graph feature fusion coefficient (optimization range 0.3 - 0.7); f k represents the gradient boosting tree base model; g represents the graph feature adapter (fully connected network); K represents the number of base models of the gradient boosting tree, k represents the index variable; f k (*) represents the prediction of the gradient boosting tree base model for the input feature *; g(*) represents the transformation of the graph feature adapter for the input graph embedding vector *;
[0305] Integrate spatio-temporal dynamic features and time decay mechanism into the Graph-Enhanced Hybrid Gradient Boosting Model, then the optimized Graph-Enhanced Hybrid Gradient Boosting Model is:
[0306]
[0307] In the formula, s t is the spatio-temporal risk weight output by ST-GAT; w t is the time decay coefficient; T represents the total number of time steps; t represents the index variable of the time step;
[0308] Composite loss function:
[0309]
[0310] In the formula, represents the composite loss function; Ω(f k) represents the complexity regularization term (number of leaf nodes / weights) of the k-th tree; W (k) represents the weight matrix of the k-th layer of GraphSAGE; α represents the regularization strength hyperparameter of the gradient boosting tree; β represents the regularization strength hyperparameter of the graph feature encoding; γ represents the regularization strength hyperparameter of the spatio-temporal dynamic feature; represents the cross-entropy loss term; represents the gradient boosting tree regularization term; represents the graph feature encoding regularization term; represents the spatio-temporal dynamic regularization term; represents the predicted value of the i-th sample; y i represents the true label of the i-th sample; n represents the total number of samples, and i represents the index variable; U (t) represents the weight matrix of the t-th layer of ST-GAT, which is used to generate spatio-temporal risk weights.
[0311] Dynamically splice structured features (such as financial data) with graph embedding features (enterprise association network vectors generated by GraphSAGE), and achieve dual-channel feature fusion through the graph-enhanced hybrid gradient boosting model (GE-HGBM), taking into account the statistical laws of numerical features and the topological characteristics of the graph structure, solving the missed detection problems caused by traditional models ignoring association relationships (such as hidden risks of isolated nodes), and improving the robustness of high-risk case prediction.
[0312] S6. Train the graph-enhanced hybrid gradient boosting model using a progressive training strategy to balance the contributions of structured features and graph features to the prediction results;
[0313] Specifically, when training the graph-enhanced hybrid gradient boosting model with a progressive training strategy, the progressive training strategy is as follows:
[0314] Phase 1:
[0315] Only use to train the model and fix these parameters {f k};
[0316] Phase 2:
[0317] Input to train the parameters of the model g(*) and GraphSAGE
[0318] Phase 3:
[0319] Unfreeze all parameters and g(*), and perform end-to-end optimization to fine-tune the overall model.
[0320] Example 2:
[0321] This embodiment provides a non-performing asset case retrieval and analysis system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program to implement the steps of the non-performing asset case retrieval and analysis method described in any one of the above.
[0322] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and all these changes and improvements fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A method for retrieving and analyzing non-performing asset cases, characterized in that: The following steps are involved: S1. Clean, standardize and extract features of the collected corporate financial data, market information and historical cases to generate structured data; S2. Perform a preliminary screening of the cleaned structured data based on the dynamic weight logistic regression rule to obtain a high-risk candidate case set, perform feature extraction on the high-risk candidate case set, and obtain the structured features of the high-risk candidate case set S3, build a multi-layer knowledge graph based on SPO triples, introduce time and space dimensions to the edge attributes of the multi-layer knowledge graph and dynamically update the edge attributes; S4. Risk propagation weights s of multi-layer knowledge graph after dynamically updating edge attributes based on spatiotemporal graph attention network calculation t , output high-risk subgraphs, locate hidden assets through subgraph search algorithms, and update node attributes of multi-layer knowledge graphs; S5. The GraphSAGE algorithm generates a graph embedding vector of the enterprise association network based on the updated graph, and integrates the graph embedding vector with the structural features of the high-risk candidate case set. Dynamically concatenate features to obtain mixed features The combined features As the input of the graph-enhanced hybrid gradient boosting model, we get the high-risk cases after fine screening; S6. Use a progressive training strategy to train the graph-enhanced hybrid gradient boosting model to balance the contribution of structured features and graph features to the prediction results.
2. The method for retrieving and analyzing non-performing asset cases according to claim 1, characterized in that: The dynamic weighted logistic regression rule extracts the initial rule set R from the structured data D based on the pre-generation rule, and uses the DWLR algorithm to perform weight assignment and threshold adjustment on the initial rule set R, and combines the Monte Carlo rule search to iteratively optimize the rule combination to obtain a high-risk candidate case set; The cleaned structured data in y i ∈{0,1}, the initial rule set generated by the DWLR algorithm is The mathematical expression of the DWLR algorithm is: Response Time <h In the formula, Represents the feature vector of the i-th sample, with dimension d; y i ∈{0,1} represents the label of the i-th sample; m represents the total number of samples; r j represents the jth rule; n represents the total number of rules; represents the initial rule space generated by the pre-generation rules; β represents the polarization degree of the control weight distribution; θ represents the comprehensive score threshold; w j Representation rule r j dynamic weight; S(x) represents the comprehensive scoring function. When S(x)>θ, it is judged as a high-risk candidate case. x represents the feature vector of the input sample; Response Time is the response time; g represents the preset value of the response time; φ j (x) represents the indicative function, which determines whether x triggers rule r j ; j represents the rule index variable; w j represents the dynamic weight of the jth rule; n represents the total number of rules; F1-score(R) represents the overall performance evaluation index of the rule set R; The weight calculation expression is: In the formula, P(r j ) represents the rule r j The accuracy rate; P(r k ) represents the rule r k The precision of ; k represents the rule index variable; exp(x) represents the natural exponential function.
3. The method for retrieving and analyzing non-performing asset cases according to claim 2 is characterized by: The specific steps involved in the Monte Carlo rule search iterative optimization rule combination to obtain a high-risk candidate case set are: S1.
1. Set Monte Carlo parameters: number of iterations T, temperature attenuation coefficient α, initial temperature τ0, F1-score threshold γ F1 , upper limit of response time h; S1.
2. Randomly select K rules from the initial rule set R to form an initial planning combination R current , calculate its comprehensive scoring function S(x), and evaluate R current 's F1-score and response time. If F1-score(R current ) ≥ γ F1 and ResponseTime < h, then save it as a candidate rule; S1.3, for the current rule combination R current Randomly do one of the following: From R\R current Randomly select a rule to add; From R current Randomly remove a rule from Randomly select R current A rule in is replaced by another unused rule in the initial rule set R; Randomly select a rule r j , fine-tune its threshold v k ; S1.4, according to the new rules combination R new , recalculate the weight w according to the DWLR algorithm j ; S1.
5. Calculate the new rule combination R new F1-score and response time; S1.
6. If R new satisfies F1 - score(R current ) ≥ γ F1 and Response Time < h, then calculate the energy difference ΔE = F1 - score(R current ) - F1 - score(R new ); By probability Accept new rule combination R new As the new current solution; Where, τ represents the current temperature parameter in the simulated annealing algorithm; If the new rule combination R new The F1-score is higher and the response time is in compliance, so it is accepted directly; S1.7, update temperature τ = α·τ0, where α represents the temperature attenuation coefficient; Repeat steps S1.1-S1.5 until the maximum number of iterations is reached; S1.
8. Traverse all candidate rule combinations and keep those that satisfy F1-score>γ F1 The combinations are sorted in descending order by F1-score; S1.
9. Select the Top-K combination, merge its rules and remove duplicates to form an optimized high-risk candidate case set R final .
4. The method for retrieving and analyzing non-performing asset cases according to claim 1, characterized in that: In S3, a multi-layer knowledge graph is constructed to introduce time and space dimensions into the edge attributes of the multi-layer knowledge graph. The specific steps involved are: S2.
1. Use the BiLSTM-CRF model to model the text sequence and define the entity recognition probability function: Among them, h t is the hidden state of BiLSTM output, A is the CRF transfer matrix, and y is the label sequence (such as BIOES annotation); Vectorize attribute values in structured data: v num Normalize the value of the numeric attribute, v cat is the category attribute embedding vector, Represents a splicing operation; S2.2, adopt the Span-based joint extraction model to define the entity-relationship joint probability, which is used to generate entity-relationship triples as the edges of the knowledge graph: Among them, e s , e o is the entity span embedding, e s,o is the context interaction feature, W r is the relationship classification weight matrix; S2.
3. Construct a concept hierarchy tree based on top-down ontology definition: Among them, c i is the entity type (such as enterprise, collateral), parent(c i ) is the parent concept; The Jaccard coefficient and the embedding cosine similarity are weightedly fused in the entity alignment similarity calculation: A e is the entity attribute set, v e is the entity embedding vector, α is the weight coefficient; Use weighted voting for conflicting attribute values from multiple sources: In the formula, w i is the confidence weight of the data source; S2.
4. Calculate the node risk propagation probability based on the random walk algorithm: Where, d ij Represents the current node v i To the target node v j The shortest path length, γ is the attenuation coefficient; Represents node v i The set of neighbor nodes; exp(-γ·d ij ) represents node v j The unnormalized propagation probability weight of exp(-γ·d ik ) represents node v k The unnormalized propagation probability weight of ; Indicates that node k belongs to v i The set of neighbors of ; S2.
5. In the entity-relationship triples generated in step S2.2, add spatiotemporal attribute fields to each edge and adjust the node risk propagation probability: If the edge in the knowledge graph is e = (s, r, o), then after adding the time and space dimension attributes, it is expressed as: And ST = <s,r,o;A time ,TO space >; In the formula, s and o represent the starting and ending entities of the edge. is the relationship type; A time Represents the time dimension attribute; A space Represents the spatial dimension attribute; e ST Represents an edge containing spatiotemporal dynamic attributes; The adjusted node risk propagation probability is: In the formula, w edge represents the spatiotemporal weight.
5. The method for retrieving and analyzing non-performing asset cases according to claim 1, characterized in that: In S4, the risk propagation weight s of the multi-layer knowledge graph after dynamically updating the edge attributes is calculated through the spatiotemporal graph attention network. t And output high-risk subgraphs, the specific steps involved are: S3.
1. Spatiotemporal encoding of node features and edge attributes: in, It represents the result obtained after spatiotemporal encoding of node features and edge attributes; is the structural feature of the high-risk candidate case set output by S2, is the time decay function value, is the spatial attribute interpolation result; W t Represents a learnable weight matrix; ReLU represents an activation function; S3.
2. Calculate the dynamic attention coefficient a ij : in, represents the spatiotemporal weight of the edge between node i and node j; γ represents a hyperparameter; f decay (t ij ) represents the time decay function; PolicyRisk(g ij ) represents the policy risk function; LeakyReLU(*) represents the activation function; W represents the learnable weight matrix; represents the spatiotemporal weight of the edge between node i and node k; h i represents the feature representation of node i after spatiotemporal encoding; h j represents the feature representation of node j after spatiotemporal encoding; a T represents the learnable attention vector; t ij represents the time factor associated with the edge between node i and node j; S3.3, according to the dynamic attention coefficient a ij Calculate the risk propagation weight s t : In the formula, represents the set of neighbor nodes of node i; a ij represents the dynamic attention coefficient; S3.4, based on risk propagation weights t Filter high-risk edges and construct candidate subgraph G risk : G risk ={(i,j)|s t (i,j)>m} Where m is the preset threshold; i and j both represent nodes in the graph; s t (i, j) represents the risk transmission intensity from node i to node j; S3.
5. For the candidate subgraph G risk The path P in (v1, v1, ..., v n ), calculate the path risk score Score(P): Where, Decay(Δt)=e -λΔt represents the time decay function, Δt represents the time interval between edges, and λ represents the control decay rate; v k represents the kth node on path P; s t (v k , v k+1 ) represents the risk transmission intensity from node k to node k+1 in path P; S3.
6. Use the greedy algorithm to select the top K paths with the highest scores, merge nodes and edges to generate a high-risk subgraph G. hight-risk .
6. The method for retrieving and analyzing non-performing asset cases according to claim 5 is characterized by: The specific steps involved in locating hidden assets and updating graph node attributes through the subgraph search algorithm are: S4.
1. Based on high-risk subgraph G hight-risk The edges and their risk propagation weights s in t , construct a weighted adjacency matrix A, where A ij =s t (i, j), the modularity Q is optimized by Louvain algorithm, and the objective function is: Among them, m is the total edge weight of the subgraph; k i is the weighted degree of node i; k j is the weighted degree of node j; δ determines community belonging; δ(c i ,c j ) represents the exponential function; A ij represents the weight of the edge between node i and node j in the adjacency matrix; S4.2, set the path score threshold θ' to filter high-risk paths; For each community C i For node v in the image, perform a depth-first search to explore paths of length ≤ L, where the path score is: In the formula, Decay(Δt uv ) represents the time decay function; Score(P) represents the risk score of path P; (u,v) represents an edge in path P, connecting node u and node v; s t (u,v) represents the risk propagation weight from node u to node v; Δt uv Represents the time interval between edges (u,v); Keep the paths with Score(P)≥θ' and store them in the candidate path set P candidate ; S4.
3. For each hidden asset node v, update its attribute vector: Where, f(*) represents the mapping function; Represents the structured feature vector of node v; Represents a vector concatenation operation; IsHiddenAsset indicates whether the node v is a hidden asset; HiddenValue represents the value attribute of a hidden asset; If the hidden asset node v is marked as a hidden asset, the upstream node attributes are updated in reverse along its associated edge: in, In the formula, represents the structured feature vector of the upstream node u; RiskExposure represents the risk exposure attribute of the node; += represents the accumulation operation; represents the set of neighbor nodes of node v; u represents an upstream node of node v; α represents the risk exposure coefficient; s t (u,v) represents the risk propagation weight from node u to node ν; S4.
4. When a new asset node joins, only the affected community C i Recalculate the modularity Q.
7. The method for retrieving and analyzing non-performing asset cases according to claim 1, characterized in that: The graph embedding vector of the enterprise association network is generated using GraphSAGE: In the formula, represents the embedding of node v at the kth layer; W (k) Represents the k-th layer trainable weight matrix; the final output The graph-enhanced hybrid gradient boosting model is an improved model based on fusion graph features, and its mathematical expression is: In the formula, Represents the predicted value of the i-th sample; Represents the structured feature vector of the i-th sample; represents the final graph embedding vector of node v; λ represents the graph feature fusion coefficient; f k represents the gradient boosting tree base model; g represents the graph feature adapter; K represents the number of base models of the gradient boosting tree, k represents the index variable; f k (*) represents the prediction of the input feature * by the gradient boosting tree-based model; g(*) represents the transformation of the input graph embedding vector * by the graph feature adapter; By integrating spatiotemporal dynamic features and time decay mechanism into the graph-enhanced hybrid gradient boosting model, the optimized graph-enhanced hybrid gradient boosting model is: In the formula, s t is the spatiotemporal risk weight output by ST-GAT; w t is the time attenuation coefficient; T represents the total number of time steps; t represents the index variable of the time step; Composite loss function: In the formula, represents the composite loss function; Ω(f k ) represents the complexity regularization term of the k-th tree (number of leaf nodes / weight); W (k) represents the weight matrix of the kth layer of GraphSAGE; α represents the regularization strength hyperparameter of the gradient boosting tree; β represents the regularization strength hyperparameter of the graph feature encoding; γ represents the regularization strength hyperparameter of the spatiotemporal dynamic features; represents the cross entropy loss term; represents the gradient boosting tree regularization term; Represents the graph feature encoding regularization term; represents the spatiotemporal dynamic regularization term; represents the predicted value of the i-th sample; y i Represents the true label of the i-th sample; n represents the total number of samples, i represents the index variable; U (t) represents the weight matrix of the tth layer of ST-GAT.
8. The method for retrieving and analyzing non-performing asset cases according to claim 7, characterized in that: The progressive training strategy trains the graph-enhanced hybrid gradient boosting model, and the progressive training strategy is: Phase 1: Use only Training the model And fix these parameters {f k }; Phase 2: enter Training model g(*) and GraphSAGE parameters Phase 3: Unfreeze all parameters and g(*), perform end-to-end optimization and fine-tune the overall model.
9. A system for searching and analyzing non-performing asset cases, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The processor executes the computer program to implement the steps of the non-performing asset case retrieval and analysis method as described in any one of claims 1-8.
Citation Information
Cited By
Multi-source data fusion-based credit business risk control decision method, apparatus and device, and medium
CN120833210A
Credit risk control decision method, device and equipment based on multi-source data fusion and medium
CN120833210B
Tube feeding nursing complication retrieval recommendation processing method and system
CN120929583A
Geological disaster risk assessment method based on mountainous area geological features and knowledge graph
CN121882217A