A risk cause identification method, device, and storage medium
By constructing the risk control object relationship map and feature variable correlation analysis, and generating a risk sub-graph, the problem that the graph neural network model is difficult to analyze the model inference results in risk management is solved, and the accuracy and comprehensiveness of risk cause identification are improved.
Patent Information
- Application Number
- CN202110763170.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-06
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-07-06
AI Technical Summary
In the field of risk management, the graph neural network model is difficult to directly analyze the impact characteristics of model inference results, resulting in inaccurate and comprehensive identification of risk causes.
By constructing a relationship map of risk control objects, extracting the feature set of specified nodes, using the graph neural network model for risk prediction, and screening the risk characteristic variables based on the correlation of feature variables, and generating a risk sub-graph for risk cause identification.
It improves the accuracy and comprehensiveness of risk cause identification, can intuitively analyze the risk influencing factors and degree of impact of risk control objects, and improves the efficiency of risk analysis.
Smart Images

Figure CN113449921B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of big data technology, and in particular, to a method, apparatus, and storage medium for identifying risk causes. Background Art
[0002] The graph neural network model realizes the combination of the graph and the neural network model. By introducing different node information and network association information between nodes into the model, the inference effect of the model is well improved. However, as a black box model, it is difficult to directly analyze the model features that have a greater impact on the model inference result. Especially in the field of risk management, not only the risk inference result is required, but also the risk causes are needed. Summary of the Invention
[0003] The purpose of the embodiments of this specification is to provide a method, apparatus, and storage medium for identifying risk causes, which can further improve the accuracy and comprehensiveness of risk cause identification.
[0004] The method, apparatus, and storage medium for identifying risk causes provided in this specification are implemented in the following manner:
[0005] A method for identifying risk causes, which is applied to a server. The method includes: obtaining a risk control object relationship graph; the risk control object relationship graph includes edges between nodes, and the nodes are used to represent risk control objects, and the edges between the nodes are used to represent the relationships between different risk control objects; taking any node in the risk control object relationship graph as a specified node, and extracting a specified feature set of the specified node, the specified feature set includes feature values corresponding to multiple feature variables; wherein, the types of the feature variables at least include the object features of the specified risk control object corresponding to the specified node, the relationship features between the specified risk control object and other risk control objects in the relationship graph, and the graph features of the specified node in the relationship graph; based on the specified feature set, performing risk prediction on the specified node to obtain a risk prediction result of the specified node; according to the correlation between each feature variable in the specified feature set and the risk prediction result, and the correlation between any two feature variables, screening out the risk feature variables of the specified node and the associated feature variables corresponding to the risk feature variables; and based on the risk feature variables of the specified node and the associated feature variables corresponding to the risk feature variables, extracting a risk sub-graph corresponding to the specified node from the relationship graph, so as to identify the risk causes of the risk control object based on the risk sub-graph.
[0006] On the other hand, an embodiment of this specification provides a risk cause identification device, which is applied to a server. The device includes: an acquisition module, configured to acquire a risk control object relationship graph; the risk control object relationship graph includes edges between nodes, where the nodes are used to represent risk control objects, and the edges between the nodes are used to represent the relationships between risk control objects; a feature extraction module, configured to use any node in the risk control object relationship graph as a specified node, and extract a specified feature set of the specified node, where the specified feature set includes feature values corresponding to multiple feature variables; where the types of the feature variables at least include object features of the specified risk control object corresponding to the specified node, relationship features between the specified risk control object and other risk control objects in the relationship graph, and graph features of the specified node in the relationship graph; a risk prediction model, configured to perform risk prediction on the specified node based on the specified feature set to obtain a risk prediction result of the specified node; a feature screening module, configured to screen risk feature variables of the specified node and associated feature variables corresponding to the risk feature variables according to the correlation between each feature variable in the specified feature set and the risk prediction result, and the correlation between any two feature variables; a risk sub-graph extraction module, configured to extract a risk sub-graph corresponding to the specified node from the relationship graph according to the risk feature variables of the specified node and the associated feature variables corresponding to the risk feature variables, so as to identify the risk cause of the risk control object based on the risk sub-graph.
[0007] On the other hand, an embodiment of this specification provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed, the steps of the method described in any one or more of the above embodiments are implemented.
[0008] The risk cause identification method, device and storage medium provided by one or more embodiments of this specification can effectively extract important feature factors that cause a risk control object to have risks by analyzing the influence of feature variables on the risk prediction result and the mutual influence between feature variables. For example, which self-features of the risk control object have an impact on the generation of risks, which other risk control objects have an impact on the risk generation of this risk control object, and which risk control object association relationships have an impact on the risk generation of this risk control object, etc. Based on this, a risk sub-graph of each risk control object can be further generated, and business personnel can more intuitively and accurately infer the cause relationship of the risk control object according to the risk sub-graph, improving the accuracy of risk cause identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:
[0010] Figure 1 It is a schematic diagram of the construction process of the risk prediction model provided by this specification;
[0011] Figure 2 It is a schematic diagram of the interpretation process of the risk prediction model provided by this specification;
[0012] Figure 3 It is a schematic diagram of the process of reasoning about the risk causes provided by this specification;
[0013] Figure 4 It is a schematic diagram of the implementation process of the risk cause identification method provided by this specification;
[0014] Figure 5 It is a schematic diagram of the module structure of the risk cause identification device provided by this specification. Detailed implementation manners
[0015] To enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in one or more embodiments of this specification in conjunction with the drawings in one or more embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of the specification, rather than all of them. Based on one or more embodiments of the specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the embodiments of the specification.
[0016] In a scenario example provided by this specification, the risk cause identification method can be applied to a single server or a server cluster composed of multiple servers. The risk cause identification can refer to the identification of the influencing factors for the generation of the risk of the risk control object and the identification of the degree of influence of the influencing factors on the generation of the risk, etc. As Figures 1 to 3 shown, the risk cause identification method can at least include three links: the construction of the risk prediction model, the interpretation of the risk prediction model, and the reasoning about the risk causes. As Figure 1 shown, the risk prediction model can be constructed by the following steps S101 to S103:
[0017] Step S101: Construction of the relationship graph.
[0018] A risk control object relationship graph can be pre-constructed. The risk control objects can be individual users, enterprise users, etc. The relationship graph G=(V, E) can include nodes V and edges E between the nodes. Among them, V is used to represent the risk control objects; E represents the relationships between different risk control objects, such as investment relationships, guarantee relationships, subordination relationships, etc.
[0019] Step S102: Feature extraction.
[0020] The types of the extracted feature variables can at least include the object features of the risk control objects, the relationship features between any two risk control objects, and the graph features of the risk control objects in the relationship graph.
[0021] The object features of the risk control object i can be expressed as The types of the feature variables can include user assets, liabilities, capital inflows, capital outflows, etc.
[0022] The relationship features of the risk control object i can be expressed as The types of the feature variables can include investment, guarantee, subordination, etc.
[0023] The graph features of the risk control object i can be expressed as The types of the feature variables can include the out-degree, in-degree, concentration degree, etc. of the node corresponding to the risk control object i in the relationship graph.
[0024] Integrate all the feature variables of the risk control object i into That is, each risk control object has K feature variables.
[0025] Assume that the target variable of the risk prediction model is expressed as y i , then when the risk control object i is a black sample, y i = 1; otherwise y i = 0.
[0026] Step S103: Model training.
[0027] Input the feature variables and the target variable into the graph neural network model for training to obtain the trained risk prediction model. The specific type of the graph neural network model selected is a binary classification node risk result prediction model. Of course, other algorithm models can also be used for risk result prediction, which is not limited here.
[0028] Such as Figure 2 shown, the risk prediction model can be explained by the following steps S201 to S203:
[0029] Step S201: Subgraph sampling.
[0030] The sampling system can generate subgraphs in three ways: BA graph, trajectory graph, and N-hop graph.
[0031] BA graph: Starting from the initial node, each time a new node is added, it is required to have an association relationship with at least three nodes in the existing nodes to form a subgraph.
[0032] Trajectory graph: Starting from the initial node, define the trajectory of node walking, and connect the nodes on the trajectory to form a subgraph.
[0033] N-hop graph: Starting from the initial node, expand the connection relationship by N layers outward to form a subgraph.
[0034] Of course, the above methods of extracting subgraphs are only preferred examples for illustration and do not constitute direct limitations. The relationship graph usually has a large amount of data. By extracting subgraphs for subsequent processing, the amount of data processing can be greatly reduced and the data processing efficiency can be improved. Of course, it is also possible not to extract subgraphs and directly identify the risk causes based on the risk control object relationship graph. The specific processing method refers to the processing method based on subgraphs below and will not be elaborated here.
[0035] N target variables can be extracted based on the above methods (that is, extract the data corresponding to N risk control objects). Starting from the N target variables, according to the subgraph generation method in the foregoing steps, the server selects the subgraph generation method according to the ratio (for example, define the ratio of BA graph, trajectory graph, and N-hop graph in the sample as 1:1:2. The ratio can be set by business personnel through the terminal device) to generate N subgraph sampling samples.
[0036] Step S202: Model prediction
[0037] Take the characteristic variables corresponding to each of the N target variables,
[0038] Input the graph neural network model trained in the foregoing steps to obtain the prediction result y i .
[0039] Defined as
[0040] Step S203: Model interpretation
[0041] Use the following first risk prediction analysis model to execute the interpretation of the risk prediction model to determine the factors that affect the risk prediction result, and the risk cause analysis can be realized based on the interpretation result.
[0042]
[0043] Among them, each element in β = [β1,…,β k ,…,β m ,…,β K represents the importance degree of each characteristic variable to the risk prediction result y, β1,…,βk , …, β m , …, β K ≥0; R K represents a K - dimensional real vector; HSIC(f k , f m ) represents the correlation metric value between the feature variables f k and f m ; f k represents the vector formed by the feature variables with the identifier k corresponding to each node; f m represents the vector formed by the feature variables with the identifier m corresponding to each node. HSIC(f k , y) represents the correlation metric value between the feature variable f k and the risk prediction result y; y represents the vector formed by the risk prediction results corresponding to each node. HSIC(y, y) is a constant; λ is an adjustment coefficient; ||·||2 represents the second - order norm.
[0044] By comprehensively analyzing multiple samples to analyze the correlation between different types of feature variables, the correlation between each type of feature variable and the risk prediction result, and the importance of each type of feature variable to the risk prediction result, and then analyzing the risk impact link, the risk link relationship between different nodes can be more accurately and comprehensively mined, improving the accuracy and comprehensiveness of risk cause analysis.
[0045] Among them, HSIC(f k , f m ) and HSIC(f k , y) can be expressed as:
[0046]
[0047]
[0048] Among them, tr(·) represents the trace of a matrix, Among them, the kernel functions of each element in M (k) , M (m) , L are:
[0049]
[0050]
[0051] If the risk prediction result is a continuous numerical value:
[0052]
[0053] If the risk prediction result is a discrete numerical value:
[0054]
[0055]
[0056] Among them, respectively represent the feature variables with identifier k corresponding to nodes i and j, respectively represent the feature variables with identifier m corresponding to nodes i and j, represents the standard deviation among the feature variables with identifier k corresponding to each node; represents the standard deviation among the feature variables with identifier m corresponding to each node; y i , y j respectively represent the risk prediction results corresponding to nodes i and j; represents the standard deviation among the risk prediction results corresponding to each node; I N is an N-dimensional unit matrix, 1 N is an N-dimensional vector, and each of its elements takes a value of 1; ||·|| F represents the Frobenius norm, i, j = 1,..., N, and N represents the number of samples.
[0057] Solving the above first risk prediction analysis model gives:
[0058] β = [β1,…,β k ,…β K
[0059] HSIC(f k , f m )
[0060] HSIC(f k , y)
[0061] β k The higher the value, the more important the feature variable corresponding to β k is to the risk prediction result. And record the feature variable corresponding to β k and the feature meaning corresponding to this feature variable, that is, whether the feature variable corresponds to a node, an edge, or a graph feature of a node.
[0062] HSIC(f k , f m ) represents the correlation between two feature variables with identifiers k and m, and the larger the value, the higher the correlation. And record the feature variables involved in HSIC(f k , f m ) and the corresponding feature meanings of the corresponding feature variables, that is, whether the corresponding feature variables correspond to nodes, edges, or graph features of nodes.
[0063] If the feature variables with identifiers k and m respectively correspond to:
[0064] Node A - Node B, then the larger the value of HSIC(f k , f m ), the closer the relationship between these two nodes;
[0065] Node C - Edge D, then the larger the value of HSIC(f k , f m ), the more important the relationship between Node C and this Edge D, indirectly indicating that the node E at the other end of this Edge D has a close relationship with Node C;
[0066] Node F - Graph features of Node F, then the larger the value of HSIC(f k , f m ), the greater the influence of Node F on the surrounding nodes;
[0067] Edge G - Graph features of Node H, then the larger the value of HSIC(f k , f m ), the stronger the conductivity of the graph features of Node H to the surrounding Edge G, indirectly indicating that the node I at the other end of this Edge G has a close relationship with Node H.
[0068] Of course, the above corresponding relationships are only for illustrative purposes, and there may be other corresponding relationships in specific implementations. Thus, through HSIC(f k , f m ), the strength of the relationship between nodes in the relationship graph and the magnitude of the influence of nodes on the surrounding nodes can be characterized more accurately and conveniently, thereby facilitating the realization of risk connection and strength analysis between nodes.
[0069] HSIC(f k , y) represents the correlation between the feature variable labeled as k and the target variable y, and the larger the value, the higher the correlation. And record the feature variables involved in HSIC(f k , y) and the corresponding feature meanings of the corresponding feature variables, that is, whether the corresponding feature variable corresponds to a node, an edge, or the graph features of a node.
[0070] The feature variables whose importance for the risk prediction result meets the specified importance condition can be used as risk feature variables, and extract the HSIC(f k , f m ) and HSIC(f k , y) corresponding to the risk feature variables, so as to determine the associated feature variables corresponding to the corresponding risk feature variables according to the extracted HSIC(f k , f m ).
[0071] For example, for β = [β1,…,β k ,…βK Sort from high to low. Set a threshold τ%, and extract the feature variables higher than τ% as risk feature variables. Of course, the first T feature variables can also be selected as risk feature variables. And extract the HSIC(f k ,f m ) and HSIC(f k ,y) corresponding to the risk feature variables.
[0072] Since the kernel function needs to be defined when calculating K (k) and L above, in order to reduce the sensitivity impact of the kernel function calculation for the parameter β. The above formula can be adjusted to:
[0073]
[0074] where β = [β1,…,β k ,…,β m ,…,β K , and each element in it represents the importance degree of each feature variable to the risk prediction result y; R K represents a K-dimensional real vector; D NOCCO (f k ,f m ) represents the correlation measurement value between the feature variables f k and f m ; D NOCCO (f k ,y) represents the correlation measurement value between the feature variable f k and the risk prediction result y; λ is an adjustment coefficient; ||·||1 represents the first-order norm. Among them, D NOCCO (f k ,f m ) and D NOCCO (f k ,y) can be expressed as:
[0075]
[0076]
[0077] where tr(·) represents the trace of a matrix, ε > 0 is a hyperparameter set by business personnel. Other parameters refer to the above formula and will not be elaborated here. D NOCCO (f k ,f m ), D NOCCO (f k ,y) have similar meanings to HSIC(f k ,f m ) and HSIC(f k ,y), and will not be elaborated here.
[0078] As shown in Figure 3 , the risk cause reasoning can be performed by the following steps S301 to S303:
[0079] Step S301: Risk sub-graph construction.
[0080] For each risk control object (i.e., node) n, after extracting the risk feature variables based on the above method, assuming the identifier of a certain risk feature variable is k, find the corresponding edge or node of the risk feature variable from the relationship graph.
[0081] Then starting from HSIC(f k , f m )(or D NOCCO (f k , f m ), take the feature variable with the identifier m as the associated feature variable of the risk feature variable with the identifier k, and find the corresponding node or edge of the associated feature variable from the relationship graph. And so on, search all the links of the risk feature variables corresponding to the risk control object (i.e., node) n.
[0082] By the above method, the nodes that are relatively closely associated with the risk control object n can be found in sequence to form the risk sub-graph corresponding to the risk control object (i.e., node) n. At the same time, HSIC(f k , y), HSIC(f k , f m )(or D NOCCO (f k , y), D NOCCO (f k , f m )) can represent the influence degree of a certain feature variable on the prediction result and the strength of the association between feature variables. When the feature variables corresponding to HSIC(f k , y), HSIC(f k , f m )(or D NOCCO (f k , y), D NOCCO (f k , f m )) are the relationship features or graph features of the risk control object, it is also possible to based on HSIC(f k , y), HSIC(f k , f m )(or D NOCCO (f k , y), D NOCCO (f k , f m)) Determine the weights of the edges involved in the corresponding feature variables in the risk sub-graph, and further determine the influence degree of the conduction relationship corresponding to the edge on the risk result based on the weight of the edge.
[0083] Each risk control object n generates a risk sub-graph. A total of N risk sub-graphs are generated.
[0084] Step S302: Risk cause reasoning.
[0085] Form a risk sub-graph according to the foregoing steps. Starting from the risk control object (i.e., node) n, record the conduction link of generating the risk sub-graph in the foregoing steps to form a risk cause reasoning relationship graph. The server provides a screening function on the terminal device, and business personnel can screen the risk link relationships to be analyzed from the risk cause reasoning relationship graph according to weights, node relationships, etc. At the same time, the server also provides a visualization method, and according to the pointing and weight relationships of the edges between nodes, uses the visualization method to display through a graphical interface. Different colors can also be marked for different weight relationships to provide more intuitive and effective visualization analysis for business personnel.
[0086] Step S303: Risk list generation.
[0087] According to the risk link relationships screened by business personnel in the foregoing steps, for the nodes marked as risks, the server can generate a risk relationship list and can include the relevant nodes on the risk link in the risk list.
[0088] For the nodes not marked as risks, the node can be marked as a non-risk node. If there are nodes marked as risks in the risk sub-graph corresponding to the non-risk node, the nodes on the shortest link between the node marked as a risk and the non-risk node can be included in the risk list.
[0089] It can be seen that through the above risk cause reasoning relationship graph, the accuracy and comprehensiveness of risk analysis and investigation can be greatly improved, and it is convenient for business personnel to more clearly and intuitively determine the influencing factors of the risk of the risk control object and the influence degree of the influencing factors on the risk generation, etc., so that the accuracy, comprehensiveness and efficiency of risk cause identification in complex business scenarios can be greatly improved.
[0090] Based on the above scenario example, this specification also provides a risk cause identification method. Figure 4 It is a schematic flowchart of an embodiment of the risk cause identification method provided in this specification. As Figure 4 shown, in an embodiment of the risk cause identification method provided in this specification, the method can be applied to a server. The method can include the following steps.
[0091] S40: Obtain the risk control object relationship graph; the risk control object relationship graph includes edges between nodes, where the nodes are used to represent risk control objects, and the edges between the nodes are used to represent the relationships between different risk control objects.
[0092] S42: Take any node in the risk control object relationship graph as the specified node, and extract the specified feature set of the specified node. The specified feature set includes the feature values corresponding to multiple feature variables. Among them, the types of the feature variables at least include the object features of the specified risk control object corresponding to the specified node, the relationship features between the specified risk control object and other risk control objects in the relationship graph, and the graph features of the specified node in the relationship graph.
[0093] S44: Based on the specified feature set, perform risk prediction on the specified node to obtain the risk prediction result of the specified node.
[0094] S46: According to the correlation between each feature variable in the specified feature set and the risk prediction result, as well as the correlation between any two feature variables, screen out the risk feature variables of the specified node and the associated feature variables corresponding to the risk feature variables.
[0095] S48: Based on the risk feature variables of the specified node and the associated feature variables corresponding to the risk feature variables, extract the risk sub-graph corresponding to the specified node from the relationship graph, so as to identify the risk causes of the risk control object based on the risk sub-graph.
[0096] In the above embodiments, by analyzing the influence of feature variables on the risk prediction result and the mutual influence between feature variables, important feature factors that cause risks for risk control objects can be effectively extracted. For example, which self-features of the risk control object have an impact on the generation of risks, which other risk control objects have an impact on the risk generation of this risk control object, and which risk control object association relationships have an impact on the risk generation of this risk control object, etc. Based on this, the risk sub-graphs of each risk control object can be further generated, and business personnel can more intuitively and accurately infer the causal relationship of the risk control object according to the risk sub-graph, improving the accuracy of risk cause identification.
[0097] In some other embodiments, the importance of each feature variable in the specified feature set to the risk prediction result may be further combined, and the risk feature variables of the specified node and the associated feature variables corresponding to the risk feature variables may be screened according to the correlation between each feature variable in the specified feature set and the risk prediction result and the correlation between any two feature variables. By further combining the importance of each feature variable in the specified feature set to the risk prediction result, feature variables with a greater impact on the result can be extracted from a large number of feature variables, and then the risk causation relationship can be sorted out based on the feature variables with a greater impact, which can make the risk causation analysis more accurate and efficient.
[0098] In some other embodiments, the server may also determine the weights of the edges between the nodes in the risk subgraph according to the correlation metric value between the risk feature variable and the risk prediction result and the correlation metric value between the risk feature variable and the corresponding associated feature variable, so as to use the weights to label the influence degree of the association relationship and the conduction relationship between each risk control object on the risk of a certain risk control object, and further improve the accuracy and comprehensiveness of risk causation identification.
[0099] In some other embodiments, the server may also construct a risk causation inference relationship graph based on the conduction links of the risk subgraphs corresponding to the nodes in the relationship graph, so as to identify the risk causation of the risk control object based on the risk causation inference relationship graph. By comprehensively integrating the risk subgraphs corresponding to each node to form a risk causation inference relationship graph, the sorting of the risk influence relationship can be made more accurate and comprehensive, and the accuracy of risk causation identification can be further improved.
[0100] In some embodiments, the risk feature variables of the specified node and the associated feature variables corresponding to the risk feature variables may be screened in the following manner:
[0101] Solve the first risk prediction analysis model to obtain the importance of each feature variable in the specified feature set to the risk prediction result, the correlation metric value between each feature variable in the specified feature set and the risk prediction result, and the correlation metric value between any two feature variables; wherein, the first risk prediction analysis model includes:
[0102]
[0103] Among them, each element in β = [β1,…,β k ,…,β m ,…,β K represents the importance of each feature variable to the risk prediction result y, and β1,…,β k ,…,β m ,…,β K ≥0; R Kdenotes a K-dimensional real vector; HSIC(f k , f m ) represents the correlation metric value between the feature variables f k and f m ; HSIC(f k , y) represents the correlation metric value between the feature variable f k and the risk prediction result y; HSIC(y, y) is a constant; λ is an adjustment coefficient; ||·||2 represents the second-order norm;
[0104] Regarding the feature variables whose importance to the risk prediction result meets the specified importance condition as risk feature variables, extract the HSIC(f k , f m ) corresponding to the risk feature variables;
[0105] According to the extracted HSIC(f k , f m ), determine the associated feature variables corresponding to the corresponding risk feature variables.
[0106] Among them, the correlation metric values between each feature variable in the specified feature set and the risk prediction result and the correlation metric values between any two feature variables are determined in the following manner:
[0107]
[0108]
[0109] Among them, tr(·) represents the trace of a matrix, Among them, the kernel functions of the elements in M (k) , M (m) , and L are:
[0110]
[0111]
[0112] If the risk prediction result is a continuous numerical value:
[0113]
[0114] If the risk prediction result is a discrete numerical value:
[0115]
[0116]
[0117] Among them, respectively represent the feature variables with the identifier k corresponding to nodes i and j, respectively represent the feature variables with the identifier m corresponding to nodes i and j, represents the standard deviation between the feature variables with the identifier k corresponding to each node; represents the standard deviation between the feature variables with the identifier m corresponding to each node; y i , y j respectively represent the risk prediction results corresponding to nodes i and j; represents the standard deviation between the risk prediction results corresponding to each node; I N is an N-dimensional unit matrix, 1 N is an N-dimensional vector, and each element of it takes the value of 1; ||·|| F represents the Frobenius norm, i, j = 1,..., N, and N represents the number of samples.
[0118] In the above manner, by comprehensively analyzing the data of multiple risk control objects to analyze the risk impact relationship, the risk impact analysis can be made more accurate and comprehensive, and further improve the accuracy and comprehensiveness of the final risk impact relationship interpretation and analysis. The multiple risk control objects comprehensively considered above can be all the risk control objects involved in the relationship graph; or they can be risk control objects selected based on a certain method, such as selecting risk control objects through the above-mentioned BA graph, trajectory graph, N-hop graph, etc. By selecting risk control objects, the data processing volume can be further reduced and the data processing efficiency can be improved.
[0119] In some other embodiments, the following method can also be used to screen the risk feature variables of the specified node and the associated feature variables corresponding to the risk feature variables:
[0120] Solve the second risk prediction analysis model to obtain the importance of each feature variable in the specified feature set to the risk prediction result, the correlation metric value between each feature variable in the specified feature set and the risk prediction result, and the correlation metric value between any two feature variables; wherein, the second risk prediction analysis model includes:
[0121]
[0122] wherein, each element in β = [β1,…,β k ,…,β m ,…,β K represents the importance of each feature variable to the risk prediction result y; R K represents a K-dimensional real vector; D NOCCO (f k , f m ) represents the correlation metric value between the feature variables f k and f m ; D NOCCO (f k, y) represents the correlation metric value between the feature variable f k and the risk prediction result y; λ is the adjustment coefficient; ||·||1 represents the first-order norm;
[0123] Take the feature variables whose importance to the risk prediction result meets the specified importance condition as risk feature variables, and extract the D corresponding to the risk feature variables NOCCO (f k , f m );
[0124] According to the extracted D NOCCO (f k , f m ) to determine the associated feature variables corresponding to the corresponding risk feature variables.
[0125] Among them, the correlation metric values between each feature variable in the specified feature set and the risk prediction result and the correlation metric values between any two feature variables are determined in the following manner:
[0126]
[0127]
[0128] Among them, tr(·) represents the trace of the matrix, Among them, Among them, M (k) , M (m) , the kernel functions of the elements in L are:
[0129]
[0130]
[0131] If the risk prediction result is a continuous numerical value:
[0132]
[0133] If the risk prediction result is a discrete numerical value:
[0134]
[0135]
[0136] Among them, respectively represent the feature variables with the identifier k corresponding to nodes i and j, respectively represent the feature variables with the identifier m corresponding to nodes i and j, represents the standard deviation between the feature variables with the identifier k corresponding to each node; represents the standard deviation between the characteristic variables with the identifier m corresponding to each node; y i and y j respectively represent the risk prediction results corresponding to nodes i and j; represents the standard deviation between the risk prediction results corresponding to each node; I N is an N-dimensional unit matrix, 1 N is an N-dimensional vector, and each element of it takes a value of 1; ||·|| F represents the Frobenius norm, i, j = 1,..., N, and N represents the number of samples.
[0137] Through the above method, the sensitivity impact of calculating the kernel function for the parameter β can be further reduced, and the accuracy of extracting the risk impact relationship can be further improved.
[0138] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. Specifically, reference can be made to the description of the relevant processing-related embodiments above, and details will not be repeated here.
[0139] The specific embodiments of this specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0140] Based on the method provided in the above embodiments, an embodiment of this specification further provides a risk cause identification device, which is applied to a server. Figure 5 represents a schematic diagram of the module structure of the risk cause identification device in an embodiment provided by this specification, as Figure 5 shown, the device includes:
[0141] An acquisition module 50, which can be used to acquire a risk control object relationship graph; the risk control object relationship graph includes edges between nodes, the nodes are used to represent risk control objects, and the edges between the nodes are used to represent the relationships between risk control objects.
[0142] The feature extraction module 52 can be used to take any node in the risk control object relationship graph as a specified node, and extract a specified feature set of the specified node, where the specified feature set includes the feature values corresponding to multiple feature variables; among them, the types of the feature variables at least include the object features of the specified risk control object corresponding to the specified node, the relationship features between the specified risk control object and other risk control objects in the relationship graph, and the graph features of the specified node in the relationship graph.
[0143] The risk prediction model 54 can be used to perform risk prediction on the specified node based on the specified feature set to obtain the risk prediction result of the specified node.
[0144] The feature screening module 56 can be used to screen the risk feature variables of the specified node and the associated feature variables corresponding to the risk feature variables according to the correlation between each feature variable in the specified feature set and the risk prediction result, and the correlation between any two feature variables.
[0145] The risk sub-graph extraction module 58 can be used to extract the risk sub-graph corresponding to the specified node from the relationship graph according to the risk feature variables of the specified node and the associated feature variables corresponding to the risk feature variables, so as to identify the risk causes of the risk control object based on the risk sub-graph.
[0146] It should be noted that the above-mentioned device may also include other implementation manners according to the description of the above embodiments. The specific implementation manners can refer to the description of the relevant method embodiments and will not be elaborated here one by one.
[0147] This specification also provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed, the steps of the method described in any one or more of the above embodiments are implemented. The storage medium may include a physical device for storing information, usually storing the information after digitization and then using media such as electricity, magnetism or optics to store it. The storage medium may include: devices that store information using electrical energy, such as various memories, such as RAM, ROM, etc.; devices that store information using magnetic energy, such as hard disks, floppy disks, magnetic tapes, magnetic core memories, bubble memories, USB drives; devices that store information using optical methods, such as CDs or DVDs. Of course, there are also other ways of readable storage media, such as quantum memories, graphene memories, and so on.
[0148] It should be noted that the embodiments of this specification are not limited to those that must conform to the standard data model / template or the situations described in the embodiments of this specification. Some industry standards or implementation schemes slightly modified on the basis of the implementation described by using a custom method or embodiment can also achieve the same, equivalent or similar, or predictable implementation effects after deformation as the above embodiments. The embodiments obtained by using these modified or deformed data acquisition, storage, judgment, processing methods, etc. still fall within the scope of the optional implementation schemes of this specification.
[0149] The embodiments in this specification are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. In the description of this specification, the description of reference terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic expression of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0150] The above is only the embodiments of this specification and is not used to limit this specification. For those skilled in the art, this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.
Claims
1. A risk cause identification method, characterized in that, Applied to a server, the method includes: Obtain a risk control object relationship graph; the risk control object relationship graph includes edges between nodes, where the nodes are used to represent risk control objects, and the edges between the nodes are used to represent the relationships between different risk control objects; Take any node in the risk control object relationship graph as a specified node, and extract a specified feature set of the specified node, where the specified feature set includes feature values corresponding to multiple feature variables; among them, the types of the feature variables at least include object features of the specified risk control object corresponding to the specified node, relationship features between the specified risk control object and other risk control objects in the relationship graph, and graph features of the specified node in the relationship graph; Perform risk prediction on the specified node based on the specified feature set to obtain a risk prediction result of the specified node; According to the correlation between each feature variable in the specified feature set and the risk prediction result, and the correlation between any two feature variables, screen out the risk feature variables of the specified node and the associated feature variables corresponding to the risk feature variables; Based on the risk feature variables of the specified node and the associated feature variables corresponding to the risk feature variables, extract a risk sub-graph corresponding to the specified node from the relationship graph, so as to identify the risk causes of the risk control object based on the risk sub-graph; Among them, the screening out of the risk feature variables of the specified node and the associated feature variables corresponding to the risk feature variables includes: Solve a first risk prediction analysis model to obtain the importance of each feature variable in the specified feature set to the risk prediction result, the correlation measurement value between each feature variable in the specified feature set and the risk prediction result, and the correlation measurement value between any two feature variables; among them, the first risk prediction analysis model includes: Among them, β = [β1, …, β k , …, β m , …, β K . Each element represents the importance of each feature variable to the risk prediction result y. β1, …, β k , …, β m , …, β K ≥0; R K represents a K-dimensional real vector; HSIC(f k , f m ) represents the correlation metric value between the feature variables f k and f m ; HSIC(f k , y) represents the correlation metric value between the feature variable f k and the risk prediction result y; HSIC(y, y) is a constant; λ is an adjustment coefficient; ||·||2 represents the second-order norm; Take the feature variables whose importance of the risk prediction results meets the specified importance conditions as risk feature variables, and extract the HSIC(f k ,f m ) corresponding to the risk feature variables; Determine the associated feature variables corresponding to the corresponding risk feature variables according to the extracted HSIC(f k ,f m ); The correlation measurement value between each feature variable in the specified feature set and the risk prediction result, and the correlation measurement value between any two feature variables are determined in the following manner: where tr(·) represents the trace of a matrix, where M (k ), M (m) , and the kernel functions of the elements in L are: If the risk prediction result is a continuous numerical value: If the risk prediction result is a discrete numerical value: Among them, respectively represent the feature variables with identifier k corresponding to nodes i and j, respectively represent the feature variables with identifier m corresponding to nodes i and j, represents the standard deviation among the feature variables with identifier k corresponding to each node; represents the standard deviation among the feature variables with identifier m corresponding to each node; y i and y j respectively represent the risk prediction results corresponding to nodes i and j; represents the standard deviation among the risk prediction results corresponding to each node; I N is an N-dimensional unit matrix, 1 N is an N-dimensional vector, and each of its elements takes a value of 1; ||·|| F represents the Frobenius norm, i, j = 1,..., N, and N represents the number of samples.
2. The method according to claim 1, wherein The method further includes: According to the correlation measurement value between the risk feature variable and the risk prediction result, and the correlation measurement value between the risk feature variable and the corresponding associated feature variable, determine the weight of the edge between the nodes in the risk sub-graph.
3. The method according to claim 1, characterized in that The method further includes: Based on the conduction link of the risk sub-graph corresponding to each node in the risk control object relationship graph, construct a risk cause inference relationship graph, so as to identify the risk causes of the risk control object based on the risk cause inference relationship graph.
4. The method according to claim 1, characterized in that, The screening out of the risk feature variables of the specified node and the associated feature variables corresponding to the risk feature variables includes: Solve a second risk prediction analysis model to obtain the importance of each feature variable in the specified feature set to the risk prediction result, the correlation measurement value between each feature variable in the specified feature set and the risk prediction result, and the correlation measurement value between any two feature variables; among them, the second risk prediction analysis model includes: Among them, each element in β = [β1, …, β k , …, β m , …, β K represents the importance of each characteristic variable to the risk prediction result y; R K represents a K-dimensional real vector; D NOCCO (f k , f m ) represents the correlation metric value between the characteristic variables f k and f m ; D NOCCO (f k , y) represents the correlation metric value between the characteristic variable f k and the risk prediction result y; λ is an adjustment coefficient; ||·||1 represents the first-order norm; The feature variables whose importance degrees of the risk prediction results meet the specified importance degree conditions are used as risk feature variables, and D corresponding to the risk feature variables is extracted NOCCO (f k ,f m ); According to the extracted D NOCCO (f k , f m ) Determine the associated feature variables corresponding to the corresponding risk feature variables.
5. The method according to claim 4, wherein The correlation metric values between each feature variable in the specified feature set and the risk prediction result, and the correlation metric values between any two feature variables are determined in the following manner: where tr(·) represents the trace of a matrix, where where the kernel functions of the elements in M (k) , M (m) , and L are: If the risk prediction result is a continuous numerical value: If the risk prediction result is a discrete numerical value: Among them, respectively represent the feature variables with identifier k corresponding to nodes i and j, respectively represent the feature variables with identifier m corresponding to nodes i and j, represents the standard deviation between the feature variables with identifier k corresponding to each node; represents the standard deviation between the feature variables with identifier m corresponding to each node; y i 、y j respectively represent the risk prediction results corresponding to nodes i and j; represents the standard deviation between the risk prediction results corresponding to each node; I N is an N-dimensional unit matrix, 1 N is an N-dimensional vector, and each element of it takes the value of 1; ||·|| F represents the Frobenius norm, i, j = 1,..., N, and N represents the number of samples.
6. A risk cause identification device, characterized in that, Applied to a server, the apparatus includes: An acquisition module, configured to acquire a risk control object relationship graph; the risk control object relationship graph includes edges between nodes, the nodes are used to represent risk control objects, and the edges between the nodes are used to represent the relationships between risk control objects; A feature extraction module, configured to use any node in the risk control object relationship graph as a specified node, and extract a specified feature set of the specified node, the specified feature set includes feature values corresponding to multiple feature variables; wherein, the types of the feature variables at least include object features of the specified risk control object corresponding to the specified node, relationship features between the specified risk control object and other risk control objects in the relationship graph, and graph features of the specified node in the relationship graph; A risk prediction model, configured to perform risk prediction on the specified node based on the specified feature set to obtain a risk prediction result of the specified node; A feature screening module, configured to screen risk feature variables of the specified node and associated feature variables corresponding to the risk feature variables according to the correlation between each feature variable in the specified feature set and the risk prediction result, and the correlation between any two feature variables; A risk sub-graph extraction module, configured to extract a risk sub-graph corresponding to the specified node from the relationship graph according to the risk feature variables of the specified node and the associated feature variables corresponding to the risk feature variables, so as to identify the risk causes of the risk control object based on the risk sub-graph; Wherein, the feature screening module is specifically configured to: Solve a first risk prediction analysis model to obtain the importance of each feature variable in the specified feature set to the risk prediction result, the correlation metric values between each feature variable in the specified feature set and the risk prediction result, and the correlation metric values between any two feature variables; wherein, the first risk prediction analysis model includes: Among them, β = [β1, …, β k , …, β m , …, β K . Each element in it represents the importance degree of each characteristic variable to the risk prediction result y. β1, …, β k , …, β m , …, β K ≥0; R K represents a K-dimensional real vector; HSIC(f k , f m ) represents the correlation measurement value between the characteristic variables f k and f m ; HSIC(f k , y) represents the correlation measurement value between the characteristic variable f k and the risk prediction result y; HSIC(y, y) is a constant; λ is an adjustment coefficient; ||·||2 represents the second-order norm; The feature variables whose importance degrees of the risk prediction results meet the specified importance degree conditions are used as risk feature variables, and the HSIC(f k ,f m ) corresponding to the risk feature variables is extracted; Determine the associated feature variables corresponding to the corresponding risk feature variables according to the extracted HSIC(f k ,f m ); The correlation metric values between each feature variable in the specified feature set and the risk prediction result, and the correlation metric values between any two feature variables are determined in the following manner: where tr(·) represents the trace of a matrix, where M (k) , M (m) , and the kernel functions of the elements in L are: If the risk prediction result is a continuous numerical value: If the risk prediction result is a discrete numerical value: Among them, respectively represent the feature variables with identifier k corresponding to nodes i and j, respectively represent the feature variables with identifier m corresponding to nodes i and j, represents the standard deviation among the feature variables with identifier k corresponding to each node; represents the standard deviation among the feature variables with identifier m corresponding to each node; y i and y j respectively represent the risk prediction results corresponding to nodes i and j; represents the standard deviation among the risk prediction results corresponding to each node; I N is an N-dimensional unit matrix, 1 N is an N-dimensional vector, and each of its elements takes a value of 1; ||·|| F represents the Frobenius norm, i, j = 1,..., N, and N represents the number of samples.
7. A computer-readable storage medium having computer instructions stored thereon, characterized in that, When the instruction is executed, the steps of the method according to any one of claims 1-5 are implemented.
Citation Information
Patent Citations
Account risk recognition method and apparatus, and electronic device
CN111340612A