Risk detection method and device, storage medium and electronic equipment
By constructing a risk propagation heterogeneous graph, the problem of the inability to detect risk diffusion and propagation in existing technologies is solved, risk diffusion detection and early warning for multiple business domains are realized, and a holistic perspective monitoring of risk diffusion is provided.
Patent Information
- Application Number
- CN202510796673.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-23
AI Technical Summary
Existing risk detection methods can only detect risks in a specific business domain, and cannot effectively detect the spread and impact of risks, nor can they conduct timely detection and early warning between different business domains.
Construct a risk propagation heterogeneous graph, obtain nodes and their relationships in multiple business domains, determine the risk propagation path based on risk cases, and perform anomaly detection on nodes in the path to achieve detection and early warning of risk spread.
It realizes the detection of risk diffusion paths and determination of impact scope in multiple business domains, and provides timely risk diffusion warning and tracing functions.
Smart Images

Figure CN120687982A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to data processing technology, and more particularly to a risk detection method, device, storage medium, and electronic device. Background Art
[0002] With the continuous development of computer technology, risk monitoring systems are being applied in various fields to identify risks, including but not limited to e-commerce and financial technology.
[0003] Currently, risk monitoring systems can evaluate business behavior based on a predefined set of rules, triggering alerts when specific rule conditions are met. Alternatively, risk monitoring systems can train a risk prediction model based on historical data. This risk prediction model is a machine learning model that is then used to assign risk scores to new business data.
[0004] In the process of implementing the present disclosure, it was found that there are at least the following technical problems in the prior art: the above-mentioned risk detection method can only perform risk detection on a specific business behavior, and cannot detect the spread of risks. Summary of the Invention
[0005] The present disclosure provides a risk detection method, device, storage medium and electronic device to detect risk diffusion paths in different business domains, and to provide early warning and tracing of risk diffusion in a timely manner.
[0006] In a first aspect, an embodiment of the present disclosure provides a risk detection method, comprising:
[0007] Obtaining a risk propagation heterogeneous graph, wherein the risk propagation heterogeneous graph includes nodes corresponding to multiple business domains, and the nodes are connected based on association relationships;
[0008] Determine a first node according to the risk case, and determine at least one risk propagation path where the first node is located in the risk propagation heterogeneous graph;
[0009] Anomaly detection is performed on each node in the risk propagation path to determine risk anomaly detection information of each node in the risk propagation path.
[0010] In a second aspect, an embodiment of the present disclosure further provides a risk detection device, comprising:
[0011] A heterogeneous graph acquisition module is used to acquire a risk propagation heterogeneous graph, wherein the risk propagation heterogeneous graph includes nodes corresponding to multiple business domains, and the nodes are connected based on association relationships;
[0012] A risk propagation path determination module, configured to determine a first node according to a risk case, and determine at least one risk propagation path where the first node is located in the risk propagation heterogeneous graph;
[0013] The risk anomaly detection module is used to perform anomaly detection on each node in the risk propagation path and determine the risk anomaly detection information of each node in the risk propagation path.
[0014] In a third aspect, an embodiment of the present disclosure further provides an electronic device, characterized in that the electronic device includes:
[0015] one or more processors;
[0016] a storage device for storing one or more programs,
[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the risk detection method provided by any embodiment of the present disclosure.
[0018] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to execute the risk detection method provided in any embodiment of the present disclosure.
[0019] The technical solution provided by the disclosed embodiments integrates and uniformly models nodes across different business domains by constructing a heterogeneous risk propagation graph comprising nodes corresponding to multiple business domains. This provides a data foundation for risk diffusion detection and holistic risk monitoring across multiple business domains. For each risk case, based on the first node corresponding to the risk case in the heterogeneous risk propagation graph, a search is performed on the nodes within the graph to identify at least one risk propagation path that meets the risk propagation criteria. This allows for risk diffusion detection across multiple business domains and the determination of the risk impact scope, providing prompts and early warnings regarding risk diffusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0021] Figure 1 A flow chart of a risk detection method provided by an embodiment of the present disclosure;
[0022] Figure 2 This is a flow chart of a risk detection method provided by an embodiment of the present disclosure;
[0023] Figure 3 is a schematic structural diagram of a risk detection device provided by an embodiment of the present disclosure;
[0024] Figure 4 It is a structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0025] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0026] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0027] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0028] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0029] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0030] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0031] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0032] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0033] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0034] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0035] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0036] Any application field may include multiple business domains, and each business domain is respectively provided with a risk strategy. The risk strategy corresponding to any business domain is used to perform risk detection on business behaviors in the business domain, so as to detect and intercept risky business behaviors. Taking the e-commerce field as an example, business domains may include but are not limited to false advertising business domains, counterfeit and shoddy business domains, and illegal and prohibited business domains. Taking the false advertising business domain as an example, the false advertising business domain corresponds to at least one risk strategy, which is used to detect whether the business behavior of the merchant contains false advertising. It is understandable that the multiple business domains here may belong to the same application field. Different application fields are set with different business domains, and different business domains can be set with different risk strategies. The specific strategy content of the risk strategy is not limited here, and it can be set according to the actual situation of the business domain, as long as it can achieve risk detection of the business domain.
[0037] Different business domains may have relationships with each other. For example, the risk strategies of different business domains may have overlapping associated data tables and fields within those tables, or overlapping models or model parameters within the risk strategies of different business domains. These relationships between different business domains lead to risk propagation between them.
[0038] Current risk detection methods can only detect business behaviors in a certain business domain, and cannot solve the problem of correlation and propagation between different business domains. They cannot also detect the risk diffusion path and impact scope from the perspective of the entire network, resulting in the inability to timely detect and warn of the spread of risks.
[0039] In response to the above technical problems, the present disclosure provides a risk detection method, see Figure 1 , Figure 1 A flow chart of a risk detection method provided by an embodiment of the present disclosure is applicable to constructing a risk propagation heterogeneous graph including multiple business domains, conducting risk diffusion detection on different business domains based on risk cases, determining the risk diffusion path and the risk detection information of each node on the risk diffusion path, so as to achieve tracing of the risk source and propagation link. The method can be executed by a risk detection device, which can be implemented in the form of software and / or hardware, and optionally, by an electronic device, which can be a mobile terminal, a PC or a server, etc.
[0040] like Figure 1 As shown, the method includes:
[0041] S110 , obtaining a risk propagation heterogeneous graph, wherein the risk propagation heterogeneous graph includes nodes corresponding to a plurality of business domains, and the nodes are connected based on association relationships.
[0042] S120: Determine a first node according to the risk case, and determine at least one risk propagation path where the first node is located in the risk propagation heterogeneous graph.
[0043] S130: Perform anomaly detection on each node in the risk propagation path to determine risk anomaly detection information of each node in the risk propagation path.
[0044] In this embodiment, the risk propagation heterogeneous graph can be understood as a graph data structure formed by nodes corresponding to multiple business domains in the application field. The risk propagation heterogeneous graph includes multiple nodes and edges between nodes, wherein the edges between nodes are determined based on the association relationship between the nodes.
[0045] For example, the risk propagation heterogeneous graph can be labeled as Among them, V represents the node set, E represents the edge set, Characterizes the node type mapping function, Representation node v i The node type, ψ represents the relationship type mapping function, ψ(v i , v j ) represents node v i To node v j The relationship type.
[0046] Among them, the nodes corresponding to the business domain can be multi-type and / or multi-level data nodes in the risk strategy of the business domain. By constructing a risk propagation heterogeneous graph based on the nodes corresponding to multiple business domains, the data relationship between multiple business domains is represented in the form of a graph data structure, providing a data basis for risk diffusion detection between different business domains.
[0047] In some embodiments of the present disclosure, the nodes in the risk propagation heterogeneous graph include at least one of the following: the policy network of each of the business domains, the risk strategy in the policy network, the data factors in the risk strategy, the algorithm model of the data factors, the features in the algorithm model, the associated data table of the features, the fields of the associated data table and the data production source of the fields.
[0048] The policy network for each business domain can be understood as a network of risk strategies corresponding to multiple business domains in the application domain, where each business domain can correspond to at least one risk strategy. Accordingly, the policy network includes multiple risk strategies. A risk strategy can be understood as a strategy for determining whether a business behavior is risky. This strategy includes data factors for risk determination. For example, a risk is determined to exist when data factor A meets a set condition. Data factor A can be an indicator of the business behavior or an associated data item. The data factors in the risk strategy can be determined by an algorithmic model, which can be a machine learning model or a data model, without limitation. The algorithmic model includes multiple features, which can be input features of the algorithmic model. The algorithmic model processes these features to obtain data factors. The business system pre-stores an associated data table. The associated data table does not include feature values corresponding to the features in the algorithmic model, that is, the field content corresponding to the fields in the associated data table. The field content corresponding to the fields in the associated data table is obtained and stored from the data source of the field.
[0049] For example, for a data field, such as the number of account logins, the data source for this field may be a monitoring algorithm / acquisition module for the number of account logins. Here, the data field is considered a node, and the algorithm / module that generates the data content of the data field is considered a node. Multiple associated data tables formed by these data fields are considered a node. At least one associated data table forms a feature set, with the feature set being a node, or each feature in the feature set being a node. An algorithmic model is formed based on the features in the feature set, with the algorithmic model being a node. Data factors are derived from the algorithmic model, with the data factors being a node. Risk strategies that include the data factors are also considered a node, and the strategy network formed by risk strategies in different business domains is also considered a node. A risk propagation heterogeneous graph is constructed by using data nodes of different types and / or different levels within multiple business domains, thereby improving the comprehensiveness of the nodes in the risk propagation heterogeneous graph. The set of node types in the risk propagation heterogeneous graph can be labeled A, and this set of node types can include multiple node types in the risk propagation heterogeneous graph. In some embodiments of the present disclosure, the nodes in the risk propagation heterogeneous graph are not limited. For example, the nodes in the risk propagation heterogeneous graph may also include the policy network of each of the business domains, the risk strategy in the policy network, the data factors in the risk strategy, the associated data table of the data factors, the fields of the associated data table and the data production source of the fields.
[0050] The above-mentioned nodes are connected by edges, and the edges between the nodes are determined by the association relationship between the nodes. An edge is set between two nodes with an association relationship, and the connection relationship between different nodes can be different. The set of relationship types in the risk propagation heterogeneous graph can be marked as R, and the set of relationship types can include multiple relationship types of the risk propagation heterogeneous graph. Optionally, the relationship types between nodes in the risk propagation heterogeneous graph include but are not limited to parallel relationships, inclusion relationships, mapping relationships, combination relationships and production relationships. Exemplarily, the relationship between different risk strategies in the policy network can be a parallel relationship; the relationship between the policy network and any risk strategy can be an inclusion relationship; the relationship between the risk strategy and the data factor can be an inclusion relationship; the relationship between the data factor and the algorithm model can be a mapping relationship, and the relationship between the algorithm model and the feature can be an inclusion relationship or a combination relationship; the relationship between the field and the associated data table can be a mapping relationship, and the relationship between the field and the data production source can be a production relationship.
[0051] In some embodiments of the present disclosure, constructing a risk propagation heterogeneous graph may include: obtaining raw data from each business domain, which may be historical data from the business system corresponding to the business domain; performing entity identification from the raw data and creating a node corresponding to each entity; determining the relationships between the entities from the raw data, and setting edges between the nodes based on the relationships to form the risk propagation heterogeneous graph. Entity identification and relationship identification may be achieved using a pre-defined algorithm, which may be a machine learning algorithm.
[0052] In some embodiments of the present disclosure, a method for constructing a risk propagation heterogeneous graph may include: obtaining raw data from each business domain, identifying entities and relationships between entities from the raw data based on a machine learning method, and dynamically constructing a risk propagation heterogeneous graph. Specifically, a machine learning model with a heterogeneous graph construction function is pre-set, and the machine learning model may be a neural network model. The raw data from each business domain is input into the above-mentioned machine learning model, and the machine learning model is used to identify entities and relationships between entities in the raw data, and output a risk propagation heterogeneous graph. The risk propagation heterogeneous graph includes nodes corresponding to the identified entities and edges representing the relationships between the entities.
[0053] A risk case can be understood as a business behavior that is not detected by the policy network of multiple business domains and is risky. The risk case can be identified by other means, such as a risk case determined by a client's complaint operation or feedback operation, or a risk case identified manually, etc., which is not limited here.
[0054] The first node can be understood as a node in the risk propagation heterogeneous graph that has an association relationship with the risk case. Optionally, the business domain to which the risk case belongs is determined, and the node corresponding to the risk strategy of the business domain to which the risk case belongs is determined as the first node. Optionally, the nodes in the risk propagation heterogeneous graph may also include nodes corresponding to historical cases, and accordingly, there is an association relationship between the nodes corresponding to the historical cases and other nodes. Determine the similarity between the risk case and the historical cases in the risk propagation heterogeneous graph, and determine the historical case with the greatest similarity to the risk case, and determine the node corresponding to the historical case with the greatest similarity as the first node. It can be understood that the risk propagation heterogeneous graph can be continuously updated, for example, the risk propagation heterogeneous graph can be updated according to a preset time interval to ensure the accuracy and real-time performance of the risk propagation heterogeneous graph. Accordingly, the nodes of the historical cases in the risk propagation heterogeneous graph are updated along with the risk propagation heterogeneous graph.
[0055] By determining the first node corresponding to the risk case, risk diffusion detection is performed on the risk propagation heterogeneous graph with the first node as the starting node, and the risk propagation path in the risk propagation heterogeneous graph is determined. The risk propagation path can be understood as the trajectory of risk propagation and diffusion between different nodes.
[0056] The risk propagation path includes multiple propagation nodes. Risk propagation and diffusion can occur between associated propagation nodes, and the propagation nodes in the risk propagation path include the first node. Two nodes connected by an edge in the risk propagation path are adjacent propagation nodes, and the risk propagation strength between the adjacent propagation nodes satisfies strength constraint information, i.e., the risk propagation condition.
[0057] In some embodiments of the present disclosure, a first node is the first propagation node, and a search is performed on nodes in a risk propagation heterogeneous graph using the first node as the starting search node to determine other nodes that meet the risk propagation conditions as propagation nodes. Nodes with associated relationships constitute a risk propagation path. There may be at least one risk propagation path determined by the first node. Nodes in the risk propagation heterogeneous graph may be searched using at least one of a breadth-first search algorithm, a random walk, or a diffusion algorithm based on a fluid dynamics model to determine the propagation nodes in the risk propagation heterogeneous graph, thereby determining the risk propagation path.
[0058] Optionally, the adjacent nodes of the first node are determined in the risk propagation heterogeneous graph, wherein the adjacent nodes are nodes connected to the first node by edges. The risk propagation strength between the first node and the adjacent nodes is determined. If the risk propagation strength between the first node and the adjacent nodes satisfies the strength constraint information, that is, satisfies the risk propagation condition, the adjacent nodes of the first node are determined as propagation nodes. Based on the determined propagation nodes, the risk propagation strength between the determined propagation nodes and the adjacent nodes is determined. When the risk propagation condition is met, a new propagation node is determined, and so on, until there are no untraversed adjacent nodes for each determined propagation node, the search for the propagation node is stopped, and at least one risk propagation path is formed based on the association relationship between each of the propagation nodes.
[0059] The intensity of risk transmission between nodes can be understood as the degree or probability of risk spreading from one node to another. A greater intensity of risk transmission indicates a greater probability of risk spreading from one node to another when a risk exists at one node.
[0060] In the above process of determining the propagation node, the method for determining the risk propagation intensity between any two adjacent nodes in the risk propagation heterogeneous graph includes: obtaining the risk value of the second node, the risk resistance factor of the third node, and the risk propagation parameters between the second node and the third node, and determining the risk propagation intensity from the second node to the third node based on the risk value of the second node, the risk resistance factor of the third node, and the risk propagation parameters. Wherein, the second node and the third node are adjacent nodes in the risk propagation heterogeneous graph, the second node is any determined propagation node, and the third node is any node to be determined among the adjacent nodes of the second node. For example, the second node can be the first node, and the third node is the adjacent node of the first node; for example, the second node can be any propagation node v i , the third node is the propagation node v i adjacent nodes.
[0061] Each node in the risk propagation heterogeneous graph is respectively provided with a risk resistance factor. The risk resistance factor of any node can characterize the resistance of the node to the risk propagated by the adjacent nodes. The larger the risk resistance factor, the stronger the resistance of the node to the risk propagated by the adjacent nodes. In other words, the larger the risk resistance factor, the more difficult it is to transmit the risk to the node. The risk resistance factor of each node can be pre-set and can be set according to the node type of the node. For the third node, the risk resistance factor of the third node is read. Optionally, a risk resistance factor data set or a risk resistance factor data matrix is pre-set, and the risk resistance factor data set or the risk resistance factor data matrix stores the risk resistance factors respectively provided for each node in the risk propagation heterogeneous graph, and the risk resistance factor of the third node is read according to the node identifier of the third node.
[0062] Each node in the risk propagation heterogeneous graph is assigned a risk value. The risk value of each node can represent the degree of risk at that node, with a larger risk value representing a greater risk at that node. The risk value of each node can change dynamically. Optionally, the risk value can be read from a business system. The business system pre-sets a risk field for each node, and the field content corresponding to the risk field is the risk value. Accordingly, the risk value of each node is read from the risk field corresponding to each node. Optionally, a node risk identification model is pre-set, and the risk value of each node is identified using the node risk identification model to obtain the risk value of each node. The risk value of each node in the risk propagation heterogeneous graph can be identified based on the risk identification cycle to obtain the risk value of each node in the current time period. The risk value of each node in the risk propagation heterogeneous graph is stored, and the storage format is not limited here. For the second node, the risk value of the second node is obtained from the stored risk values based on the second node's identifier.
[0063] Risk propagation parameters between various node types are pre-set. These risk propagation parameters can represent the risk propagation coefficient between nodes of any node type. The magnitude of the risk propagation parameters is related to the node type for risk propagation. For example, the risk propagation parameters can be stored in the form of a node type propagation parameter matrix. The node type propagation parameter matrix includes risk propagation parameters between any node types, that is, risk propagation parameters between nodes of the same type and risk propagation parameters between different node types.
[0064] The risk propagation parameter between the second node and the third node is related to the node types of the adjacent second node and the third node. The node types of the second node and the third node are obtained, and the risk propagation parameter between the second node and the third node is obtained based on the matching of the node types of the second node and the third node in a node type propagation parameter matrix.
[0065] Optionally, the risk propagation intensity from the second node to the third node may be determined based on the risk value of the second node, the risk resistance factor of the third node, and the risk propagation parameter between the second node and the third node. Specifically, based on the risk resistance factor RR (v j ) Determine the propagation factor of the third node after resisting risk propagation, that is, 1-RR(v j ), based on the risk value R(v i ), the third node's transmission factor after resisting risk transmission is 1-RR(v j ) and the risk propagation parameter between the second node and the third node The product of determines the risk propagation intensity from the second node to the third node. i Represents the second node; v j Characterize the third node; Characterizes the node type of the second node, Characterizes the node type of the third node; 1-RR(v j ) represents the remaining propagable portion after the third node resists the risk. Optionally, the risk propagation intensity can also be adjusted using a global propagation adjustment coefficient α, where α is a hyperparameter and an adjustable parameter used to uniformly control the propagation amplitude.
[0066] The risk resistance factor of any node can be determined based on the node's inherent security numerical representation and redundant backup capability numerical representation. The node's inherent security numerical representation can reflect the node's ability to protect and isolate itself, and is a value between 0 and 1. The node's redundant backup capability numerical representation can reflect the robustness of the node's backup mechanism, and is a value between 0 and 1. The node's inherent security numerical representation and redundant backup capability numerical representation can be set based on the node's actual state, and can be based on pre-set data values. Alternatively, the node's inherent security numerical representation and redundant backup capability numerical representation can be dynamically adjusted based on node state information. A first indicator that can represent inherent security is periodically obtained, and the inherent security numerical representation is determined based on the first indicator. A second indicator that can represent redundant backup capability is periodically obtained, and the redundant backup capability numerical representation is determined based on the second indicator. For example, the second indicator can be whether the node performs data backup, and can be represented by an identifier such as 0 or 1. Different identifiers can correspond to different numerical representations of redundant backup capability. Similarly, the first indicator can be whether the node is equipped with isolation and / or isolation level. Different values of the first indicator correspond to different numerical representations of inherent security.
[0067] Optionally, the risk resistance factor of any node can be obtained by weighted calculation of the numerical representation of the node's own security and the numerical representation of its redundant backup capability, where the weight of the numerical representation of the node's own security is β, and the weight of the numerical representation of its redundant backup capability can be 1-β, where β is a weighting coefficient used to balance the effects between security and redundant backup.
[0068] In some embodiments of the present disclosure, the method for obtaining the risk propagation intensity may also include: obtaining the adaptive edge weight between the second node and the third node at the current moment, and determining the risk propagation intensity from the second node to the third node based on the adaptive edge weight between the second node and the third node at the current moment, the risk value of the second node, the risk resistance factor of the third node, and the risk propagation parameter; accordingly, the risk propagation intensity of the risk propagation from the second node to the third node can be obtained by the adaptive edge weight W between the second node and the third node at the current moment. ij , the risk value R(v i ), the third node's transmission factor after resisting risk transmission is 1-RR(v j ) and the risk propagation parameter between the second node and the third node The adaptive edge weight W between the second node and the third node at the current moment is determined by the product of ij It can represent the strength of the propagation relationship between the second node and the third node.
[0069] Among them, the adaptive edge weight between the second node and the third node at the current moment is obtained by updating the adaptive edge weight at the previous moment based on the correlation data of the risk values of the second node and the third node in the first time window. Specifically, the risk values of the second node and the third node in the first time window are obtained respectively, and the correlation data of the second node and the third node at the current moment and the correlation data change rate are determined by the risk values of the second node and the third node in the first time window respectively. The adaptive edge weight at the previous moment is updated based on the correlation data change rate of the second node and the third node at the current moment to obtain the adaptive edge weight between the second node and the third node at the current moment.
[0070] The first time window can be a time window from time tw to time t, and the risk value of the second node at each moment in the first time window and the average risk value of the second node in the first time window are obtained; the risk value of the third node at each moment in the first time window and the average risk value of the third node in the first time window are obtained; based on the risk value of the second node at each moment in the first time window and the average risk value of the second node in the first time window, the deviation of the risk value of the second node at any moment in the first time window from the risk value average and the standard deviation of the deviation are determined, and the risk value of the third node at each moment in the first time window and the risk value average of the third node in the first time window are determined, and the correlation data of the second node and the third node at the current moment are determined based on the deviation of the risk value of the second node at any moment in the first time window from the risk value average and the standard deviation of the deviation, and the deviation of the risk value of the third node at any moment in the first time window from the risk value average and the standard deviation of the deviation. Among them, the sum of the products of the risk numerical deviation of the second node (that is, the deviation of the risk numerical value at any moment from the risk numerical mean) and the risk numerical deviation of the third node at each moment in the first time window is determined, as well as the standard deviation of the risk numerical deviation of the third node and the product of the standard deviation of the risk numerical deviation of the second node are determined, and the correlation data of the second node and the third node at the current moment are determined based on the ratio of the sum of the products of the above-mentioned risk numerical deviations to the product of the standard deviations.
[0071] Based on the correlation data RC(v i ,v j ,t), and the historical correlation data mean RC of the second and third nodes avg (v i ,v j ) determines the correlation data change rate RD(v i ,vj ,t). Specifically, the product of the difference and the sensitivity parameter is used as the input parameter of the hyperbolic tangent function, and the output function value based on the hyperbolic tangent function is used as the correlation data change rate, which can be a value between -1 and 1. The mean of the historical correlation data of the second node and the third node can be, for example, the mean of the correlation data within a preset time period. The sensitivity parameter is a hyperparameter used to control the magnitude of the correlation change, and is usually a value greater than zero. The hyperbolic tangent function is used to map the input parameter to a value between -1 and 1.
[0072] Correspondingly, the updating method of the adaptive edge weight between the second node and the third node at the current time t can be: obtaining the correlation data change rate RD (v i ,v j ,t-1), and RD(v i ,v j ,t) is determined in the same way and will not be described here; the correlation data change rate RD(v i ,v j ,t-1) to determine the weight adjustment coefficient, that is, 1+λ·RD(v i ,v j ,t-1), the adaptive edge weight W between the second node and the third node at the previous time t-1 is adjusted by the weight adjustment coefficient ij (t-1) is adjusted to obtain the adaptive edge weight W between the second node and the third node at the current time t ij (t), for example, can be obtained by adjusting the weight coefficient and W ij The product of (t-1) is used as the adaptive edge weight W at the current time t ij (t). λ is the learning rate, which is a hyperparameter.
[0073] In the disclosed embodiment, by obtaining the adaptive edge weight between the second node and the third node at the current moment, the risk propagation intensity from the second node to the third node is adaptively adjusted to improve the accuracy of the risk propagation intensity.
[0074] There may be at least one third node adjacent to the second node, and the risk propagation intensity of the risk propagated from the second node to each third node is determined. A determination is made as to whether the risk propagation intensity satisfies intensity constraint information, the intensity constraint information including an intensity threshold. If the risk propagation intensity is greater than or equal to the intensity threshold, the third node is determined to be the propagation node, and further propagation nodes are determined based on the third node. If the risk propagation intensity is less than the intensity threshold, the third node is determined not to be a propagation node.
[0075] When multiple propagation nodes are identified, at least one risk propagation path is formed based on the relationships between them. Specifically, if an edge exists between two propagation nodes, the two propagation nodes are determined to be in the same risk propagation path. The identified propagation nodes are then traversed to determine the risk propagation path to which each propagation node belongs. By starting the search in the risk propagation heterogeneous graph with the first node identified by the risk case as the starting node and using the risk propagation strength between nodes as the judgment criterion, at least one risk propagation path is determined, thereby enabling the detection of risk diffusion.
[0076] It can be understood that since the risk propagation heterogeneous graph includes nodes corresponding to different business domains, the same risk propagation path may include nodes corresponding to the same business domain or nodes corresponding to different business domains, thereby realizing risk diffusion detection between different business domains.
[0077] Each risk propagation path may include multiple propagation nodes. Multiple propagation nodes in at least one risk propagation path may include abnormal nodes, resulting in inaccurate risk case detection. These abnormal nodes may be one or more. By performing anomaly detection on multiple propagation nodes in at least one risk propagation path, risk anomaly detection information for each node in the risk propagation path is determined. The abnormal nodes are identified based on the risk anomaly detection information for each node, and timely adjustments or warnings are performed on the abnormal nodes to reduce the impact of risks on different business domains.
[0078] In some embodiments of the present disclosure, a pre-configured anomaly detection model can be used to perform anomaly detection on multiple nodes in a risk propagation path, outputting risk anomaly detection information for each node in the risk propagation path. This risk anomaly detection information can be in the form of a data value, where a larger data value indicates a greater probability of an anomaly at that node.
[0079] Visually display at least one risk propagation path and the risk anomaly detection information of each node in the risk propagation path to provide risk warnings to operators.
[0080] The technical solution of the disclosed embodiments integrates and uniformly models nodes across different business domains by constructing a risk propagation heterogeneous graph comprising nodes corresponding to multiple business domains. This provides a data foundation for risk diffusion detection and holistic risk monitoring across multiple business domains. For each risk case, based on the first node corresponding to the risk case in the risk propagation heterogeneous graph, a search is performed on the nodes within the graph to identify at least one risk propagation path that meets the risk propagation criteria. This allows for risk diffusion detection across multiple business domains and the determination of the risk impact scope, providing prompts and early warnings regarding risk diffusion.
[0081] In some embodiments of the present disclosure, each of the risk propagation paths is evaluated for importance in at least one dimension to obtain an evaluation value of the risk propagation path in each dimension; the importance evaluation result of the risk propagation path is determined based on the evaluation value of the risk propagation path in each dimension; wherein, the at least one dimension of the importance evaluation includes at least one of the following: path sensitivity, path diversity, and path criticality.
[0082] The risk propagation path can be understood as a path formed by multiple propagation nodes and the edges between them. For example, the node sequence included in the risk propagation path is v1, v2…v |p| Where |p| is the number of nodes in the risk propagation path.
[0083] Optionally, a path evaluation model can be set up to transmit the node sequence included in each risk propagation path to the path evaluation model to obtain the evaluation values of the risk propagation path in different dimensions. The evaluation values of the risk propagation path in each dimension are then fused to obtain the risk propagation path importance evaluation result. The fusion process here can be weighted fusion.
[0084] Among them, path sensitivity can be understood as the degree of risk transmission sensitivity in each node on the risk transmission path, path diversity can be understood as the degree of diversity of node types and relationship types on the risk transmission path; path criticality can be understood as the degree of importance of each node on the risk transmission path in the risk transmission heterogeneous graph.
[0085] In some embodiments of the present disclosure, the evaluation value of the risk propagation path in the path sensitivity dimension is determined based on the length between the propagation nodes in the risk propagation path and the risk propagation intensity between the propagation nodes in the risk propagation path; the risk propagation path can be understood as multiple path segments, each path segment may include two adjacent propagation nodes, and accordingly, the evaluation value of the risk propagation path in the path sensitivity dimension is positively correlated with the risk propagation intensity of each path segment (i.e., the risk propagation intensity between two adjacent propagation nodes), and negatively correlated with the length between the propagation nodes in the risk propagation path.
[0086] Optionally, the evaluation value of the risk propagation path in the path sensitivity dimension is determined based on the product of the risk propagation intensities of each path segment and the attenuation impact data determined based on the length of the path segment (i.e., the path segment between two adjacent propagation nodes). Exemplarily, the evaluation value of the risk propagation path in the path sensitivity dimension is obtained as follows: for path segment i in the risk propagation path, the attenuation impact data of path segment i is determined based on the length or distance of path segment i in the risk propagation path and the distance penalty coefficient, i.e., 1-γ·d i , where di is the length or distance of path segment i in the risk transmission path, d i It can be based on node v i With node v i+1 The similarity between the nodes is determined, the greater the similarity, d i The smaller. Or, d i It can be based on node v i With node v i+1 The Euclidean distance between nodes; γ is the distance penalty coefficient, which can be a coefficient between 0 and 1; the path segment i can be understood as the node v i With node v i+1 propagation path segments between nodes v i With node v i+1 The risk propagation intensity between can be used as the risk propagation intensity of path segment i in the risk propagation path; based on the attenuation impact data of path segment i, the risk propagation intensity RI (v i →v i+1 ) and the edge reliability coefficient β of path segment i i The product of is used as the evaluation value corresponding to the path segment i. The risk propagation path p may include multiple path segments i. The product of the evaluation values corresponding to the multiple path segments i is used as the evaluation value in the path sensitivity dimension. The edge reliability coefficient β i It can be based on node v i With node v i+1 The node type and / or relationship type of the risk determines the risk propagation path.
[0087] In some embodiments of the present disclosure, the evaluation value of the risk propagation path in the path diversity dimension is determined based on the proportion of each node type and the proportion of each relationship type in the risk propagation path. Specifically, the proportion of each node type in the risk propagation path can be determined based on the number of different node types in the risk propagation path and the total number of all node types in the risk propagation heterogeneous graph, the proportion of each relationship type can be determined based on the number of different relationship nodes in the risk propagation path and the total number of all relationship types in the risk propagation heterogeneous graph, and the evaluation value of the risk propagation path in the path diversity dimension can be determined based on the product of the proportion of each node type in the risk propagation path and the proportion of each relationship type.
[0088] In some embodiments of the present disclosure, the evaluation value of the risk propagation path in the path criticality dimension is determined based on the centrality value representation and the bridging value representation of each propagation node in the risk propagation path. The centrality value representation of the propagation node can reflect the core degree or influence of the propagation node in the risk propagation path, and the bridging value of the propagation node can reflect the importance of the connection between the propagation node and other propagation nodes. The centrality value of the propagation node can be determined based on the position of the propagation node in the risk propagation path; the bridging value of the propagation node can be determined based on the number of edges connected to the propagation node. Exemplarily, the evaluation value of the risk propagation path in the path criticality dimension can be determined by determining the criticality evaluation data of multiple propagation nodes in the risk propagation path, and obtaining the evaluation value of the risk propagation path in the path criticality dimension based on the normalized values of the criticality evaluation data of multiple propagation nodes, wherein the criticality evaluation data of each propagation node can be obtained based on the weighted calculation of the centrality value representation, centrality value weight, bridging value representation and bridging value weight of the propagation node. Determine the sum of the critical evaluation data of multiple propagation nodes, and use the ratio of the sum of the critical evaluation data of multiple propagation nodes to the number of nodes in the risk propagation path as the evaluation value of the risk propagation path in the critical dimension of the path.
[0089] Based on the evaluation data of risk transmission paths in various dimensions such as path sensitivity, path diversity and path criticality, the evaluation data of path sensitivity, path diversity and path criticality are weighted and fused to obtain the evaluation results of the importance of risk transmission paths, among which the sum of the weight coefficients of path sensitivity, path diversity and path criticality is 1.
[0090] In the disclosed embodiment, each risk propagation path is evaluated quantitatively in at least one dimension, such as path sensitivity, path diversity, and path criticality, to determine the importance evaluation results of the risk propagation path across multiple dimensions. The risk propagation paths are ranked based on the importance evaluation results, and the importance evaluation results and ranking results of the risk propagation paths are displayed. This provides prompt information on the importance of risk diffusion along each risk propagation path, facilitating operators to prioritize and adjust risk propagation paths ranked higher in importance evaluation results, thereby ensuring the safety of important paths.
[0091] Figure 2 This is a flow chart of a risk detection method provided by an embodiment of the present disclosure. Based on the above embodiment, the anomaly detection of each node in the risk propagation path is optimized. The method specifically includes:
[0092] S210 , obtaining a risk propagation heterogeneous graph, wherein the risk propagation heterogeneous graph includes nodes corresponding to a plurality of business domains, and the nodes are connected based on association relationships.
[0093] S220: Determine a first node according to the risk case, and determine at least one risk propagation path where the first node is located in the risk propagation heterogeneous graph.
[0094] S230. Perform at least one dimension of anomaly detection on each node in the risk propagation path to obtain an anomaly detection value of each dimension, and obtain risk anomaly detection information of the node based on the anomaly detection value of at least one dimension.
[0095] The at least one dimension of anomaly detection includes at least one of the following: a structural dimension, a timing dimension, and a propagation dimension.
[0096] In some embodiments of the present disclosure, anomaly detection in the structural dimension can be understood as a detection to determine whether there is an anomaly in the number of edges between a node and nodes of the same type. The anomaly detection value of any node in the structural dimension can represent the degree of anomaly of the node in the structural dimension. Optionally, the anomaly detection value of any node in the structural dimension is determined based on the number of edges of the node and the average number of edges of nodes of the same type as the node; specifically, for the node to be detected v i , get the node v to be detected i The number of edges E(v i ), and obtain the risk propagation heterogeneous graph with the node v to be detected i The average number of edges of nodes of the same type Get the node v to be detected i The number of edges E(v i ) and the average number of edges of the same type of nodes, and the standard deviation of the number of edges of the same type of nodes. The node to be detected v is determined based on the ratio of the square of the above difference to the standard deviation of the number of edges of the same type of nodes. i Anomaly detection value in the structural dimension.
[0097] In some embodiments of the present disclosure, anomaly detection in the time series dimension can be understood as detecting whether the characteristic value of a node is abnormal in the second time window. The anomaly detection value of any node in the time series dimension can represent the degree of abnormality of the node in the time series dimension. The characteristic value of a node can be understood as the field data value of the node itself. For example, if the node is the number of historical logins, the characteristic value of the node can be a specific value of the number of historical logins, such as 100.
[0098] Optionally, the anomaly detection value of any of the nodes in the time series dimension is determined based on the observed characteristic value and the predicted characteristic value of the node at each time point in the second time window, wherein the second time window may be a time window formed from time t-τ to time t, and τ is the size of the second time window. Specifically, the anomaly detection value of the node in the time series dimension is determined based on the average of the absolute deviations between the observed characteristic value and the predicted characteristic value corresponding to each time point in the second time window of any node, wherein for any time point k in the second time window, the absolute value of the deviation between the observed characteristic value and the predicted characteristic value corresponding to the time point k, i.e., the absolute deviation, is obtained, and the ratio of the sum of the above absolute deviations corresponding to multiple time points k to the size τ of the second time window is determined as the anomaly detection value of the node in the time series dimension. Among them, the observed characteristic value is the characteristic value of the node read from the business system; the predicted characteristic value is obtained based on the change trend of the historical observed characteristic value of the node; when the change trend of the historical observed characteristic value remains stable, the predicted characteristic value is equal to the historical observed characteristic value; when the change trend of the historical observed characteristic value is increasing or decreasing, the first change rate is determined based on the change trend of the historical observed characteristic value, and the predicted characteristic value is determined based on the historical observed characteristic value and the first change rate.
[0099] In some embodiments of the present disclosure, anomaly detection in the propagation dimension can be understood as detecting whether the cumulative risk value of a node is abnormal. The anomaly detection value of any node in the propagation dimension can represent the degree of abnormality of the node in the propagation dimension.
[0100] Optionally, the anomaly detection value of any node in the propagation dimension is determined based on the actual cumulative risk value and the predicted cumulative risk value of the node, wherein the actual cumulative risk value of the node is determined based on the risk value of the node and the risk propagation strength of the adjacent nodes, wherein the adjacent nodes can be understood as all nodes pointing to the node v j The predecessor node.
[0101] The predicted cumulative risk value is obtained based on a change trend prediction of the historical actual cumulative risk value of the node. The historical actual cumulative risk value is obtained, and the change trend of the historical actual cumulative risk value is determined based on the historical actual cumulative risk value.
[0102] When the changing trend of the historical actual cumulative risk value remains stable, the predicted cumulative risk value is equal to the historical actual cumulative risk value; when the changing trend of the historical actual cumulative risk value is increasing or decreasing, the second change rate is determined based on the changing trend of the historical actual cumulative risk value, and the predicted cumulative risk value is determined based on the historical actual cumulative risk value and the second change rate.
[0103] The anomaly detection value of any node in the propagation dimension may be determined based on the ratio of the absolute deviation between the actual cumulative risk value and the predicted cumulative risk value to the predicted cumulative risk value. The absolute deviation between the actual cumulative risk value and the predicted cumulative risk value may be understood as the absolute value of the deviation between the actual cumulative risk value and the predicted cumulative risk value.
[0104] For any node, after determining the anomaly detection values corresponding to the node in the structural dimension, timing dimension and propagation dimension respectively, the anomaly detection values corresponding to the node in the structural dimension, timing dimension and propagation dimension respectively are fused to obtain the risk anomaly detection information of the node, where the sum of the weight coefficients of the structural dimension, timing dimension and propagation dimension is 1.
[0105] In some embodiments of the present disclosure, risk anomaly detection information of the node is obtained based on the anomaly detection value of at least one dimension, including: fusing the anomaly detection value of the at least one dimension to obtain first risk anomaly detection information; determining a synergistic enhancement coefficient based on similarity data between the node and adjacent nodes, and enhancing the first risk anomaly detection information based on the synergistic enhancement coefficient to obtain second risk anomaly detection information of the node.
[0106] The first risk anomaly detection information can be obtained by fusing the anomaly detection values of the structural dimension, the time series dimension, and the propagation dimension to obtain a fusion value. The synergistic enhancement coefficient can be understood as a coefficient for enhancing the first risk anomaly detection information, and enhancing the similarity between the nodes to improve the accuracy of the risk anomaly detection information of the node. The method for determining the synergistic enhancement coefficient includes: obtaining the node v i The adjacent node v j The risk anomaly detection information of node v is based on the activation function i The adjacent node v j The risk anomaly detection information of node v is mapped and processed to obtain i The adjacent node v j The risk anomaly detection information of the node v is mapped to the value of the activation function, which includes but is not limited to the sigmoid function. i and adjacent node v j Similarity data for node v i At least one corresponding adjacent node v j , determine the node v i The adjacent node v j The sum of the products of the mapping value of the activation function and the similarity data of the risk anomaly detection information is M, and the corresponding synergistic enhancement coefficient is 1+M.
[0107] The adjacent nodes here can be understood as all nodes pointing to node vj For a predecessor node, the second risk anomaly detection information may be the product of the synergy enhancement coefficient and the first risk anomaly detection information.
[0108] The collaborative enhancement coefficient is determined by the risk anomaly detection information of adjacent nodes and the similar data between adjacent nodes, and the risk anomaly detection information of the detected node is coordinated and enhanced. The mutual influence between nodes is comprehensively considered, and the accuracy of the risk anomaly detection information of the node is improved.
[0109] The technical solution provided by the embodiments of the present disclosure improves the comprehensiveness and diversity of anomaly detection by performing multi-dimensional anomaly detection on each node in the risk propagation path. The risk anomaly detection information for the node obtained by fusing the multi-dimensional anomaly detection values improves the accuracy of node anomaly detection. On this basis, the risk anomaly detection information for each node in the risk propagation path is displayed, and the risk anomaly levels of the nodes are displayed and compared to determine the source of the risk, facilitate timely warning and adjustment of abnormal nodes, and improve safety.
[0110] Figure 3 This is a schematic diagram of the structure of a risk detection device provided by an embodiment of the present disclosure, such as Figure 3 As shown, the device includes: a heterogeneous graph acquisition module 310, a risk propagation path determination module 320 and a risk anomaly detection module 330.
[0111] A heterogeneous graph acquisition module 310 is used to acquire a risk propagation heterogeneous graph, wherein the risk propagation heterogeneous graph includes nodes corresponding to multiple business domains, and the nodes are connected based on association relationships;
[0112] A risk propagation path determination module 320 is configured to determine a first node according to a risk case, and determine at least one risk propagation path where the first node is located in the risk propagation heterogeneous graph;
[0113] The risk anomaly detection module 330 is configured to perform anomaly detection on each node in the risk propagation path and determine risk anomaly detection information of each node in the risk propagation path.
[0114] The technical solution provided by the disclosed embodiments integrates and uniformly models nodes across different business domains by constructing a risk propagation heterogeneous graph comprising nodes corresponding to multiple business domains. This provides a data foundation for risk diffusion detection and holistic risk monitoring across multiple business domains. For each risk case, based on the first node corresponding to the risk case in the risk propagation heterogeneous graph, a search is performed on the nodes within the graph to identify at least one risk propagation path that meets the risk propagation criteria. This allows for risk diffusion detection across multiple business domains and the determination of the risk impact scope, providing prompts and early warnings regarding risk diffusion.
[0115] Based on the above embodiments, optionally, the nodes in the risk propagation heterogeneous graph include at least one of the following: the policy network of each business domain, the risk strategy in the policy network, the data factors in the risk strategy, the algorithm model of the data factors, the features in the algorithm model, the associated data table of the features, the fields of the associated data table and the data production source of the fields.
[0116] Based on the above embodiment, optionally, the risk propagation path includes multiple propagation nodes, and the propagation nodes include the first node; the risk propagation strength between adjacent propagation nodes satisfies strength constraint information.
[0117] Optionally, the risk propagation path determination module 320 is used to: obtain the risk value of the second node, the risk resistance factor of the third node, and the risk propagation parameters between the second node and the third node, wherein the second node and the third node are adjacent nodes in the risk propagation heterogeneous graph, the second node is the propagation node, and the risk propagation parameters are related to the node types of the adjacent second node and the third node; determine the risk propagation intensity from the second node to the third node based on the risk value of the second node, the risk resistance factor of the third node, and the risk propagation parameters.
[0118] Optionally, the strength constraint information includes a strength threshold;
[0119] The risk propagation path determination module 320 is also used to: when the risk propagation intensity is greater than or equal to the intensity threshold, determine the third node as the propagation node; wherein, the first node is the first propagation node, and other propagation nodes are determined with the first node as the starting search node; each of the propagation nodes forms at least one risk propagation path based on the association relationship between nodes.
[0120] Optionally, the risk propagation path determination module 320 is further configured to: obtain an adaptive edge weight between the second node and the third node at the current moment, and determine the risk propagation intensity from the second node to the third node based on the adaptive edge weight between the second node and the third node at the current moment, the risk value of the second node, the risk resistance factor of the third node, and the risk propagation parameter;
[0121] The adaptive edge weight between the second node and the third node at the current moment is obtained by updating the adaptive edge weight at the previous moment based on correlation data of risk values of the second node and the third node within the first time window.
[0122] Based on the above embodiment, optionally, the risk anomaly detection module 330 is used to: perform at least one dimension of anomaly detection on each node in the risk propagation path, obtain an anomaly detection value of each dimension, and obtain risk anomaly detection information of the node based on the anomaly detection value of at least one dimension; wherein, the at least one dimension of anomaly detection includes at least one of the following: structural dimension, timing dimension, and propagation dimension.
[0123] Optionally, the anomaly detection value of any of the nodes in the structural dimension is determined based on the number of edges of the node and the average number of edges of nodes of the same type as the node;
[0124] The anomaly detection value of any of the nodes in the time series dimension is determined based on the observed characteristic value and the predicted characteristic value of the node at each time point in the second time window, wherein the predicted characteristic value is obtained based on the change trend prediction of the historical observed characteristic value of the node;
[0125] The anomaly detection value of any node in the propagation dimension is determined based on the actual cumulative risk value and the predicted cumulative risk value of the node, wherein the actual cumulative risk value of the node is determined based on the risk value of the node and the risk propagation intensity of the adjacent nodes, and the predicted cumulative risk value is obtained based on the change trend prediction of the historical actual cumulative risk value of the node.
[0126] Optionally, the risk anomaly detection module 330 is also used to: fuse the anomaly detection values of the at least one dimension to obtain first risk anomaly detection information; determine a synergistic enhancement coefficient based on the similarity data between the node and the adjacent nodes, and enhance the first risk anomaly detection information based on the synergistic enhancement coefficient to obtain second risk anomaly detection information of the node.
[0127] Based on the above embodiment, the device may optionally further include a path evaluation module, which is used to: perform importance evaluation on each of the risk propagation paths in at least one dimension to obtain an evaluation value of the risk propagation path in each dimension; determine an importance evaluation result of the risk propagation path based on the evaluation value of the risk propagation path in each dimension; wherein, at least one dimension of the importance evaluation includes at least one of the following: path sensitivity, path diversity, and path criticality.
[0128] Optionally, the evaluation value of the risk propagation path in the path sensitivity dimension is determined based on the length between propagation nodes in the risk propagation path and the risk propagation intensity between propagation nodes in the risk propagation path;
[0129] The evaluation value of the risk propagation path in the path diversity dimension is determined based on the proportion of each node type and the proportion of each relationship type in the risk propagation path;
[0130] The evaluation value of the risk propagation path in the path criticality dimension is determined based on the centrality numerical representation and the bridging numerical representation of each propagation node in the risk propagation path.
[0131] The risk detection device provided in the embodiments of the present disclosure can execute the risk detection method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0132] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.
[0133] Figure 4 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Figure 4 , which shows an electronic device (eg Figure 4 The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0134] like Figure 4 As shown, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the electronic device 500 are also stored in the RAM 503. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An edit / output (I / O) interface 505 is also connected to the bus 504.
[0135] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 4 The electronic device 500 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0136] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0137] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0138] The electronic device provided in the embodiment of the present disclosure and the risk detection method provided in the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0139] An embodiment of the present disclosure provides a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the risk detection method provided in the above embodiment.
[0140] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0141] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0142] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0143] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:
[0144] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device is enabled to: obtain a risk propagation heterogeneous graph, wherein the risk propagation heterogeneous graph includes nodes corresponding to multiple business domains, and each of the nodes is connected based on an association relationship; determine a first node based on a risk case, and determine at least one risk propagation path where the first node is located in the risk propagation heterogeneous graph; perform anomaly detection on each node in the risk propagation path, and determine risk anomaly detection information of each node in the risk propagation path.
[0145] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0146] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0147] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."
[0148] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0149] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0150] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0151] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0152] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A risk detection method, characterized in that: include: Obtaining a risk propagation heterogeneous graph, wherein the risk propagation heterogeneous graph includes nodes corresponding to multiple business domains, and the nodes are connected based on association relationships; Determine a first node according to the risk case, and determine at least one risk propagation path where the first node is located in the risk propagation heterogeneous graph; Anomaly detection is performed on each node in the risk propagation path to determine risk anomaly detection information of each node in the risk propagation path.
2. The method according to claim 1, characterized in that The nodes in the risk propagation heterogeneous graph include at least one of the following: the policy network of each business domain, the risk strategy in the policy network, the data factors in the risk strategy, the algorithm model of the data factors, the features in the algorithm model, the associated data table of the features, the fields of the associated data table and the data production source of the fields.
3. The method according to claim 1, characterized in that The risk propagation path includes multiple propagation nodes, including the first node; the risk propagation intensity between adjacent propagation nodes satisfies intensity constraint information.
4. The method according to claim 3, characterized in that The method for determining the risk propagation intensity between any two adjacent nodes in the risk propagation heterogeneous graph includes: Obtaining a risk value of a second node, a risk resistance factor of a third node, and a risk propagation parameter between the second node and the third node, wherein the second node and the third node are adjacent nodes in the risk propagation heterogeneous graph, the second node is the propagation node, and the risk propagation parameter is related to the node type of the adjacent second node and the third node; The risk propagation intensity from the second node to the third node is determined based on the risk value of the second node, the risk resistance factor of the third node, and the risk propagation parameter.
5. The method according to claim 4, characterized in that The intensity constraint information includes an intensity threshold; When the risk propagation intensity is greater than or equal to the intensity threshold, the third node is determined as the propagation node; wherein, the first node is the first propagation node, and other propagation nodes are determined with the first node as the starting search node; each of the propagation nodes forms at least one risk propagation path based on the association relationship between nodes.
6. The method according to claim 4, characterized in that The method further comprises: Obtaining an adaptive edge weight between the second node and the third node at the current moment, and determining a risk propagation intensity from the second node to the third node based on the adaptive edge weight between the second node and the third node at the current moment, a risk value of the second node, a risk resistance factor of the third node, and the risk propagation parameter; The adaptive edge weight between the second node and the third node at the current moment is obtained by updating the adaptive edge weight at the previous moment based on correlation data of risk values of the second node and the third node within the first time window.
7. The method according to claim 1, characterized in that Perform anomaly detection on each node in the risk propagation path, including: Performing at least one dimension of anomaly detection on each node in the risk propagation path to obtain an anomaly detection value of each dimension, and obtaining risk anomaly detection information of the node based on the anomaly detection value of the at least one dimension; The at least one dimension of anomaly detection includes at least one of the following: a structural dimension, a timing dimension, and a propagation dimension.
8. The method according to claim 7, characterized in that An anomaly detection value of any of the nodes in the structural dimension is determined based on the number of edges of the node and the average number of edges of nodes of the same type as the node; The anomaly detection value of any of the nodes in the time series dimension is determined based on the observed characteristic value and the predicted characteristic value of the node at each time point in the second time window, wherein the predicted characteristic value is obtained based on the change trend prediction of the historical observed characteristic value of the node; The anomaly detection value of any node in the propagation dimension is determined based on the actual cumulative risk value and the predicted cumulative risk value of the node, wherein the actual cumulative risk value of the node is determined based on the risk value of the node and the risk propagation intensity of the adjacent nodes, and the predicted cumulative risk value is obtained based on the change trend prediction of the historical actual cumulative risk value of the node.
9. The method according to claim 7, characterized in that Obtaining risk anomaly detection information of the node based on an anomaly detection value of at least one dimension includes: Performing fusion processing on the anomaly detection values of the at least one dimension to obtain first risk anomaly detection information; A synergistic enhancement coefficient is determined based on similarity data between the node and adjacent nodes, and the first risk anomaly detection information is enhanced based on the synergistic enhancement coefficient to obtain second risk anomaly detection information of the node.
10. The method according to claim 1, characterized in that The method further comprises: Performing an importance evaluation of at least one dimension on each of the risk propagation paths to obtain an evaluation value of the risk propagation path in each dimension; and determining an importance evaluation result of the risk propagation path based on the evaluation value of the risk propagation path in each dimension; The at least one dimension of the importance evaluation includes at least one of the following: path sensitivity, path diversity, and path criticality.
11. The method according to claim 10, characterized in that The evaluation value of the risk propagation path in the path sensitivity dimension is determined based on the length between the propagation nodes in the risk propagation path and the risk propagation intensity between the propagation nodes in the risk propagation path; The evaluation value of the risk propagation path in the path diversity dimension is determined based on the proportion of each node type and the proportion of each relationship type in the risk propagation path; The evaluation value of the risk propagation path in the path criticality dimension is determined based on the centrality numerical representation and the bridging numerical representation of each propagation node in the risk propagation path.
12. A risk detection device, characterized in that: include: A heterogeneous graph acquisition module is used to acquire a risk propagation heterogeneous graph, wherein the risk propagation heterogeneous graph includes nodes corresponding to multiple business domains, and the nodes are connected based on association relationships; A risk propagation path determination module, configured to determine a first node according to a risk case, and determine at least one risk propagation path where the first node is located in the risk propagation heterogeneous graph; The risk anomaly detection module is used to perform anomaly detection on each node in the risk propagation path and determine the risk anomaly detection information of each node in the risk propagation path.
13. An electronic device, characterized in that: The electronic device comprises: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the risk detection method according to any one of claims 1 to 11.
14. A storage medium comprising computer executable instructions, wherein the computer executable instructions are used to perform the risk detection method according to any one of claims 1 to 11 when executed by a computer processor.