Risk prediction method and device, equipment, storage medium and product
By constructing a target graph and utilizing adaptive neighborhood sampling and multi-relation graph convolution strategies, entities are aggregated and risk features are analyzed, solving the problem of inefficient identification of risk groups for illegal transactions in existing technologies and achieving efficient identification of illegal transaction risks.
Patent Information
- Application Number
- CN202511223567.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-10-21
AI Technical Summary
Existing technologies are unable to efficiently identify groups at risk of illegal transactions based on massive amounts of data, and manual screening methods are inefficient and cannot effectively uncover illegal transaction behaviors.
By constructing a target graph, and using an adaptive neighborhood sampling strategy and a multi-relationship graph convolution strategy to aggregate entities, combined with risk feature analysis, risk clusters are identified.
It enables efficient identification of groups at risk of illegal transactions, fully explores the hidden information in big data, and improves the efficiency of identifying illegal transaction behavior.
Smart Images

Figure CN120823041A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of big data, and in particular to a risk prediction method, device, equipment, storage medium and product. Background Art
[0002] With the development of the Internet, transaction methods (interactive behaviors in the financial field) have become increasingly diversified, and the participants and scenarios of transactions have become more complicated. However, various new illegal trading behaviors have also emerged. Among them, illegal trading behaviors in the form of groups will have a greater impact on the healthy development of the financial industry.
[0003] In this context, to ensure the healthy development of the financial industry, rigorous manual screening is often deployed to identify groups engaging in illegal trading. However, manual screening is a relatively backward and inefficient method, unable to effectively identify groups at risk of illegal trading based on massive amounts of big data. Summary of the Invention
[0004] The present application provides a risk prediction method, apparatus, device, storage medium and product to solve the problem of being unable to efficiently identify groups with risks of illegal transactions based on massive big data.
[0005] In a first aspect, the present application provides a risk prediction method, comprising:
[0006] Obtain interaction data of multiple objects and build a target graph based on the interaction data of multiple objects;
[0007] Based on the adaptive neighborhood sampling strategy and multi-relation graph convolution strategy, multiple entities in the target graph are aggregated to obtain multiple initial clusters;
[0008] Risk characteristic analysis is performed on each initial cluster to obtain the risk probability of each initial cluster, and the initial cluster whose risk probability reaches the probability threshold is regarded as the risk cluster.
[0009] In one embodiment, a target graph is constructed based on interaction data of multiple objects, including:
[0010] Extract multiple entities from the interaction data of multiple objects, and obtain the association relationships between the entities by mining the association relationships of the interaction data of multiple objects;
[0011] Generate graph data based on multiple entities and the relationships between them;
[0012] Obtain a multi-layer message passing network architecture, combine the multi-layer message passing network architecture with graph data, and obtain an initial graph;
[0013] The initial map is optimized to obtain the target map.
[0014] In one embodiment, performing a map optimization on the initial map to obtain a target map includes:
[0015] Performing topological enhancement and feature enhancement on the initial atlas to obtain a first atlas;
[0016] Performing progressive training and adversarial regularization training on the first graph to obtain a second graph;
[0017] The second atlas is compressed to obtain a target atlas.
[0018] In one embodiment, based on the adaptive neighborhood sampling strategy and the multi-relation graph convolution strategy, multiple entities in the target graph are aggregated to obtain multiple initial clusters, including:
[0019] Based on the adaptive neighborhood sampling strategy, multiple entities in the target graph are aggregated to obtain multiple first clusters;
[0020] Based on the multi-relationship graph convolution strategy, multiple entities in the target graph are aggregated to obtain multiple second clusters;
[0021] A plurality of first clusters and a plurality of second clusters are fused to obtain a plurality of initial clusters.
[0022] In one embodiment, risk characteristics analysis is performed on each initial cluster to obtain the risk probability of each initial cluster, including:
[0023] For each initial cluster, perform node-level risk feature analysis and graph-level risk feature analysis to obtain the cluster features of the initial cluster;
[0024] Score the cluster features to obtain the cluster score of the initial cluster;
[0025] Based on the cluster scores, the risk probability of the initial cluster is determined.
[0026] In one embodiment, node-level risk feature analysis and graph-level risk feature analysis are performed on each initial cluster to obtain cluster features of the initial cluster, including:
[0027] For each initial cluster, node-level risk feature analysis and graph-level risk feature analysis are performed on each entity in the initial cluster to obtain the entity-level individual features and graph-level individual features of each entity;
[0028] Perform feature aggregation on the entity-level individual features and the map-level individual features of each entity to obtain the individual features of each entity;
[0029] The individual features of each entity in the initial cluster are aggregated to obtain the cluster features of the initial cluster.
[0030] In a second aspect, the present application provides a risk prediction device, comprising:
[0031] A target graph construction module is used to obtain the interaction data of multiple objects and construct a target graph based on the interaction data of multiple objects;
[0032] The initial cluster acquisition module is used to aggregate multiple entities in the target graph based on the adaptive neighborhood sampling strategy and the multi-relationship graph convolution strategy to obtain multiple initial clusters;
[0033] The risk cluster prediction module is used to analyze the risk characteristics of each initial cluster, obtain the risk probability of each initial cluster, and take the initial cluster whose risk probability reaches the probability threshold as the risk cluster.
[0034] In a third aspect, the present application provides an electronic device, comprising: a memory, a processor;
[0035] Memory stores computer-executable instructions;
[0036] The processor executes the computer-executable instructions stored in the memory, so that the processor implements the steps of the method in any one of the above embodiments when executing the instructions.
[0037] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the method in any one of the above embodiments when the computer program is executed by a processor.
[0038] In a fifth aspect, the present application provides a computer program product, comprising a computer program, which implements the steps of the method in any one of the above embodiments when executed by a processor.
[0039] The risk prediction method, device, equipment, storage medium and product provided by the present application first obtain the interaction data of multiple objects, and construct a target graph based on the interaction data of multiple objects, so as to effectively mine and learn the information carried by the interaction data of multiple objects. Furthermore, based on the adaptive neighborhood sampling strategy and the multi-relationship graph convolution strategy, the multiple entities in the target graph can be aggregated to achieve preliminary grouping of the multiple entities and obtain multiple initial clusters. Furthermore, the risk characteristics are analyzed for each initial cluster to obtain the risk probability of each initial cluster, and the initial cluster whose risk probability reaches the probability threshold is used as the risk cluster. By adopting the above method, the target graph can be used to fully mine the implicit information in the big data. On this basis, the risk clusters (groups) with higher risks can be efficiently determined by grouping the entities in the target graph and then performing risk probability assessment on the initial clusters obtained by the grouping. That is, the group with the risk of illegal transactions can be efficiently identified based on massive big data. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0041] Figure 1 A flowchart of the steps of a risk prediction method in one embodiment;
[0042] Figure 2 A flowchart of the steps for constructing a target map in one embodiment;
[0043] Figure 3 A flowchart of the steps for constructing a target graph based on graph optimization in one embodiment;
[0044] Figure 4 A flowchart of the steps of dividing clusters by aggregation processing in one embodiment;
[0045] Figure 5 A flowchart of steps for evaluating initial cluster risk probability in one embodiment;
[0046] Figure 6 A flowchart of the steps for determining the initial cluster risk probability through risk feature analysis in one embodiment;
[0047] Figure 7 is a structural diagram of a risk prediction device in one embodiment;
[0048] Figure 8 FIG. 1 is a schematic structural diagram of an electronic device in an embodiment.
[0049] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0050] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0051] It should be noted that the risk prediction method and device of the present application can be used in the field of big data, and can also be used in any field other than big data. The application field of the risk prediction method and risk prediction device of the present application is not limited.
[0052] In the past, the identification of groups at risk of illegal transactions mostly relied on rigorous manual screening methods, without fully leveraging big data analysis to enhance and empower identification capabilities, resulting in low identification efficiency. Therefore, how to efficiently identify groups at risk of illegal transactions based on massive amounts of big data has become a technical challenge that needs to be addressed. Based on this, this application proposes a risk prediction method.
[0053] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0054] In an exemplary embodiment, Figure 1 As shown, a risk prediction method is provided. Taking the method applied to a server as an example, the method may include the following steps 102 to 106. Among them:
[0055] Step 102: Acquire interaction data of multiple objects, and construct a target graph based on the interaction data of the multiple objects.
[0056] Specifically, the interaction data between multiple objects can refer to the transaction data of a large number of clients of a financial institution, which can include the following information: client account number, recent transaction frequency, transaction time, transaction notes, transaction type, transaction resource volume, the account number of the other party to the transaction, etc. In every financial transaction, there is a resource transferor and a resource recipient. Clients of a financial institution can act as both resource transferors and resource recipients.
[0057] Among them, the target graph is a structured semantic knowledge base used to describe entities and the relationships between them. It is essentially a way to organize and represent knowledge using graph structures (nodes and edges). Entities are nodes, and the relationships between entities are edges.
[0058] Optionally, the server can obtain the interaction data of multiple objects from the stored data. Furthermore, the server can parse the interaction data of multiple objects to determine the interaction process represented by each piece of interaction data, and then take both parties involved in each interaction process as entities to obtain multiple entities. That is, all the interacting parties involved in the interaction data are regarded as entities. Furthermore, the server can perform association mining on the interaction data of multiple objects to obtain the association relationship between entities, and then construct a target graph based on the multiple entities and the association relationship between entities using the graph neural network framework. Among them, there is an association between interacting entities. The more frequent the interaction between two entities and the more resources are interacted, the closer the association relationship between the entities. Graph Neural Networks (GNN) is a deep learning model specially designed to process graph structured data.
[0059] Step 104: Based on the adaptive neighborhood sampling strategy and the multi-relationship graph convolution strategy, multiple entities in the target graph are aggregated to obtain multiple initial clusters.
[0060] Among them, when aggregating the target graph, it is necessary to aggregate the information of the entity and its neighboring entities (adjacent entities). In the traditional neighbor sampling method, when aggregating, the same number of neighboring entities are usually fixedly sampled for each entity. However, the adaptive neighborhood sampling adopted in this embodiment abandons this approach and instead dynamically determines the number of neighboring entities to be sampled based on the degree of each entity itself (the number of neighboring entities). Based on this, the adaptive neighborhood sampling strategy has the following advantages: (1) more balanced information, so that entities with more neighboring entities can aggregate more diverse information, and entities with fewer neighboring entities can avoid information dilution or the introduction of too much noise; (2) more efficient calculation, reducing redundant sampling and computational overhead on entities with fewer neighboring entities; (3) improved expressiveness, which can better capture the structural information of entities at different positions in the target graph; (4) alleviate over-smoothing, more sufficient sampling of entities with more neighboring entities helps to obtain deeper and more differentiated neighboring entity information, which is conducive to alleviating the over-smoothing problem in the graph.
[0061] Among them, there are association relationships between different entities in the target graph. In this embodiment, a multi-relationship graph convolution strategy is adopted. When aggregating multiple entities in the target graph, different types of association relationships in the target graph can be explicitly distinguished and utilized.
[0062] Optionally, the server can aggregate multiple entities in the target graph based on an adaptive neighborhood sampling strategy to obtain multiple first clusters. Simultaneously, the server can also aggregate multiple entities in the target graph based on a multi-relationship graph convolution strategy to obtain multiple second clusters. Based on this, the server can fuse and reference multiple first clusters and multiple second clusters to obtain multiple initial clusters. Each initial cluster includes at least one entity, each representing an interacting party, which can be a customer of a financial institution or another party interacting with a customer of the financial institution.
[0063] Step 106 , performing risk feature analysis on each initial cluster to obtain the risk probability of each initial cluster, and taking the initial cluster whose risk probability reaches the probability threshold as the risk cluster.
[0064] The initial cluster's risk probability reaching a probability threshold indicates that the probability of the interacting parties represented by this initial cluster engaging in illegal transactions has reached the probability threshold. A risk cluster is a group that presents a risk of illegal transactions. The probability threshold can be flexibly configured based on the needs of the actual application scenario.
[0065] Optionally, for each initial cluster, the server performs four risk profile analyses: historical risk indicator analysis, behavioral profile analysis, attribute profile analysis, and correlation profile analysis. The server then aggregates the risk probabilities from each risk profile analysis to determine the risk probability for that initial cluster. Furthermore, the server may select the initial cluster whose risk probability reaches a threshold as the risk cluster.
[0066] For example, the four risk feature analyses performed on a certain initial cluster can be specifically as follows: (1) Historical risk indicator analysis: Analyze the frequency or proportion of risk events (such as fraud, default, failure, security incidents, etc.) that occurred in the past for each entity in the initial cluster. (2) Behavioral feature analysis: Analyze whether each entity in the initial cluster has abnormal behavior. (3) Attribute feature analysis: Analyze whether each entity in the initial cluster has high-risk attributes, such as whether it is a newly registered account, belongs to a high-risk industry, or is located in a high-risk area. (4) Association feature analysis: Analyze the degree of association between each entity in the initial cluster and known high-risk entities, and whether there are known high-risk entities in the initial cluster.
[0067] In the above embodiment, the interaction data of multiple objects are first obtained, and based on the interaction data of multiple objects, a target graph is constructed to effectively mine and learn the information carried by the interaction data of multiple objects. Furthermore, based on the adaptive neighborhood sampling strategy and the multi-relationship graph convolution strategy, multiple entities in the target graph can be aggregated to achieve preliminary grouping of multiple entities and obtain multiple initial clusters. Furthermore, risk feature analysis is performed on each initial cluster to obtain the risk probability of each initial cluster, and the initial cluster whose risk probability reaches the probability threshold is used as the risk cluster. Using the above method, the target graph can be used to fully mine the implicit information in big data. On this basis, by grouping the entities in the target graph and then performing risk probability assessment on the initial clusters obtained by clustering, risk clusters (groups) with higher risks can be efficiently determined. That is, groups with illegal transaction risks can be efficiently identified based on massive big data.
[0068] In some embodiments, as Figure 2 As shown, in step 102, constructing a target graph based on the interaction data of multiple objects may specifically include the following steps:
[0069] Step 202 : extract multiple entities from the interaction data of multiple objects, and obtain the association relationships between the entities by performing association mining on the interaction data of the multiple objects.
[0070] The multiple entities include: all interacting parties in the interaction process represented by all interaction data.
[0071] Alternatively, natural language processing techniques, such as named entity recognition, keyword extraction, dictionary matching, rule matching, or deep learning models, can be used to identify fields representing interacting parties from the interaction data of multiple objects, thereby determining multiple interacting parties and, in turn, obtaining multiple entities. Furthermore, the server can parse the interaction data of multiple objects to obtain interaction records between each two entities. The more interaction records between entities and the more resources interacted, the closer the relationship between the entities. For example, different weights can be assigned to the relationship between entities based on the number of interactions between the two entities and the total amount of resources interacted. The closer the relationship, the higher the weight.
[0072] For example, before extracting entities from interaction data, preprocessing can be performed on the interaction data lines of multiple objects, such as data cleaning, denoising, and standardization, to more efficiently and accurately extract entities using natural language processing techniques. During the entity extraction process, existing knowledge graph libraries can also be used to assist in defining entity types and standardizing relationship types.
[0073] Step 204: Generate graph data based on the multiple entities and the relationships between the entities.
[0074] Optionally, the server may respectively use multiple entities as nodes of a graph and use association relationships as edges connecting the nodes, thereby generating graph data.
[0075] For example, both nodes and edges can have their own attributes. For example, a "user" node can have attributes such as "age," "gender," and "industry," while an edge can have attributes such as "timestamp" and "number of interactive resources." Furthermore, the generated graph data can be stored in a dedicated graph database within the server, or in the server's memory or files for subsequent use.
[0076] Step 206: Obtain a multi-layer message passing network architecture, combine the multi-layer message passing network architecture and graph data, and obtain an initial graph.
[0077] The multi-layer message passing network architecture can include the following layers: graph convolution layer, graph attention layer, and skip connection layer. The multi-layer message passing network architecture is a model architecture that can learn nodes and node relationships. The graph convolutional network (GCN) mainly learns the representation of entities in graph data by aggregating the feature information of entities and their neighboring entities. The graph attention network (GAT) can apply the attention mechanism to multi-layer network structures. The skip connection layer can directly pass the output of a layer in the network to the input of a deeper layer (usually one or several layers later).
[0078] Optionally, after obtaining the multi-layer message passing network architecture, the server can input the graph data into the multi-layer message passing network architecture in the form of a node feature matrix (describing the entities in the graph data) and an adjacency matrix (describing the association relationship between entities), and perform local neighborhood aggregation on the graph data through the multi-layer message passing network architecture to further learn and mine the graph data, thereby realizing the fusion of the multi-layer message passing network architecture and the graph data.
[0079] Step 208: Optimize the initial map to obtain a target map.
[0080] Optionally, in order to more effectively mine the information carried by the initial graph, the server can optimize the initial graph, for example, perform enhancement processing, training optimization, compression processing and other optimization operations to obtain the target graph.
[0081] In the above embodiment, by combining graph data with a multi-layer message passing network architecture and then optimizing the graph to obtain a target graph, it is possible to effectively mine and learn the information carried by the interaction data of multiple objects, so that risk clusters with illegal transaction risks can be efficiently identified based on the target graph.
[0082] In one embodiment, Figure 2 On the basis of Figure 3 As shown, the above step 208 optimizes the initial map to obtain the target map, including:
[0083] Step 302: Perform topology enhancement and feature enhancement on the initial graph to obtain a first graph.
[0084] Among them, topology enhancement can improve the structural robustness of the graph and learn the structural invariance of the graph. Feature enhancement can improve the feature robustness of the graph and make up for the feature loss of the graph.
[0085] Optionally, the server can enhance the topology of the initial graph by performing operations such as edge dropping, node dropping, and subgraph sampling, thereby obtaining a topologically enhanced initial graph. Furthermore, the server can enhance the topologically enhanced initial graph by performing operations such as feature masking, feature shuffling, noise addition, and manual feature enhancement to obtain a first graph.
[0086] For example, edge discarding can include, with a certain probability, deleting loose associations in the initial graph, i.e., deleting edges representing loose associations. Node discarding can include, with a certain probability, deleting entities (nodes) in the initial graph that have no abnormal transaction records and no associations (no connected edges) with other entities (other nodes).
[0087] For example, feature masking can involve randomly setting certain dimensions of the feature vectors of entities in the graph to zero with a certain probability. Feature shuffling can involve randomly shuffling a certain feature dimension across entities in the graph. For example, the first feature vectors associated with all entities are collected, randomly shuffled, and then redistributed. Adding noise can involve adding Gaussian or uniformly distributed noise to the feature vectors of entities in the graph.
[0088] Based on this, the common sparsity and incompleteness problems in the initial graph can be alleviated, so that the graph structure in the obtained first graph can better reflect the real complex relationships between entities.
[0089] Step 304: Perform progressive training and adversarial regularization training on the first atlas to obtain a second atlas.
[0090] Among them, progressive training can solve the problem that graphs may be large in scale and complex in structure, which makes one-time training difficult and ineffective. Adversarial regularization training can improve the robustness and generalization of graphs.
[0091] Optionally, progressive training may specifically include: (1) subgraph sampling: sampling different subgraphs from the first graph in rounds, for example, generating multiple entity sequences as subgraphs through random walks; (2) learning: first sampling simple, dense subgraphs, and then gradually introducing sparser, more complex subgraphs. After completing progressive training, adversarial regularization training may specifically include: introducing a discriminator, and using the discriminator to perform adversarial training on the graph after progressive training, so that the graph can learn a smoother and more stable representation space. After completing progressive training and adversarial regularization training in sequence, a second graph with stronger generalization ability can be obtained.
[0092] Step 306: compress the second atlas to obtain a target atlas.
[0093] Among them, compression processing can reduce the complexity, size and computational overhead of the graph while retaining its core information and performance as much as possible, making the graph more suitable for deployment and application, reducing graph storage space, accelerating inference calculations, and preventing potential overfitting.
[0094] Optionally, compression processing can be achieved by sequentially performing the following operations on the second graph to obtain the target graph: (1) Knowledge distillation: Using the second graph as a reference, a graph with a smaller structure and fewer parameters is trained. The training goal is to imitate the output of the second graph so that the graph after knowledge distillation can achieve performance similar to that of the second graph, but with a greatly reduced volume and computational complexity. (2) Parameter quantization: After knowledge distillation, the high-precision floating-point parameters (such as 32-bit floating-point numbers) of the obtained graph are converted to low-precision formats (such as 16-bit floating-point numbers, 8-bit integers), so that the volume of the graph after parameter quantization is significantly reduced. (3) Model pruning: After parameter quantization, redundant and unimportant parameters in the graph are identified and removed, so that the volume of the graph is further reduced and the inference speed is accelerated.
[0095] In the above embodiment, by performing enhancement processing, training optimization, compression processing and other optimization operations on the initial graph, a target graph is obtained, which can effectively mine and learn the information carried by the interaction data of multiple objects, so that risk clusters with illegal transaction risks can be efficiently identified based on the target graph.
[0096] In one embodiment, Figure 1 On the basis of Figure 4 As shown, step 104 aggregates multiple entities in the target graph based on the adaptive neighborhood sampling strategy and the multi-relationship graph convolution strategy to obtain multiple initial clusters, which may include the following steps:
[0097] Step 402: Based on the adaptive neighborhood sampling strategy, multiple entities in the target graph are aggregated to obtain multiple first clusters.
[0098] Each first cluster may include at least one entity.
[0099] Optionally, based on an adaptive neighborhood sampling strategy, the number of sampled neighbor entities, aggregation function, number of layers, etc. can be flexibly selected to aggregate connected and similar entities (nodes) in the target graph, that is, multiple entities in the target graph are aggregated to obtain multiple first clusters.
[0100] For example, the number of sampled neighbor entities can be adjusted (e.g., 10 for one-hop sampling, 5 for two-hop sampling, etc.), different aggregation functions can be selected (e.g., mean aggregation, pooling aggregation, LSTM aggregation, etc.), and the number of sampling layers (i.e., the K value for K-hops) can be selected. Based on this, the radius of the graph that the aggregated information can cover can be adaptively adjusted.
[0101] Step 404: Based on the multi-relationship graph convolution strategy, multiple entities in the target graph are aggregated to obtain multiple second clusters.
[0102] There are various types of relationships in the target graph, such as "upper-lower relationship", "consumer-service provider relationship", "investor-investee relationship", "family relationship", "friend relationship", etc. Each second cluster may include at least one entity.
[0103] Optionally, the server can configure different weights for different types of relationships based on a multi-relationship graph convolution strategy, and then construct a weight matrix based on the configured multiple weights. Subsequently, the weight matrix can be used to aggregate multiple entities in the target graph to distinguish different types of relationships between entities, thereby obtaining multiple second clusters. For example, if the relationship between entity 1 and entity 2, entity 3, and entity 4 is "relationship A", and all correspond to weight a, then entity 2, entity 3, and entity 4 can be aggregated into the same second cluster. That is, the second cluster not only considers the relationship between entities in the target graph, but also deeply considers the specific type of relationship.
[0104] Step 406: Merge the multiple first clusters and the multiple second clusters to obtain multiple initial clusters.
[0105] Optionally, a consensus mechanism may be introduced to merge multiple first clusters and multiple second clusters to obtain multiple initial clusters.
[0106] For example, using a consensus mechanism based on voting, for each entity, the first cluster and / or second cluster to which the entity belongs can be determined. If the first and second clusters intersect, the intersection can be used as the initial cluster to which the entity belongs. If no intersection exists, a cluster that retains the first or second cluster as much as possible can be selected from the first or second cluster as the initial cluster to which the entity belongs.
[0107] For example, using a consensus mechanism based on hypergraph partitioning, for each entity, all clusters in the first and second clusters are considered "meta-clusters," and a hypergraph is constructed: the entity is considered an entity in the hypergraph, and the hyperedges in the hypergraph are each "meta-cluster" (i.e., all entities belonging to the same "meta-cluster" are connected by a hyperedge). The hypergraph is then partitioned so that as many hyperedges as possible are retained intact in the same final cluster. This is equivalent to finding a cluster partitioning result that can simultaneously satisfy multiple first clusters and multiple second clusters to the greatest extent possible, and based on this, multiple initial clusters are obtained.
[0108] In the above embodiment, based on the two strategies of adaptive neighborhood sampling and multi-relationship graph convolution, multiple entities in the target graph can be aggregated to accurately achieve preliminary clustering of entities in the graph, so that subsequent analysis can be carried out on each initial cluster separately to accurately predict risk clusters with risks of illegal transactions.
[0109] In one embodiment, Figure 5 As shown, in step 106, risk characteristics analysis is performed on each initial cluster to obtain the risk probability of each initial cluster, including:
[0110] Step 502 : For each initial cluster, node-level risk feature analysis and graph-level risk feature analysis are performed to obtain cluster features of the initial cluster.
[0111] Among them, risk feature analysis can include the following: historical risk indicator analysis, behavioral feature analysis, attribute feature analysis and correlation feature analysis.
[0112] Optionally, for each initial cluster, node-level risk feature analysis (historical risk indicator analysis, behavioral feature analysis, attribute feature analysis, and association feature analysis) can be performed on each entity in the initial cluster to obtain the entity-level individual features of the entity (the risk probability of the entity). Furthermore, for each initial cluster, the mean of the entity-level individual features of all entities in the initial cluster can be used as the common map-level individual features of each entity in the initial cluster. Furthermore, for each entity, the entity-level individual features and map-level individual features of the entity can be feature aggregated (such as averaging) to obtain the individual features of the entity. Then, the cluster features of the initial cluster (in the numerical form of risk probability) can be obtained by feature aggregation of the individual features of each entity in the initial cluster, averaging, selecting the mode, or selecting the median.
[0113] Step 504: Score the cluster features to obtain the cluster score of the initial cluster.
[0114] Optionally, for each initial cluster, the higher the probability value represented by the cluster characteristics, the higher the cluster score of the initial cluster. To normalize and standardize the risk probabilities of all initial clusters, the cluster characteristics of all initial clusters can be aggregated and divided into probability value intervals accordingly. Initial clusters belonging to the same probability value interval are assigned the same cluster score.
[0115] For example, if the cluster characteristics of initial cluster A are 35% risk probability, the cluster characteristics of initial cluster B are 70% risk probability, the cluster characteristics of initial cluster C are 100% risk probability, the cluster characteristics of initial cluster D are 38% risk probability, etc., after aggregating the cluster characteristics of all initial clusters, the probability values can be divided into intervals: first interval 0-10%; second interval 10-25%; third interval 25%-50%; fourth interval 50%-75%; fifth interval 75%-100%. Initial clusters A and D are assigned a cluster score of 2, initial cluster B is assigned a cluster score of 4, and initial cluster C is assigned a cluster score of 5.
[0116] Step 506: Determine the risk probability of the initial cluster based on the cluster score.
[0117] A mapping relationship between the cluster score and the final risk probability of the initial cluster may be pre-configured. The higher the cluster score, the higher the final risk probability of the initial cluster.
[0118] Optionally, for each initial cluster, the risk probability of the initial cluster may be determined according to the cluster score by querying a mapping relationship between the cluster score and the risk probability.
[0119] For example, the mapping relationship between the cluster score and the final risk probability of the initial cluster can be: cluster score 1 corresponds to a risk probability of 5%; cluster score 2 corresponds to a risk probability of 18%; cluster score 3 corresponds to a risk probability of 38%; cluster score 4 corresponds to a risk probability of 63%; cluster score 5 corresponds to a risk probability of 88%.
[0120] In the above embodiment, the cluster characteristics of each initial cluster can be accurately mined and learned through node-level and graph-level risk feature analysis, thereby accurately evaluating the risk probability of each initial cluster.
[0121] In one possible implementation, Figure 5 On the basis of Figure 6 As shown, step 502 includes steps 602 to 606. Among them:
[0122] Step 602 : For each initial cluster, node-level risk feature analysis and graph-level risk feature analysis are performed on each entity in the initial cluster to obtain entity-level individual features and graph-level individual features of each entity.
[0123] Optionally, for each initial cluster, node-level risk feature analysis (including four types of analysis: historical risk indicator analysis, behavioral feature analysis, attribute feature analysis, and association feature analysis) can be performed on each entity in the initial cluster to obtain the entity's risk probability under the four analyses. The entity's risk probabilities under the four analyses are then averaged to obtain the entity's entity-level individual features. Furthermore, the graph-level individual features of each entity can be determined by averaging the entity-level individual features of all entities in the initial cluster and using this average as the common graph-level individual feature for all entities in the initial cluster.
[0124] For example, four risk feature analyses are performed on a certain entity to obtain the risk probabilities under the four analyses. Specifically, the following can be used: (1) Historical risk indicator analysis: Analyze the frequency or proportion of risk events (such as fraud, default, failure, security incidents, etc.) that occurred in the past. If the frequency or proportion of a certain risk event for the entity exceeds the historical risk indicator probability threshold, the risk probability corresponding to the historical risk indicator analysis is determined to be 20%. (2) Behavioral feature analysis: Analyze whether the entity has abnormal behavior, such as whether there is a sudden increase in transaction amount, abnormal login location, abnormal operation frequency. If the entity hits a certain abnormal behavior, the risk probability corresponding to the behavioral feature analysis is determined to be 20%. (3) Attribute feature analysis: Analyze whether the entity has high-risk attributes, such as whether it is a newly registered account, belongs to a high-risk industry, or is located in a high-risk area. If the entity hits any high-risk attribute, the risk probability corresponding to the attribute feature analysis is determined to be 20%. (4) Association feature analysis: Analyze the closeness of the association between the entity and known high-risk entities and whether it is a high-risk entity. If the number of interactions between the entity and a known high-risk entity reaches a preset number, or the entity is a known high-risk entity, the risk probability corresponding to the association feature analysis is determined to be 20%.
[0125] Step 604 : performing feature aggregation on the entity-level individual features and the map-level individual features of each entity to obtain the individual features of each entity.
[0126] Optionally, for each entity in the initial cluster, feature aggregation may be performed on the entity-level individual features and the map-level individual features of each entity. For example, the entity-level individual features and the map-level individual features may be averaged to obtain the individual features of the entity.
[0127] Step 606 : performing feature aggregation on the individual features of each entity in the initial cluster to obtain cluster features of the initial cluster.
[0128] Optionally, feature aggregation may be performed on the individual features of each entity in the initial cluster. For example, the individual features of each entity in the initial cluster may be averaged, the mode may be selected, or the median may be selected to obtain the cluster features of the initial cluster.
[0129] In the above embodiment, risk feature analysis can be performed at both the node and graph levels for each entity. Feature aggregation is then used to obtain the individual features of each entity, which in turn aggregates the cluster features of each initial cluster. This allows accurate mining and learning of the cluster features of each initial cluster, allowing for accurate assessment and determination of the risk probability of each initial cluster.
[0130] In one embodiment, an evaluation indicator matrix (including multiple auxiliary indicators) can also be constructed to verify the risk prediction method involved above. For example, the precision and recall of the risk prediction method can be verified by the auxiliary indicator F1-Macro (F1 score); the hit rate of the risk prediction method can also be verified by the auxiliary indicator Hit@10 (hit rate); and the reliability of the risk prediction method can also be verified by ROC-AUC (Receiver Operating Characteristic Curve, ROC curve).
[0131] Based on the same inventive concept, the present application also provides a device for implementing the aforementioned risk prediction. The solution provided by the device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more risk prediction device embodiments provided below can be found in the above-mentioned limitations of the risk prediction method, and will not be repeated here.
[0132] In an exemplary embodiment, Figure 7 As shown, a risk prediction device is provided, comprising: a target map construction module 702, an initial cluster acquisition module 704 and a risk cluster prediction module 706, wherein:
[0133] A target graph construction module is used to obtain the interaction data of multiple objects and construct a target graph based on the interaction data of multiple objects;
[0134] The initial cluster acquisition module is used to aggregate multiple entities in the target graph based on the adaptive neighborhood sampling strategy and the multi-relationship graph convolution strategy to obtain multiple initial clusters;
[0135] The risk cluster prediction module is used to analyze the risk characteristics of each initial cluster, obtain the risk probability of each initial cluster, and take the initial cluster whose risk probability reaches the probability threshold as the risk cluster.
[0136] In one embodiment, the target map construction module includes:
[0137] An entity extraction unit is used to extract multiple entities from the interaction data of multiple objects, and obtain the association relationship between the entities by performing association mining on the interaction data of multiple objects;
[0138] A graph data generating unit, configured to generate graph data based on a plurality of entities and association relationships between the entities;
[0139] An initial graph construction unit is used to obtain a multi-layer message passing network architecture, combine the multi-layer message passing network architecture with graph data, and obtain an initial graph;
[0140] The target map construction unit is used to optimize the initial map to obtain the target map.
[0141] In one embodiment, the target map construction unit is specifically used to:
[0142] Performing topological enhancement and feature enhancement on the initial atlas to obtain a first atlas;
[0143] Performing progressive training and adversarial regularization training on the first graph to obtain a second graph;
[0144] The second atlas is compressed to obtain a target atlas.
[0145] In one embodiment, the initial cluster acquisition module includes:
[0146] A first cluster obtaining unit is configured to perform aggregation processing on multiple entities in the target graph based on an adaptive neighborhood sampling strategy to obtain multiple first clusters;
[0147] A second cluster obtaining unit is used to aggregate multiple entities in the target graph based on a multi-relationship graph convolution strategy to obtain multiple second clusters;
[0148] The cluster fusion unit is used to fuse multiple first clusters and multiple second clusters to obtain multiple initial clusters.
[0149] In one embodiment, the risk cluster prediction module includes:
[0150] A risk feature analysis unit is used to perform node-level risk feature analysis and graph-level risk feature analysis on each initial cluster to obtain the cluster features of the initial cluster;
[0151] A cluster score obtaining unit is used to score the cluster features and obtain the cluster score of the initial cluster;
[0152] The risk probability prediction unit is used to determine the risk probability of the initial cluster based on the cluster score.
[0153] In one embodiment, the risk signature analysis unit is specifically configured to:
[0154] For each initial cluster, node-level risk feature analysis and graph-level risk feature analysis are performed on each entity in the initial cluster to obtain the entity-level individual features and graph-level individual features of each entity;
[0155] Perform feature aggregation on the entity-level individual features and the map-level individual features of each entity to obtain the individual features of each entity;
[0156] The individual features of each entity in the initial cluster are aggregated to obtain the cluster features of the initial cluster.
[0157] In the above embodiment, the risk prediction method provided by the present application can also be applied to the electronic device 800, such as Figure 8 As shown, it should be understood that processor 801 can be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention can be directly implemented by a hardware processor or performed by a combination of hardware and software modules in the processor.
[0158] The memory 802 may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory.
[0159] The communication component 803 is used to communicate and interconnect with other devices.
[0160] The bus involved in electronic device 800 may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of illustration, the buses in the drawings of this application are not limited to only one bus or one type of bus.
[0161] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method of any one of the above embodiments are implemented.
[0162] In one embodiment, a computer program product is provided, comprising a computer program, which implements the steps of the method of any one of the above embodiments when executed by a processor.
[0163] It should be noted that the information collected by this application is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of relevant data comply with the relevant laws, regulations and standards of relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0164] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0165] It should be understood that the present application is not limited to the exact structure described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A risk prediction method, characterized in that: The method comprises: Acquire interaction data of multiple objects, and construct a target graph based on the interaction data of the multiple objects; Based on the adaptive neighborhood sampling strategy and the multi-relationship graph convolution strategy, multiple entities in the target graph are aggregated to obtain multiple initial clusters; A risk characteristic analysis is performed on each of the initial clusters to obtain a risk probability of each of the initial clusters, and the initial cluster whose risk probability reaches a probability threshold is taken as a risk cluster.
2. The method according to claim 1, characterized in that The constructing a target graph based on the interaction data of the multiple objects includes: Extracting multiple entities from the interaction data of the multiple objects, and obtaining association relationships between the entities by performing association mining on the interaction data of the multiple objects; generating graph data according to the plurality of entities and the association relationships between the entities; Obtaining a multi-layer message passing network architecture, and combining the multi-layer message passing network architecture with the graph data to obtain an initial graph; The initial map is optimized to obtain a target map.
3. The method according to claim 2, characterized in that The performing of spectrum optimization on the initial spectrum to obtain a target spectrum comprises: Performing topological enhancement and feature enhancement on the initial map to obtain a first map; Performing progressive training and adversarial regularization training on the first atlas to obtain a second atlas; The second atlas is compressed to obtain the target atlas.
4. The method according to claim 1, wherein Based on the adaptive neighborhood sampling strategy and the multi-relationship graph convolution strategy, multiple entities in the target graph are aggregated to obtain multiple initial clusters, including: Based on the adaptive neighborhood sampling strategy, aggregating multiple entities in the target graph to obtain multiple first clusters; Based on the multi-relationship graph convolution strategy, aggregating multiple entities in the target graph to obtain multiple second clusters; The multiple first clusters and the multiple second clusters are merged to obtain multiple initial clusters.
5. The method according to claim 1, wherein The risk characteristic analysis is performed on each of the initial clusters to obtain the risk probability of each of the initial clusters, including: For each of the initial clusters, performing node-level risk feature analysis and graph-level risk feature analysis to obtain cluster features of the initial cluster; Scoring the cluster features to obtain a cluster score of the initial cluster; The risk probability of the initial cluster is determined according to the cluster score.
6. The method according to claim 5, characterized in that The node-level risk feature analysis and graph-level risk feature analysis are performed on each of the initial clusters to obtain cluster features of the initial cluster, including: For each of the initial clusters, performing node-level risk feature analysis and graph-level risk feature analysis on each entity in the initial cluster to obtain entity-level individual features and graph-level individual features of each entity; Performing feature aggregation on the entity-level individual features and the map-level individual features of each of the entities to obtain the individual features of each of the entities; Perform feature aggregation on the individual features of each entity in the initial cluster to obtain cluster features of the initial cluster.
7. A risk prediction device, characterized in that: The device comprises: A target graph construction module is used to obtain interaction data of multiple objects and construct a target graph based on the interaction data of the multiple objects; An initial cluster acquisition module is used to aggregate multiple entities in the target graph based on an adaptive neighborhood sampling strategy and a multi-relationship graph convolution strategy to obtain multiple initial clusters; The risk cluster prediction module is used to perform risk feature analysis on each of the initial clusters to obtain the risk probability of each of the initial clusters, and to take the initial cluster whose risk probability reaches a probability threshold as the risk cluster.
8. An electronic device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.