Abnormal node determination method, device, storage medium and electronic device

By determining the subgraph of nodes in the target relationship diagram of credit card business and calculating the embedded characterization vector, the problem of inefficient determination of abnormal nodes in the prior art is solved, and more efficient and accurate credit card business abnormal detection is achieved.

CN117313008BActive Publication Date: 2025-06-06CHINA CONSTR BANK CORP (DALIAN BRANCH)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311198679.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-15
Publication Date
2025-06-06
Estimated Expiration
2043-09-15

AI Technical Summary

Technical Problem

The determination efficiency of credit card business abnormal nodes in the prior art is low. It is mainly because neo4j stores the graph structure in memory, which requires more hardware resources when the data volume increases. In addition, traditional machine learning algorithms require manual feature engineering, which is complex and time-consuming.

Method used

By determining the subgraph corresponding to each node in the target relationship graph, the embedded representation vector is calculated using node information and edge parameters, and then determining whether the node is abnormal. This method avoids information loss caused by constructing homogeneous graphs and reduces dependence on artificial feature engineering.

Benefits of technology

It improves the determination efficiency of abnormal nodes, reduces the hardware resource requirements, and simplifies the feature engineering process, achieving faster and more accurate credit card business abnormality detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117313008B_ABST
    Figure CN117313008B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention provides a method, device, storage medium and electronic device for determining abnormal nodes, wherein the method comprises: determining the subgraph corresponding to each node of N nodes of the first node type from the target relationship graph, respectively, to obtain N subgraphs, the target relationship graph comprising N nodes, nodes of the second node type, nodes of the third node type, and edges of the first edge type and edges of the second edge type; determining an embedded representation vector for representing each node of the N nodes according to the node information of each node in each subgraph of the N subgraphs and the parameters on each edge, to obtain N embedded representation vectors; determining whether each node of the N nodes is abnormal according to the N embedded representation vectors and the N subgraphs. Through the embodiment of the present invention, the technical problem of low efficiency in determining abnormal nodes existing in the related art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of data processing technology, and in particular, to a method, device, storage medium and electronic device for determining an abnormal node. Background Art

[0002] With the rapid development of bank credit card business, abnormal behavior detection of credit card business has also become a focus of attention. For example, credit card team fraud has also become one of the serious problems faced by banks. Usually, credit card business involves several different types of nodes, such as credit cards, repayers, merchants, etc. Anti-fraud is an identification service. Its core is to collect, analyze and process big data information, use anti-fraud evaluation indicators to determine user fraud behavior, establish anti-fraud models, and solve risk problems in different scenarios. In related technologies, it is mainly to build a graph database through neo4j, use machine learning algorithms to identify fraudulent customer groups based on graph structures such as connectivity and neighbor nodes, and calculate the article rank score of the customer group, the number of overdue customers and other indicators to analyze the customer group, and then locate suspicious customer groups, determine customers, repayers and merchants suspected of fraud, and finally export the information bills of customers with fraud risks for manual investigation.

[0003] Methods in related technologies Since neo4j stores graph structures in memory, this means that when the amount of data increases, more hardware resources are needed to support high-performance processing, which increases costs and complexity; when performing credit card business anomaly detection, it is necessary to detect whether there are abnormal nodes. Related technologies mainly consider the structural nature of the graph and use machine learning algorithms to model graph data as homogeneous graphs. Nodes and edges are usually regarded as the same type, and node neighbor information is used to perform tasks such as node classification and link prediction on the graph. In fact, they only model part of the information in the interactive system, or do not distinguish the heterogeneity of entities and relationships, resulting in irreversible information loss; at the same time, traditional machine learning algorithms usually need to rely on manual feature engineering, that is, extracting and selecting features. In the detection of credit card abnormal behavior (or node anomaly), feature engineering can be a complex and time-consuming task that requires the knowledge and experience of domain experts. It can be seen that the methods for determining abnormal nodes in related technologies have the problem of low efficiency.

[0004] With respect to the technical problem of low efficiency in determining abnormal nodes existing in the related technologies, no effective solution has been proposed so far. Summary of the invention

[0005] The embodiments of the present invention provide a method, device, storage medium and electronic device for determining an abnormal node, so as to at least solve the technical problem of low efficiency in determining an abnormal node existing in the related art.

[0006] According to one embodiment of the present invention, a method for determining an abnormal node is provided, comprising: in a case where there are N nodes of a first node type in a target relationship graph, determining a subgraph corresponding to each of the N nodes from the target relationship graph to obtain N subgraphs, wherein the i-th subgraph among the N subgraphs includes the i-th node among the N nodes and nodes whose hop counts between the i-th node and the i-th node in the target relationship graph are less than or equal to M, N and M are positive integers greater than or equal to 2, i is a positive integer greater than or equal to 1 and less than or equal to N, the target relationship graph includes the N nodes, nodes of a second node type, nodes of a third node type, and edges of a first edge type and edges of a second edge type, the edges of the first edge type being edges between nodes of the first node type and nodes of the second node type, and the edges of the second edge type being edges between nodes of the second node type and nodes of the third node type The edges between nodes, the parameters on the edges of the first edge type include event parameters of an event of a first event type, the event of the first event type represents an event performed by an account of a first account type represented by a node of the first node type on an account of a second account type represented by a node of the second node type, the parameters on the edges of the second edge type include event parameters of an event of a second event type, the event of the second event type represents an event performed by an account of the second account type represented by a node of the second node type on an account of a third account type represented by a node of the third node type; according to the node information of each node in each subgraph of the N subgraphs and the parameters on each edge, determine an embedded representation vector used to represent each node in the N nodes to obtain N embedded representation vectors; according to the N embedded representation vectors and the N subgraphs, determine whether each of the N nodes has an abnormality.

[0007] In an exemplary embodiment, the method of determining an embedded representation vector for representing each node in the N nodes according to the node information of each node in each subgraph of the N subgraphs and the parameters on each edge to obtain N embedded representation vectors includes: determining the embedded representation vector of the ith node by the following steps to obtain the ith embedded representation vector, wherein the subgraph corresponding to the ith node is the ith subgraph in the N subgraphs: in a case where the ith subgraph includes K nodes of the first node type, determining a feature vector for representing each node in the K nodes according to the K node information corresponding to the K nodes and the parameters on the edges of the first edge type connecting each node in the K nodes to obtain K feature vectors, wherein the K node information includes account information of an account of the first account type represented by each node in the K nodes, wherein K is a positive integer greater than or equal to 1 and less than N; in a case where the ith subgraph includes P nodes of the second node type, determining a feature vector for representing each node in the K nodes according to the P node information corresponding to the P nodes and the parameters on the edges of the first edge type connecting each node in the K nodes to obtain K feature vectors. The method comprises: determining a feature vector for representing each of the P nodes according to the K feature vectors, wherein the edge connecting each of the P nodes includes an edge of the first edge type and / or an edge of the second edge type, and the P node information includes account information of an account of the second account type represented by each of the P nodes, and P is a positive integer greater than or equal to 1; in the case where the i-th subgraph includes Q nodes of the third node type, determining a feature vector for representing each of the Q nodes according to the Q node information corresponding to the Q nodes and the parameters of the edge of the second edge type connecting each of the Q nodes, and obtaining Q feature vectors, wherein the Q node information includes account information of an account of the third account type represented by each of the Q nodes, and Q is a positive integer greater than or equal to 1; determining the i-th embedded representation vector according to the K feature vectors, the P feature vectors and the Q feature vectors.

[0008] In an exemplary embodiment, determining a feature vector for representing each of the K nodes according to K node information corresponding to the K nodes and parameters on the edge of the first edge type connecting each of the K nodes to obtain K feature vectors includes: determining a feature vector for representing a j-th node among the K nodes by the following steps to obtain a j-th feature vector among the K feature vectors, wherein j is a positive integer greater than or equal to 1 and less than or equal to K; when the node information corresponding to the j-th node and the parameters on the edge of the first edge type connecting the j-th node include a total of X parameters, converting each of the X parameters into a corresponding vector to obtain X first vectors, wherein X is a positive integer greater than or equal to 1; inputting the X first vectors into a fully connected layer respectively to obtain X second vectors; inputting the X second vectors into a long short-term memory network respectively to obtain X encoding vectors; and performing an average pooling operation on the X encoding vectors respectively to obtain the j-th feature vector.

[0009] In an exemplary embodiment, determining the i-th embedded representation vector based on the K feature vectors, the P feature vectors, and the Q feature vectors includes: processing the K feature vectors to obtain a first fusion vector; processing the P feature vectors to obtain a second fusion vector; processing the Q feature vectors to obtain a third fusion vector; and determining the i-th embedded representation vector based on the first fusion vector, the second fusion vector, and the third fusion vector.

[0010] In an exemplary embodiment, the processing of the K feature vectors to obtain a first fused vector includes: inputting the K feature vectors into a long short-term memory network respectively to obtain K encoding vectors respectively; and performing an average pooling operation on the K encoding vectors respectively to obtain the first fused vector.

[0011] In an exemplary embodiment, determining the i-th embedded representation vector based on the first fusion vector, the second fusion vector and the third fusion vector includes: determining a feature vector corresponding to the i-th node among the K feature vectors; determining the i-th embedded representation vector based on the feature vector corresponding to the i-th node, the first fusion vector, the second fusion vector and the third fusion vector.

[0012] In an exemplary embodiment, the subgraph corresponding to each of the N nodes is determined from the target relationship graph to obtain N subgraphs, including: determining the i-th subgraph corresponding to the i-th node through the following steps: determining the nodes whose hop numbers to the i-th node are 1 to M in sequence in the target relationship graph to obtain the i-th group of nodes; determining the edges between the i-th group of nodes and each two nodes in the i-th node in the target relationship graph to obtain the i-th group of edges; and determining the subgraph consisting of the i-th node, the i-th group of nodes and the i-th group of edges as the i-th subgraph.

[0013] In an exemplary embodiment, the node of the first node type is used to represent a repayment account, the node of the second node type is used to represent a credit card account, and the node of the third node type is used to represent a merchant account. The account of the first account type is the repayment account, the account of the second account type is the credit card account, and the account of the third account type is the merchant account: the event of the first event type is a repayment event performed by the repayment account on the credit card account, and the event of the second event type is a consumption event performed by the credit card account on the merchant account.

[0014] In an exemplary embodiment, determining whether each of the N nodes is abnormal based on the N embedded representation vectors and the N subgraphs includes: determining whether the i-th node is abnormal through the following steps, wherein the subgraph corresponding to the i-th node is the i-th subgraph in the N subgraphs: when there is a k-th node of the first node type in the i-th subgraph and the k-th node is different from the i-th node, obtaining the i-th embedded representation vector for representing the i-th node and the k-th embedded representation vector for representing the k-th node from the N embedded representation vectors, wherein k is a positive integer greater than or equal to 1 and less than or equal to N; when the similarity between the i-th embedded representation vector and the k-th embedded representation vector is less than or equal to a preset similarity threshold, determining that the i-th node is abnormal.

[0015] In an exemplary embodiment, before determining the subgraph corresponding to each of the N nodes in the target relationship graph and obtaining the N subgraphs, the method further includes: deleting parameters that are not related to determining whether an event of the first event type occurs abnormally from the parameters on the edges of the first edge type in the initial relationship graph, and deleting parameters that are not related to determining whether an event of the second event type occurs abnormally from the parameters on the edges of the second edge type in the initial relationship graph, to obtain the target relationship graph; wherein the initial relationship graph includes the N nodes, nodes of the second node type, nodes of the third node type, and edges of the first edge type and edges of the second edge type.

[0016] According to another embodiment of the present invention, a device for determining an abnormal node is also provided, comprising: a first acquisition module, for determining, when there are N nodes of a first node type in a target relationship graph, a subgraph corresponding to each of the N nodes from the target relationship graph, to obtain N subgraphs, wherein the i-th subgraph among the N subgraphs includes the i-th node among the N nodes and nodes whose number of hops between the i-th node and the target relationship graph is less than or equal to M, N and M are positive integers greater than or equal to 2, i is a positive integer greater than or equal to 1 and less than or equal to N, the target relationship graph includes the N nodes, nodes of a second node type, nodes of a third node type, and edges of a first edge type and edges of a second edge type, the edges of the first edge type being edges between nodes of the first node type and nodes of the second node type, and the edges of the second edge type being edges between nodes of the second node type and nodes of the third node type. The parameters on the edges of the first edge type include event parameters of an event of a first event type, the event of the first event type represents an event executed by an account of a first account type represented by a node of the first node type on an account of a second account type represented by a node of the second node type, the parameters on the edges of the second edge type include event parameters of an event of a second event type, the event of the second event type represents an event executed by an account of the second account type represented by a node of the second node type on an account of a third account type represented by a node of the third node type; a second obtaining module is used to determine an embedded representation vector used to represent each node in the N nodes according to node information of each node in each subgraph of the N subgraphs and parameters on each edge, so as to obtain N embedded representation vectors; a determining module is used to determine whether each of the N nodes is abnormal according to the N embedded representation vectors and the N subgraphs.

[0017] According to yet another embodiment of the present invention, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the steps of any one of the above method embodiments when run.

[0018] According to yet another embodiment of the present invention, there is provided an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0019] According to the present invention, when there are N nodes of the first node type in the target relationship graph, a subgraph corresponding to each of the N nodes is determined from the target relationship graph to obtain N subgraphs, wherein the i-th subgraph in the N subgraphs includes the i-th node in the N nodes and a node whose hop count between the i-th node and the target relationship graph is less than or equal to M, and the target relationship graph includes N nodes, nodes of the second node type, nodes of the third node type, and edges of the first edge type and edges of the second edge type, wherein the parameters on the edges of the first edge type include event parameters of an event of the first event type, and the event of the first event type represents the first account represented by the node of the first node type The method comprises the following steps: determining an event executed by an account of the second type of node on an account of the second type of node represented by a node of the second node type, the parameters on the edge of the second edge type include the event parameters of the event of the second event type, and the event of the second event type represents an event executed by an account of the second type of node represented by the node of the second node type on an account of the third type of account represented by a node of the third node type; and determining an embedded representation vector for representing each of the N nodes based on the node information of each node and the parameters on each edge included in each of the N subgraphs, obtaining N embedded representation vectors, and then determining whether each of the N nodes is abnormal based on the N embedded representation vectors and the N subgraphs. The target relationship graph includes different types of nodes and different types of edges, which can effectively distinguish the heterogeneity of entities and relationships, thereby avoiding the problem of information loss caused by constructing a homogeneous graph adopted in the related art, and obtaining N embedded representation vectors by determining an embedded representation vector for representing each of the N nodes, and determining whether each of the N nodes is abnormal based on the N embedded representation vectors and the N subgraphs, thereby avoiding the low efficiency of determining abnormal nodes due to the need to rely on manual feature engineering in the related art. Therefore, the technical problem of low efficiency in determining abnormal nodes in the related art is solved, and the effect of improving the efficiency in determining abnormal nodes is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a hardware structure block diagram of a mobile terminal of a method for determining an abnormal node according to an embodiment of the present invention;

[0021] Figure 2 is a flow chart of a method for determining an abnormal node according to an embodiment of the present invention;

[0022] Figure 3 is an overall framework diagram of abnormal node detection according to an embodiment of the present invention;

[0023] Figure 4 is a flowchart of abnormal node detection according to an embodiment of the present invention;

[0024] Figure 5 is an example diagram of a heterogeneous graph neural network model according to an embodiment of the present invention;

[0025] Figure 6 4 is a structural block diagram of a device for determining an abnormal node according to an embodiment of the present invention. DETAILED DESCRIPTION

[0026] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings and in combination with the embodiments.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0028] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 FIG. 1 is a block diagram of the hardware structure of a mobile terminal of the method for determining an abnormal node according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.

[0029] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method for determining abnormal nodes in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, to implement the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the mobile terminal via a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0030] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0031] In this embodiment, a method for determining an abnormal node is provided. Figure 2 is a flow chart of a method for determining an abnormal node according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:

[0032] Step S202, in the case where there are N nodes of the first node type in the target relationship graph, determine the subgraph corresponding to each node in the N nodes from the target relationship graph to obtain N subgraphs, wherein the i-th subgraph in the N subgraphs includes the i-th node in the N nodes and the nodes whose hop counts between the i-th node and the i-th node in the target relationship graph are less than or equal to M, N and M are positive integers greater than or equal to 2, i is a positive integer greater than or equal to 1 and less than or equal to N, and the target relationship graph includes the N nodes, nodes of the second node type, nodes of the third node type, and edges of the first edge type and the second edge type, and the edges of the first edge type are the nodes of the first node type and the i-th node. an edge between nodes of the second node type, the edge of the second edge type being an edge between a node of the second node type and a node of the third node type, the parameters on the edge of the first edge type comprising event parameters of an event of a first event type, the event of the first event type representing an event executed by an account of a first account type represented by the node of the first node type on an account of a second account type represented by the node of the second node type, the parameters on the edge of the second edge type comprising event parameters of an event of a second event type, the event of the second event type representing an event executed by an account of the second account type represented by the node of the second node type on an account of a third account type represented by the node of the third node type;

[0033] Step S204, determining an embedded representation vector for representing each of the N nodes according to the node information of each node in each of the N subgraphs and the parameters on each edge, to obtain N embedded representation vectors;

[0034] Step S206: Determine whether each of the N nodes is abnormal according to the N embedded representation vectors and the N subgraphs.

[0035] Through the above steps, when there are N nodes of the first node type in the target relationship graph, a subgraph corresponding to each of the N nodes is determined from the target relationship graph to obtain N subgraphs, wherein the i-th subgraph in the N subgraphs includes the i-th node in the N nodes and a node whose hop number between the i-th node and the target relationship graph is less than or equal to M, and the target relationship graph includes N nodes, nodes of the second node type, nodes of the third node type, and edges of the first edge type and edges of the second edge type, wherein the parameters on the edges of the first edge type include event parameters of an event of the first event type, and the event of the first event type indicates the first account represented by the node of the first node type. The method comprises the following steps: determining an event executed by an account of a second type of number on an account of a second type of account represented by a node of a second node type, the parameters on the edge of the second edge type include the event parameters of the event of the second event type, and the event of the second event type represents an event executed by an account of a second type of account represented by a node of a second node type on an account of a third type of account represented by a node of a third node type; and determining an embedded representation vector for representing each of the N nodes based on the node information of each node and the parameters on each edge included in each of the N subgraphs, obtaining N embedded representation vectors, and then determining whether each of the N nodes is abnormal based on the N embedded representation vectors and the N subgraphs. The target relationship graph includes different types of nodes and different types of edges, which can effectively distinguish the heterogeneity of entities and relationships, thereby avoiding the problem of information loss caused by constructing a homogeneous graph adopted in the related art, and obtaining N embedded representation vectors by determining an embedded representation vector for representing each of the N nodes, and determining whether each of the N nodes is abnormal based on the N embedded representation vectors and the N subgraphs, thereby avoiding the low efficiency of determining abnormal nodes due to the need to rely on manual feature engineering in the related art. Therefore, the technical problem of low efficiency in determining abnormal nodes in the related art is solved, and the effect of improving the efficiency in determining abnormal nodes is achieved.

[0036] Among them, the executor of the above steps can be a processor, or a device end, such as a computer terminal, or an application control program in the device, or a processor with human-computer interaction capabilities configured on a storage device, or a processing device or processing unit with similar processing capabilities, etc., but not limited to this.

[0037] In the above embodiment, taking the anomaly detection of the credit card business as an example, the general credit card business involves several different types of nodes, such as credit card, repayer, merchant and other nodes, that is, the node of the first node type, the node of the second node type, and the node of the third node type can respectively represent one of them. For example, the node of the first node type can represent the repayer, the node of the second node type can represent the credit card, and the node of the third node type can represent the merchant; the edge of the first edge type can be used to represent the edge between the repayer node and the credit card node, that is, the edge of the first edge type can be used to represent the repayment relationship or repayment event in the credit card business (corresponding to the event of the first event type mentioned above), and the edge of the second edge type can be used to represent the edge between the credit card node and the merchant node, that is, the edge of the second edge type can be used to represent the consumption relationship or consumption event in the credit card business ( corresponding to the above-mentioned event of the second event type); the parameters on the edge of the first edge type and the parameters on the edge of the second edge type can be used to represent the characteristics of the edge, such as the parameters on the edge can include credit card contract number, merchant number, merchant name, minimum transaction time, maximum transaction time, total transaction amount, total number of transactions, number of transaction days, average number of transactions per month, monthly number variance, average number of transactions per month range, average monthly amount, monthly consumption amount range, whether to pay in installments, number of transaction channel types, the mean of the number of days between two transactions and the variance of the number of days between two transactions, among which the credit card contract number, merchant number, merchant name, average number of transactions per month range, monthly consumption amount range, number of transaction channel types and variance of the number of days between two transactions can be used to represent the event parameters of the event of the above-mentioned first event type or the event parameters of the event of the second event type.

[0038] The target relationship graph can be a graph data structure constructed using TigerGragh, or a graph data structure obtained by converting TigerGragh graph data into the graph data structure used by PyG. The graph database constructed in this embodiment is a heterogeneous graph, which includes three types of nodes: credit cards, payees, and merchants, as well as two types of edges: consumption relationship between credit cards and merchants and repayment relationship between credit cards and payees. This is different from the related art in which graph data is constructed as a homogeneous graph, nodes and edges are regarded as the same type, and the heterogeneity of entities and relationships is not distinguished.

[0039] Then, according to the node information of each node in each subgraph of the N subgraphs and the parameters on each edge, the embedded representation vector corresponding to each node in the N nodes is determined. Each subgraph may include nodes of the first node type, nodes of the second node type, and nodes of the third node type. Taking the subgraph corresponding to the ith node in the above N nodes (the ith subgraph) as an example, the ith subgraph also includes other nodes that satisfy a predetermined hop number M (such as 2, or 3, or other) relationship with the ith node, that is, the hop number with the ith node in the target relationship graph is within M, which can be a node of the first node type, or a node of the second node type or a node of the third node type. The above predetermined hop number M is settable, that is, the ith subgraph is equivalent to containing information on all other nodes and edges associated with the ith node, and the ith subgraph is equivalent to the neighbor sampling obtained with the ith node as the center node. graph (that is, the number of hops between all other nodes in the subgraph and the i-th node is within M); similarly, for any node of the above N nodes of the first node type, the corresponding subgraph can be obtained, and a total of N subgraphs are obtained, and each subgraph contains the information of all nodes and edges associated with the corresponding node of the first node type, so that each of the N embedded representation vectors obtained also incorporates the information of all nodes and edges associated with the corresponding node; then determine whether each of the N nodes is abnormal based on the N embedded representation vectors and the N subgraphs. For example, assuming that there are other nodes of the same type as the i-th node in the i-th subgraph (such as the k-th node in the above N nodes), the i-th embedded representation vector corresponding to the i-th node can be compared with the embedded representation vector corresponding to the k-th node. If the similarity between the two is low, it can be considered that the i-th node may be abnormal. Through this embodiment, it is possible to more accurately capture the deep feature information between nodes, and thus more accurately identify whether there are abnormal nodes. The target relationship graph includes different types of nodes and different types of edges, which can effectively distinguish the heterogeneity of entities and relationships, thereby avoiding the problem of information loss caused by constructing homogeneous graphs used in related technologies, and obtaining N embedded representation vectors by determining the embedded representation vector used to represent each of the N nodes, and determining whether each of the N nodes is abnormal based on the N embedded representation vectors and N subgraphs, thereby avoiding the need to rely on manual feature engineering in related technologies, resulting in low efficiency in determining abnormal nodes. Therefore, the technical problem of low efficiency in determining abnormal nodes in related technologies is solved, and the effect of improving the efficiency of determining abnormal nodes is achieved.

[0040] In an optional embodiment, determining an embedded representation vector for representing each node in the N nodes according to node information of each node in each subgraph of the N subgraphs and parameters on each edge to obtain N embedded representation vectors includes: determining the embedded representation vector of the ith node through the following steps to obtain the ith embedded representation vector, wherein the subgraph corresponding to the ith node is the ith subgraph in the N subgraphs: in a case where the ith subgraph includes K nodes of the first node type, determining a feature vector for representing each node in the K nodes according to K node information corresponding to the K nodes and parameters on edges of the first edge type connecting each node in the K nodes to obtain K feature vectors, wherein the K node information includes account information of an account of the first account type represented by each node in the K nodes, wherein K is a positive integer greater than or equal to 1 and less than N; in a case where the ith subgraph includes P nodes of the second node type, determining a feature vector for representing each node in the K nodes according to P node information corresponding to the P nodes and parameters on edges of the first edge type connecting each node in the K nodes to obtain K feature vectors. The method comprises: determining a feature vector for representing each of the P nodes according to the K feature vectors, wherein the edge connecting each of the P nodes includes an edge of the first edge type and / or an edge of the second edge type, and the P node information includes account information of an account of the second account type represented by each of the P nodes, and P is a positive integer greater than or equal to 1; in the case where the i-th subgraph includes Q nodes of the third node type, determining a feature vector for representing each of the Q nodes according to the Q node information corresponding to the Q nodes and the parameters of the edge of the second edge type connecting each of the Q nodes, and obtaining Q feature vectors, wherein the Q node information includes account information of an account of the third account type represented by each of the Q nodes, and Q is a positive integer greater than or equal to 1; determining the i-th embedded representation vector according to the K feature vectors, the P feature vectors and the Q feature vectors.

[0041] In the above embodiment, taking any one of the N nodes (such as the i-th node) as an example, the subgraph corresponding to the i-th node is the i-th subgraph in the N subgraphs, and the i-th subgraph may include the nodes of the first node type, the nodes of the second node type, and the nodes of the third node type. For example, assuming that the i-th subgraph includes K nodes of the first node type, P nodes of the second node type, and Q nodes of the third node type, that is, all nodes included in the i-th subgraph are divided into three categories according to the node type, and the K nodes of the first node type include the i-th node of the N nodes. Then, according to the node information of K nodes in the K nodes of the first node type and the parameters on the edge of the first edge type connected to the K nodes (referring to the parameters contained in the i-th subgraph connected to the K nodes) ), obtain K feature vectors, that is, obtain K feature vectors corresponding to K nodes of the first node type; similarly, obtain P feature vectors corresponding to P nodes of the second node type, and obtain Q feature vectors corresponding to Q nodes of the third node type, wherein the edge connecting each of the P nodes may be an edge of the first edge type and / or an edge of the second edge type, and the edge connecting each of the Q nodes may be an edge of the second edge type; then, determine the i-th embedded representation vector based on the K feature vectors (feature vectors corresponding to nodes of the first node type), the P feature vectors (feature vectors corresponding to nodes of the second node type) and the Q feature vectors (feature vectors corresponding to nodes of the third node type). For example, the K feature vectors, the P feature vectors, and the Q feature vectors may be processed separately to obtain a first fusion vector, a second fusion vector, and a third fusion vector. Then, the first fusion vector, the second fusion vector, and the third fusion vector may be fused to obtain the i-th embedded representation vector, that is, the i-th embedded representation vector obtained incorporates the node information and edge information of neighboring nodes of the same type as the i-th node, as well as the node information of other types of nodes that have neighbor relationships with the i-th node and the edge information between them.

[0042] In an optional embodiment, determining a feature vector for representing each of the K nodes according to K node information corresponding to the K nodes and parameters on the edge of the first edge type connecting each of the K nodes to obtain K feature vectors includes: determining a feature vector for representing a j-th node among the K nodes by the following steps to obtain a j-th feature vector among the K feature vectors, wherein j is a positive integer greater than or equal to 1 and less than or equal to K; when the node information corresponding to the j-th node and the parameters on the edge of the first edge type connecting the j-th node include a total of X parameters, converting each of the X parameters into a corresponding vector to obtain X first vectors, wherein X is a positive integer greater than or equal to 1; inputting the X first vectors into a fully connected layer respectively to obtain X second vectors; inputting the X second vectors into a long short-term memory network respectively to obtain X encoding vectors; and performing an average pooling operation on the X encoding vectors respectively to obtain the j-th feature vector.

[0043] In the above embodiment, taking the jth node among the K nodes of the first node type included in the above i-th subgraph as an example, assuming that the node information associated with the jth node and the parameters on the edge of the first edge type connecting the jth node contain a total of X parameters, each of the X parameters is converted into a vector to obtain X first vectors, and then the X first vectors are input into the fully connected layer to obtain X second vectors, and then the X second vectors are respectively input into the long short-term memory network (such as LSTM, or Bi-LSTM) to obtain X encoding vectors, and then the X encoding vectors are average pooled to obtain the jth feature vector; according to a similar method, each feature vector in the above K feature vectors can be obtained. Similarly, according to the same method as in this embodiment, each feature vector in the P feature vectors in the above embodiment and each feature vector in the Q feature vectors in the above embodiment can also be obtained.

[0044] In an optional embodiment, determining the i-th embedded representation vector based on the K feature vectors, the P feature vectors and the Q feature vectors includes: processing the K feature vectors to obtain a first fusion vector; processing the P feature vectors to obtain a second fusion vector; processing the Q feature vectors to obtain a third fusion vector; determining the i-th embedded representation vector based on the first fusion vector, the second fusion vector and the third fusion vector.

[0045] In the above embodiment, K feature vectors can be processed to obtain a first fusion vector, P feature vectors can be processed to obtain a second fusion vector, and Q feature vectors can be processed to obtain a third fusion vector, that is, the first fusion vector fuses the node information of K nodes of the first node type (such as repayer nodes) included in the i-th subgraph and the edge information related to the K nodes, while the second fusion vector fuses the node information of P nodes of the second node type (such as credit card nodes) included in the i-th subgraph and the edge information related to the P nodes, and the third fusion vector fuses the node information of Q nodes of the third node type (such as merchant nodes) included in the i-th subgraph and the edge information related to the Q nodes. Then, the i-th embedded representation vector is determined according to the first fusion vector, the second fusion vector and the third fusion vector, for example, the first fusion vector, the second fusion vector and the third fusion vector are processed using an attention mechanism module to obtain the i-th embedded representation vector.

[0046] In an optional embodiment, the processing of the K feature vectors to obtain a first fusion vector includes: inputting the K feature vectors into a long short-term memory network respectively to obtain K encoding vectors respectively; performing an average pooling operation on the K encoding vectors respectively to obtain the first fusion vector.

[0047] In the above embodiment, K feature vectors can be respectively input into a long short-term memory network (such as LSTM or Bi-LSTM) to obtain K encoding vectors, and then the K encoding vectors are averaged and pooled to obtain a first fused vector. Similarly, the second fused capacity and the third fused vector in the above embodiment can be obtained by the same method as in this embodiment.

[0048] In an optional embodiment, determining the i-th embedded representation vector based on the first fusion vector, the second fusion vector and the third fusion vector includes: determining a feature vector corresponding to the i-th node among the K feature vectors; determining the i-th embedded representation vector based on the feature vector corresponding to the i-th node, the first fusion vector, the second fusion vector and the third fusion vector.

[0049] In the above embodiment, the feature vector corresponding to the i-th node can be determined from the K feature vectors. It can be understood that the feature vector corresponding to the i-th node in the K feature vectors contains all the information of the i-th node itself and the edges associated with the i-th node. Then, based on the feature vector corresponding to the i-th node, the first fusion vector, the second fusion vector and the third fusion vector, the i-th embedded representation vector is determined. That is, the final i-th embedded representation vector not only fuses all the information associated with the i-th node itself, but also aggregates all the information associated with the neighbor nodes of the same type as the i-th node (including node information and edge information), and also aggregates all the information associated with the neighbor nodes of different types of the i-th node. It should be noted that the neighbor nodes here refer to the nodes included in the i-th subgraph corresponding to the i-th node, that is, the nodes whose number of hops to the i-th node is within M.

[0050] In an optional embodiment, the subgraph corresponding to each of the N nodes is determined from the target relationship graph to obtain N subgraphs, including: determining the i-th subgraph corresponding to the i-th node through the following steps: determining the nodes whose hop numbers to the i-th node are 1 to M in sequence in the target relationship graph to obtain the i-th group of nodes; determining the edges between the i-th group of nodes and each two nodes in the i-th node in the target relationship graph to obtain the i-th group of edges; and determining the subgraph consisting of the i-th node, the i-th group of nodes and the i-th group of edges as the i-th subgraph.

[0051] In the above embodiment, taking the i-th node among the N nodes of the first node type as an example, a node whose hop count to the i-th node is within M (such as M=2, or 3, or other) is determined in the target relationship graph, so that a group of nodes (for example, it can be called the i-th group of nodes) can be obtained, and then the edges between the i-th group of nodes and any two nodes in the i-th node are determined in the target relationship graph to obtain a group of edges (for example, it can be called the i-th group of edges). In this way, the i-th node, the i-th group of nodes and the i-th group of edges can constitute the i-th subgraph. For each subgraph in the N subgraphs, the same method as above can be used to obtain it, and the above hop count M can be set. Through this implementation, the purpose of determining the subgraph corresponding to any node from the target relationship graph is achieved.

[0052] In an optional embodiment, the node of the first node type is used to represent a repayment account, the node of the second node type is used to represent a credit card account, and the node of the third node type is used to represent a merchant account. The account of the first account type is the repayment account, the account of the second account type is the credit card account, and the account of the third account type is the merchant account: the event of the first event type is a repayment event performed by the repayment account on the credit card account, and the event of the second event type is a consumption event performed by the credit card account on the merchant account.

[0053] In the above embodiment, in combination with the specific application scenario, the node of the first node type mentioned above (or called the first type node) can be used to represent the repayment account, the node of the second node type (or called the second type node) can be used to represent the credit card account, and the node of the third node type (or called the third type node) can be used to represent the merchant account; the event of the first event type mentioned above can be used to represent the repayment event executed by the repayment account on the credit card account, and the event of the second event type can be used to represent the consumption event executed by the credit card account on the merchant account.

[0054] In an optional embodiment, determining whether each of the N nodes is abnormal based on the N embedded representation vectors and the N subgraphs includes: determining whether the i-th node is abnormal through the following steps, wherein the subgraph corresponding to the i-th node is the i-th subgraph in the N subgraphs: when there is a k-th node of the first node type in the i-th subgraph and the k-th node is different from the i-th node, obtaining the i-th embedded representation vector for representing the i-th node and the k-th embedded representation vector for representing the k-th node from the N embedded representation vectors, wherein k is a positive integer greater than or equal to 1 and less than or equal to N; when the similarity between the i-th embedded representation vector and the k-th embedded representation vector is less than or equal to a preset similarity threshold, determining that the i-th node is abnormal.

[0055] In the above embodiment, taking the above i-th node as an example, when there is a k-th node (not the same node as the i-th node) of the first node type (i.e., the same type as the i-th node) in the i-th subgraph, the i-th embedded representation vector corresponding to the i-th node and the k-th embedded representation vector corresponding to the k-th node can be obtained in the N embedded representation vectors, and then the similarity between the i-th embedded representation vector and the k-th embedded representation vector is compared. When the similarity between the two is less than or equal to the preset similarity threshold, it is determined that the i-th node is abnormal; using the same method as above, it can be determined whether other nodes (such as other nodes of the same type as the i-th node, or other nodes of different types from the i-th node) are abnormal; optionally, when it is determined that the i-th node is abnormal according to the i-th subgraph, it can be further determined whether other nodes included in the i-th subgraph are abnormal, so that it can be further determined whether the edges of the first edge type and / or the second edge type associated with the i-th node in the i-th subgraph are abnormal. Thereby achieving the purpose of more accurately determining whether the events corresponding to the edges contained in the subgraph are abnormal.

[0056] In an optional embodiment, before determining the subgraph corresponding to each of the N nodes in the target relationship graph and obtaining the N subgraphs, the method further includes: deleting parameters that are not related to determining whether an event of the first event type occurs abnormally from the parameters on the edges of the first edge type in the initial relationship graph, and deleting parameters that are not related to determining whether an event of the second event type occurs abnormally from the parameters on the edges of the second edge type in the initial relationship graph, to obtain the target relationship graph; wherein the initial relationship graph includes the N nodes, nodes of the second node type, nodes of the third node type, and edges of the first edge type and edges of the second edge type.

[0057] In the above embodiment, some parameters of the parameters on the edge of the first edge type in the initial relationship graph can be deleted. For example, some parameters that are irrelevant to determining whether an event of the first event type (such as a repayment event) is abnormal can be deleted. Similarly, some parameters of the parameters on the edge of the second edge type in the initial relationship graph can be deleted. For example, some parameters that are irrelevant to determining whether an event of the second event type (such as a consumption event) is abnormal can be deleted. That is, in this embodiment, parameters that are more important for determining whether a repayment or consumption event is abnormal can be screened out from the edge parameters included in the initial relationship graph. For example, the credit card contract number, merchant number, merchant name, monthly average number of extremes, monthly consumption amount extremes, number of transaction channel types, and variance of the number of days between two transactions can be screened out.

[0058] Obviously, the above described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The present invention will be specifically described below in conjunction with the embodiments.

[0059] An embodiment of the present invention provides a method for detecting abnormal behavior in credit card business. In the embodiment of the present invention, TigerGraph is first used to build a graph database, and data is processed and stored in parallel on multiple nodes to achieve high-performance processing on large-scale data such as credit card transactions; then PyTorch Geometric (PyG) is used to perform heterogeneous graph neural network modeling on credit card data, fully utilize multimodal features and graph structure information, automatically learn feature representation, learn the most useful features from the original data, and capture complex transaction patterns and abnormal behaviors (such as fraudulent behavior).

[0060] Figure 3 This is an overall framework diagram of abnormal node detection according to an embodiment of the present invention, which mainly includes the following parts:

[0061] Graph database construction, that is, building a credit card customer relationship graph data structure;

[0062] Heterogeneous graph neural network modeling: input graph data into heterogeneous graph neural network and use heterogeneous graph neural network to model;

[0063] Node anomaly detection, which uses information about node embedding and graph structure to detect anomalies;

[0064] Result display, such as output corresponding graph display and bill flow display.

[0065] The main steps of the embodiment of the present invention are described below. Figure 4 is a flowchart of abnormal node detection according to an embodiment of the present invention, including:

[0066] S1, the first step is to build the credit card customer relationship graph data structure.

[0067] Step S1 specifically includes:

[0068] S11, first, obtain a csv format file to be processed, and the data to be processed includes credit card data, merchant data and repayer data. Use TigerGraph to build a graph data structure, including three types of nodes: credit card, repayer and merchant, and two types of edges: credit card-merchant consumption relationship and credit card-repayer repayment relationship. Among them, the credit card entity features include the credit card contract number, the repayer entity features include the repayer account number, the merchant entity features include the merchant account number and the merchant name, the repayment relationship and the consumption relationship have the same edge feature content, but different code values, and the features include credit card contract number, merchant number, merchant name, minimum transaction time, maximum transaction time, total transaction amount, total number of transactions, number of transaction days, average number of transactions per month, monthly number variance, average number of transactions range, average monthly amount, monthly consumption amount range, whether to install, number of transaction channel types, mean of the number of days between the two transactions, and variance of the number of days between the two transactions.

[0069] S12, then import the data in csv format. Take the credit card data, merchant data and repayer data as node data, take the repayment relationship data of the repayer corresponding to the credit card data and the consumption relationship of the merchant corresponding to the credit card data as edge data, and mark the characteristic data in the repayment relationship and consumption relationship on the corresponding edge.

[0070] S2, the second step, data input conversion. Input the graph data into the heterogeneous graph neural network, that is, convert the Tiger Graph graph data into the data format used by PyG. First, use the Python third-party library of PyTigerGraph to connect to the TigerGraph graph database to export nodes, edges and corresponding attributes. Then build the Data object of PyG to represent the graph data. From a business perspective, credit card cash withdrawal will have frequent large transactions, so the present invention focuses more on the monitoring of abnormal large transactions and frequent transactions, and selects some features as the input of the heterogeneous graph neural network, which can more accurately predict fraudulent behavior. The present invention screens the credit card contract number, merchant number, merchant name, monthly average number of extremes, monthly consumption amount extremes, number of transaction channel types and variance of the number of days between the dates of the previous and next transactions in the edge features. Add the feature attributes of the node to the x attribute of the Data object, create a unique index mapping for each node, and the unique indexes of the credit card, merchant and repayer are the credit card number, merchant number and repayer number, respectively, and store them in three dictionaries. According to the repayment relationship edge and consumption relationship edge and corresponding attributes in TigerGraph, the edge data is added to the edge_index and edge_attr attributes of the Data object. Finally, the constructed Data object is returned as the input format required by PyG.

[0071] S3, the third step, uses heterogeneous graph neural network modeling. The model diagram is as follows Figure 5 shown.

[0072] Step S3 specifically includes:

[0073] S31, each time a single repayer is selected as the central node, and heterogeneous neighbors are sampled. The repayer is the central node because the repayer is the most core role in credit card transactions, and they are directly involved in the use and repayment of credit cards. Fraudulent behavior usually leads to changes in transaction patterns, such as abnormal transaction locations, frequent transactions, etc. Using the repayer as the central node can help discover abnormal transaction patterns related to the repayer, rather than detecting a single transaction in isolation, and can more accurately identify potential fraudulent behavior. The present invention sets the number of neighbor hops to 2. In addition to directly adjacent neighbor nodes, two-hop neighbor nodes are also sampled, that is, the neighbors of the neighbor nodes of the node, which helps to more comprehensively understand the contextual information of the node. The sampled subgraphs can be processed separately according to the entity type and processed through the neural network f 1 Encode it into a fixed-size embedding. Specifically, the features of the present invention are divided into string and numerical forms. We do not convert them into a unified vector linearly, but use par2vec to encode the string features, retain the original characteristics of the numerical features, and use a bidirectional LSTM (Bi-LSTM) architecture to capture "deep" feature interactions and obtain greater expressive power. This process obtains the node content C from the node ν ν ; Its feature embedding f 1 (ν) is calculated as follows:

[0074] FC is the feature conversion method, which can be identity mapping, fully connected neural network, etc., where i represents the node content C ν The i-th content in x i Represents the vector obtained by converting the i-th content. When converting between string and numeric features, the FC here is Par2vec and identity conversion respectively; operator Represents a cascade operation; the input gate f of LSTM i , forget the door i , output gate i , candidate memory state Memory state c i , hidden state h i as follows:

[0075]

[0076] Where σ represents the sigmoid activation function, tanh represents the hyperbolic tangent function, represents the Hadamard inner product, which is the corresponding multiplication of two matrix elements. i is the output hidden state of the ith content, U i (As mentioned above f , U z , U o )、W i (As mentioned above W f , W z , W o ), b i (As mentioned above f , b z , b o ) is a learnable parameter, f i 、z i , o i are the input gate vector, forget gate vector, and output gate vector of the first content feature, respectively. More specifically, the above architecture first uses different FC layers to transform different content features, then uses Bi-LSTM to capture the “deep” feature interactions and accumulate the expressive power of all content features, and finally uses the average pooling layer (MP) to obtain the content embedding v in all hidden states.

[0077] S32, aggregated neighbors’ content embedding. It consists of two consecutive steps: (1) aggregation of neighbors of the same type; (2) combination of neighbors of different types.

[0078] S32 includes the following steps:

[0079] S321, same type of neighbor aggregation. We use a neural network f 2 To aggregate the content embedding of ν′. Formally, the embedding formula for ν aggregating t types of neighbors is as follows:

[0080]

[0081] Among them, LSTM is the same as formula (2) except for the input and parameter settings. We use Bi-LSTM to aggregate the content embeddings of all t types of neighbors and use the average of all hidden states to represent the aggregate embedding of the same type. t (v) represents the content of the neighbor node, and v′ represents the content of the neighbor node N t (v) Any of the following.

[0082] S322, type aggregation. Because different types of neighbors will make different contributions to the final representation of ν, the present invention adopts an attention mechanism to output the embedding ε v for:

[0083]

[0084] where ε v , α v,* represents the importance of different embeddings, f 1 (v) is the content embedding of ν obtained by formula (1), is the type-based aggregate embedding obtained by formula (3). We denote the set of embeddings as LeakyReLU represents the leaky version of the ReLU activation function, and u is the attention mechanism parameter.

[0085] Thus, we get the embedding vector representation ε of the repayer entity a’s aggregated neighbor and edge feature information v .

[0086] S4, using node embedding and graph structure information for anomaly detection. In the subgraph formed by aggregating neighbor information starting from each repayer entity a, the cosine similarity between repayer entity a and other repayer entity nodes in the subgraph in the embedding space is calculated. From a business perspective, when consumption and repayment suddenly have large and frequent transactions, the greater the suspicion of fraud, that is, if the embedding of a node's neighbor node has a low similarity with the embedding of the node, then the node can be regarded as abnormal data.

[0087] S5, output the corresponding graph display and bill flow display.

[0088] In the above embodiment, a heterogeneous graph neural network is used to model the transaction data. The credit card transaction data is converted into a heterogeneous graph. The graph convolution layer in the model propagates information from the cardholder node to the transaction node, while considering the heterogeneous features and relationships of the edges, synchronizing and making full use of the graph structure information and feature information, and effectively avoiding the problem of information loss affecting the prediction effect. Tiger Graph is used to store and display the graph structure of credit card data. The unstructured data of credit card transactions is converted into TigerGraph graph structure data, which can better perform real-time data processing and high-performance computing in the credit card transaction business scenario with large amounts of data and frequent update operations, thereby supporting the efficient execution of various graph algorithms and analysis tasks on large-scale credit card transaction graphs.

[0089] The embodiment of the present invention uses a heterogeneous graph neural network to model credit card transaction data. The heterogeneous graph neural network can effectively fuse multiple types of feature information. By fusing these different types of features into a unified heterogeneous graph, the complex associations and features of credit card transactions can be more comprehensively described; the heterogeneous graph neural network can learn the embedded representation of nodes and edges, and can aggregate global information on nodes to capture the complex relationships and potential patterns between transactions. It does not rely solely on the local features of a single node. This enables the model to take into account factors such as the context of the transaction and the user's historical behavior, thereby improving the accuracy of detecting fraudulent behavior; by learning the differences between normal transaction patterns and fraudulent behavior, the heterogeneous graph neural network can perform anomaly detection and prediction. The model can automatically identify potential fraud patterns and provide real-time anti-fraud decisions and risk assessments.

[0090] The embodiment of the present invention uses the TigerGraph graph database to store credit card transaction data and display the fraud relationship network. The relationship between entities is represented by a graph structure, so as to better capture the transaction patterns and abnormal behaviors. TigerGraph supports distributed storage and computing, can quickly process large amounts of credit card fraud data, and provide responses in real time or near real time. It supports the storage and query of multimodal data, so that different types of data can be associated and analyzed in the same graph database. At the same time, staff can more intuitively understand the associations and abnormal behaviors in fraud data through the visual interface displayed by the platform, which facilitates decision-making and early warning.

[0091] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the method described in each embodiment of the present invention.

[0092] In this embodiment, a device for determining an abnormal node is also provided. Figure 6 is a structural block diagram of an apparatus for determining an abnormal node according to an embodiment of the present invention, such as Figure 6 As shown, the device comprises:

[0093] The first acquisition module 602 is used to determine the subgraph corresponding to each node in the target relationship graph respectively from the target relationship graph when there are N nodes of the first node type, so as to obtain N subgraphs, wherein the i-th subgraph in the N subgraphs includes the i-th node in the N nodes and the nodes whose hop counts between the i-th node and the target relationship graph are less than or equal to M, N and M are positive integers greater than or equal to 2, i is a positive integer greater than or equal to 1 and less than or equal to N, and the target relationship graph includes the N nodes, nodes of the second node type, nodes of the third node type, and edges of the first edge type and the second edge type, and the edges of the first edge type are nodes of the first node type and nodes of the third node type. an edge between nodes of the second node type, the edge of the second edge type being an edge between a node of the second node type and a node of the third node type, the parameters on the edge of the first edge type comprising an event parameter of an event of a first event type, the event of the first event type representing an event executed by an account of a first account type represented by the node of the first node type on an account of a second account type represented by the node of the second node type, the parameters on the edge of the second edge type comprising an event parameter of an event of a second event type, the event of the second event type representing an event executed by an account of the second account type represented by the node of the second node type on an account of a third account type represented by the node of the third node type;

[0094] A second obtaining module 604 is used to determine an embedded representation vector for representing each of the N nodes according to the node information of each node in each of the N subgraphs and the parameters on each edge, so as to obtain N embedded representation vectors;

[0095] The determination module 606 is used to determine whether each of the N nodes is abnormal according to the N embedded representation vectors and the N subgraphs.

[0096] In an optional embodiment, the second acquisition module 604 includes: a first acquisition unit, configured to determine the embedded representation vector of the i-th node through the following steps to obtain the i-th embedded representation vector, wherein the subgraph corresponding to the i-th node is the i-th subgraph in the N subgraphs: in the case where the i-th subgraph includes K nodes of the first node type, according to the K node information corresponding to the K nodes and the parameters on the edges of the first edge type connecting each of the K nodes, determine the feature vector used to represent each of the K nodes to obtain K feature vectors, wherein the K node information includes the account information of the account of the first account type represented by each of the K nodes, wherein K is a positive integer greater than or equal to 1 and less than N; in the case where the i-th subgraph includes P nodes of the second node type, according to the P node information corresponding to the P nodes and the parameters on the edges connecting each of the P nodes , determine a feature vector for representing each of the P nodes to obtain P feature vectors, wherein the edge connecting each of the P nodes includes an edge of the first edge type and / or an edge of the second edge type, the P node information includes account information of an account of the second account type represented by each of the P nodes, and P is a positive integer greater than or equal to 1; in the case where the i-th subgraph includes Q nodes of the third node type, determine a feature vector for representing each of the Q nodes according to the Q node information corresponding to the Q nodes and parameters on the edge of the second edge type connecting each of the Q nodes to obtain Q feature vectors, wherein the Q node information includes account information of an account of the third account type represented by each of the Q nodes, and Q is a positive integer greater than or equal to 1; determine the i-th embedded representation vector according to the K feature vectors, the P feature vectors and the Q feature vectors.

[0097] In an optional embodiment, the first obtaining unit is used to obtain K feature vectors in the following manner: determine a feature vector representing a j-th node among the K nodes by the following steps to obtain a j-th feature vector among the K feature vectors, where j is a positive integer greater than or equal to 1 and less than or equal to K: when the node information corresponding to the j-th node and the parameters on the edge of the first edge type connecting the j-th node include a total of X parameters, convert each of the X parameters into a corresponding vector to obtain X first vectors, where X is a positive integer greater than or equal to 1; input the X first vectors into a fully connected layer respectively to obtain X second vectors; input the X second vectors into a long short-term memory network respectively to obtain X encoding vectors; and perform an average pooling operation on the X encoding vectors respectively to obtain the j-th feature vector.

[0098] In an optional embodiment, the first obtaining unit is used to determine the i-th embedded representation vector according to the K feature vectors, the P feature vectors and the Q feature vectors in the following manner: processing the K feature vectors to obtain a first fusion vector; processing the P feature vectors to obtain a second fusion vector; processing the Q feature vectors to obtain a third fusion vector; determining the i-th embedded representation vector according to the first fusion vector, the second fusion vector and the third fusion vector.

[0099] In an optional embodiment, the first obtaining unit is used to process the K feature vectors in the following manner to obtain a first fusion vector: input the K feature vectors into a long short-term memory network respectively to obtain K encoding vectors respectively; perform an average pooling operation on the K encoding vectors respectively to obtain the first fusion vector.

[0100] In an optional embodiment, the above-mentioned first obtaining unit is used to determine the i-th embedded representation vector according to the first fusion vector, the second fusion vector and the third fusion vector in the following manner: determine the feature vector corresponding to the i-th node among the K feature vectors; determine the i-th embedded representation vector according to the feature vector corresponding to the i-th node, the first fusion vector, the second fusion vector and the third fusion vector.

[0101] In an optional embodiment, the above-mentioned first acquisition module 602 includes: a first determination unit, used to determine the i-th subgraph corresponding to the i-th node through the following steps: determine the nodes whose hop numbers to the i-th node are 1 to M in sequence in the target relationship graph, and obtain the i-th group of nodes; determine the edges between the i-th group of nodes and each two nodes in the i-th node in the target relationship graph, and obtain the i-th group of edges; determine the subgraph consisting of the i-th node, the i-th group of nodes and the i-th group of edges as the i-th subgraph.

[0102] In an optional embodiment, the node of the first node type is used to represent a repayment account, the node of the second node type is used to represent a credit card account, and the node of the third node type is used to represent a merchant account. The account of the first account type is the repayment account, the account of the second account type is the credit card account, and the account of the third account type is the merchant account: the event of the first event type is a repayment event performed by the repayment account on the credit card account, and the event of the second event type is a consumption event performed by the credit card account on the merchant account.

[0103] In an optional embodiment, the above-mentioned determination module 606 includes: a second determination unit, used to determine whether the i-th node has an abnormality through the following steps, wherein the subgraph corresponding to the i-th node is the i-th subgraph in the N subgraphs: when there is a k-th node of the first node type in the i-th subgraph, and the k-th node is different from the i-th node, obtain the i-th embedded representation vector used to represent the i-th node and the k-th embedded representation vector used to represent the k-th node from the N embedded representation vectors, wherein k is a positive integer greater than or equal to 1 and less than or equal to N; when the similarity between the i-th embedded representation vector and the k-th embedded representation vector is less than or equal to a preset similarity threshold, determine that the i-th node has an abnormality.

[0104] In an optional embodiment, the above-mentioned device also includes: a processing module, which is used to determine the subgraph corresponding to each of the N nodes in the target relationship graph respectively, and before obtaining the N subgraphs, delete the parameters that are not related to judging whether an event of the first event type occurs abnormally from the parameters of the edges of the first edge type in the initial relationship graph, and delete the parameters that are not related to judging whether an event of the second event type occurs abnormally from the parameters of the edges of the second edge type in the initial relationship graph, so as to obtain the target relationship graph; wherein the initial relationship graph includes the N nodes, the nodes of the second node type, the nodes of the third node type, and the edges of the first edge type and the edges of the second edge type.

[0105] It should be noted that each of the above modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or each of the above modules is located in different processors in any combination.

[0106] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.

[0107] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0108] An embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0109] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0110] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail herein.

[0111] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order than here, or they can be made into each integrated circuit module separately, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0112] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for determining abnormal nodes. It is characterized in that include: In the case where there are N nodes of the first node type in the target relationship graph, a subgraph corresponding to each of the N nodes is determined from the target relationship graph to obtain N subgraphs, wherein the i-th subgraph in the N subgraphs includes the i-th node in the N nodes and nodes whose hop counts between the i-th node and the i-th node in the target relationship graph are less than or equal to M, N and M are positive integers greater than or equal to 2, i is a positive integer greater than or equal to 1 and less than or equal to N, and the target relationship graph includes the N nodes, nodes of the second node type, nodes of the third node type, and edges of the first edge type and edges of the second edge type, and the edges of the first edge type are nodes of the first node type and nodes of the second node type. an edge between nodes of the first node type, the edge of the second edge type is an edge between a node of the second node type and a node of the third node type, the parameters on the edge of the first edge type include event parameters of an event of the first event type, the event of the first event type represents an event executed by an account of the first account type represented by the node of the first node type on an account of the second account type represented by the node of the second node type, the parameters on the edge of the second edge type include event parameters of an event of the second event type, the event of the second event type represents an event executed by the account of the second account type represented by the node of the second node type on an account of the third account type represented by the node of the third node type; Determine, according to the node information of each node in each subgraph of the N subgraphs and the parameters on each edge, an embedded representation vector for representing each node in the N nodes, to obtain N embedded representation vectors; Determining whether each of the N nodes is abnormal according to the N embedded representation vectors and the N subgraphs; Among them, the node of the first node type is used to represent the repayment account, the node of the second node type is used to represent the credit card account, and the node of the third node type is used to represent the merchant account. The account of the first account type is the repayment account, the account of the second account type is the credit card account, and the account of the third account type is the merchant account; the event of the first event type is the repayment event performed by the repayment account on the credit card account, and the event of the second event type is the consumption event performed by the credit card account on the merchant account.

2. The method according to claim 1, It is characterized in that The step of determining an embedded representation vector for representing each of the N nodes according to the node information of each node in each of the N subgraphs and the parameters on each edge, and obtaining the N embedded representation vectors, comprises: The embedded representation vector of the i-th node is determined by the following steps to obtain the i-th embedded representation vector, wherein the subgraph corresponding to the i-th node is the i-th subgraph in the N subgraphs: In a case where the i-th subgraph includes K nodes of the first node type, determining a feature vector for representing each of the K nodes according to K node information corresponding to the K nodes and parameters on the edge of the first edge type connecting each of the K nodes, to obtain K feature vectors, wherein the K node information includes account information of an account of the first account type represented by each of the K nodes, wherein K is a positive integer greater than or equal to 1 and less than N; In a case where the i-th subgraph includes P nodes of the second node type, determining a feature vector for representing each of the P nodes according to P node information corresponding to the P nodes and parameters on an edge connecting each of the P nodes, to obtain P feature vectors, wherein the edge connecting each of the P nodes includes an edge of the first edge type and / or an edge of the second edge type, the P node information includes account information of an account of the second account type represented by each of the P nodes, and P is a positive integer greater than or equal to 1; In a case where the i-th subgraph includes Q nodes of the third node type, determining a feature vector for representing each of the Q nodes according to Q node information corresponding to the Q nodes and parameters on the edge of the second edge type connecting each of the Q nodes, to obtain Q feature vectors, wherein the Q node information includes account information of an account of the third account type represented by each of the Q nodes, and Q is a positive integer greater than or equal to 1; The i-th embedded representation vector is determined according to the K feature vectors, the P feature vectors and the Q feature vectors.

3. The method according to claim 2, It is characterized in that The step of determining a feature vector for representing each of the K nodes according to the K node information corresponding to the K nodes and the parameters of the edge of the first edge type connecting each of the K nodes to obtain the K feature vectors includes: The feature vector representing the j-th node among the K nodes is determined by the following steps to obtain the j-th feature vector among the K feature vectors, where j is a positive integer greater than or equal to 1 and less than or equal to K: In a case where the node information corresponding to the j-th node and the parameters of the edge of the first edge type connecting the j-th node include X parameters in total, converting each parameter of the X parameters into a corresponding vector to obtain X first vectors, where X is a positive integer greater than or equal to 1; Input the X first vectors into a fully connected layer respectively to obtain X second vectors; Inputting the X second vectors into a long short-term memory network respectively to obtain X encoding vectors; An average pooling operation is performed on each of the X encoding vectors to obtain the j-th feature vector.

4. The method according to claim 2, It is characterized in that The determining the i-th embedded representation vector according to the K feature vectors, the P feature vectors and the Q feature vectors includes: Processing the K feature vectors to obtain a first fusion vector; Processing the P feature vectors to obtain a second fusion vector; Processing the Q feature vectors to obtain a third fusion vector; Determine the i-th embedded representation vector according to the first fusion vector, the second fusion vector and the third fusion vector.

5. The method according to claim 4, It is characterized in that The step of processing the K feature vectors to obtain a first fusion vector includes: Input the K feature vectors into the long short-term memory network respectively to obtain K encoding vectors respectively; An average pooling operation is performed on each of the K encoding vectors to obtain the first fused vector.

6. The method according to claim 5, It is characterized in that The determining the i-th embedded representation vector according to the first fusion vector, the second fusion vector and the third fusion vector comprises: Determine a feature vector corresponding to the i-th node among the K feature vectors; Determine the i-th embedded representation vector according to the feature vector corresponding to the i-th node, the first fusion vector, the second fusion vector and the third fusion vector.

7. The method according to any one of claims 1 to 6, It is characterized in that Determining a subgraph corresponding to each of the N nodes from the target relationship graph to obtain N subgraphs includes: The i-th subgraph corresponding to the i-th node is determined by the following steps: Determine in the target relationship graph the nodes whose hop numbers to the i-th node are 1 to M in sequence, to obtain the i-th group of nodes; Determine the edges between the i-th group of nodes and every two nodes in the i-th node in the target relationship graph to obtain the i-th group of edges; A subgraph formed by the i-th node, the i-th group of nodes, and the i-th group of edges is determined as the i-th subgraph.

8. The method according to any one of claims 1 to 6, It is characterized in that The determining, according to the N embedded representation vectors and the N subgraphs, whether each of the N nodes is abnormal includes: Whether the i-th node is abnormal is determined by the following steps, wherein the subgraph corresponding to the i-th node is the i-th subgraph in the N subgraphs: In a case where there is a kth node of the first node type in the i-th subgraph, and the kth node is different from the i-th node, obtaining an i-th embedding representation vector for representing the i-th node and a k-th embedding representation vector for representing the k-th node from the N embedding representation vectors, where k is a positive integer greater than or equal to 1 and less than or equal to N; When the similarity between the i-th embedded representation vector and the k-th embedded representation vector is less than or equal to a preset similarity threshold, it is determined that an abnormality occurs in the i-th node.

9. The method according to any one of claims 1 to 6, It is characterized in that Before respectively determining the subgraph corresponding to each of the N nodes in the target relationship graph to obtain the N subgraphs, the method further includes: Delete parameters irrelevant to determining whether an event of the first event type occurs abnormally from parameters on edges of the first edge type in the initial relationship graph, and delete parameters irrelevant to determining whether an event of the second event type occurs abnormally from parameters on edges of the second edge type in the initial relationship graph, to obtain the target relationship graph; The initial relationship graph includes the N nodes, nodes of the second node type, nodes of the third node type, edges of the first edge type, and edges of the second edge type.

10. A device for determining an abnormal node, It is characterized in that include: A first obtaining module is used to determine, in the case where there are N nodes of the first node type in the target relationship graph, the subgraph corresponding to each node in the target relationship graph, to obtain N subgraphs, wherein the i-th subgraph in the N subgraphs includes the i-th node in the N nodes and the nodes whose hop counts between the i-th node and the i-th node in the target relationship graph are less than or equal to M, N and M are positive integers greater than or equal to 2, i is a positive integer greater than or equal to 1 and less than or equal to N, and the target relationship graph includes the N nodes, nodes of the second node type, nodes of the third node type, and edges of the first edge type and the second edge type, and the edges of the first edge type are the nodes of the first node type and the an edge between nodes of a second node type, the edge of the second edge type being an edge between a node of the second node type and a node of the third node type, the parameters on the edge of the first edge type comprising event parameters of an event of a first event type, the event of the first event type representing an event executed by an account of a first account type represented by the node of the first node type on an account of a second account type represented by the node of the second node type, the parameters on the edge of the second edge type comprising event parameters of an event of a second event type, the event of the second event type representing an event executed by an account of the second account type represented by the node of the second node type on an account of a third account type represented by the node of the third node type; A second obtaining module is used to determine an embedded representation vector for representing each of the N nodes according to the node information of each node in each of the N subgraphs and the parameters on each edge, so as to obtain N embedded representation vectors; A determination module, configured to determine whether each of the N nodes is abnormal according to the N embedded representation vectors and the N subgraphs; Among them, the node of the first node type is used to represent the repayment account, the node of the second node type is used to represent the credit card account, and the node of the third node type is used to represent the merchant account. The account of the first account type is the repayment account, the account of the second account type is the credit card account, and the account of the third account type is the merchant account; the event of the first event type is the repayment event performed by the repayment account on the credit card account, and the event of the second event type is the consumption event performed by the credit card account on the merchant account.

11. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method described in any one of claims 1 to 9 when executed by a processor.

12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, the steps of the method described in any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Account relationship identification method and device

    CN112215500A

  • Account detection method, device and equipment based on graph neural network

    CN112818257A