Methods, systems, and computer program products for generic depth map neural networks
By introducing trainable depth parameters into the graph neural network and redefining the power of the eigenvalues of the Laplace matrix of the graph, the problem that graph neural networks in the prior art are difficult to automatically tune depth is solved, and better performance on heterogeneous graphs is achieved.
Patent Information
- Application Number
- CN202380066698.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-27
- Filing Date
- 2023-09-27
- Publication Date
- 2025-05-27
AI Technical Summary
Existing graph neural networks (GNNs) have challenges in dealing with heterogeneous graphs, and their depth usually requires manual setting, making it difficult to automatically tune to adapt to the homogeneous or heterogeneous properties of the graph.
By introducing trainable depth parameters, redefine the power of the eigenvalues of the Laplace matrix of the graph, a new symmetric normalized Laplace matrix is generated, thereby automatically tuning the depth of the graph convolution network during training.
It realizes that the heterogeneous or homogeneous characteristics of the graph are automatically detected without manually setting the depth and find the best feature map weight distribution, thereby improving the performance of the graph neural network when processing heterogeneous graphs.
Smart Images

Figure CN120051773A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 410,437, filed on September 27, 2022, the entire disclosure of which is incorporated herein by reference. Technical Field
[0003] This disclosure relates to graph neural networks and, in some non - limiting embodiments or aspects, to methods, systems, and computer program products for understanding the depth of graph neural networks. Background Art
[0004] Existing graph neural networks (GNNs) can be designed with fixed integer layers, which may limit the performance of GNNs. In addition, existing GNNs may depend on homogeneous assumptions, and heterogeneous graphs may still pose challenges to GNNs. Summary of the Invention
[0005] Accordingly, improved systems, devices, products, apparatuses, and / or methods for general - depth graph neural networks are provided.
[0006] According to some non - limiting embodiments or aspects, a method is provided, the method comprising: obtaining, by at least one processor, a graph G including an adjacency matrix A and an attribute matrix X, wherein an augmented adjacency matrix is determined based on the adjacency matrix A and the identity matrix I wherein based on the augmented adjacency matrix the identity matrix I, and the degree matrix D of the adjacency matrix X, a symmetric normalized Laplacian matrix L for the augmented adjacency matrix is determined sym and wherein the symmetric normalized Laplacian matrix L for the augmented adjacency matrix is determined sym and the eigen - decomposition of the symmetric normalized Laplacian matrix L for the augmented adjacency matrix provides an eigen - vector matrix U and an eigen - value matrix Λ; and training a graph convolutional network by the at least one processor by: applying, by the at least one processor, a transformation function to the symmetric normalized Laplacian matrix L for the augmented adjacency matrix sym to generate a filter obtaining, by the at least one processor, the eigen - vector matrix U and the eigen - value matrix Λ from the eigen - decomposition of the symmetric normalized Laplacian matrix L for the augmented adjacency matrix setting, by the at least one processor, a trainable depth parameter d to the power of each eigenvalue of the eigen - value matrix Λ of the symmetric normalized Laplacian matrix L sym to generate a new symmetric normalized Laplacian matrix S sym d ; and using the at least one processor, using a filter Apply graph convolution to generate an output embedding matrix H.
[0007] In some non-limiting embodiments or aspects, apply graph convolution according to the following update equation:
[0008]
[0009] where H is the output embedding matrix, σ(·) is a non-linear activation function, is the filter, X is the attribute matrix, W is the trainable weight matrix, and d is the trainable depth parameter.
[0010] In some non-limiting embodiments or aspects, the trainable depth parameter d includes a first trainable depth parameter d h and a second trainable depth parameter d l , and wherein apply graph convolution according to the following update equation:
[0011]
[0012] where H is the output embedding matrix, σ(·) is a non-linear activation function, is the filter, X is the attribute matrix, W is the trainable weight matrix, and K is a hyperparameter.
[0013] In some non-limiting embodiments or aspects, the method further includes: using the at least one processor, using a machine learning classifier to process the output embedding matrix H to generate at least one predicted label for at least one node of the graph G.
[0014] In some non-limiting embodiments or aspects, the graph G includes a plurality of edges and a plurality of nodes for the plurality of edges, wherein the plurality of nodes are associated with a plurality of accounts in an electronic payment network, and wherein the plurality of edges are associated with a plurality of transactions between the plurality of accounts in the electronic payment network.
[0015] In some non-limiting embodiments or aspects, the at least one predicted label for at least one node of the graph G includes a prediction as to whether at least one account associated with the at least one node is a malicious account.
[0016] According to some non-limiting embodiments or aspects, there is provided a system, the system comprising: at least one processor programmed and / or configured to: obtain a graph G including an adjacency matrix A and an attribute matrix X, wherein an augmented adjacency matrix is determined based on the adjacency matrix A and the identity matrix I wherein based on the augmented adjacency matrix the identity matrix I and the degree matrix D of the adjacency matrix X determine the augmented adjacency matrix The symmetric normalized Laplacian matrix L of sym , and wherein for the augmented adjacency matrix The symmetric normalized Laplacian matrix L of sym The eigen - decomposition of provides an eigen - vector matrix U and an eigenvalue matrix Λ; and training a graph convolutional network by: applying a transformation function to the symmetric normalized Laplacian matrix L of the augmented adjacency matrix The symmetric normalized Laplacian matrix L of sym to generate a filter from the eigen - decomposition of the symmetric normalized Laplacian matrix L of the augmented adjacency matrix The symmetric normalized Laplacian matrix L of sym obtaining an eigen - vector matrix U and an eigenvalue matrix Λ; setting a trainable depth parameter d to the power of each eigenvalue of the eigenvalue matrix Λ of the symmetric normalized Laplacian matrix L sym to generate a new symmetric normalized Laplacian matrix S d ; and applying graph convolution using the filter to generate an output embedding matrix H.
[0017] In some non - limiting embodiments or aspects, graph convolution is applied according to the following update equation:
[0018]
[0019] where H is the output embedding matrix, σ(·) is a non - linear activation function, is the filter, X is the attribute matrix, W is the trainable weight matrix, and d is the trainable depth parameter.
[0020] In some non - limiting embodiments or aspects, the trainable depth parameter d includes a first trainable depth parameter d h and a second trainable depth parameter d l , and wherein graph convolution is applied according to the following update equation:
[0021]
[0022] where H is the output embedding matrix, σ(·) is a non - linear activation function, is the filter, X is the attribute matrix, W is the trainable weight matrix, and K is a hyper - parameter.
[0023] In some non - limiting embodiments or aspects, the at least one processor is further programmed and / or configured to: process the output embedding matrix H using a machine - learning classifier to generate at least one predicted label for at least one node of the graph G.
[0024] In some non - limiting embodiments or aspects, graph G includes a plurality of edges and a plurality of nodes for the plurality of edges, wherein the plurality of nodes are associated with a plurality of accounts in an electronic payment network, and wherein the plurality of edges are associated with a plurality of transactions between the plurality of accounts in the electronic payment network.
[0025] In some non - limiting embodiments or aspects, the at least one prediction label for the at least one node of graph G includes a prediction as to whether at least one account associated with the at least one node is a malicious account.
[0026] According to some non - limiting embodiments or aspects, there is provided a computer program product including a non - transitory computer - readable medium, the non - transitory computer - readable medium including program instructions that, when executed by at least one processor, cause the at least one processor to: obtain graph G including an adjacency matrix A and an attribute matrix X, wherein an augmented adjacency matrix is determined based on the adjacency matrix A and the identity matrix I wherein based on the augmented adjacency matrix the degree matrix D of the identity matrix I and the adjacency matrix X to determine for the augmented adjacency matrix the symmetric normalized Laplacian matrix L sym , and wherein for the augmented adjacency matrix the symmetric normalized Laplacian matrix L sym the eigen - decomposition of provides an eigen - vector matrix U and an eigen - value matrix Λ; and train a graph convolutional network by: applying a transformation function to the symmetric normalized Laplacian matrix L for the augmented adjacency matrix sym to generate a filter obtain the eigen - vector matrix U and the eigen - value matrix Λ from the eigen - decomposition of the symmetric normalized Laplacian matrix L for the augmented adjacency matrix sym ; set a trainable depth parameter d to the power of each eigenvalue of the eigen - value matrix Λ of the symmetric normalized Laplacian matrix L sym to generate a new symmetric normalized Laplacian matrix S d ; and apply graph convolution using the filter to generate an output embedding matrix H.
[0027] In some non - limiting embodiments or aspects, graph convolution is applied according to the following update equation:
[0028]
[0029] where H is the output embedding matrix, σ(·) is a non - linear activation function, is a filter, X is an attribute matrix, W is a trainable weight matrix, and d is a trainable depth parameter.
[0030] In some non-limiting embodiments or aspects, the trainable depth parameter d includes a first trainable depth parameter d h and a second trainable depth parameter d l , and wherein graph convolution is applied according to the following update equation:
[0031]
[0032] where H is an output embedding matrix, σ(·) is a non-linear activation function, is a filter, X is an attribute matrix, W is a trainable weight matrix, and K is a hyperparameter.
[0033] In some non-limiting embodiments or aspects, the program instructions, when executed by at least one processor, further cause the at least one processor to: process the output embedding matrix H using a machine learning classifier to generate at least one predicted label for at least one node of the graph G.
[0034] In some non-limiting embodiments or aspects, the graph G includes a plurality of edges and a plurality of nodes for the plurality of edges, wherein the plurality of nodes are associated with a plurality of accounts in an electronic payment network, and wherein the plurality of edges are associated with a plurality of transactions between the plurality of accounts in the electronic payment network.
[0035] In some non-limiting embodiments or aspects, the at least one predicted label for at least one node of the graph G includes a prediction as to whether at least one account associated with the at least one node is a malicious account.
[0036] Additional embodiments or aspects are set forth in the following numbered clauses:
[0037] Clause 1: A method, comprising: obtaining, using at least one processor, a graph G including an adjacency matrix A and an attribute matrix X, wherein an augmented adjacency matrix is determined based on the adjacency matrix A and the identity matrix I wherein based on the augmented adjacency matrix the identity matrix I and the degree matrix D of the adjacency matrix X, a symmetric normalized Laplacian matrix L for the augmented adjacency matrix is determined, and wherein an eigen decomposition of the symmetric normalized Laplacian matrix L for the augmented adjacency matrix sym provides an eigenvector matrix U and an eigenvalue matrix Λ; and training a graph convolutional network using the at least one processor by: applying, using the at least one processor, a transformation function to the augmented adjacency matrix of the symmetric normalized Laplacian matrix L sym and the eigenvector matrix U to obtain a transformed eigenvector matrix; and using the at least one processor, training a graph convolutional network by: using the at least one processor, applying a transformation function to the augmented adjacency matrix The symmetric normalized Laplacian matrix L sym to generate a filter Using the at least one processor, from the symmetric normalized Laplacian matrix L for the augmented adjacency matrix to obtain an eigenvector matrix U and an eigenvalue matrix Λ; using the at least one processor, setting the trainable depth parameter d to the power of each eigenvalue of the eigenvalue matrix Λ of the symmetric normalized Laplacian matrix L sym to generate a new symmetric normalized Laplacian matrix S sym ; and using the at least one processor, applying graph convolution using the filter d to generate an output embedding matrix H.
[0038] Clause 2: The method according to clause 1, wherein the graph convolution is applied according to the following update equation:
[0039]
[0040] where H is the output embedding matrix, σ(·) is a non-linear activation function, is the filter, X is the attribute matrix, W is the trainable weight matrix, and d is the trainable depth parameter.
[0041] Clause 3: The method according to any one of clauses 1 or 2, wherein the trainable depth parameter d includes a first trainable depth parameter d h and a second trainable depth parameter d l , and wherein the graph convolution is applied according to the following update equation:
[0042]
[0043] where H is the output embedding matrix, σ(·) is a non-linear activation function, is the filter, X is the attribute matrix, W is the trainable weight matrix, and K is a hyperparameter.
[0044] Clause 4: The method according to any one of clauses 1-3, further comprising: using the at least one processor, processing the output embedding matrix H using a machine learning classifier to generate at least one predicted label for at least one node of the graph G.
[0045] Clause 5: The method according to any one of clauses 1-4, wherein the graph G includes a plurality of edges and a plurality of nodes for the plurality of edges, wherein the plurality of nodes are associated with a plurality of accounts in an electronic payment network, and wherein the plurality of edges are associated with a plurality of transactions between the plurality of accounts in the electronic payment network.
[0046] Clause 6: The method according to any one of Clauses 1-5, wherein the at least one predicted label for the at least one node of the graph G includes a prediction as to whether at least one account associated with the at least one node is a malicious account.
[0047] Clause 7: A system comprising: at least one processor programmed and / or configured to: obtain a graph G comprising an adjacency matrix A and an attribute matrix X, wherein an augmented adjacency matrix is determined based on the adjacency matrix A and the identity matrix I wherein based on the augmented adjacency matrix the identity matrix I and the degree matrix D of the adjacency matrix X to determine a symmetric normalized Laplacian matrix L for the augmented adjacency matrix and wherein the eigen-decomposition of the symmetric normalized Laplacian matrix L for the augmented adjacency matrix sym provides an eigenvector matrix U and an eigenvalue matrix Λ; and train a graph convolutional network by: applying a transformation function to the symmetric normalized Laplacian matrix L for the augmented adjacency matrix to generate a filter sym obtaining an eigenvector matrix U and an eigenvalue matrix Λ from the eigen-decomposition of the symmetric normalized Laplacian matrix L for the augmented adjacency matrix ; setting a trainable depth parameter d to the power of each eigenvalue of the eigenvalue matrix Λ of the symmetric normalized Laplacian matrix L sym to generate a new symmetric normalized Laplacian matrix S ; and applying graph convolution using the filter to generate an output embedding matrix H. sym sym d ; and applying graph convolution using the filter d to generate an output embedding matrix H. ; and applying graph convolution using the filter
[0048] Clause 8: The system according to Clause 7, wherein the graph convolution is applied according to the following update equation:
[0049]
[0050] where H is the output embedding matrix, σ(·) is a non-linear activation function, is the filter, X is the attribute matrix, W is a trainable weight matrix, and d is a trainable depth parameter.
[0051] Clause 9: The system according to any one of Clauses 7 or 8, wherein the trainable depth parameter d comprises a first trainable depth parameter d h and a second trainable depth parameter d l and wherein the graph convolution is applied according to the following update equation:
[0052]
[0053] where H is the output embedding matrix, σ(·) is a non-linear activation function, is a filter, X is an attribute matrix, W is a trainable weight matrix, and K is a hyperparameter.
[0054] Clause 10: The system according to any one of Clauses 7-9, wherein the at least one processor is further programmed and / or configured to: process the output embedding matrix H using a machine learning classifier to generate at least one predicted label for at least one node of the graph G.
[0055] Clause 11: The system according to any one of Clauses 7-10, wherein the graph G includes a plurality of edges and a plurality of nodes for the plurality of edges, wherein the plurality of nodes are associated with a plurality of accounts in an electronic payment network, and wherein the plurality of edges are associated with a plurality of transactions between the plurality of accounts in the electronic payment network.
[0056] Clause 12: The system according to any one of Clauses 7-11, wherein the at least one predicted label for the at least one node of the graph G includes a prediction as to whether at least one account associated with the at least one node is a malicious account.
[0057] Clause 13: A computer program product comprising a non-transitory computer-readable medium having program instructions that, when executed by at least one processor, cause the at least one processor to: obtain a graph G including an adjacency matrix A and an attribute matrix X, wherein an augmented adjacency matrix is determined based on the adjacency matrix A and the identity matrix I wherein based on the augmented adjacency matrix the identity matrix I and the degree matrix D of the adjacency matrix X to determine for the augmented adjacency matrix the symmetric normalized Laplacian matrix L sym and wherein the eigen-decomposition of the symmetric normalized Laplacian matrix L for the augmented adjacency matrix provides an eigenvector matrix U and an eigenvalue matrix Λ; and train a graph convolutional network by: applying a transformation function to the symmetric normalized Laplacian matrix L for the augmented adjacency matrix sym to generate a filter the symmetric normalized Laplacian matrix L for the augmented adjacency matrix sym to generate a filter obtaining an eigenvector matrix U and an eigenvalue matrix Λ from the eigen-decomposition of the symmetric normalized Laplacian matrix L for the augmented adjacency matrix the symmetric normalized Laplacian matrix L for the augmented adjacency matrix sym setting a trainable depth parameter d to the symmetric normalized Laplacian matrix L symThe power of each eigenvalue of the eigenvalue matrix Λ to generate a new symmetric normalized Laplacian matrix S d ; and applying graph convolution using a filter to generate an output embedding matrix H.
[0058] Clause 14: The computer program product according to clause 13, wherein the graph convolution is applied according to the following update equation:
[0059]
[0060] where H is the output embedding matrix, σ(·) is a non-linear activation function, is the filter, X is the attribute matrix, W is the trainable weight matrix, and d is the trainable depth parameter.
[0061] Clause 15: The computer program product according to any one of clauses 13 or 14, wherein the trainable depth parameter d includes a first trainable depth parameter d h and a second trainable depth parameter d l , and wherein the graph convolution is applied according to the following update equation:
[0062]
[0063] where H is the output embedding matrix, σ(·) is a non-linear activation function, is the filter, X is the attribute matrix, W is the trainable weight matrix, and K is a hyperparameter.
[0064] Clause 16: The computer program product according to any one of clauses 13 - 15, wherein the program instructions, when executed by at least one processor, further cause the at least one processor to: process the output embedding matrix H using a machine learning classifier to generate at least one predicted label for at least one node of the graph G.
[0065] Clause 17: The computer program product according to any one of clauses 13 - 16, wherein the graph G includes a plurality of edges and a plurality of nodes for the plurality of edges, wherein the plurality of nodes are associated with a plurality of accounts in an electronic payment network, and wherein the plurality of edges are associated with a plurality of transactions between the plurality of accounts in the electronic payment network.
[0066] Clause 18: The computer program product according to any one of clauses 13 - 17, wherein the at least one predicted label for the at least one node of the graph G includes a prediction as to whether at least one account associated with the at least one node is a malicious account.
[0067] These and other features and characteristics of the present disclosure, as well as the methods of operation and functions of the related structural elements and combinations of the various parts, and the manufacturing economy, will become more apparent when considering the following description and the appended claims in conjunction with the accompanying drawings, all of which form a part of this specification, where like reference numerals designate corresponding parts in the various figures. It should be understood, however, that the drawings are for illustrative and descriptive purposes only and are not intended as a definition of the limits. Unless the context clearly dictates otherwise, the singular forms "a" and "the" as used in this specification and the claims include plural referents. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Additional advantages and details are explained in more detail below with reference to the exemplary embodiments shown in the schematic drawings, in which:
[0069] Figure 1 is a diagram of a non-limiting embodiment or aspect of an environment in which the systems, devices, products, apparatuses, and / or methods described herein can be implemented;
[0070] Figure 2 is Figure 1 a diagram of a non-limiting embodiment or aspect of one or more devices and / or components of one or more systems;
[0071] Figure 3A is a flowchart of a non-limiting embodiment or aspect of a process for a general deep graph neural network;
[0072] Figure 3B is a flowchart of a non-limiting embodiment or aspect of a process for a general deep graph neural network;
[0073] Figure 4 shows decomposing an example symmetric normalized adjacency matrix into three feature maps;
[0074] Figure 5 is a graph showing the feature map weights compared to the eigenvalues of non-limiting embodiments or aspects of a simplified graph convolutional network (SGC) and a graph convolutional network (GCN) at different depths;
[0075] Figure 6 shows the difference between the filters for GCN / SGC and the filters for a non-limiting embodiment or aspect of GCN;
[0076] Figure 7 is a table showing the semi-supervised node classification performance regarding homogeneous graphs and heterogeneous graphs;
[0077] Figure 8 is a graph showing the node classification accuracy of a non-limiting embodiment or aspect of GCN with respect to the trainable depth parameters on an example dataset;
[0078] Figure 9 is a table showing the performance of a single layer of conventional GCN on the augmented diffusion matrix;
[0079] Figure 10 is a graph showing the weights of the eigenmaps of the eigenvalues on the augmented diffusion matrix and the original diffusion matrix relative to the example dataset; and
[0080] Figure 11 is a heatmap showing the difference between the augmented diffusion matrix and the original diffusion matrix for the example dataset. DETAILED DESCRIPTION
[0081] It should be understood that, unless expressly specified to the contrary, the present disclosure may assume various alternative variations and step sequences. It should also be understood that the specific apparatus and processes shown in the drawings and described in the following specification are merely exemplary and non-limiting embodiments or aspects. Accordingly, the specific dimensions and other physical characteristics related to the embodiments or aspects disclosed herein should not be considered limiting.
[0082] Aspects, components, elements, structures, acts, steps, functions, instructions, etc. used herein should not be construed as critical or essential unless expressly so described. Also, as used herein, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more" and "at least one". Further, as used herein, the term "set" is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and may be used interchangeably with "one or more" or "at least one". In cases where only one item is intended, the term "one" or similar language is used. Moreover, as used herein, the term "having" and / or its like is intended to be an open-ended term. Additionally, unless otherwise expressly stated, the phrase "based on" is intended to mean "at least partially based on".
[0083] As used herein, the term "communicate" can refer to the receipt, acceptance, transmission, conveyance, provision, etc. of data (e.g., information, signals, messages, instructions, commands, etc.). A unit (e.g., a device, a system, a component of a device or system, a combination thereof, and / or the like) communicating with another unit means that the one unit is capable of receiving information from and / or sending information to the other unit, directly or indirectly. This can refer to a direct or indirect connection, in essence, wired and / or wireless (e.g., a direct communication connection, an indirect communication connection, and / or the like). Additionally, although the information sent can be modified, processed, relayed, and / or routed between a first unit and a second unit, the two units can still communicate with each other. For example, even if a first unit receives information passively and does not actively send information to a second unit, the first unit can still communicate with the second unit. As another example, if at least one intermediate unit processes the information received from a first unit and conveys the processed information to a second unit, then the first unit can communicate with the second unit.
[0084] Obviously, the systems and / or methods described herein can be implemented in different forms of hardware, software, or a combination of hardware and software. The actual specific control hardware or software code for implementing these systems and / or methods does not limit the implementation. Therefore, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, and it should be understood that the software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0085] In this document, some non-limiting embodiments or aspects are described in connection with thresholds. As used herein, meeting a threshold can refer to a value that is greater than the threshold, more than the threshold, higher than the threshold, greater than or equal to the threshold, less than the threshold, fewer than the threshold, lower than the threshold, less than or equal to the threshold, equal to the threshold, etc.
[0086] As used herein, the term "transaction service provider" can refer to an entity that receives a transaction authorization request from a merchant or other entity and, in some cases, provides payment assurance through an agreement between the transaction service provider and the issuer institution. For example, a transaction service provider can include, for example, a payment network, or any other entity that processes transactions. The term "transaction processing system" can refer to one or more computing devices operated by or on behalf of a transaction service provider, such as a transaction processing server that executes one or more software applications. The transaction processing system can include one or more processors and, in some non-limiting embodiments, can be operated by or on behalf of a transaction service provider.
[0087] As used herein, the term "account identifier" can include one or more primary account numbers (PANs), tokens, or other identifiers associated with a customer account. The term "token" can refer to an identifier used as an alternative or replacement for an original account identifier such as a PAN. An account identifier can be alphanumeric or any combination of characters and / or symbols. A token can be associated with a PAN or other original account identifier in one or more data structures (e.g., one or more databases, etc.) such that the token can be used to conduct a transaction without directly using the original account identifier. In some examples, an original account identifier such as a PAN can be associated with multiple tokens for different individuals or purposes.
[0088] As used herein, the terms "issuer institution", "portable financial device issuer", "issuer", or "issuer bank" can refer to one or more entities that provide one or more accounts to a user (e.g., a customer, consumer, organization, etc.) for conducting transactions (e.g., payment transactions), such as initiating credit card payment transactions and / or debit card payment transactions. For example, an issuer institution can provide a user with an account identifier such as a PAN that uniquely identifies one or more accounts associated with the user. The account identifier can be included on a portable financial device such as a physical financial instrument (e.g., a payment card), and / or can be electronic and used for electronic payments. In some non-limiting embodiments or aspects, an issuer institution can be associated with a bank identification number (BIN) that uniquely identifies the issuer institution. As used herein, the term "issuer institution system" can refer to one or more computer systems operated by or on behalf of an issuer institution, such as a server computer executing one or more software applications. For example, an issuer institution system can include one or more authorization servers for authorizing payment transactions.
[0089] As used herein, the term "merchant" can refer to an individual or entity that provides goods and / or services or access to goods and / or services to a user (e.g., a customer) based on a transaction (e.g., a payment transaction). As used herein, the term "merchant" or "merchant system" can also refer to one or more computer systems, computing devices, and / or software applications operated by or on behalf of a merchant, such as a server computer that executes one or more software applications. As used herein, a "point of sale (POS) system" can refer to one or more computers and / or peripheral devices used by a merchant to conduct a payment transaction with a user, including one or more card readers, near field communication (NFC) receivers, radio frequency identification (RFID) receivers, and / or other contactless transceivers or receivers, contact-based receivers, payment terminals, computers, servers, input devices, and / or other similar devices that can be used to initiate a payment transaction. A POS system can be part of a merchant system. A merchant system can also include a merchant plug-in for facilitating online Internet-based transactions via a merchant web page or software application. A merchant plug-in can include software that runs on a merchant server or is hosted by a third party to facilitate such online transactions.
[0090] As used herein, the term "mobile device" can refer to one or more portable electronic devices configured to communicate with one or more networks. By way of example, a mobile device can include a cellular phone (e.g., a smart phone or a standard cellular phone), a portable computer (e.g., a tablet computer, a laptop computer, etc.), a wearable device (e.g., a watch, glasses, lenses, clothing, etc.), a personal digital assistant (PDA), and / or other similar devices. As used herein, the terms "client device" and "user device" refer to any electronic device configured to communicate with one or more servers or remote devices and / or systems. A client device or user device can include a mobile device, a network-enabled appliance (e.g., a network-enabled television, refrigerator, thermostat, etc.), a computer, a POS system, and / or any other device or system capable of communicating with a network.
[0091] As used herein, the term "computing device" can refer to one or more electronic devices configured to process data. In some examples, a computing device can include the necessary components for receiving, processing, and outputting data, such as a processor, a display, a memory, an input device, a network interface, and / or the like. A computing device can be a mobile device. By way of example, a mobile device can include a cellular phone (e.g., a smart phone or a standard cellular phone), a portable computer, a wearable device (e.g., a watch, glasses, lenses, clothing, and / or the like), a PDA, and / or other similar devices. A computing device can also be a desktop computer or other form of non-mobile computer.
[0092] As used herein, the term "payment device" can refer to a portable financial device, an electronic payment device, a payment card (e.g., a credit card or debit card), a gift card, a smart card, a smart medium, a payroll card, a healthcare card, a wristband, a machine-readable medium containing account information, a keychain device or fob, an RFID transponder, a retailer discount or membership card, a cellular phone, an electronic wallet mobile application, a PDA, a pager, a security card, a computer, an access card, a wireless terminal, a transponder, etc. In some non-limiting embodiments or aspects, the payment device may include volatile or non-volatile memory for storing information (e.g., an account identifier, an account holder name, etc.).
[0093] As used herein, the term "server" and / or "processor" can refer to or include one or more computing devices operated by multiple parties in a network environment such as the Internet or facilitating communication and processing among the multiple parties, but it should be understood that communication can be facilitated through one or more public or private network environments and there may be various other arrangements. Additionally, multiple computing devices (e.g., servers, POS devices, mobile devices, etc.) communicating directly or indirectly in a network environment can constitute a "system". As used herein, a reference to a "server" or "processor" can refer to the previously described server and / or processor stated to implement a previous step or function, a different server and / or processor, and / or a combination of servers and / or processors. For example, as used in the specification and claims, a first server and / or a first processor stated to perform a first step or function can refer to the same or a different server and / or processor stated to perform a second step or function.
[0094] As used herein, the term "acquirer" can refer to an entity licensed and / or approved by a transaction service provider to initiate transactions using the transaction service provider's portable financial device. An acquirer can also refer to one or more computer systems operated by or on behalf of the acquirer, such as a server computer (e.g., an "acquirer server") executing one or more software applications. The "acquirer" can be a merchant bank, or in some cases, the merchant system can be the acquirer. The transaction can include an original credit transaction (OCT) and an account funding transaction (AFT). The transaction service provider can authorize the acquirer to sign up merchants of the service provider to initiate transactions using the transaction service provider's portable financial device. The acquirer can contract with a payment service provider to enable the service provider to sponsor merchants. The acquirer can monitor the compliance of the payment service provider according to the regulations of the transaction service provider. The acquirer can conduct due diligence on the payment service provider and ensure appropriate due diligence is conducted before signing up sponsored merchants. The acquirer can be responsible for all transaction service provider programs they operate or sponsor. The acquirer can be responsible for the actions of their payment service providers and the merchants sponsored by them or their payment service providers.
[0095] As used herein, the term "payment gateway" may refer to an entity and / or a payment processing system operated by or on behalf of such entity, where the entity (e.g., merchant service provider, payment service provider, payment service facilitator, payment service facilitator under contract with an acquirer, payment aggregator, etc.) provides payment services (e.g., transaction service provider payment services, payment processing services, etc.) to one or more merchants. The payment services may be associated with the use of a portable financial device managed by a transaction service provider. As used herein, the term "payment gateway system" may refer to one or more computer systems, computer devices, servers, server groups, etc., operated by or on behalf of a payment gateway.
[0096] As used herein, the term "authentication system" may refer to one or more computing devices that authenticate users and / or accounts, such as but not limited to transaction processing systems, merchant systems, issuer systems, payment gateways, third-party authentication services, etc.
[0097] As used herein, the terms "request", "response", "request message", and "response message" may refer to one or more messages, data packets, signals, and / or data structures used to transfer data between two or more components or units.
[0098] As used herein, the term "application programming interface" (API) may refer to computer code that allows communication between different systems or system components (hardware and / or software). For example, an API may include function calls, functions, subroutines, communication protocols, fields, etc., that can be used and / or accessed by other systems or other system components (hardware and / or software).
[0099] As used herein, the term "user interface" or "graphical user interface" refers to a generated display, such as one or more graphical user interfaces (GUIs) with which a user can interact directly or indirectly (e.g., via a keyboard, mouse, touch screen, etc.).
[0100] Graph Convolutional Networks (GCNs) have shown strong capabilities in various graph learning tasks such as node classification, community detection, etc. Since the representational ability of GCNs is largely determined by their depth (e.g., the number of graph convolutional layers, etc.), a great deal of research work has been carried out to find the optimal depth that enhances the model's ability for downstream tasks. After increasing the depth, an over-smoothing problem emerges: if the depth of the GCN exceeds an uncertain threshold, the performance of the GCN deteriorates. It has been revealed that graph convolutional operations are a special form of Laplacian smoothing. Thus, the similarity between graph node embeddings grows with depth, making these embeddings eventually indistinguishable. Various techniques have been developed to alleviate this problem. For example, applying pairwise normalization can make distant nodes dissimilar, and discarding sampled edges during training slows down the growth of embedding smoothness with depth.
[0101] In addition to the over-smoothing problem caused by large GCN depths, another fundamental phenomenon widely present in real-world graphs is homophily and heterophily. In homophilic graphs, nodes with similar labels or attributes tend to be interconnected, while in heterophilic graphs, connected nodes usually have different labels or dissimilar attributes. Many Graph Neural Networks (GNNs) are developed based on the homophily assumption, and models that can perform well on heterophilic graphs usually require special handling and complex designs. Despite the achievements of these methods, little correlation has been found between the depth of the GNN models employed and their ability to represent graph heterophily.
[0102] For many (if not all) GNNs, the depth is manually set as a hyperparameter before training, and finding the appropriate depth usually requires a large number of trials or good prior knowledge of the graph dataset. Since the depth represents the number of graph convolutional operations and naturally only takes positive integer values, little attention has been paid to the question of whether non-integer depths are achievable, and if so, whether non-integer depths are actually meaningful, and whether non-integer depths can bring unique advantages to current graph learning mechanisms.
[0103] Non-limiting embodiments or aspects of the present disclosure can revisit the depth of GCNs from spectral and spatial perspectives and elucidate the interdependencies among the following components in graph learning: the depth of GCNs, the spectrum of graph signals, and the homophily / heterophily of the underlying graph. First, via the eigen-decomposition of the symmetric normalized graph Laplacian, non-limiting embodiments or aspects of the present disclosure can present the correlation between graph homophily / heterophily and eigenvector frequencies. Second, by introducing the concept of eigen-graphs, non-limiting embodiments or aspects of the present disclosure can show that graph topology is equivalent to a weighted linear combination of eigen-graphs, and the weight values determine the ability of GCNs to capture homophilous / heterophilous graph signals. Third, non-limiting embodiments or aspects of the present disclosure can reveal that the eigen-graph weights can be controlled by the depth of GCNs, such that an automatically tunable depth parameter is required to adjust the eigen-graph weights to a specified distribution to match the underlying graph homophily / heterophily.
[0104] Non-limiting embodiments or aspects of the present disclosure can provide a method, system, and / or computer program product that obtains a graph G including an adjacency matrix A and an attribute matrix X, where an augmented adjacency matrix is determined based on the adjacency matrix A and the identity matrix I where based on the augmented adjacency matrix the degree matrix D of the identity matrix I and the adjacency matrix X, a symmetric normalized Laplacian matrix L for the augmented adjacency matrix is determined and where the eigen-decomposition of the symmetric normalized Laplacian matrix L for the augmented adjacency matrix sym provides an eigenvector matrix U and an eigenvalue matrix Λ; and trains a graph convolutional network by: applying a transformation function to the symmetric normalized Laplacian matrix L for the augmented adjacency matrix to generate a filter sym obtaining the eigenvector matrix U and the eigenvalue matrix Λ from the eigen-decomposition of the symmetric normalized Laplacian matrix L for the augmented adjacency matrix setting a trainable depth parameter d to the power of each eigenvalue of the eigenvalue matrix Λ of the symmetric normalized Laplacian matrix L sym to generate a new symmetric normalized Laplacian matrix S from the augmented adjacency matrix the symmetric normalized Laplacian matrix L sym and applying graph convolution using the filter sym to generate an output embedding matrix H. For example, graph convolution is applied according to the following update equation: d where H is the output embedding matrix, σ(·) is a non-linear activation function, is the filter, X is the attribute matrix, W is a trainable weight matrix, and d is a trainable depth parameter. where H is the output embedding matrix, σ(·) is a non-linear activation function, is the filter, X is the attribute matrix, W is a trainable weight matrix, and d is a trainable depth parameter.
[0105] In this way, non-limiting embodiments or aspects of the present disclosure can achieve adaptive GCN depth by extending its definition from positive integers to arbitrary real numbers, which has theoretical feasibility guarantees from functional calculus. Using the trainable depth parameter, non-limiting embodiments or aspects of the present disclosure can provide a simple and powerful model, for example, a redefined depth-GCN (ReD-GCN) with two variants (such as RED-GCN-S and Red-GCN-D, etc.). For example, non-limiting embodiments or aspects of the present disclosure can perform eigenvalue decomposition of the normalized adjacency matrix to obtain the basis matrix and the eigenvalue matrix, add the depth parameter as the power of the eigenvalue matrix to construct a new adjacency matrix, and multiply the new trainable adjacency matrix by the given feature matrix for downstream tasks. Accordingly, non-limiting embodiments or aspects of the present disclosure can extend the depth of the GNN from the integer domain to the real domain and search for the optimal depth and / or automatically detect the heterogeneous / homogeneous characteristics of a given graph. For example, if the depth parameter is less than 0, the input graph can be a heterogeneous graph. Otherwise, the input graph can be a homogeneous graph.
[0106] A large number of experiments described herein demonstrate the automatic optimal depth search ability and find that negative depth plays a role in processing heterogeneous graphs. A systematic study of the optimal depth is carried out in both the spectral domain and the spatial domain, which in turn provides a novel graph augmentation method with clear geometric interpretability, which has the greatest advantage over the original input topology, especially for graphs with heterogeneity.
[0107] Now refer to Figure 1 , Figure 1 is a diagram of an example environment 100 in which the devices, systems, methods, and / or products described herein can be implemented. As Figure 1 shown, the environment 100 includes a transaction processing network 101, a user device 112, and / or a communication network 116. The transaction processing network can include a merchant system 102, a payment gateway system 104, an acquirer system 106, a transaction service provider system 108, and an issuer system 110. The transaction processing network 101, the merchant system 102, the payment gateway system 104, the acquirer system 106, the transaction service provider system 108, the issuer system 110, and / or the user device 112 can be interconnected by a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection (e.g., establish a connection for communication, etc.).
[0108] The merchant system 102 may include one or more devices that are capable of receiving information and / or data (e.g., via a communication network 116, etc.) from the payment gateway system 104, the acquirer system 106, the transaction service provider system 108, the issuer system 110, and / or the user device 112, and / or transmitting information and / or data (e.g., via a communication network 116, etc.) to the payment gateway system 104, the acquirer system 106, the transaction service provider system 108, the issuer system 110, and / or the user device 112. The merchant system 102 may include devices capable of receiving information and / or data from the user device 112 and / or transmitting information and / or data to the user device 112 via a communication connection with the user device 112 (e.g., an NFC communication connection, an RFID communication connection, a communication connection, etc.). For example, the merchant system 102 may include computing devices such as servers, server clusters, client devices, client device clusters, and / or other similar devices. In some non-limiting embodiments or aspects, the merchant system 102 may be associated with the merchant described herein. In some non-limiting embodiments or aspects, the merchant system 102 may include one or more devices that can be used by the merchant to conduct payment transactions with users, such as computers, computer systems, and / or peripheral devices. For example, the merchant system 102 may include a POS device and / or a POS system.
[0109] The payment gateway system 104 may include one or more devices that are capable of receiving information and / or data (e.g., via a communication network 116, etc.) from the merchant system 102, the acquirer system 106, the transaction service provider system 108, the issuer system 110, and / or the user device 112, and / or transmitting information and / or data (e.g., via a communication network 116, etc.) to the merchant system 102, the acquirer system 106, the transaction service provider system 108, the issuer system 110, and / or the user device 112. For example, the payment gateway system 104 may include computing devices such as servers, server clusters, and / or other similar devices. In some non-limiting embodiments or aspects, the payment gateway system 104 is associated with the payment gateway described herein.
[0110] The acquirer system 106 may include one or more devices capable of receiving information and / or data (e.g., via communication network 116, etc.) from the merchant system 102, the payment gateway system 104, the transaction service provider system 108, the issuer system 110, and / or the user device 112, and / or transmitting information and / or data (e.g., via communication network 116, etc.) to the merchant system 102, the payment gateway system 104, the transaction service provider system 108, the issuer system 110, and / or the user device 112. For example, the acquirer system 106 may include computing devices such as servers, server clusters, and / or other similar devices. In some non-limiting embodiments or aspects, the acquirer system 106 may be associated with the acquirer described herein.
[0111] The transaction service provider system 108 may include one or more devices capable of receiving information and / or data (e.g., via communication network 116, etc.) from the merchant system 102, the payment gateway system 104, the acquirer system 106, the issuer system 110, and / or the user device 112, and / or transmitting information and / or data (e.g., via communication network 116, etc.) to the merchant system 102, the payment gateway system 104, the acquirer system 106, the issuer system 110, and / or the user device 112. For example, the transaction service provider system 108 may include computing devices such as servers (e.g., transaction processing servers, etc.), server clusters, and / or other similar devices. In some non-limiting embodiments or aspects, the transaction service provider system 108 may be associated with the transaction service provider described herein. In some non-limiting embodiments or aspects, the transaction service provider 108 may include and / or access one or more internal and / or external databases including transaction data.
[0112] The issuer system 110 may include one or more devices capable of receiving information and / or data (e.g., via communication network 116, etc.) from the merchant system 102, the payment gateway system 104, the acquirer system 106, the transaction service provider system 108, and / or the user device 112, and / or transmitting information and / or data (e.g., via communication network 116, etc.) to the merchant system 102, the payment gateway system 104, the acquirer system 106, the transaction service provider system 108, and / or the user device 112. For example, the issuer system 110 may include computing devices such as servers, server clusters, and / or other similar devices. In some non-limiting embodiments or aspects, the issuer system 110 may be associated with the issuer institution described herein. For example, the issuer system 110 may be associated with an issuer institution that issues payment accounts or instruments (e.g., credit accounts, debit accounts, credit cards, debit cards, etc.) to users (e.g., users associated with the user device 112, etc.).
[0113] In some non - limiting embodiments or aspects, the transaction processing network 101 includes multiple systems in a communication path for processing transactions. For example, the transaction processing network 101 may include a merchant system 102, a payment gateway system 104, an acquirer system 106, a transaction service provider system 108, and / or an issuer system 110 in a communication path (e.g., a communication path, a communication channel, a communication network, etc.) for processing electronic payment transactions. For example, the transaction processing network 101 may process (e.g., initiate, conduct, authorize, etc.) electronic payment transactions via a communication path among the merchant system 102, the payment gateway system 104, the acquirer system 106, the transaction service provider system 108, and / or the issuer system 110.
[0114] The user device 112 may include one or more devices that are capable of receiving information and / or data (e.g., via a communication network 116, etc.) from the merchant system 102, the payment gateway system 104, the acquirer system 106, the transaction service provider system 108, and / or the issuer system 110, and / or of transmitting information and / or data (e.g., via a communication network 116, etc.) to the merchant system 102, the payment gateway system 104, the acquirer system 106, the transaction service provider system 108, and / or the issuer system 110. For example, the user device 112 may include a client device, etc. In some non - limiting embodiments or aspects, the user device 112 is capable of receiving information (e.g., from the merchant system 102, etc.) via a short - range wireless communication connection (e.g., an NFC communication connection, an RFID communication connection, a communication connection, etc.), and / or of transmitting information via a short - range wireless communication connection (e.g., to the merchant system 102). In some non - limiting embodiments or aspects, the user device 112 may include an application associated with the user device 112, such as an application stored on the user device 112, a mobile application stored and / or executed on the user device 112 (e.g., a mobile device application, a native application for a mobile device, a mobile cloud application for a mobile device, an electronic wallet application, an issuer bank application, etc.). In some non - limiting embodiments or aspects, for one or more transactions in the payment network, the user device 112 may be associated with a sender account and / or a receiver account in the payment network.
[0115] The communication network 116 may include one or more wired and / or wireless networks. For example, the communication network 116 may include a cellular network (such as a Long-Term Evolution (LTE) network, a Third Generation (3G) network, a Fourth Generation (4G) network, a Fifth Generation (5G) network, a Code Division Multiple Access (CDMA) network, etc.), a Public Land Mobile Network (PLMN), a Local Area Network (LAN), a Wide Area Network (WAN), a Metropolitan Area Network (MAN), a telephone network (such as a Public Switched Telephone Network (PSTN)), a private network, an ad-hoc network, an intranet, the Internet, a fiber-optic based network, a cloud computing network, etc., and / or a combination of these or other types of networks.
[0116] Provide Figure 1 The number and arrangement of the devices and systems shown are for example purposes. There may be additional devices and / or systems, fewer devices and / or systems, different devices and / or systems, and / or differently arranged devices and / or systems than those shown. Additionally, two or more of the devices and / or systems shown in Figure 1 may be implemented within a single device and / or system, or a single device and / or system shown in Figure 1 may be implemented as multiple distributed devices and / or systems. Additionally or alternatively, a set of devices and / or systems of the environment 100 (e.g., one or more devices or systems) may perform one or more functions described as being performed by another set of devices and / or systems of the environment 100. Figure 1 Now referring to
[0117] is a diagram of example components of the device 200. The device 200 may correspond to one or more devices of the merchant system 102, one or more devices of the payment gateway system 104, one or more devices of the acquirer system 106, one or more devices of the transaction service provider system 108, one or more devices of the issuer system 110, and / or the user device 112 (e.g., one or more devices of the system of the user device 112, etc.). In some non-limiting embodiments or aspects, one or more devices of the merchant system 102, one or more devices of the payment gateway system 104, one or more devices of the acquirer system 106, one or more devices of the transaction service provider system 108, one or more devices of the issuer system 110, and / or the user device 112 (e.g., one or more devices of the system of the user device 112, etc.) may include at least one device 200 and / or at least one component of the device 200. As Figure 2 , Figure 2 shown, the device 200 may include a bus 202, a processor 204, a memory 206, a storage component 208, an input component 210, an output component 212, and a communication interface 214. Figure 2
[0118] Bus 202 may include components that permit communication among components of device 200. In some non-limiting embodiments or aspects, processor 204 may be implemented in hardware, software, or a combination of hardware and software. For example, processor 204 may include a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), etc.), a microprocessor, a digital signal processor (DSP), and / or any processing component that can be programmed to perform functions (e.g., a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc.). Memory 206 may include random access memory (RAM), read only memory (ROM), and / or another type of dynamic or static storage device that stores information and / or instructions for use by processor 204 (e.g., flash memory, magnetic memory, optical memory, etc.).
[0119] Storage component 208 may store information and / or software associated with the operation and use of device 200. For example, storage component 208 may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, a solid state disk, etc.), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cartridge, a magnetic tape, and / or another type of computer-readable medium, as well as corresponding drives.
[0120] Input component 210 may include components that permit device 200 to receive information, e.g., via user input (such as a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, a microphone, etc.). Additionally or alternatively, input component 210 may include sensors for sensing information (e.g., a global positioning system (GPS) component, an accelerometer, a gyroscope, an actuator, etc.). Output component 212 may include components that provide output information from device 200 (e.g., a display, a speaker, one or more light emitting diodes (LEDs), etc.).
[0121] Communication interface 214 may include transceiver-like components (e.g., a transceiver, a separate receiver and transmitter, etc.) that enable device 200 to communicate with other devices, e.g., via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection. Communication interface 214 may permit device 200 to receive information from another device and / or provide information to another device. For example, communication interface 214 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, an interface, a cellular network interface, etc.
[0122] Device 200 may perform one or more of the processes described herein. Device 200 may perform these processes based on software instructions stored by a computer-readable medium such as memory 206 and / or storage component 208 and executed by a processor 204 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), etc.). A computer-readable medium (e.g., a non-transitory computer-readable medium) is defined herein as a non-transitory memory device. A non-transitory memory device includes a memory space located within a single physical storage device or a memory space spread across multiple physical storage devices.
[0123] The software instructions may be read into memory 206 and / or storage component 208 from another computer-readable medium or from another device via communication interface 214. When executed, the software instructions stored in memory 206 and / or storage component 208 may cause processor 204 to perform one or more of the processes described herein. Additionally or alternatively, hardwired circuitry may be used in place of or in combination with the software instructions to perform one or more of the processes described herein. Accordingly, the embodiments or aspects described herein are not limited to any particular combination of hardware circuitry and software.
[0124] Memory 206 and / or storage component 208 may include a data storage device or one or more data structures (e.g., a database, etc.). Device 200 is capable of receiving information from the data storage device or one or more data structures in memory 206 and / or storage component 208, storing the information in the data storage device or one or more data structures, transmitting information to the data storage device or one or more data structures, or searching for information stored therein.
[0125] Provide Figure 2 The number and arrangement of the components shown in Figure 2 are provided as an example. In some non-limiting embodiments or aspects, device 200 may include additional components, fewer components, different components, or components arranged differently compared to the components shown in
[0126] Now refer to Figure 3A , Figure 3Ais a flowchart of a non - limiting embodiment or aspect of process 300 for ordinal encoding based on target label concentration. In some non - limiting embodiments or aspects, one or more steps in process 300 may be performed (e.g., fully, partially, etc.) by the transaction service provider system 108 (e.g., one or more devices of the transaction service provider system 108). In some non - limiting embodiments or aspects, one or more steps of process 300 may be performed (e.g., fully, partially, etc.) by another device or group of devices independent of or including the transaction service provider system 108, such as the merchant system 102 (e.g., one or more devices of the merchant system 102), the payment gateway system 104 (e.g., one or more devices of the payment gateway system 104), the acquirer system 106 (e.g., one or more devices of the acquirer system 106), the issuer system 110 (e.g., one or more devices of the issuer system 110), and / or the user device 112 (e.g., one or more devices of the system of the user device 112).
[0127] As Figure 3A shown, at step 302, process 300 includes obtaining a graph including an adjacency matrix and an attribute matrix. For example, the transaction service provider system 108 may obtain a graph G including an adjacency matrix A and an attribute matrix X. An augmented adjacency matrix can be determined based on the adjacency matrix A and the identity matrix I An augmented adjacency matrix can be based on the identity matrix I and the degree matrix D of the adjacency matrix X to determine the symmetric normalized Laplacian matrix L for the augmented adjacency matrix The symmetric normalized Laplacian matrix L for the augmented adjacency matrix sym . The symmetric normalized Laplacian matrix L for the augmented adjacency matrix The symmetric normalized Laplacian matrix L sym The eigen - decomposition of the symmetric normalized Laplacian matrix L can provide an eigen - vector matrix U and an eigen - value matrix Λ.
[0128] For example, it should be noted that in this disclosure, matrices may be represented by bold uppercase letters (e.g., A), column vectors by bold lowercase letters (e.g., u), and scalars by lowercase letters (e.g., α). Superscripts can be used for the transpose of matrices and vectors (e.g., and ). The attributed undirected graph G = {A, X} contains an adjacency matrix and an attribute matrix where the number of nodes is n and the dimension of node attributes is q. D represents the diagonal degree matrix of A. The adjacency matrix with self - loops is given by (I is the identity matrix), and all variables derived from are modified with the symbol ~, e.g., represents The diagonal degree matrix. M d represents the d-th power of matrix M, and the parameters and node embedding matrix in the d-th layer of the GCN are represented by W (d) and H (d) respectively.
[0129] In some non-limiting embodiments or aspects, the graph G includes a plurality of edges and a plurality of nodes for the plurality of edges. The plurality of nodes may be associated with a plurality of accounts in an electronic payment network. The plurality of edges may be associated with a plurality of transactions between the plurality of accounts in the electronic payment network.
[0130] The layer-by-layer message passing and aggregation of the GCN can be defined according to the following equation (1):
[0131]
[0132] where H (d) / H (d+1) represents the embedding matrix in the d / (d + 1)-th layer (H (0) = X); W (d) is a trainable parameter matrix; and σ(·) is a non-linear activation function. In the case of removing σ(·) in each layer, the simplified graph convolution (SGC) can be obtained according to the following equation (2):
[0133]
[0134] where and the parameter W (i) in each layer is compressed into a trainable
[0135] In graph theory, the graph Laplacian and its symmetric normalized counterpart L sym = I – D -1 / 2 AD -1 / 2 have the properties of the underlying graph G. L sym has eigenvalues [λ 1 , λ 2 ,..., λ n , where λ i ∈ [0, 2), Here, the eigenvalues can be placed in ascending order: 0 = λ 1 ≤ λ 2 ≤ ··· ≤ λ n < 2, and it can be eigen-decomposed as: where U = [u 1 , u 2 ,..., u n is an eigenvector matrix (u i ⊥ uj , ) and Λ is a diagonal eigenvalue matrix defined according to the following equation (3):
[0136]
[0137] For each eigenvector u i , there exists As described below, this n×n matrix can be regarded as the weighted adjacency matrix of a graph with possible negative edges, and the graph can be called the i-th eigen-graph of G. Accordingly, L sym can be written as a linear combination of all eigen-graphs weighted by the corresponding eigenvalues according to the following equation (4):
[0138]
[0139] where the first eigenvalue λ 1 = 0, and the corresponding eigen-graph has the same value 1 / n for all entries. Therefore, the SGC can be defined according to the following equation (5):
[0140]
[0141] The SGC with d layers can use d consecutive graph convolution operations, which involve multiplying times. Since the orthogonality of, that is the SGC can be defined according to the following equation (6):
[0142]
[0143] where and the depth d of the SGC serves as the power of the eigenvalues of. can be regarded as the sum of eigen-graphs weighted by the coefficient .
[0144] Graph homogeneity describes to what extent edges tend to link nodes with the same labels and similar features. Non-limiting embodiments or aspects of the present disclosure may focus on edge homogeneity: where if x is true, then <x>= 1, otherwise 0. The graph is more homogeneous for h(G) closer to 1, or more heterogeneous for h(G) closer to 0.
[0145] As Figure 3A shown, at step 304, process 300 includes training a GCN. For example, transaction service provider system 108 can train a GCN. Additional details regarding step 304 of process 300 are provided below with respect to Figure 3B
[0146] An inherent connection can be established between a feature map with small / large weights, a graph signal with high / low frequencies, and a graph with homogeneous / heterogeneous characteristics. The positive / negative depth d of a GCN can be shown to affect the feature map weights and thus determine the expressive power of the algorithm to process homogeneous / heterogeneous graph signals. By means of functional calculus, the theoretical feasibility of extending the domain of d from to can be proposed. By making d a trainable parameter, non-limiting embodiments or aspects of the model can be provided, such as RED-GCN and its variants, which can automatically detect the homogeneity / heterogeneity of the input graph and find the corresponding optimal depth.
[0147] The eigenvectors of the graph Laplacian form a complete set of basis vectors in the n-dimensional space, which can express the original node attributes X as a linear combination. From the perspective of spectral graph analysis, the frequency of the eigenvector u i reflects the degree of deviation between the j-th entry u j and the k-th entry u k of each connected node pair v i and v i in G. This deviation is measured by the zero-crossing set of u i : Z(u i ):={e=(v j ,v k )∈ε:u i [j]u i [k]<0}, where ε is the set of edges in graph G. A larger / smaller |Z(u i )| indicates a higher / lower eigenvector frequency. The zero-crossing also corresponds to the negatively weighted edges in the feature map. Due to the widespread positive correlation between λ i and |Z(u i )|, large / small eigenvalues mainly correspond to high / low frequencies of the associated eigenvectors. As Figure 4 shown, for a fictional instance with n = 3, where λ 1 = 0, |Z(u 1 )| = 0, and the feature map Closely related to the same edge weight 1 / n; negative edge weights exist in the second and third feature maps, which indicates more zero crossings (|Z(u 2 )| = 1 and |Z(u 3 )| = 2) and higher eigenvector frequencies.
[0148] Since node labels are related to their attributes, and node attribute similarity indicates the degree of smoothness / homogeneity, plus node attributes can be represented by eigenvectors, the deviation between eigenvector entries naturally implies the degree of heterogeneity. Obviously, high-frequency eigenvectors and their corresponding feature maps have advantages in graph heterogeneity capture. Therefore, high-frequency feature maps should adopt larger weights when modeling heterogeneous graphs, while low-frequency feature maps should adopt larger weights when dealing with homogeneous graphs. Furthermore, the feature map weights are controlled by the depth d of GCN / SGC (for example, for SGC with depth d, the weight of the i-th feature map is etc.), and changing the layer d of SGC will adjust the weights of different feature maps. Therefore, the depth d controls the expressive power of the model to effectively filter low / high-frequency signals of graph homogeneity / heterogeneity.
[0149] Now refer to Figure 3B , Figure 3B which is a flowchart of a non-limiting embodiment or aspect of process 350 for ordinal encoding based on target label concentration. In some non-limiting embodiments or aspects, one or more steps of process 350 may be performed (e.g., fully, partially, etc.) by the transaction service provider system 108 (e.g., one or more devices of the transaction service provider system 108). In some non-limiting embodiments or aspects, one or more steps of process 300 may be performed (e.g., fully, partially, etc.) by another device or group of devices independent of or including the transaction service provider system 108, such as the merchant system 102 (e.g., one or more devices of the merchant system 102), the payment gateway system 104 (e.g., one or more devices of the payment gateway system 104), the acquirer system 106 (e.g., one or more devices of the acquirer system 106), the issuer system 110 (e.g., one or more devices of the issuer system 110), and / or the user device 112 (e.g., one or more devices of the system of the user device 112).
[0150] As Figure 3B shown, at step 352, process 300 includes applying a transformation function to the symmetric normalized Laplacian matrix for the augmented adjacency matrix to generate a filter. For example, the transaction service provider system 108 may apply the transformation function to the symmetric normalized Laplacian matrix L of the augmented adjacency matrix sym to generate a filter
[0151] Non - limiting embodiments or aspects of the present disclosure recognize that instead of manually setting the depth d, the depth d can be built into the model as a trainable parameter, such that an appropriate set of feature map weights for homogeneous / heterogeneous matching with the graph can be automatically achieved by finding the optimal d in an end - to - end manner during training. Differentiable variables require continuity, which requires extending the depth d from the discrete positive integer domain to the continuous real number domain According to function calculus, applying an arbitrary function f to the graph Laplacian L sym is equivalent to applying the same function only to the eigenvalue matrix Λ defined according to the following equation (7):
[0152]
[0153] This also applies to and Therefore, any depth SGC can be implemented via a power function according to the following equation (8):
[0154]
[0155] However, since where and (1 - λ i ) takes on zero or negative values, there are no well - defined or involved complex - number - based calculations involving real - valued d (e.g., (-0.5) 3 / 8 etc.). Additionally, even for integer values ds, the behavior of is complex compared to , and for negative ds, it diverges when , as shown in the graph (a) of Figure 5 . Therefore, it may be difficult to obtain a favorable weight distribution by tuning d.
[0156] As Figure 3B shows, at step 354, process 300 includes obtaining an eigenvector matrix and an eigenvalue matrix from the eigendecomposition of the symmetric normalized Laplacian matrix for the augmented adjacency matrix. For example, the transaction service provider system 108 can obtain an eigenvector matrix U and an eigenvalue matrix Λ from the eigendecomposition of the symmetric normalized Laplacian matrix L for the augmented adjacency matrix sym .
[0157] To avoid this complexity and alleviate the difficulty of manipulating feature map weights, one can use the graph Laplacian L sym or The transformation function g(·) of the operation shifts g(λ i ) or to an appropriate value range so that the power of its real value d is easily obtained and performs well compared to λ i or . Without loss of generality, non-limiting embodiments or aspects of the present disclosure may mainly focus on L sym and λ i . There may be multiple choices for g(·) that meet the requirements. Non-limiting embodiments or aspects of the present disclosure may use the following equation (9):
[0158]
[0159] This choice of g(·) can have three characteristics: (1) Positive eigenvalues. Since the i-th eigenvalue λ sym of L i ∈[0, 2), the corresponding eigenvalue of is g(λ i ) = 1 / 2(2 - λ i ) ∈ (0, 1]. Therefore, for any d ∈ R, the d-th power of g(λ i ) is computable. (2) Monotonicity compared to the eigenvalue λ. As shown in the curve graph (b) of Figure 5 , g(λ i ) d = (1 - 1 λ) d is monotonically increasing / decreasing at negative / positive depths, where λ varies between 0 and 2. (3) Geometric interpretability. The filter can be expressed according to the following equation (10):
[0160]
[0161] As Figure 3B shown, at step 358, process 300 includes setting the trainable depth parameter to the power of each eigenvalue of the eigenvalue matrix of the symmetric normalized Laplacian matrix to generate a new symmetric normalized Laplacian matrix. For example, the transaction service provider system 108 can set the trainable depth parameter d to the power of each eigenvalue of the eigenvalue matrix Λ of the symmetric normalized Laplacian matrix L sym to generate a new symmetric normalized Laplacian matrix S d .
[0162] As Figure 6 shown, in the spatial domain, Normalization, adding self-loops, and scaling all edge weights by 1 / 2 (e.g., a lazy random walk, etc.), while in typical GCN / SGC contains two operations: adding self-loops and normalization.
[0163] With the help of the transformation g, the depth d is redefined in the real number domain, and the message propagation process of depth d can be achieved through the following steps: (1) Perform eigen-decomposition on L sym ; (2) Calculate through the weights g(λ i ) d and the weighted sum of all feature maps (3) Multiply by the original node attribute X.
[0164] When d takes integer values, negative d can be intuitively explained from the perspective of matrix inversion and information diffusion process. Since Therefore is 's inverse matrix. In the diffusion dynamics, X can be regarded as the intermediate state generated in a series of message propagation steps. Effectively propagates messages one step forward, while can eliminate 's influence on X and restore the original message by moving backward: Therefore, traces back to the previous state of the message in the series. However, due to its non-positive eigenvalues, neither A nor L has an inverse. Non-integer d indicates that backward or forward propagation can be a continuous process.
[0165] As Figure 3B shown, at step 358, process 300 includes using a filter for graph convolution to generate an output embedding matrix. For example, the transaction service provider system 108 can use the filter to apply graph convolution to generate the output embedding matrix H.
[0166] In some non-limiting embodiments or aspects, by further making d a trainable parameter, the non-limiting embodiments or aspects of the present disclosure can provide a redefined depth-GCN-single (RED-GCN-S), whose final node embedding matrix can be given by the following equation (11):
[0167]
[0168] where σ(·) is a non-linear activation function; W is a trainable parameter matrix; and d is a trainable depth parameter. For example, graph convolution can be applied according to equation (11). As Figure 5 As shown in the curve graph (b), the weight distribution of different frequency / feature maps can be tuned via d: (1) For d = 0, the weights are evenly distributed among all frequency components (g(λ i ) d = 1), which means that the graph signal does not prefer a specific frequency; (2) For d > 0, the weights g(λ i ) d decrease with the corresponding frequency, indicating that low-frequency components are favored, making RED-GCN-S effectively act as a low-pass filter and thus capture graph homogeneity; (3) For d < 0, high-frequency components obtain amplified weights, making RED-GCN-S act as a high-pass filter to capture graph heterogeneity. During training, RED-GCN-S can tune its frequency filtering functionality by automatically finding the optimal d to adapt to the underlying graph signal.
[0169] During optimization, RED-GCN-S includes a single depth d unified for all feature maps and selects its preference for homogeneity or heterogeneity. However, RED-GCN-S can use the complete eigen-decomposition of L sym , which may be costly for large graphs. Additionally, the high-frequency and low-frequency components in the graph signal may not be mutually exclusive, i.e., there is a possibility that the graph has both homogeneous and heterogeneous corresponding parts. Therefore, non-limiting embodiments or aspects of the present disclosure may provide another variant of RED-GCN: RED-GCN-D (dual), which introduces two separate trainable depths, the first trainable depth parameter d h and the second trainable depth parameter d l , to obtain more flexible weighting of high-frequency and low-frequency related feature maps. The Arnoldi method can be used to perform EVD on L sym and obtain the top K largest and smallest eigenpairs (λ i , u i ). By denoting U l = U[:, 0:K] and U h = U[:, n-K:n] (U l and ), non-limiting embodiments or aspects of the present disclosure can define a new diffusion matrix according to the following equation (12)
[0170]
[0171] where and are diagonal matrices of the top K smallest and largest eigenvalues, where K is a hyperparameter. λ i = 2 can be excluded (i.e., g(λ i ) = 0), since it corresponds to the presence of bipartite components. The final node embedding of RED-GCN-D can be defined according to the following equation (13):
[0172]
[0173] where the depth d l and d h are trainable; and W is a trainable parameter matrix. RED-GCN-D can be scaled on large graphs by choosing K << n such that approximate the full diffusion matrix by covering only a small subset of all the feature maps. For small graphs, can be used to include all the feature maps and thus higher flexibility can be obtained by means of two separate depth parameters rather than a unified depth parameter than .
[0174] Non-limiting embodiments or aspects of the present disclosure can provide various differences from ODE-based GNNs. Previous attempts at GNNs with continuous diffusion were mainly inspired by the graph diffusion equation, which is an ordinary differential equation (ODE) that characterizes the time-varying dynamic message propagation process. In contrast, non-limiting embodiments or aspects of the present disclosure can start from discrete graph convolution operations without explicitly involving ODEs. The goal of the Crystal Graph Neural Network (CGNN) is to build a depth GNN resistant to over-smoothing by adopting the framework of ordinary differential equations (ODEs), but its time parameter t is a predefined non-trainable hyperparameter in the positive domain, which is different from RED-GCN. In non-limiting embodiments or aspects of the present disclosure, CGNN components for preventing over-smoothing, restarting the distribution (e.g., skip connections from the first layer, etc.) are not required. In addition, CGNN applies the same depth to all frequency components, while RED-GCN-D has the flexibility to adopt two different depths respectively to adapt to high-frequency components and low-frequency components. The Graph Neural Definition (GRAND) introduces a non-Euler multi-step scheme with an adaptive step size to obtain a more accurate solution of the diffusion equation, and its depth (total integration time) is continuous, but still predefined / non-trainable and only takes positive values. Decoupled Graph Convolution (DGC) decouples the SGC depth into two predefined non-trainable hyperparameters: a positive real value T that controls the total time and a positive integer value K dgc . However, implementing negative depths in DGC is not applicable because the implementation propagates through K dgc times, rather than through an arbitrary real-valued exponent d on the feature map weights in RED-GCN.
[0175] As Figure 3A As shown, at step 302, process 300 includes processing the output embedding matrix using a machine learning classifier to generate at least one predicted label for at least one node of the graph. For example, the transaction service provider system 108 can process the output embedding matrix H using a machine learning classifier to generate at least one predicted label for at least one node of graph G. As an example, graph G may include multiple edges and multiple nodes for the multiple edges, the multiple nodes may be associated with multiple accounts in an electronic payment network, and / or the multiple edges may be associated with multiple transactions between the multiple accounts in the electronic payment network. In this example, the at least one predicted label for the at least one node of graph G includes a prediction as to whether at least one account associated with the at least one node is a malicious account (e.g., money mule, etc.).
[0176] Example experiment
[0177] In this section, non-limiting embodiments or aspects of RED-GCN are evaluated on semi-supervised node classification tasks on both homogeneous and heterogeneous graphs.
[0178] Eleven datasets are used for evaluation, including four homogeneous graphs: Cora, Citeseer, Pubmed, and DBLP, and seven heterogeneous graphs: Cornell, Texas, Wisconsin, Actor, Chameleon, Squirrel, and cornell5. All datasets are collected from the public GCN platform Pytorch-Geometric. For Cora, Citeseer, and Pubmed which have data splits in Pytorch-Geometric, the training / validation / test set splits are maintained as in GCN. For the remaining eight datasets, each dataset is randomly split into 20 / 20 / 60% for training, validation, and testing.
[0179] Non-limiting embodiments or aspects of RED-GCN are compared with seven baseline methods, including four classic GNNs: GCN, SGC, APPNP, and ChebNet, and three GNNs customized for heterogeneous graphs: FAGCN, GPRGNN, and H2GCN. Accuracy (ACC) is used as the evaluation metric. The average ACC is reported with the standard deviation (std) of all methods, which are each obtained through five runs with different initializations.
[0180] The semi-supervised node classification performance for homogeneous and heterogeneous graphs is shown respectively in Figure 7 Tables 1 and 2 below.
[0181] As observed from Table 1, different methods have similar performance on homogeneous graphs. RED-GCN-S achieves the best accuracy on two datasets: Cora and DBLP. On the remaining two datasets, RED-GCN-S is only 1.1% and 0.4% lower than the best baselines (APPNP on Citeseer and SGC on Pubmed). For RED-GCN-D, RED-GCN-D obtains similar performance to other methods, but RED-GCN-D only uses the top K largest / smallest eigenpairs.
[0182] RED-GCN-S / RED-GCN-D outperforms every baseline on all heterogeneous graphs, as shown in Table 2. These results indicate that RED-GCN has the ability to automatically detect the underlying graph heterogeneity without manually setting the model depth and without prior knowledge of the input graph. On the three large datasets Squirrel, Chameleon, and cornell5, even with only a small fraction of the feature maps, RED-GCN-D is able to achieve better performance than RED-GCN-S with the full set of feature maps. This shows that the graph signals in some real-world graphs may be dominated by a few low-frequency components and high-frequency components, and allowing two independent depth parameters in RED-GCN-D brings the flexibility to capture both low frequencies and high frequencies simultaneously.
[0183] A systematic study was also conducted on the node classification performance with respect to the trainable depth d.
[0184] As Figure 8 shown, the optimal depth and its corresponding classification accuracy are annotated. For the two homogeneous graphs Cora and Citeseer, the optimal depths are positive (5.029 and 3.735) in terms of the best ACC, while for the two heterogeneous graphs Actor and Squirrel, the optimal depths are negative (-0.027 and -3.751). These results demonstrate that the non-limiting embodiments or aspects of the present disclosure actually automatically capture graph heterogeneity / homogeneity by finding the appropriate depth to suppress or amplify the relative weights of the corresponding frequency components. That is, the high-frequency components / low-frequency components are suppressed for homogeneous / heterogeneous graphs, respectively.
[0185] For Figure 8 the two homogeneous graphs in plots (a) and (b), a sharp performance drop is observed when the depth d approaches 0 because the feature maps obtain nearly uniform weights. For the heterogeneous Actor dataset, its optimal depth -0.027 is close to 0, as Figure 8 as shown in the curve graph (c). In addition, the performance of RED-GCN-D (27.6%) is similar to that of GCN (27.5%), and both are much worse than RED-GCN-S (35.3%). This result indicates that the Actor is a special graph in which all frequency components have similar importance. Since there is no intermediate frequency component between the high-end component and the low-end component, the performance of RED-GCN-D is severely affected. For the conventional GCN, the suppressed weights of the high-frequency components deviate from the nearly uniform spectrum, resulting in a low ACC on this dataset.
[0186] It is particularly interesting to analyze the changes brought about by negative depth in the spatial domain and how this change affects subsequent model performance.
[0187] Graph augmentation. By selecting the optimal depth d according to the best performance on the validation set, a new diffusion matrix is obtained With the optimal d fixed, use to replace the normalized adjacency matrix in Equation (1) This is equivalent to applying the conventional GCN to the new topology. This topology effectively acts as a structural augmentation of the original graph. The impact of this augmentation on performance is tested on three heterogeneous graphs: Texas, Cornell, and Wisconsin, as shown in the table of the graph. Obviously, for the conventional GCN, the performance obtained using this new topology is better than that obtained using the original input graph: it significantly brings a 20%-30% improvement in ACC. In addition, the augmented topology also enables the conventional GCN to outperform RED-GCN-S and RED-GCN-D in two of the three datasets. Essentially, the augmented graph is a reweighted linear combination of feature graphs, and its topological structure inherently assigns higher weights to the feature graphs corresponding to higher frequencies, as Figure 10 shown.
[0188] To further understand how the topology of with negative optimal d is different from that of and why the performance is significantly improved, Figure 11 presents the heatmap of for the Cornell dataset in Figure 11 There are horizontal and vertical lines (light lines marked by dashed ellipses) in the heatmap in , which correspond to the hub nodes in the graph, i.e., the nodes with the largest degree. Interestingly, the connections between this node and most other nodes in the graph experience negative weight changes. Therefore, the influence of the hub node on most other nodes is systematically reduced. Thus, amplification magnifies the deviation between node embeddings and promotes the representation of graph heterogeneity.
[0189] Therefore, non-limiting embodiments or aspects of the present disclosure can provide a deep trainable GCN by redefining GCN over the real number field and reveal the interdependence between negative GCN depth and graph heterogeneity. A new problem of automatic GCN depth tuning for graph homophily / heterophily detection is proposed, and non-limiting embodiments or aspects of the present disclosure provide a simple and powerful solution, RED-GCN, which has two variants, RED-GCN-S and RED-GCN-D. An effective graph amplification method is also achieved through a new understanding of the message propagation mechanism generated by negative depth. The excellent performance of non-limiting embodiments or aspects of the present disclosure is demonstrated by example experiments of semi-supervised node classification on eleven graph datasets.
[0190] Although the embodiments or aspects have been described in detail for purposes of illustration and description, it should be understood that such details are for those purposes only and that the embodiments or aspects are not limited to the disclosed embodiments or aspects, but rather are intended to cover modifications and equivalent arrangements within the spirit and scope of the appended claims. For example, it should be understood that the present disclosure contemplates, to the extent possible, that one or more features of any embodiment or aspect can be combined with one or more features of any other embodiment or aspect. In fact, any of these features can be combined in a manner not specifically recited in the claims and / or disclosed in the specification. Although each of the dependent claims listed below may directly depend on only one claim, the disclosure of possible embodiments includes each dependent claim combined with each other claim in the claim set.< / x>
Claims
1. A method, which comprises: Obtain a graph G including an adjacency matrix A and an attribute matrix X using at least one processor, wherein an augmented adjacency matrix is determined based on the adjacency matrix A and the identity matrix I wherein based on the augmented adjacency matrix the identity matrix I and the degree matrix D of the adjacency matrix X are used to determine, for the augmented adjacency matrix a symmetric normalized Laplacian matrix L sym , and wherein an eigen-decomposition of the symmetric normalized Laplacian matrix L for the augmented adjacency matrix sym provides an eigenvector matrix U and an eigenvalue matrix Λ; and using the at least one processor to train a graph convolutional network in the following manner: Apply a transformation function to the symmetric normalized Laplacian matrix L for the augmented adjacency matrix using the at least one processor thereof sym to generate a filter Using the at least one processor, obtaining the eigenvector matrix U and the eigenvalue matrix Λ from an eigen decomposition of the symmetric normalized Laplacian matrix L for the augmented adjacency matrix ; sym Set the trainable depth parameter d to the power of each eigenvalue of the eigenvalue matrix Λ of the symmetric normalized Laplacian matrix L using the at least one processor sym to generate a new symmetric normalized Laplacian matrix S d ; and Using the at least one processor, use the filter Apply graph convolution to generate an output embedding matrix H.
2. The method according to claim 1, wherein the graph convolution is applied according to the following update equation: where H is the output embedding matrix, σ(·) is a non-linear activation function, is the filter, X is the attribute matrix, W is a trainable weight matrix, and d is the trainable depth parameter.
3. The method according to claim 1, wherein the trainable depth parameter d comprises a first trainable depth parameter d h and a second trainable depth parameter d l , and wherein the graph convolution is applied according to the following update equation: where H is the output embedding matrix, σ(·) is a non-linear activation function, is the filter, X is the attribute matrix, W is a trainable weight matrix, and K is a hyperparameter.
4. The method according to claim 1, which further comprises: using the at least one processor to process the output embedding matrix H using a machine learning classifier to generate at least one predicted label for at least one node of the graph G.
5. The method according to claim 4, wherein the graph G includes a plurality of edges and a plurality of nodes for the plurality of edges, wherein the plurality of nodes are associated with a plurality of accounts in an electronic payment network, and wherein the plurality of edges are associated with a plurality of transactions between the plurality of accounts in the electronic payment network.
6. The method according to claim 5, wherein the at least one predicted label for the at least one node of the graph G includes a prediction as to whether at least one account associated with the at least one node is a malicious account.
7. A system, which comprises: at least one processor, which is programmed and / or configured to: Obtain a graph G including an adjacency matrix A and an attribute matrix X, where an augmented adjacency matrix is determined based on the adjacency matrix A and an identity matrix I where based on the augmented adjacency matrix the identity matrix I and the degree matrix D of the adjacency matrix X are used to determine a symmetric normalized Laplacian matrix L for the augmented adjacency matrix ; and where an eigen - decomposition of the symmetric normalized Laplacian matrix L for the augmented adjacency matrix sym provides an eigen - vector matrix U and an eigen - value matrix Λ; and the symmetric normalized Laplacian matrix L for the augmented adjacency matrix sym ; and an eigen - decomposition of the symmetric normalized Laplacian matrix L for the augmented adjacency matrix provides an eigen - vector matrix U and an eigen - value matrix Λ; and train a graph convolutional network in the following manner: Apply a transformation function to the augmented adjacency matrix of the symmetric normalized Laplacian matrix L sym to generate a filter from the symmetric normalized Laplacian matrix L for the amplified adjacency matrix sym obtain the eigenvector matrix U and the eigenvalue matrix Λ from the eigenvalue decomposition; Set the trainable depth parameter d to the power of each eigenvalue of the eigenvalue matrix Λ of the symmetric normalized Laplacian matrix L to generate a new symmetric normalized Laplacian matrix S sym ; d ; and Use the filter Apply graph convolution to generate an output embedding matrix H.
8. The system according to claim 7, wherein the graph convolution is applied according to the following update equation: where H is the output embedding matrix, σ(·) is a non-linear activation function, is the filter, X is the attribute matrix, W is a trainable weight matrix, and d is the trainable depth parameter.
9. The system according to claim 7, wherein the trainable depth parameter d comprises a first trainable depth parameter d h and a second trainable depth parameter d l , and wherein the graph convolution is applied according to the following update equation: where H is the output embedding matrix, σ(·) is a non-linear activation function, is the filter, X is the attribute matrix, W is a trainable weight matrix, and K is a hyperparameter.
10. The system according to claim 7, wherein the at least one processor is further programmed and / or configured to: process the output embedding matrix H using a machine learning classifier to generate at least one predicted label for at least one node of the graph G.
11. The system according to claim 10, wherein the graph G includes a plurality of edges and a plurality of nodes for the plurality of edges, wherein the plurality of nodes are associated with a plurality of accounts in an electronic payment network, and wherein the plurality of edges are associated with a plurality of transactions between the plurality of accounts in the electronic payment network.
12. The system according to claim 11, wherein the at least one predicted label for the at least one node of the graph G includes a prediction as to whether at least one account associated with the at least one node is a malicious account.
13. A computer program product comprising a non-transitory computer-readable medium, the non-transitory computer-readable medium including program instructions that, when executed by at least one processor, cause the at least one processor to: Obtain a graph G including an adjacency matrix A and an attribute matrix X, wherein an augmented adjacency matrix is determined based on the adjacency matrix A and the identity matrix I wherein based on the augmented adjacency matrix the identity matrix I and the degree matrix D of the adjacency matrix X are used to determine the symmetric normalized Laplacian matrix L for the augmented adjacency matrix ; and sym wherein the eigen - decomposition of the symmetric normalized Laplacian matrix L for the augmented adjacency matrix provides an eigen - vector matrix U and an eigen - value matrix Λ; and sym train a graph convolutional network in the following manner: Apply a transformation function to the augmented adjacency matrix of the symmetric normalized Laplacian matrix L sym to generate a filter From the symmetric normalized Laplacian matrix L for the amplified adjacency matrix sym obtain the eigenvector matrix U and the eigenvalue matrix Λ from the eigenvalue decomposition; Set the trainable depth parameter d to the power of each eigenvalue of the eigenvalue matrix Λ of the symmetric normalized Laplacian matrix L to generate a new symmetric normalized Laplacian matrix S sym ; and d ; and Use the filter Apply graph convolution to generate an output embedding matrix H.
14. The computer program product according to claim 13, wherein the graph convolution is applied according to the following update equation: where H is the output embedding matrix, σ(·) is a non-linear activation function, is the filter, X is the attribute matrix, W is the trainable weight matrix, and d is the trainable depth parameter.
15. The computer program product according to claim 13, wherein the trainable depth parameter d includes a first trainable depth parameter d h and a second trainable depth parameter d l , and wherein the graph convolution is applied according to the following update equation: where H is the output embedding matrix, σ(·) is the non-linear activation function, is the filter, X is the attribute matrix, W is the trainable weight matrix, and K is the hyperparameter.
16. The computer program product according to claim 13, wherein the program instructions, when executed by at least one processor, further cause the at least one processor to: process the output embedding matrix H using a machine learning classifier to generate at least one predicted label for at least one node of the graph G.
17. The computer program product according to claim 16, wherein the graph G includes a plurality of edges and a plurality of nodes for the plurality of edges, wherein the plurality of nodes are associated with a plurality of accounts in an electronic payment network, and wherein the plurality of edges are associated with a plurality of transactions between the plurality of accounts in the electronic payment network.
18. The computer program product according to claim 17, wherein the at least one prediction label for the at least one node of the graph G includes a prediction as to whether at least one account associated with the at least one node is a malicious account.
Citation Information
Patent Citations
Tree-shaped risk account identification method and device, server and storage medium
CN110473083A
Systems and methods for improved anomaly detection in attributed networks
US20200065292A1