Group based anomaly detection

By generating and analyzing connected graphs of suspicious entities, the method effectively identifies and responds to anomalous groups, addressing imprecision in current detection methods and enhancing security against coordinated attacks and fraud.

WO2025160733A1PCT designated stage Publication Date: 2025-08-07VISA INTERNATIONAL SERVICE ASSOCIATION +1

Patent Information

Application Number
PCT/CN2024/074631
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-30
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Current tools for identifying group-based anomalies, such as coordinated denial-of-service attacks or cash-out fraud, are imprecise due to simplified analysis and reliance on crude groupings or human judgment, necessitating more effective detection methods.

Method used

A method involving determining a set of nodes representing suspicious entities, generating a connected graph, splitting it into sub-graphs, ranking potentially anomalous groups, and taking actions on high-ranked subsets using a computer.

Benefits of technology

Enhances the detection of anomalous groups by providing precise identification and enabling automated responses to potential fraudulent activities, reducing false positives and improving security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024074631_07082025_PF_FP_ABST
    Figure CN2024074631_07082025_PF_FP_ABST
Patent Text Reader

Abstract

A method is disclosed. In one example, the method includes determining a set of nodes representing suspicious entities. The method also includes generating a connected graph using the set of nodes, where the connected graph comprises edges between nodes in the set of nodes. The method also includes determining sub-graphs of the connected graph by splitting the connected graph, determining a set of potentially anomalous groups using the sub-graphs, ranking the potentially anomalous groups in the set of potentially anomalous groups, and determining a subset of the potentially anomalous groups that are higher ranked than other potentially anomalous subgraphs in the set of potentially anomalous groups. The method further includes taking one or more actions with respect to the subset of potentially anomalous group.
Need to check novelty before this filing date? Find Prior Art

Description

GROUP BASED ANOMALY DETECTION

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] None.BACKGROUND

[0003] Certain patterns of transactions require coordination between groups. Such transactions are examples of group-based anomalies. One example of anomalous group behavior can be a set of computers that are working together to access a system through denial or service attacks or the like. Another such example are improper “cash-out” transactions. Cash-out transactions refer to the withdrawal of cash by individuals from entities such as merchants using credit cards and then using that cash (e.g., investing in cryptocurrency) in a different manner during the interest free period through which transactions are to be paid back. In some cases, an individual and a merchant can collaborate to conduct a fake transaction using a credit card. The merchant will give the individual cash in an amount less than the amount of the fake transaction. The merchant will eventually be paid by the issuer of the credit card, and the issuer will ultimately be left with the debt if the individual does not pay the issuer back for the fake transaction. The nature of this activity requires a group of individuals that are acting together in an inappropriate way. Such transactions are considered fraudulent in some jurisdictions. Current tools for the identification of such are imprecise, as they often rely on simplified analysis, crude groupings, or rely on imprecise human judgment.

[0004] Better ways to detect and prevent such transactions are needed Embodiments address these and other problems.

[0005] BRIEF SUMMARY

[0006] One embodiment includes a method of detecting anomalous groups, the method comprising: determining, by a computer, a set of nodes representing suspicious entities; generating, by the computer, a connected graph using the set of nodes, wherein the connected graph comprises edges between nodes in the set of nodes; determining, by the computer, sub-graphs of the connected graph by splitting the connected graph; determining, by the computer, a set of potentially anomalous groups using the sub-graphs; ranking, by the computer, the potentially anomalous groups in the set of potentially anomalous groups; determining, by the computer, a subset of the potentially anomalous groups that are higher ranked than other potentially anomalous subgraphs in the set of potentially anomalous groups; and taking, by the computer, one or more actions with respect to the subset of potentially anomalous groups.

[0007] Another embodiment of the invention includes a computer comprising: a processor; and a non-transitory computer readable medium, the non-transitory computer readable medium comprising code executable by the processor, to perform operations comprising: determining a set of nodes representing suspicious entities; generating a connected graph using the set of nodes, wherein the connected graph comprises edges between nodes in the set of nodes; determining sub-graphs of the connected graph by splitting the connected graph; determining a set of potentially anomalous groups using the sub-graphs; ranking the potentially anomalous groups in the set of potentially anomalous groups; determining a subset of the potentially anomalous groups that are higher ranked than other potentially anomalous subgraphs in the set of potentially anomalous groups; and taking, by the computer, one or more actions with respect to the subset of potentially anomalous groups.

[0008] These and other embodiments of the invention are described in further detail below.BRIEF DESCRIPTION OF THE DRAWINGS

[0009] FIG. 1 shows a flow diagram of an overall process flow according to embodiments.

[0010] FIG. 2 shows a data preparation process flow according to embodiments.

[0011] FIG. 3A shows a process for identifying suspicious entities.

[0012] FIG. 3B shows a graph of nodes representing suspicious entities and interactions associated with the nodes.

[0013] FIG. 3C shows another table of transactions conducted by different entities.

[0014] FIG. 3D shows a table where transactions are split into two.

[0015] FIG. 3E shows a table where data are aggregated to obtain a number of intersections between two entities.

[0016] FIG. 3F shows matching consumers and the transactions when they intersect.

[0017] FIG. 3G shows a table with an overall similarity score calculation.

[0018] FIG. 4A shows process for determining potentially anomalous groups using a graph splitting algorithm.

[0019] FIG. 4B shows a connected graph and splitting of the connected graph.

[0020] FIG. 5 shows a process flow for determining scores for different groups.

[0021] FIG. 6 shows a graphic illustrating different risk levels for different group anomaly scores.

[0022] FIG. 7 shows a block diagram of a transaction processing system.

[0023] FIG. 8 shows a block diagram of a processing computer.

[0024] TERMS

[0025] Prior to discussing embodiments of the disclosure, some terms can be described in further detail.

[0026] A “user” may refer to an individual. In some embodiments, a user may be associated with data. The data may be associated with one or more personal accounts and / or user devices. A user can be identified by his or her data, personal accounts, and / or devices.

[0027] A "user device" may include any device that can be operated by a user. A user device can be referred to as a communication device, which can allow for communication to one or more computers. A communication device can be referred to as a mobile device if the mobile device has the ability to communicate data remotely.

[0028] A “mobile device” may comprise any suitable electronic device that may be transported and operated by a user, which may also provide remote communication capabilities over a network. Examples of remote communication capabilities include using a mobile phone (wireless) network, wireless data network (e.g., 3G, 4G or similar networks) , Wi-Fi, Wi-Max, or any other communication medium that may provide access to a network such as the Internet or a private network. Examples of mobile devices include mobile phones (e.g., cellular phones) , PDAs, tablet computers, net books, laptop computers, personal music players, hand-held specialized readers, etc. Further examples of mobile devices include wearable devices, such as smart watches, fitness bands, ankle bracelets, rings, earrings, etc., as well as automobiles with remote communication capabilities. A mobile device may comprise any suitable hardware and software for performing such functions, and may also include multiple devices or components (e.g., when a device has remote access to a network by tethering to another device -i.e., using the other device as a modem –both devices taken together may be considered a single mobile device) . A mobile device may further comprise means for determining / generating location data. For example, a mobile device may comprise means for communicating with a global positioning system (e.g., GPS) .

[0029] An “application” may be computer code or other data stored on a computer readable medium (e.g., memory element or secure element) that may be executable by a processor to complete a task.

[0030] A “resource provider” may be an entity that can provide a resource such as goods, services, information, and / or access. Examples of resource providers include merchants, access devices, secure data access points, etc. A “merchant” may typically be an entity that engages in transactions and can sell goods or services, or provide access to goods or services.

[0031] An “access device” may be any suitable device that provides access to a remote system. An access device may also be used for communicating with a merchant computer, a transaction processing computer, an authentication computer, or any other suitable system. An access device may generally be located in any suitable location, such as at the location of a merchant. An access device may be in any suitable form. Some examples of access devices include POS or point of sale devices (e.g., POS terminals) , cellular phones, PDAs, personal computers (PCs) , tablet PCs, hand-held specialized readers, set-top boxes, electronic cash registers (ECRs) , automated teller machines (ATMs) , virtual cash registers (VCRs) , kiosks, security systems, access systems, and the like. An access device may use any suitable contact or contactless mode of operation to send or receive data from, or associated with, a user mobile device. In some embodiments, where an access device may comprise a POS terminal, any suitable POS terminal may be used and may include a reader, a processor, and a computer-readable medium. A reader may include any suitable contact or contactless mode of operation. For example, exemplary card readers can include radio frequency (RF) antennas, optical scanners, bar code readers, or magnetic stripe readers to interact with a payment device and / or mobile device. In some embodiments, a cellular phone, tablet, or other dedicated wireless device used as a POS terminal may be referred to as a mobile point of sale or an “mPOS” terminal.

[0032] A “transport computer” may refer to an intermediary computer that can transport data. A transport computer can be a computer of an acquirer. An "acquirer" may be an entity that can process interactions on behalf of a resource provider. For example, the acquirer can be a business entity (e.g., a commercial bank) that establishes relationships with resource providers, such that the resource providers can meet transaction processing requirements. Some entities can perform both issuer and acquirer functions. Some embodiments may encompass such single entity issuer-acquirers.

[0033] An “authorizing computer” may be a computer of an authorizing entity. An “authorizing entity” may be an entity that can authorize interactions. Examples of an authorizing entity may be an issuer, a governmental agency, a document repository, an access administrator, etc. An “issuer” may typically refer to a business entity (e.g., a bank) that maintains an account for a user. An issuer may also issue credentials to a user, such as a user account.

[0034] An “authorization request message” may be an electronic message that requests authorization for an interaction. An authorization request message according to some embodiments may comply with ISO 8583, which is a standard for systems that exchange electronic interaction information associated with a user using an issued user account. The authorization request message may include an issuer account identifier that may be associated with the user’s account. An authorization request message can also comprise additional data elements corresponding to “identification information” including, by way of example only: a service code, a CVV (card verification value) , a primary account number (PAN) , a token, a username, an expiration date, etc. An authorization request message may also comprise “interaction information, ” such as any information associated with a current interaction, such as an interaction location, transaction amount, resource provider identifier, resource provider location, bank identification number (BIN) , merchant category code (MCC) , information identifying resources being provided / exchanged, etc., as well as any other information that may be utilized in determining whether to identify and / or authorize an interaction.

[0035] An “authorization response message” may be a message that responds to an authorization request. The authorization response message may include, by way of example only, one or more of the following status indicators: Approval --transaction was approved; Decline --transaction was not approved; or Call Center --response pending more information, merchant calls the toll-free authorization phone number. The authorization response message may also include an authorization code. The code may serve as proof of authorization for an interaction.

[0036] A “server computer” may include a powerful computer or cluster of computers. For example, the server computer can be a large mainframe, a minicomputer cluster, or a group of servers functioning as a unit. In one example, the server computer may be a database server coupled to a Web server. The server computer may be coupled to a database and may include any hardware, software, other logic, or combination of the preceding for servicing the requests from one or more client computers. The server computer may comprise one or more computational apparatuses and may use any of a variety of computing structures, arrangements, and compilations for servicing the requests from one or more client computers.

[0037] A “graphics processing unit” or “GPU” may refer to an electronic circuit designed for the creation of images intended for output to a display device. The display device may be a screen, and the GPU may accelerate the creation of images in a frame buffer by rapidly manipulating and altering memory. GPUs may have a parallel structure that make them more efficient than general-purpose CPUs for algorithms where the processing of large blocks of data is done in parallel. Examples of GPUs may include RadeonTM HD 6000 Series, PolarisTM 11, NVIDIA GeForceTM 900 Series, NVIDIA PascalTM, etc.

[0038] A “connected graph” may refer to a representation of a graph in a plane of distinct vertices connected by edges. The distinct vertices in a topological graph may be referred to as “nodes. ” Each node may represent specific information for an event or may represent specific information for a profile of an entity or object. The nodes may be related to one another by a set of edges, E. An “edge” may be described as an unordered pair composed of two nodes as a subset of the graph G = (V, E) , where is G is a graph comprising a set V of vertices (nodes) connected by a set of edges E. For example, a topological graph may represent a transaction network in which a node representing a transaction may be connected by edges to one or more nodes that are related to the transaction, such as nodes representing information of a device, a user, a transaction type, etc. An edge may be associated with a numerical value, referred to as a “weight, ” that may be assigned to the pairwise connection between the two nodes. The edge weight may be identified as a strength of connectivity between two nodes and / or may be related to a cost or distance, as it often represents a quantity that is required to move from one node to the next.

[0039] The term “artificial intelligence model” or “AI model” may refer to a model that may be used to predict outcomes in order achieve a pre-defined goal. The AI model may be developed using a learning algorithm, in which training data is classified based on known or inferred patterns. An AI model may also be referred to as a “machine learning model” or “predictive model. ”

[0040] A “subgraph” or “sub-graph” may refer to a graph formed from a subset of elements of a larger graph. The elements may include vertices and connecting edges, and the subset may be a set of nodes and edges selected amongst the entire set of nodes and edges for the larger graph. For example, a plurality of subgraph can be formed by randomly sampling graph data, wherein each of the random samples can be a subgraph. Each subgraph can overlap another subgraph formed from the same larger graph.

[0041] A “community” may refer to a group / collection of nodes in a graph that are densely connected within the group. A community may be a subgraph or a portion / derivative thereof and a subgraph may or may not be a community and / or comprise one or more communities.

[0042] A “data set” may refer to a collection of related sets of information composed of separate elements that can be manipulated as a unit by a computer. A data set may comprise known data, which may be seen as past data or “historical data. ” Data that is yet to be collected, may be referred to as future data or “unknown data. ” When future data is received at a later point it time and recorded, it can be referred to as “new known data” or “recently known” data, and can be combined with initial known data to form a larger history.

[0043] “Unsupervised learning” may refer to a type of learning algorithm used to classify information in a dataset by labeling inputs and / or groups of inputs. One method of unsupervised learning can be cluster analysis, which can be used to find hidden patterns or grouping in data. The clusters may be modeled using a measure of similarity, which can be defined using one or metrics, such as Euclidean distance.

[0044] “Machine learning” may refer to an artificial intelligence process in which software applications may be trained to make accurate predictions through learning. The predictions can be generated by applying input data to a predictive model formed from performing statistical analysis on aggregated data.

[0045] A “transaction” may an interaction or exchange between two entities (e.g., a consumer and a merchant, a client and a server, etc. ) . In some embodiments, a transaction may occur through the use of a financial instrument (e.g., a credit card or other payment card) . A transaction may have associated transaction data, which may include for example, a transaction time, method used, transaction amount, geographical location, etc.DETAILED DESCRIPTION

[0046] Embodiments of the invention provide for methods to discover groups of entities (e.g., a group of credit card holders and / or a group of merchants) that may be performing anomalous activities within the ordinary activities of other entities. Embodiments allow for the detection of entities meeting one or more criteria for being included in a potential group, the formation of a graph (or other informationally equivalent structure) representing the relationships between the entities, detection of groups from within the graph, assignment of risk scores for the detected groups, and taking actions with respect to the detected groups. In some cases, visualization techniques can allow users to visualize aspects of the detected groups.

[0047] I. OVERVIEW OF GROUP DETECTION

[0048] FIG. 1 illustrates an example method according to embodiments of the disclosed invention. FIG. 1 illustrates a method 100. Additional details of various steps of method 100 are further provided below. Method 100 may be used to receive as an input transaction data and provide as an output one or more graphs, where each graph represents a group of individuals believed to be acting in concert to conduct a pattern of transactions (e.g., a cash out transaction) .

[0049] At step 110, data can be prepared for analysis. At step 110, a dataset can be received. The received dataset may be a raw dataset related to transaction data. The dataset may be information related to historical transaction data. For example, all transactions within a specific geogrpahic region may be received at this step. The dataset may also be time based (e.g., the last month or week) . The data may also be sampled or processed to filter transactions meeting specific critera (e.g., above a specific dollar amount) . Other information may include features such as transaction times, transaction time intervals, transaction amounts, transaction frequency, or cumulative spend amount. At step 110, a reduced dataset may be created from the received dataset following screening and processing data. The reduced dataset may include information related to a user (e.g., a credit card holder) , the time of the transaction, and a merchant related to the transaction. Additional details related to step 110 are described below with respect to FIGS. 2A–2B.

[0050] At step 120, a relationship graph or connected graph which provides relations between entitiies may be generated. At step 120, the reduced dataset obtained from step 110 may analyzed to detect associations between entities such as individuals and merchants. At step 120, weaker associations may be removed from the dataset while stronger assocations may be retained based on algorithmically analyzing the reduced dataset. For instance, the reduced dataset may be used to generate a graph data by using similarity measurements between the data instances of the reduced dataset. In some embodiments, a SynchroTrap algorithm (see Qiang Cao, Christopher Palow, Christopher Palow, and Christopher Palow. 2014. Uncovering Large Groups of Active Malicious Accounts in Online Social Networks. In ACM CCS. 477–488. ) may be used to determine relationships between data instances (e.g., entities) of the reduced dataset. Additional details related to step 120 are described with repsect to FIGs. 3A–3F. Stated differently, at step 120, a connected graph may be generated in which each node of the graph represents a suspicious entity and an edge of the graph represents an association between the users connected by the edge. An example connected graph is illustrated below with respect to FIG. 3B.

[0051] The suspicious entities can include consumers, merchants, computers, IP addresses, or any other entity. A suspicious entity may or may not be actually fraudulent or malevoent, but may be suspected of being so. In some embodiments, the suspicious entities may be those that may be suspected of being in a cash out group or gang. In other embodiments, suspicious entities can include computers in a network that be potentialy fraudulent or malevoent.

[0052] At step 130, communities may be determined by “splitting” the connected graph generated at step 120. The splitting of the graph may occur through the use of one or more community detection algorithms (e.g., clustering algorithms) . As one example, a Louvain algorithm may be used (See Blondel, Vincent D; Guillaume, Jean-Loup; Lambiotte, Renaud; Lefebvre, Etienne (9 October 2008) . "Fast unfolding of communities in large networks" . Journal of Statistical Mechanics: Theory and Experiment. 2008 (10) : P10008) (arXiv: 0803.0476) . The Louvian algorithm is a hierarchical clustering algorithm based on graph theory. Additional details regarding the Louvain algorithm are provided below. The communities may be outputted at this step. In some instances, each community which is determined at this step may be considered a separate anomalous group. Additional aspects of step 130 are illustrated below with respect to FIGS. 3A–3F.

[0053] At step 140, the determined anomalous group (s) may be validated. In this step, the data in each of the determined anomalous group (s) can be analyzed to determine if they are indicative of cash out gangs. For instance, the identifies of the consumers and / or merchants can be examined to determine if they have been associated with such activity in the past.

[0054] At step 150, the anomalous groups may be provided with risk scores and the groups may be ranked by the risk scores. The generation of the risk score for each group may be be done by any suitable method. As one example, the generation of the risk metric may be based on an entropy weight method (EWM) .

[0055] In some embodiments, in step 160, after determining the highest ranked anomalous groups, an action may be taken with respect to them. Also, in step 160, visualization techniques may be used to visualize the groups in the context of one or more connected graphs.

[0056] II. PREPARING DATASETS FROM TRANSACTION DATA

[0057] FIG. 2 illustrates aspects of data preperation. The data once prepared may be used for analysis with graph techniques. Any of the steps illustrated with respect to FIG. 2 may be used to prepare data for analysis, which may be used for generation of graphs. FIG. 2 illustrates additional aspects of step 110. However, any of these steps may be performed to prepare data, and steps may be combined, skipped, or modified.

[0058] At step 110A, transaction data may be obtained. The transaction data can be obtained by requesting the transaction data or generating it. The transaction data may include unprocessed data related a plurality of transactions. Transaction data can be generated from transactions conducted between computers. Exemplary transactions can be payment transactions, such as those described with respect to FIG. 7. In some embodiments, the transaction data requested may be based on time, geographical scope, or periodicity.

[0059] At step 110B, transactions may be filtered to form filtered transaction data. The filtered transaction data can then be used to identify potential suspicious entities that may be involved in anomalous transactions. For example, transactions in a set can be filtered to remove transactions conducted by known legitimate individuals (e.g., users or consumers) and merchants. The transactions in the subset that remains can be conducted by potentially suspicious individuals or merchants.

[0060] In some embodiments, the filtering may be based on one or more categories related to a specific type of behavior which is to be investigated with respect to the transactions. For example, to determine if the transactions are related to improper “cash-out” type transactions, a specific library of features may be used to filter for transactions which are likely to be relevant to this type of activity. In some embodiments, the features may be related to categories (e.. g, transaction features, billing features, and user information features) with specific sub-categories for each category. Based on the features, rules may be created to determine which transactions to filter. Exemplary transaction features may include transaction time intervals, suspicious merchants, suspicious consumers, transaction amounts, transaction frequencies, suspicious MCCs (merchant category codes) , etc. Billing features may include billing information, repayment rates, credit limit utilization percentages, and historical usage of credit, etc.. Additionally, the number of features being considered may be restricted to obtain better performance on the filtering process. Various rules related to cash out detection may be created based on the set of features considered.

[0061] Additionally, additional rules and / or distinguishing features may be generated or synthesized through the use of machine learning techniques or other statistical techniques. For example, a binary classification machine learning model may be used on a feature set provided to the model. The classification may be whether a specific transaction is or is not relevant for a specific type of behavior or pattern of behavior (e.g., cash-out fraud) . In other examples, the machine learning model may also generate additional features for consideration. After the training of the machine learning model, the most statistically significant (e.g., with high predictive power) rules (which may be based on one feature or multiple features together) may be chosen to be placed into a production environment to filter transaction data.

[0062] At step 110C, additional filtering of transactions may occur based on positive lists. For example, a consumer positive list and a merchant positive list may be used. A consumer positive list may be maintained which may be based on consumers that are known to be legitimate consumers that would not perform the improper behavior that is to be detected. A merchant positive list may be based on trusted merchants and / or popular merchants (e.g., merchants above a certain size or spend) . Transaction data related to members of a positive list may be removed from the dataset and thus not analyzed further.

[0063] At step 110D, data cleaning may occur. At this step, extraneous data which is not requried to generate graphs may be removed from the dataset. For example, geographic data in the transaction data which is not relevant (e.g., extraneous) to detection of a particular group anomaly may be removed from the dataset of transaction data. The cleansed data can transaction data for a number of transactions. The transaction data for each transaction can include a transaction time for a transaction, and a merchant and consumer that conducted the transaction.

[0064] The set of transactions in subset of transactions that remain can be conducted by potentially suspicious individuals and / or merchants. These individuals and / or merchants can nodes in a connected graph.

[0065] III. DETERMINING RELATIONSHIPS BETWEEN ENTITIES

[0066] FIG. 3A illustrates aspects of generating a graph. Any of the steps illustrated with respect to FIG. 3A may be used to generate a graph. FIG. 3A illustrates additional aspects of step 120. The various steps discussed with FIG. 3A may be used in conjunction or as a part of a SynchroTrap algorithm.

[0067] Prior to a discussion of FIG. 3A, a discussion of regarding connected graphs may be useful. A graph may be a relationship graph (or connected graph) which illustarted relationships between various consumers. In a relationship graph, each node of a graph may represent a consumer and an edge between two nodes may represent a relationship between the two consumers. In some examples the edge may be unweighted while in others the edge may be weighted. For additional context, an example relationship graph 300 is illustrated in FIG. 3B. Relationship graph 300 contains a number of nodes and vertices which are not labeled for simplicity. For example, a graph may be represented as a matheamtical pair (V, E) , where V is a set of vertices (also referred to as nodes) and E is a set of edges between the vertices. The relationshp graph may also be undirected (where there is not directionality between the edges) or a directed graph (where there is a directionality assocaited with an edge between two nodes) .

[0068] As one example of an equivalent data structure to a tree, a matrix may be used, which may be stored in computer memory. A positive number (e.g., 1, 2, 5) in the matrix represents a connection and the weight of the connection between two nodes while a “0” represents no connection. As the network trees described herein are undirected (i.e., the edges do not have directionality associated with them) , the matrix is symmetric around its diagonal. A data structure equivalent to the matrix may be stored in memory of a computer to reference a created network tree. Additionally, references to metadata may be included within the matrix as an array or with a pointer to another database.

[0069] Turning back to FIG. 3A, at step 120A, the previously filtered transaction data may be discretized. Discretization may refer to a process of turning a continuous or near-contious function or relationship to one that may only occur at discrete intervals or values (e.g., natural numbers, every minute, hour, second, or other defined interval) . The discretization analysis can be useful in determine if entities (e.g., consumers and / or merchants) are operating as an anomalous group (e.g., a gang) in an improper cash-out scheme, because they typically conduct cash-out transactions in short time windows, sometimes in regular time intervals. Other behavious such as denial of service attacks can also occur within short intervals.

[0070] At step 120A, the time may be discretized based on, for example, an hourly interval. A process of discretization may be seen in table 310 of FIG. 3C. Table 310 contains columns labeled Consumer ID, consumption time, Merchant ID, and Tsim (a time window) . The consumption time (or transaction time) may be the time when a transaction occurred. In table 310, an hourly interval may be used to cluster the activities of consumers, where from an arbitrary time, a window with a time period “T” (Tsim) (in this example, an hour) before and after the arbitrary time may be chosen, to cluster the activities of a specific consumer with a specific merchant. This may be represented in the Tsim column of table 310, where two time intervals are asscoaited with a transaction conducted by a consumer. For example, in the first row, Tsim is “20211216 (09#10) ; 20211216 (10#11) . ” “09#10” can correspond to one time interval (e.g., one hour) on one side of the time 10: 03: 32, while “10#11) can correspond to another time interval (e.g., one hour) on the other side of the time 10: 03: 32.

[0071] At step 120B, the discretized data may be grouped or split to allow for easier analysis. FIG. 3D illustrates table 320. Similar to table 310, table 320 contains labeled Consumer ID 302, transaction date and time 304, Merchant ID 306, and time intervals 308, Tsim. However, each transaction is split into two data entries. For example, Tsim “20211216 (09#10) ; 20211216 (10#11) ” had one entry for “CUST_03” in the Table 310 in FIG. 3C, but has two entries Tsim “20211216 (09#10) ” and Tsim “20211216 (10#11) ” for “CUST_03” in FIG. 3D.

[0072] At step 120C, a time difference can be calculated and two data instances (e.g., data rows in table 320) may be self-matched when the calculated time difference is below a threshold. The self-matching may be restricted by a time constraint and a merchant constraint. Stated alternatively, data table matches itself and takes the difference, with the time limit set to be less than the threshold. An example of a table showing time differences between consumers is shown in FIG. 3E. FIG. 3E shows a table 110 with columns including: a first customer identifier 322, a transaction date and time associated with a customer with the first customer identifier and a merchant with the merchant identifier in column 234 in the same row, a merchant identifier column 326, a time interval column 328, a second consumer identifier column 330, a transaction date and time column 332 of transactions conducted between the merchants and second consumers, and a time difference column 334.

[0073] At step 120D, the results from step 120C may be aggregated based on a constraint similarity. Step 120C may yield detailed data which may be aggregated to obtain the the number of intersections between two consumers. The number of intersections between two consumers may allow for the association between the consumers to be determined. FIG. 3F illustrates table 330 which represents an aggregated table, along with a number of intersections between any two consumers 330. For instance, the first row represents 2 intersections between consumer 1 and consumer 2.

[0074] The information from table 330 may be used to generate a graph (e.g., graph 300 illustrated in FIG. 3B) . For instance, any two consumers in table 330 (columns 1 and 2) may represent nodes of a graph and column 3 may represent a weight to an edge between the two consumers. In this manner, only consumers which are connected to one another (and the corresponding strength of that connection) may be represented.

[0075] At step 120E, an overall similarity calculation may be generated. FIG. 3G illustrates table 340 which may be referred to to illustrate this step.

[0076] At this step, for an “X” value, and a “Y” value (which may be the first and second column of table 340, and which may be pairwise consumers (e.g., consumer 1 and consumer 2, consumer 2 and consumer 3, and consumer 1 and consumer 3) ) , various metrics may be calculated. In table 340, X∩Y can refer to the number of times that transactions conducted by consumers X and Y are conducted that the same merchant (s) . X_Cnt and Y_Cnt can refer to the number of transactions conducted by consumer X and consumer Y the merchant (s) , respectively. XUY may refer to the total number of transactions that X and Y may have been conducted at different merchants.

[0077] At this step, a Jaccard index may also be calculated. The Jaccard index (also known as the Jaccard similarity coefficient and represented by the letter “J” ) is a statistic used for gauging the similarity and diversity of sample sets. The Jaccard index may only be between the values of 0 and 1, which may also be represented as a percentage.

[0078] J for two sets, sets A and B, may be calculated as: J (A, B) = (the number in elements in both sets)  /  (the number of the elements in either set) = J (A, B) = |A∩B|  / |A∪B|. The Jaccard index can be exemplified by the values in the column labeled SIM (F2) .

[0079] For example, in FIG. 3G the similarity score for certain rows is 100%. However, for the 5th row illustrated, the similarity score is only 5.26%, which indicates that there is not much overlap between the two consumers.

[0080] IV. DETERMINING ANOMALOUS GROUPS

[0081] Upon the generation of a connected graph, sub-graphs of the connected graphs may be determined. These may form “communities” of nodes. The communities and the subgraphs can be used to identify anomalous groups. This may be performed using the method 130 in FIG. 4A.

[0082] At step 130A, a connected graph may be obtained. The graph may be a connected graph, such as connected graph 300 illustrated with respect to FIG. 3B. The connected graph may be split into one or more groups as further explained below. At this step, the connected graph may be decomposed based on a decomposition algorithm. The decomposition algorithm may be based on an edge-separator or a vertex-separator algorithm.

[0083] At step 130B, a graph splitting algorithm (which may also be referred to a graph partition algorithm) may be chosen. The graph splitting algorithm may be a community detection algorithm. Mathematically, a split of a graph is a partition of the vertices of the graph into mutually exclusive subsets. The mutually exclusive subsets may be equivalent to sub-graphs of the input graph. The choice of a splitting algorithm may be chosen based on the size of the graph. the required size of the outputted sub-graphs of the input graph, the computational complexity, and time-constraints related to an analysis of a graph.

[0084] A community detection method may be used to locate communities based on a graph structure, such as strongly connected nodes. The community detection method may ignore the properties of the nodes but focus on the strength of the edges between the nodes. A community may be formed by nodes of similar types in a network. Community detection algorithms may be used to evaluate how groups of nodes are clustered or partitioned, and the tendency of those groups to strengthen or break apart (such as for example, over time) . Example community detection methods may include both agglomerative methods (where edges are added one by one to a graph which only contains nodes) and divisive methods (where edges are removed one by one from a complete graph) .

[0085] Example algorithms which may be chosen from at this step include: a Louvian Algorithm, a Surprise Detection Algorithm, a Leiden Community Detection Algorithm, a Walktrap Community Detection Algorithm, a Modularity Optimization Algorithm, a Label Propagation Algorithm, a Weakly Connected Components Algorithm, a Strongly Connected Components Algorithm, a Triangle Counting Algorithm, a Local Clustering Coefficient Algorithm, and the K-1 Coloring Algorithm.

[0086] At step 130C, the connected graph may be split into groups (also referred to as communities) by applying a chosen graph splitting algorithm. At this step, the chosen algorithm may be implemented on an input graph and an output graph, or multiple output graphs, may be obtained. An example output graph with split groups may be observed in FIG. 4B. The relationships between the nodes can be re-drawn so that the subgroups identified in graph 400 in FIG. 4B can be redrawn or re-visualized as the new separate graphs 410 in FIG. 4B.

[0087] At step 130D, the groups may be separately stored and / or visualized. These groups may be anomalous groups, and may also be referred to as cliques, crews, or gangs. FIG. 4B also illustrates graphs 410 of groups, where the groups are visualized independent of graph 400. Each group (whether viewed in graph 400 or 410) may represent an anomalous group, gang, or crew, which may be further analyzed for risk or other metrics.

[0088] V. ASSIGNING RISK TO GROUPS

[0089] [Rectified under Rule 91, 20.02.2024]In step 150, a risk score may be assigned to each group determined from the connected graph. Additional aspects of step 150 are described with respect to FIG. 5, which is further discussed below.

[0090] At step 150A, a community rating score process can be determined. Any suitable scoring method can be used. In some embodiments, an entropy weight method (EWM) calculation may be used.

[0091] At step 150B, the entropy weight method (EWM) calculation may be performed for each group. The entropy weight method (EWM) is a known weighting method that measures value dispersion in decision-making. The greater the degree of dispersion, the greater the degree of differentiation, and more information can be derived. EWM calculated scores can be used to assign a risk metric to each group of entities and to differentiate higher risk groups of entities.

[0092] In information theory, entropy is a measure of uncertainty. The greater the amount of information, the smaller the uncertainty and the smaller the entropy. Similarly, the less information, the greater the uncertainty and the entropy. The entropy value can be used to judge the randomness and disorder of an event. It can also be used to judge the dispersion of an indicator. The greater the dispersion of the indicator, the greater the impact (weight) of the indicator on the comprehensive evaluation, and the smaller its entropy.

[0093] Calculation of an entropy weight method may include calculation from a list of indicators related to a particular type of group behavior (e.g., cash-out fraud) . The indicator list may include for example, a cash-out merchant ratio, cash-out consumer ratio, total transaction amount, number of cash-out merchants, and number of cash-out consumers. The cash-out merchant ratio can be a ratio of known cash out merchants (e.g., identified via prior analyses) in a group to the total number of merchants in the group. The cash-out customer ratio can be a ratio of known cash out customers (e.g., identified via prior analyses) in a group to the total number of consumers in the group.

[0094] In discussion of the EWM, there may be m indicators and n samples. The samples may be each anomalous group.

[0095] For the m indicators and n samples are set in the evaluation, and the measured value of the ith indicator in the jth sample is recorded as xij. The first step is the standardization or normalization of measured values. Normalization of the measured indicator values may be performed by

[0096] The second step is to calculate the proportion, pij of the ith sample value under the jth indicator to that of the indicator according to the following formula:

[0097] In the third step, the entropy value of the jth indicator, ej, may be calculated as:

[0098] In the fourth step, the redundancy of the information entropy, dj, may be calculated as:dj=1-ej, j=1, …, m

[0099] In the fifth step, the weight of each calculator may be designated as wj.

[0100] In a sixth step, the comprehensive score each sample may be calculated as and be designated as si. The comprehensive score may be the risk rating related to that particular sample.

[0101] At step 150C, a score rating for each community may be generated, based on for example, the entropy weight method.

[0102] At step 150D, the graph (s) may be restored if desired.

[0103] At step 150E, after the communities with the highest scores are determined, the entities associated with those communities can be identified as anomalous groups. The groups with the highest scores can be validated as being cash out gangs (e.g., through an investigation) . This can be performed manually (e.g., using past knowledge of cash out gangs) or automatically. In some embodiments, signals may be sent to disable or impair the suspicious entities in the subset of the potentially anomalous groups from performing certain actions, or one or more alerts can be provided to an authority entity (e.g., a bank) regarding the suspicious entities in the subset of the potentially anomalous groups Further, one or more other automated actions can be taken with respect to the anomalous groups. In an example, the accounts associated with account numbers involved in a possible cash-out scheme can be automatically frozen so that further transactions cannot be conducted. In yet other embodiments, transactions coming from merchants and consumers in suspected gangs can be automatically declined. Further, the graph can be displayed to a user if desired.

[0104] Note that although the identification of cash out gangs is discussed in detail, embodiments of the invention are not limited thereto. For example, the suspicious entities can be computers in a network that might be suspected of attacking honest computers. If the suspicious entities are a group of computers that are suspected performing malevolent attacks, then preventative actions such as automatically preventing those computers from accessing a targeted system can be performed.

[0105] FIG. 6 shows the risk level associated with different anomalous groups. The scores can be determined using the entropy weight processes described above.

[0106] The transaction data that is initially obtained in step 110 in FIG. 1 can be obtained in any suitable manner. In some embodiments, the transaction data are generated from payment card transaction data. Such data may be present in authorization request messages, authorization response messages, and / or clearing and settlement messages.

[0107] FIG. 7 shows a flow diagram of a resource provider processing a transaction according to embodiments. The data produced by the system and flow in FIG. 7 can be examples of the transaction data described above. FIG. 7 shows a user device 706, which may be a payment card interacting with a resource provider computer 706 which may be a merchant computer such as a POS terminal. The resource provider computer 706 is in communication with an authorizing entity computer 720 via a transport computer 710 and a processing computer 716.

[0108] In step S702, after interacting with the user device 706, the resource provider computer 706 can generate an authorization request message comprising a transaction amount and a credential such as a primary account number, or a payment token. The resource provider computer 706 can then transmit the authorization request message to the transport computer 710.

[0109] In step S704, after the transport computer 710 receives the authorization request message, the transport computer 710 can forward it to the processing computer 716.

[0110] In step S706, after receiving the authorization request message, the processing computer 716 can transmit the authorization request message to the authorizing entity computer 720.

[0111] After the authorizing entity computer 720 receives the authorization request message, it can make a determination as to whether or not the transaction is authorized. It can determine if the account associated with the credential or token has sufficient funds for the transaction. It can also determine if the transaction is potentially fraudulent by analyzing data elements of the authorization request.

[0112] In step S710, the authorizing entity computer 720 can then generate an authorization response message. The authorizing entity computer 720 can then transmit it to the processing computer 716.

[0113] In step S712, the processing computer 716 can transmit the authorization response message to the transport computer 710.

[0114] In step S714, the transport computer can transmit the authorization response message to the resource provider computer 706.

[0115] Later, a clearing and settlement process can occur between the transport computer 710, the processing computer 716, and the authorizing entity computer 720.

[0116] FIG. 8 shows a block diagram of an authentication server computer 800 according to an embodiment. The authentication server computer 800 may comprise a processor 802, which may be coupled to a computer readable medium 804, a database 806, and a network interface 808. The database 806 may contain transaction data.

[0117] The computer readable medium 804 may comprise a number of software modules including a transaction processing module 804A, a filtering module 804B, a graph module 804C, a scoring module 304D, and an action module 804E.

[0118] The transaction processing module 804A and the processor 802 can perform transaction processing such as authorization and clearing and settlement processing. They can perform data translation, data augmentation, and data security processes.

[0119] The filtering module 804B and the processor 802 filter transaction data as described above.

[0120] The graph module 804C and the processor 802 can generate graphs and sub-graphs as described above.

[0121] The scoring module 804C and the processor 802 can score communities of nodes and sub-graphs as described above.

[0122] The action module 804C and the processor 802 can perform actions in response to the determination of anomalous groups. Such actions can include shutting down accounts, preventing access to computers, displaying graphs highlighting the groups, etc.

[0123] The computer readable medium 804 may also comprise code, executable by the processor 802 for performing a method comprising: determining a set of nodes representing suspicious entities; generating a connected graph using the set of nodes, wherein the connected graph comprises edges between the nodes; determining sub-graphs of the connected graph by splitting the connected graph; determining a set of potentially anomalous groups using the sub-graphs; ranking the potentially anomalous groups; determining a subset of the potentially anomalous groups that are higher ranked than other potentially anomalous subgraphs in the set; and taking, by the computer, one or more actions with respect to the subset of potentially anomalous groups.

[0124] It should be understood that any of the embodiments can be implemented in the form of control logic using hardware (e.g., an application specific integrated circuit or field programmable gate array) and / or using computer software with a generally programmable processor in a modular or integrated manner. As user herein, a processor includes a multi-core processor on a same integrated chip, or multiple processing units on a single circuit board or networked. Based on the disclosure and teachings provided herein, a person of ordinary skill in the art will know and appreciate other ways and / or methods to implement embodiments using hardware and a combination of hardware and software.

[0125] Any of the software components or functions described in this application may be implemented as software code to be executed by a processor using any suitable computer language such as, for example, Java, C, C++, C#or scripting language such as Perl or Python using, for example, conventional or object-oriented techniques. The software code may be stored as a series of instructions or commands on a computer readable medium for storage and / or transmission, suitable media include random access memory (RAM) , a read only memory (ROM) , a magnetic medium such as a hard-drive or a floppy disk, or an optical medium such as a compact disk (CD) or DVD (digital versatile disk) , flash memory, and the like. The computer readable medium may be any combination of such storage or transmission devices.

[0126] Such programs may also be encoded and transmitted using carrier signals adapted for transmission via wired, optical, and / or wireless networks conforming to a variety of protocols, including the Internet. As such, a computer readable medium according to an embodiment may be created using a data signal encoded with such programs. Computer readable media encoded with the program code may be packaged with a compatible device or provided separately from other devices (e.g., via Internet download) . Any such computer readable medium may reside on or within a single computer product (e.g., a hard drive, a CD, or an entire computer system) , and may be present on or within different computer products within a system or network. A computer system may include a monitor, printer, or other suitable display for providing any of the results mentioned herein to a user.

[0127] Any of the methods described herein may be totally or partially performed with a computer system including one or more processors, which can be configured to perform the steps. Thus, embodiments can be directed to computer systems configured to perform the steps of any of the methods described herein, potentially with different components performing a respective steps or a respective group of steps. Although presented as numbered steps, steps of methods herein can be performed at a same time or in a different order. Additionally, portions of these steps may be used with portions of other steps from other methods. Also, all or portions of a step may be optional. Additionally, any of the steps of any of the methods can be performed with modules, circuits, or other means for performing these steps.

[0128] The specific details of particular embodiments may be combined in any suitable manner without departing from the spirit and scope of embodiments of the disclosure. However, other embodiments may be directed to specific embodiments relating to each individual aspect, or specific combinations of these individual aspects.

[0129] For the purposes of explanation, specific details are set forth in order to provide a thorough understanding of the exemplary embodiments. However, it will be apparent that various embodiments may be practiced without these specific details. For example, circuits, systems, algorithms, structures, techniques, networks, processes, and other components may be shown as components in block diagram form in order not to obscure the embodiments in unnecessary detail.

[0130] The above description is illustrative and is not restrictive. Many variations will become apparent to those skilled in the art upon review of the disclosure. The scope should, therefore, be determined not with reference to the above description, but instead should be determined with reference to the pending claims along with their full scope or equivalents.

[0131] A recitation of "a" , "an" or "the" is intended to mean "one or more" unless specifically indicated to the contrary.

[0132] All patents, patent applications, publications, and descriptions mentioned above are herein incorporated by reference in their entirety for all purposes. None is admitted to be prior art.

Claims

1.A method of detecting anomalous groups, the method comprising:determining, by a computer, a set of nodes representing suspicious entities;generating, by the computer, a connected graph using the set of nodes, wherein the connected graph comprises edges between nodes in the set of nodes;determining, by the computer, sub-graphs of the connected graph by splitting the connected graph;determining, by the computer, a set of potentially anomalous groups using the sub-graphs;ranking, by the computer, the potentially anomalous groups in the set of potentially anomalous groups;determining, by the computer, a subset of the potentially anomalous groups that are higher ranked than other potentially anomalous subgraphs in the set of potentially anomalous groups; andtaking, by the computer, one or more actions with respect to the subset of the potentially anomalous groups.2.The method of claim 1, further comprising, prior to determining the set of nodes:obtaining, by the computer, transaction data associated with a plurality transactions;filtering, by the computer using one or more filters, the transaction data to form filtered transaction data; anddetermining, by the computer, the suspicious entities from the filtered transaction data.3.The method of claim 2, further comprising:cleaning the transaction data to remove extraneous data.4.The method of claim 2, further comprising:performing a similarity calculation between two suspicious entities using the filtered transaction data.5.The method of claim 4, wherein the similarity calculation is compares whether the two suspicious entities performed transactions with another entity within a similar time period.6.The method of claim 1, wherein determining, by the computer, the sub-graphs of the connected graph by splitting the connected graph, comprises using a graph split algorithm to split the connected graph.7.The method of claim 6, wherein the graph split algorithm is a Louvain algorithm.8.The method of claim 1, wherein the suspicious entities are suspicious individuals, computers, IP addresses, or merchants.9.The method of claim 1, wherein the subset of the potentially anomalous groups further comprising groups that are suspected of being involved with fraud.10.The method of claim 1, wherein the suspicious entities are individuals.11.The method of claim 1, wherein taking the one or more actions comprises:automatically taking preventative action with respect to the suspicious entities in the subset of the subset of the potentially anomalous groups.12.The method of claim 11, wherein the preventative action comprises performing an investigation of the suspicious entities in the subset of the potentially anomalous groups, sending signals to disable the suspicious entities in the subset of the potentially anomalous groups, and providing one or more alerts an authority entity regarding the suspicious entities in the subset of the potentially anomalous groups.13.A computer comprising:a processor; anda non-transitory computer readable medium, the non-transitory computer readable medium comprising code executable by the processor, to perform operations comprising:determining a set of nodes representing suspicious entities;generating a connected graph using the set of nodes, wherein the connected graph comprises edges between nodes in the set of nodes;determining sub-graphs of the connected graph by splitting the connected graph;determining a set of potentially anomalous groups using the sub-graphs;ranking the potentially anomalous groups in the set of potentially anomalous groups;determining a subset of the potentially anomalous groups that are higher ranked than other potentially anomalous subgraphs in the set of potentially anomalous groups; andtaking, by the computer, one or more actions with respect to the subset of the potentially anomalous groups.14.The computer of claim 13, wherein the operations further comprise, prior to determining the set of nodes:obtaining, by the computer, transaction data associated with a plurality transactions;filtering, by the computer using one or more filters, the transaction data to form filtered transaction data; anddetermining, by the computer, the suspicious entities from the filtered transaction data.15.The computer of claim 14, wherein a filter in the one or more filters is based on a feature set obtained using training data.16.The computer of claim 15, wherein the feature set is generated using a supervised machine learning model.17.The computer of claim 15, wherein the feature set is based on features obtained from a logistic regression model.18.The computer of claim 13, wherein determining, by the computer, the sub-graphs of the connected graph by splitting the connected graph, comprises using a graph split algorithm to split the connected graph.19.The computer of claim 13, wherein the suspicious entities are suspicious individuals, computers, IP addresses, or merchants.20.The computer of claim 13, wherein taking the one or more actions comprises:automatically taking preventative action with respect to the suspicious entities in the subset of the subset of the potentially anomalous groups.

Citation Information

Patent Citations

  • Abnormal user identification method and device, storage medium and electronic equipment

    CN111612039A

  • Abnormal behavior detection processing method and device

    CN113935832A

  • Abnormal transaction entity identification method, apparatus and device, and readable storage medium

    CN116012152A

  • Abnormal transaction identification method and device, computer equipment and storage medium

    CN117436882A

  • Level of network suspicion detection

    US11190534B1

Cited By

  • Integration link resonance early warning method and early warning device

    CN121724683A