Probabilistic account linking
Through the probabilistic account linking method, the account data and clustering information of the online trading platform are utilized to dynamically adjust the account linking probability, which solves the problem of multi-account fraud identification on the online trading platform and achieves effective fraud detection and satisfaction of anti-money laundering requirements.
Patent Information
- Application Number
- CN202510319378.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-19
- Filing Date
- 2025-03-18
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies make it difficult to effectively identify and prevent fraudulent activities conducted through multiple accounts on online trading platforms, especially financial fraud. Traditional deterministic and subjective methods are prone to misjudgments and uncertainties and cannot meet the requirements of anti-money laundering regulations.
A probabilistic account linking method is adopted. By defining linking strategies and generating the average linking probability for each strategy, account data and account clustering information are used to evaluate whether accounts belong to the same entity, and the linking probability is dynamically adjusted to identify potential fraudulent behavior.
It improves the ability of online trading platforms to identify accounts to be linked, can effectively detect and prevent fraudulent activities, meet anti-money laundering requirements, and provide a reproducible and unified application method.
Smart Images

Figure CN120707230A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to probabilistic account linking. Background Art
[0002] The digital world of e-commerce, payments, banking, and other systems that support online transactions between entities has provided numerous opportunities for fraud. Specifically, bad actors can establish multiple accounts on these digital platforms to carry out a variety of fraudulent activities. For example, a particular concern is financial activities associated with collusion or self-payment. This type of activity is often directly linked to money laundering, which is subject to anti-money laundering (AML) regulations, which require online trading platforms to monitor and potentially report fraudulent activity. Other fraudulent activities conducted through multiple accounts in the context of e-commerce platforms include: sellers using multiple accounts to create duplicate listings or profit from chargebacks; fraudsters with multiple active accounts attempting to fraudulently redeem gift cards; and fake bidding, in which sellers use a second account to bid in their own auctions to inflate prices. Summary of the Invention
[0003] Some aspects of the present technology relate, in particular, to probabilistic account linking of online trading platforms to identify accounts belonging to the same entity. According to some configurations, linking strategies are defined for linking accounts on the online trading platform. Each linking strategy identifies one or more account attributes for which accounts can share the same attribute values. An average linking probability for each linking strategy is generated using account data for accounts on the online trading platform. The average linking probability for a linking strategy represents the likelihood that two accounts belong to the same entity if they share a linking strategy. In some aspects, the average linking probability for a given linking strategy is generated based on account clustering information, in which each account cluster uses, for example, an entity resolution method to identify one or more accounts that have been identified as corresponding to the same entity. The average linking probabilities for the linking strategies are stored so that they can be used for account linking purposes.
[0004] When evaluating whether to link two accounts, a linking policy shared by the two accounts is identified. This shared linking policy is one in which the two accounts share the same attribute value for the account attribute of the shared linking policy. Average link probabilities for the shared policy are retrieved, and based on these average linking probabilities, account linking probabilities for the two accounts are generated. This account linking probability represents the likelihood that the two accounts belong to the same entity and is used to determine whether to link the two accounts. In some aspects, the account linking probability is compared to a threshold, and if the account linking probability meets the threshold, the two accounts are linked. One or more actions may be taken in response to linking the two accounts, such as blocking transactions between the two accounts or reporting past transactions between the two accounts.
[0005] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] The present technology is described in detail below with reference to the accompanying drawings, in which:
[0007] Figure 1 is a block diagram illustrating an exemplary system according to some embodiments of the present disclosure;
[0008] Figure 2 is a table showing examples of account clustering information that can be used to generate link probabilities according to some embodiments of the present disclosure;
[0009] Figure 3 is a table illustrating an example of average link probabilities that can be used to generate account link probabilities for two accounts according to some embodiments of the present disclosure;
[0010] Figure 4 is a flow chart illustrating an overall method for generating an average link probability according to some embodiments of the present disclosure;
[0011] Figure 5 is a flow chart illustrating a method for generating an average link probability for a selected linking strategy according to some embodiments of the present disclosure;
[0012] Figure 6 is a flow chart illustrating a method for performing fraud detection using an account linking probability between two accounts according to some embodiments of the present disclosure;
[0013] Figure 7 is a flow chart illustrating a method for generating an account linking probability between two accounts according to some embodiments of the present disclosure;
[0014] Figure 8 is a flow chart illustrating a method for performing fraud detection when attempting a transaction according to some embodiments of the present disclosure; and
[0015] Figure 9 is a block diagram of an exemplary computing environment suitable for use in embodiments of the present disclosure. DETAILED DESCRIPTION
[0016] Overview
[0017] The ease with which users can create accounts and conduct electronic transactions on online trading platforms (including, for example, e-commerce, payment, and banking systems) presents unique challenges for identifying and combating fraud, challenges that did not exist before the advent of these platforms. To combat fraud and comply with anti-money laundering (AML) requirements, online trading platforms have developed entity resolution technologies that identify when multiple accounts belong to the same entity. Typically, deterministic techniques are employed to definitively identify accounts belonging to the same entity. However, these deterministic methods are often overly restrictive and miss instances where accounts should be linked. Furthermore, this can expose online trading platforms to liability for failing to comply with AML requirements. To address this shortcoming of deterministic entity resolution, some online trading platforms have adopted subjective methods for linking accounts. For example, a method can be used that assigns weights to different criteria used to link two accounts based on human judgment. However, such subjective methods are often difficult to implement at scale and introduce excessive uncertainty.
[0018] Aspects of the technology described herein improve the ability of online transaction platforms to identify accounts to be linked, for example, to detect and / or prevent fraud in online transactions conducted by an entity using multiple accounts. Rather than relying on deterministic or subjective methods for linking accounts, the technology described herein employs a probabilistic linking approach for determining the likelihood that accounts belong to the same entity.
[0019] According to some aspects of the technology described herein, linking policies are defined for evaluating whether to link accounts on an online trading platform. Each linking policy identifies one or more account attributes that can share the same attribute value. For example, a linking policy may correspond to an IP address, such that if two accounts have the same IP address value, the two accounts share the IP address linking policy. As another example, a linking policy may correspond to a mailing address and a credit card number, such that if two accounts share the same mailing address value and the same credit card number, the two accounts share the mailing address-credit card linking policy.
[0020] An average link probability for each linking strategy is generated using account data for accounts on the online trading platform, and the average link probabilities are stored so that they can be used for account linking purposes. The average link probability for a linking strategy represents the likelihood that two accounts belong to the same entity if they share a linking strategy. According to some aspects, multiple average link probabilities are generated for a given linking strategy (where each average link probability for a linking strategy corresponds to the number of accounts that share the same linking strategy attribute value). For example, for an IP address linking strategy, a first average link probability may be generated when there are two accounts sharing the same IP address, a second average link probability may be generated when there are three accounts sharing the same IP address, a third average link probability may be generated when there are four accounts sharing the same IP address, and so on. Thus, for a given linking strategy, the likelihood that two accounts belong to the same entity may vary based on the total number of accounts that share the same linking strategy attribute value with the two accounts.
[0021] In some aspects, in order to generate an average link probability for a given linking strategy and a given number of accounts that share a linking strategy attribute value, a link probability is determined for each of a plurality of different attribute values of an account attribute of the linking strategy, wherein the given number of accounts share each of these attribute values in the linking strategy. The average link probability for the linking strategy is generated by the given number of accounts based on these link probabilities. For example, assume that an average link probability for an IP address linking strategy is generated when there are six accounts that share the same IP address value. In this case, the IP address values of the six accounts that share each of these IP address values are identified in the account data. The link probability for each of these IP address values is determined, and the average link probability for the IP address linking strategy when there are six accounts that share the same IP address is generated as the average of these link probabilities.
[0022] As will be described in further detail below, the link probability for a given attribute value of a linking policy can be determined by identifying accounts having the attribute value, identifying the account clusters to which each of these accounts belongs, and generating a link probability based on the number of accounts in each cluster and the total number of accounts having the attribute value of the linking policy. Each account cluster can include one or more accounts that have been identified as belonging to the same entity, for example, using a deterministic entity resolution method.
[0023] Once generated, the average linking probability is used to evaluate whether to link accounts on an online trading platform. Pairs of accounts to be linked can be evaluated for a variety of different applications. By way of example only, and not limitation, two accounts can be evaluated at the time of a proposed transaction between them to determine whether to allow or block the transaction. As another example, the two accounts can be evaluated based on past transactions between them, for example, to determine whether to report the transaction for AML purposes. As another example, accounts that share one or more linking policies with a fraudulent account (i.e., accounts previously identified as engaging in fraudulent activity) can be evaluated against a fraudulent account to determine whether the account should be considered fraudulent.
[0024] In order to evaluate two accounts to be linked, a link policy shared by the two accounts is identified. The shared link policy is a link policy in which the two accounts share the same attribute value for the account attribute of the shared link policy. For example, if two accounts have the same IP address value, the two accounts share the IP address link policy. The average link probability for the shared policy is retrieved, and the account link probability of the two accounts is generated based on these average link probabilities. In some aspects, the average link probability retrieved for each shared policy corresponds to the total number of accounts that share the same attribute value for the shared policy. For example, if the two accounts being evaluated share the same IP address value, and a total of six accounts share the same IP address value, the average link probability when the six accounts share the same IP address value will be retrieved and used to generate the account link probability of the two accounts.
[0025] In some aspects, when multiple shared linking strategies exist, a recursive function is used to generate the account linking probabilities for two accounts. In some cases, the shared linking strategies of two accounts may have overlapping account attributes. For example, two accounts may share an IP address linking strategy and an IP address-credit card linking strategy. In this case, a subset of the shared linking strategies with the highest average linking probabilities is selected, such that the subset does not include any overlapping account attributes.
[0026] The account linking probability generated for the two accounts is used to determine whether to link the two accounts. In some aspects, the account linking probability is compared to a threshold, and if the account linking probability meets the threshold, the two accounts are linked. In response to linking the two accounts, one or more actions may be performed. For example, a notification may be generated and sent to an administrator's computing device, who may then perform a task based on the notification. As another example, transactions between the two accounts may be blocked. As another example, past transactions between the two accounts may be reported based on the linking.
[0027] While much of the description provided herein focuses on fraud detection / prevention, in other aspects, the probabilistic account linking of the present technology can be used for applications other than fraud detection / prevention. By way of example only, and not limitation, in some configurations, a recommendation system can utilize the probabilistic account linking described herein for recommendation purposes. For example, accounts on different platforms can be compared to link these accounts, and information about users associated with these accounts can be aggregated to select recommendations on one of the platforms to provide to the user. Thus, the action taken in response to linking accounts using aspects of the technology described herein can be to generate recommendations based on the account linking.
[0028] The techniques described herein offer several improvements over existing technologies. For example, the probabilistic account linking method described herein identifies accounts that should be linked but are not deterministically linked. Consequently, this technique enables the detection of instances of fraud that would otherwise be missed by deterministic methods. Furthermore, the linking method described herein can be applied retroactively not only to identify fraudulent transactions that have already occurred, but also to assess whether to prevent them before they occur. Furthermore, the techniques described herein provide a reproducible method that can be uniformly applied to any account created on an online trading platform.
[0029] Example System for Probabilistic Account Linking
[0030] Referring now to the accompanying drawings, Figure 1 is a block diagram illustrating an exemplary system 100 for performing fraud detection through probabilistic account linking according to an embodiment of the present disclosure. It should be understood that this arrangement and other arrangements described herein are set forth only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, commands, and groupings of functions) can be used to supplement or replace the arrangements and elements shown, and some elements can be omitted entirely. In addition, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in combination with other components and in any suitable combination and location. The various functions performed by one or more entities described herein can be performed by hardware, firmware, and / or software. For example, the various functions can be performed by a processor executing instructions stored in a memory.
[0031] System 100 is an example of a suitable architecture for implementing some aspects of the present disclosure. System 100 includes, among other components not shown, any number of user devices 102A-102N, an online transaction platform 104, and a fraud detection system 106. Figure 1 Each of the illustrated user devices 102A to 102N, the online transaction platform 104, and the fraud detection system 106 may include one or more computer devices, such as the ones discussed below. Figure 9 The computing device 900. Figure 1 As shown, user devices 102A through 102N, online transaction platform 104, and fraud detection system 106 can communicate via a network 108, which may include, but is not limited to, one or more local area networks (LANs) and / or wide area networks (WANs). Such network environments are common in offices, enterprise-wide computer networks, intranets, and the Internet. It should be understood that any number of user devices and servers may be employed within system 100 within the scope of the present technology. Each may comprise a single device or multiple devices collaborating in a distributed environment. For example, online transaction platform 104 and fraud detection system 106 may each be provided by multiple server devices, which collectively provide the functionality of online transaction platform 104 and fraud detection system 106 as described herein. Furthermore, other components not shown may also be included within the network environment.
[0032] Each of user devices 102A through 102N may be a client device on the client side of operating environment 100, while online trading platform 104 and fraud detection system 106 may be on the server side of operating environment 100. Online trading platform 104 and / or fraud detection system 106 may each include server-side software designed to work in conjunction with the client-side software on user devices 102A through 102N to implement any combination of features and functionality discussed herein. For example, each of user devices 102A through 102N may include an application (not shown) for interacting with online trading platform 104 and / or fraud detection system 106. This application may be, for example, a web browser or a dedicated application for providing functionality such as interacting with online trading platform 104 and / or fraud detection system 106. This division of operating environment 100 is provided to illustrate one example of a suitable environment, and does not require that any combination of online trading platform 104 and fraud detection system 106 be maintained as separate entities for every implementation. For example, in some aspects, fraud detection system 106 is part of online trading platform 104. While operating environment 100 illustrates a configuration in a networked environment with separate user devices, online transaction platforms, and fraud detection systems, it should be understood that other configurations may be employed in which aspects of the various components are combined.
[0033] Each of the user devices 102A to 102N may include any type of computing device capable of being used by a user. For example, in one aspect, the user device may be a computer system described herein. Figure 9The type of computing device 900 described. By way of example and not limitation, user devices 102A-102N may each be embodied as a personal computer (PC), a laptop computer, a mobile or mobile device, a smartphone, a tablet computer, a smartwatch, a wearable computer, a personal digital assistant (PDA), an MP3 player, a global positioning system (GPS) or device, a video player, a handheld communication device, a gaming device or system, an entertainment system, an in-vehicle computer system, an embedded system controller, a remote control, an appliance, a consumer electronic device, a workstation, or any combination of these described devices, or any other suitable device. Different users may be associated with each of user devices 102A-102N and may interact with online transaction platform 104 and / or fraud detection system 106 via user devices 102A-102N.
[0034] The online transaction platform 104 may be implemented using one or more server devices, one or more platforms with corresponding application programming interfaces, cloud infrastructure, and the like. The online transaction platform 104 generally comprises any computer-based system that enables entities to establish accounts and use these accounts to conduct electronic transactions via the user devices 102A-102N over the network 108. In some aspects, the online transaction platform 104 comprises a listing platform (e.g., an e-commerce platform) that typically provides item listings to the user devices 102A-102N and facilitates electronic purchase transactions of items between seller accounts and buyer accounts. These item listings describe (physical or digital) items available for purchase, rental, streaming, downloading, etc. In other aspects, the online transaction platform 104 comprises a payment platform that facilitates electronic payment transactions between two accounts. In yet other aspects, the online transaction platform comprises a banking platform that facilitates electronic transfers of funds between accounts.
[0035] As previously indicated, the online trading platform 104 employs accounts to enable and track user interactions with the online trading platform, including electronic transactions between accounts. For example, in the context of a listing platform, a buyer account can be used to browse, search, and purchase items, while a seller account can be used to market and sell items, manage inventory, set prices, and track sales performance. The account data storage device 110 stores information about each account on the online trading platform 104. The account data for each account maintained in the account data storage device 110 can include values for each of a number of different account attributes, such as the user's name, email address, mailing address, IP address, credit card number, phone number, bank account number, and device identifier. In some aspects, the account data is stored as attribute-value pairs (e.g., name: John Smith; email: jsmith@email.com; etc.). Furthermore, the account data for an account can be stored along with an account identifier that uniquely identifies the account.
[0036] At a high level, the fraud detection system 106 uses a probabilistic approach to determine whether to link accounts. Figure 1 As shown, the fraud detection system 106 includes an average link probability component 112, an account link probability component 114, and an account link component 116. The components of the fraud detection system 106 may be in addition to other components that provide additional functionality beyond the features described herein. The fraud detection system 106 may be implemented using one or more server devices, one or more platforms with corresponding application programming interfaces, cloud infrastructure, etc. Figure 1 In the configuration of FIG, the fraud detection system 106 is shown as being separate from each of the user devices 102A to 102N and the online transaction platform 104, but it should be understood that in other configurations, some functions of the fraud detection system 106 may be provided on the online transaction platform 104 and / or the user devices 102A to 102N.
[0037] In some aspects, the functions performed by the components of the fraud detection system 106 are associated with one or more applications, services, or routines. Specifically, such applications, services, or routines may run on one or more user devices, servers, be distributed across one or more user devices and servers, or be implemented in the cloud. Furthermore, in some aspects, these components of the fraud detection system 106 may be distributed across a network (including one or more servers and client devices in the cloud) and / or may reside on a user device. Furthermore, these components, the functions performed by these components, or the services performed by these components may be implemented at an appropriate abstraction layer (e.g., operating system layer, application layer, hardware layer, etc.) of a computing system. Alternatively or additionally, the functions of these components and / or aspects of the techniques described herein may be performed at least in part by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like. Furthermore, while functionality is described herein with respect to specific components shown in the example system 100, it is contemplated that in some aspects the functionality of these components may be shared or distributed across other components.
[0038] The average link probability component 112 of the fraud detection system 106 uses the account data from the account data storage device 110 to generate an average link probability for a link policy. A link policy can be defined as one or more account attributes corresponding to the same value used to evaluate whether to link accounts based on the accounts sharing the same link policy for each account attribute. In some cases, a link policy identifies a single account attribute. For example, a link policy can consist of an email address attribute, such that two accounts share the link policy based on the two accounts having the same email address value. In other cases, a link policy identifies two or more account attributes. For example, a link policy can consist of a mailing address attribute and an email address attribute, such that two accounts share the link policy based on the two accounts having the same mailing address value and the same email address value.
[0039] The average link probability for a link policy represents the probability that two accounts correspond to the same entity based on the fact that the two accounts share the link policy. For example, if the average link probability for a link policy consisting of an email address attribute is 20%, then two accounts that share the same email address value can be considered to have a 20% chance of belonging to the same entity based solely on the link policy. A different average link probability can be provided for each link policy based on the number of accounts that share the same link policy attribute value. For example, the average link probability for an email address link policy can be 20% when there are four accounts that share the same email address value, 18% when there are five accounts that share the same email address value, and 16% when there are six accounts that share the same email address value.
[0040] For a given linking strategy, the average link probability component 112 determines a link probability for each unique attribute value of the account attributes of the linking strategy, or, when the linking strategy is based on two or more account attributes, determines a link probability for each unique combination of attribute values. The average link probability component 112 then generates an average link probability for the linking strategy based on the link probabilities of the unique attribute values and based on the number of accounts that share each attribute value.
[0041] For example, given an email address-based linking strategy, the average link probability component 112 calculates a link probability for each of a plurality of unique email address values. To generate an average link probability when four accounts share an email address value, the average link probability component 112 identifies a link probability for each of the unique email address values for the four accounts that share the same email address value. The average link probability component 112 then averages these link probabilities to provide an average link probability for the email address linking strategy when the four accounts share the same email address value. Similarly, to generate an average link probability when five accounts share an email address value, the average link probability component 112 identifies a link probability for each of the unique email address values for the five accounts that share the same email address value. The average link probability component 112 then averages these link probabilities to provide an average link probability for the email address linking strategy when the five accounts share the same email address value.
[0042] As another example, given a linking strategy based on a combination of name and mailing address, average link probability component 112 calculates a link probability for each of multiple unique combinations of name and mailing address values. To generate an average link probability when four accounts share the same name and mailing address values, average link probability component 112 identifies a link probability for each of the unique combinations of name and mailing address values for the four accounts that share the same combination of name and mailing address values. Average link probability component 112 then averages these link probabilities to provide an average link probability for the name-mailing address linking strategy when the four accounts share the same name value and the same mailing address value. Similarly, to generate an average link probability when five accounts share the same name and mailing address value, average link probability component 112 identifies a link probability for each of the unique combinations of name and mailing address values for the five accounts that share the same combination of name and mailing address values. Average link probability component 112 then averages these link probabilities to provide an average link probability for the name-mailing address linking strategy when the five accounts share the same name value and the same mailing address value.
[0043] According to some aspects of the techniques described herein, the average link probability component 112 determines the average link probability for a linking strategy by utilizing account cluster information for accounts stored in the account data store 110. Each account cluster in the account cluster information identifies a group of accounts that has been identified with a high degree of certainty as corresponding to the same entity (e.g., by account identifier). The average link probability component 112 can generate account clusters or otherwise access account cluster information generated by another component. Account clusters can be generated using any of a variety of entity resolution techniques, such as rule-based matching, exact matching of account attributes, fuzzy matching of account attributes, and machine learning-based approaches. In some aspects, account clustering is provided through a deterministic linking process of accounts, in which multiple accounts are deterministically determined to belong to the same entity or cluster. In some cases, the clustering information includes clusters that group accounts together using transitive account links. For example, if account A is linked to account B and account B is linked to account C, then account A will be linked to account C, such that accounts A, B, and C will form an account cluster.
[0044] The account cluster information serves as a ground truth for calculating the link probability for a unique attribute value of a link strategy. To calculate the link probability for a given attribute value of a link strategy, the average link probability component 112 accesses account data from the account data storage device 110 to identify the accounts with that particular attribute value (e.g., by account identifier) and also identifies the account cluster to which each of these accounts belongs (e.g., by cluster identifier). In some aspects, the average link probability component 112 uses the following function to calculate the link probability for a given attribute value of a link strategy:
[0045] (1)
[0046] Among them, m i Indicates belonging to the cluster with the identifier c i is the number of accounts in the same cluster, and N is the total number of accounts with the linking strategy attribute value.
[0047] For example, suppose we are determining the link probability of a linking strategy based on IP addresses. Given a specific IP address value, we identify the accounts with that IP address value and the account clusters to which each of these accounts belongs. Figure 2 A table with example information is provided showing accounts with a given IP address value (by account identifier 202) and the account clusters for each of these accounts (by cluster identifier 204). In this example, there are a total of six accounts (A, B, C, D, E, F) with this IP address value, and these six accounts fall into three clusters (C1, C2, C3). Here: when i=1, the account with cluster identifier ci m i (number of accounts) is 3 (i.e., accounts A, B, C in cluster C1); when i=2, with cluster identifier c i m i (number of accounts) is 2 (i.e., accounts D and E in cluster C2); when i=3, with cluster identifier c i m i (Number of accounts) is 1 (i.e., account F in cluster C3). Given equation (1) above, the link probability for this IP address value is 27%:
[0048] P=((3*(3-1))+(2*(2-1))+(1*(1-1))) / (6*(6-1))=0.27.
[0049] In other words, if two of the six accounts are picked at random (no repeats), the link probability provides the chance that the pair will have the same cluster identifier. In this example, the chance is 0.27 (or less than 1 in 3 attempts will result in the same cluster pair). This process can be repeated to determine the link probability for each of the multiple unique IP address values for the six accounts that also share the same IP address value, and the average link probability for the IP address linking strategy when the six accounts share the same IP address is generated as the average of these link probabilities.
[0050] Reference again Figure 1 The average link probabilities for the different link strategies generated by the average link probability component 112 are stored in the link probability data storage device 118. In some aspects, each average link probability is stored in association with a link strategy identifier identifying the associated link strategy and an indicator identifying the number of accounts that share a link strategy attribute value.
[0051] The account linking probability component 114 uses the average linking probability stored in the linking probability data store 118 to determine the account linking probability between each pair of accounts. The account linking probability for a pair of accounts represents the likelihood that the two accounts belong to the same entity and is used to determine whether to link the accounts, as described in further detail below. The account linking probability for two accounts can be determined at different times for different fraud detection and prevention purposes. For example, in some aspects, the existence of a past transaction between two accounts (with or without certain conditions, such as transactions exceeding a certain amount (e.g., over $2,000)) can trigger the generation of an account linking probability to determine whether to link the accounts. As another example, the account linking probability between two accounts can be generated when an anticipated transaction is made to determine whether to prevent or otherwise block the transaction. In yet another example, the account linking probability for all pairs of accounts in the account data store 110 that share one or more linking policies can be periodically generated to determine whether to link the accounts. As another example, the account linking probability can be generated for any accounts that share one or more linking policies with a fraudulent account (i.e., accounts that have been associated with fraudulent activity) to determine whether to also identify these accounts as fraudulent.
[0052] Given a pair of accounts, the account link probability component 114 identifies a link policy shared by the two accounts. A shared link policy for two accounts includes a link policy where both accounts have the same attribute value (or the same attribute value for a combination of attributes) for the link policy. For example, if the two accounts share the same IP address value, the IP address link policy will be identified as a shared link policy. As another example, if the two accounts share the same name value and the same mailing address value, the name-mailing address link policy will be identified as a shared link policy.
[0053] For each shared link policy for the pair of accounts, the account link probability component 114 retrieves an average link probability based on the total number of accounts that share the same attribute value for that shared link policy. For example, the account link probability component 114 can perform a lookup in the link probability data store 118 based on the link policy identifier for each shared link policy and the total number of accounts that share the link policy attribute value to retrieve the corresponding average link probability. For example, if two accounts share the same IP address value (among other shared link policies), and a total of six accounts share the same IP address value, then the average link probability for the six accounts sharing the same IP address value will be retrieved.
[0054] The account link probability component 114 uses the retrieved average link probability (if any) to generate an account link probability for the pair of accounts. If no shared link policy exists, the account link probability is zero. If only a single shared link policy exists, the account link probability is equal to the average link probability for that shared link policy. If multiple shared link policies exist, the account link probability is generated based on those shared link policies.
[0055] In some aspects, the account linking probability component 114 uses recursion to determine the account linking probability of each pair of accounts with multiple shared linking strategies. For example, the following recursive function can be used:
[0056] P[n]=P[n-1]+(1-P[n-1])*b[n] (2)
[0057] Wherein, P[n] represents the account link probability of n kinds of shared link strategies, and b[n] represents the average link probability of the nth shared link strategy.
[0058] For example, suppose a pair of accounts has three shared linking strategies with the following average linking probabilities: phone = 0.5; bank account = 0.33; and IP address = 0.02. The account linking probability of these two accounts will be recursively calculated using equation (2) above as follows:
[0059] P[0]=0
[0060] P[1]=P[0]+(1-P[0])*b[1]=0+(1-0)*0.5=0.5
[0061] P[2]=P[2-1]+(1-P[2-1])*b[2]=0.5+(1-0.5)*0.33=0.665
[0062] P[3]=P[3-1]+(1-P[3-1])*b[3]=0.665+(1-0.665)*.02=0.6717
[0063] Therefore, in this example, the probability of account linking between the two accounts is 67.17%, based on the fact that they share three linking strategies (phone, bank account, and IP address). This calculation treats the linking probabilities between shared linking strategies as independent events.
[0064] In the case where a linking policy includes multiple account attributes, the linking policies shared by two accounts can have overlapping account attributes. For example, suppose two accounts have the following shared linking policies: email address - mailing address; email address; IP address; mailing address - credit card; mailing address - IP address; phone number, device identifier; and bank account. In this example, there are shared linking policies with overlapping account attributes—three linking policies include mailing address, two linking policies include email address, and two linking policies include IP address. In this case, a subset of shared linking policies is selected such that no account attributes overlap within the subset, and the selected subset of shared linking policies is used to determine the account linking probability for the two accounts.
[0065] In some aspects, a subset of shared link policies can be selected from a total set of shared link policies with overlapping account attributes by selecting the link policies with the highest average link probability that covers each account attribute included in the shared link policies, without any overlapping account attributes in the selected subset. This can, for example, include performing a primary sort of the set of shared link policies in descending order of average link probability, followed by a secondary sort of the link policies in lexicographic order of their account attributes. The subset of link policies can then be determined by traversing the ordered list starting with the highest average link probability and iteratively selecting link policies with account attributes not already included in previously selected link policies, while excluding link policies that include account attributes from previously selected link policies.
[0066] For example, Figure 3An example of shared linking strategies 306 and corresponding average linking probabilities 304 for a pair of accounts 302 is provided. In this example, there are eight linking strategies shared by the account pair AB, including five linking strategies based on individual account attributes and three linking strategies based on a combination of two account attributes. Performing a first sort based on average linking probabilities and a second sort based on lexicographic order provides the following ordered set: S1 = [(0.97, email address), (0.5, phone), (0.4, address-credit card), (0.33, bank account), (0.3, address-IP address), (0.02, IP address), (0.02, machine ID), (0.1, email)]. Using this ordered set, the following linking strategies are selected: email address; phone; bank account; IP address; and machine ID. Because the email, address-credit card, and address-IP address linking strategies each include account attributes included in the previously selected linking strategies, they are excluded. For example, because the previously selected email address linking strategy includes the email attribute, the email linking strategy is excluded. In some configurations, between the IP address and machine ID linking strategies, when the IP address is ranked higher than the machine ID in lexicographic order, the IP address will be selected. The account linking probability will then be determined based on the average linking probability of the selected subset of linking strategies, for example, using the recursive method discussed above.
[0067] An example algorithm that can be used by the account linking probability component 114 to determine the account linking probability of a pair of accounts is provided below:
[0068] deffinal_probability_calculator (S1):
[0069] token=[]##Define an empty array for collecting link strategies
[0070] token_prb=[]##Define an empty array for collecting average link probabilities
[0071] ##Find the strategies and probabilities that should be used for the final combination
[0072] ##Probability calculation.
[0073] for x in S1:
[0074] key = str(x[1])
[0075] value = Decimal(str(x[0]))
[0076] print ('This is my key(strategy) and value(average probability)pair:')
[0077] print ('key:'+ str(key) + 'value:'+ str(x[0]))
[0078] ##If it is a single strategy, directly check whether it exists in the token array,
[0079] ## Otherwise, if it is a dual policy, split it into a single policy and check if it exists
[0080] ##In the token array. If it does not exist, add the key to the token array
[0081] ##And add the probability to the token_prb array
[0082] if (key in token) or ((len(key.split('|'))==2 and key.split('|')[0]in token) or
[0083] (len(key.split('|'))==2 and key.split('|')[1] in token)) :
[0084] print (key+' exists in token array, cannot use '+key)
[0085] elif len(key.split('|'))==2:
[0086] token.append(key)
[0087] token.append(key.split('|')[0])
[0088] token.append(key.split('|')[1])
[0089] token_prb.append(Decimal(value))
[0090] elif len(key.split('|'))==1:
[0091] token.append(key)
[0092] token_prb.append(value)
[0093] Resultant set:S2=[ (0.97,email-address), (0.5,phone), (0.33,bankaccount), (0.02,IP address) ]
[0094] ##Calculate the combined probability based on the set S2 stored in the token_prb array
[0095] comb_prb=Decimal('0.0000')
[0096] i=0
[0097] j=0
[0098] if len(token_prb) > 1:
[0099] print('token_prb:'+ str(token_prb))
[0100] for i in range(len(token_prb)):
[0101] for j in range(1,len(token_prb)):
[0102] if j==i+1:
[0103] if i==0:
[0104] comb_prb=(token_prb[i]+token_prb[j])-(token_prb[i]*token_prb[j])
[0105] else:
[0106] comb_prb=(comb_prb+token_prb[j])-(comb_prb*token_prb[j])
[0107] elif len(token_prb) == 1 :
[0108] comb_prb=token_prb[0]
[0109] ##Return the combination probability comb_prb
[0110] if format(comb_prb, '.2f')=='1.00':
[0111] comb_prb=Decimal(.99)
[0112] return format(comb_prb, '.2f')
[0113] The account linking component 116 determines whether to link each pair of accounts based on their account linking probabilities determined by the account linking probability component 114. Linking accounts generally involves associating the two accounts and causing an action to be performed. Linking may include storing information associating the two accounts to identify the accounts as linked (e.g., storing the information in the account data storage device 110 or other database).
[0114] A number of different actions can be taken with respect to linked accounts, and the type of action taken can depend on the specific fraud detection application. For example, in some aspects, the action can include generating a notification via one or more user interfaces that identifies the linked accounts and can provide data about the accounts. The notification can be transmitted via network 108 to a device associated with an administrator. This notification allows the administrator to review the linked accounts and determine whether to manually close one or both accounts, report transactions between the two accounts, or perform some other fraud-related task. In other aspects, automatic actions are taken based on the two accounts being linked. For example, in the case of an anticipated transaction, the account linking component 116 can cause the online trading platform 104 to block or otherwise prevent the transaction from completing. In the case of a past transaction, the account linking component 116 can automatically trigger a fraudulent transaction report for AML or other purposes. As another example, the account linking component 116 can cause the online trading platform 104 to close one or both accounts.
[0115] In some aspects, the account linking component 116 uses thresholds to determine whether to link accounts. For a given pair of accounts, the account linking component 116 compares the determined probability of account linking for the accounts with a threshold. If the probability of account linking meets the threshold, the accounts are linked. The threshold used by the account linking component 116 can vary based on the fraud detection application. For example, an application that automatically prevents a potential transaction from completing might use one threshold, while an application that automatically closes an account might use a different threshold. In other aspects, multiple thresholds can be used to enable different actions to be taken. This allows for advanced types of actions to be taken based on the strength of the probability of account linking for a pair of accounts. For example, if a first, lower threshold is met, the accounts are linked and a notification is provided to an administrator; and if a second, higher threshold is met, the accounts are linked and automatically closed. In some cases, all types of actions can be taken for a pair of accounts where the probability of account linking meets a threshold. For example, if two thresholds are met for a given pair of accounts, two related actions can be taken (e.g., reporting to an administrator and automatically closing one or both accounts). In other aspects, a single threshold can enable multiple actions to be taken.
[0116] Example Method for Probabilistic Account Linking
[0117] Now refer to Figure 4 , a flow chart showing an overall method 400 for generating an average link probability is provided. The method 400 may be performed, for example, by Figure 1 The method 400 and any other method described herein may be performed by the average link probability component 112 of the fraud detection system 106. Each block of the method 400 and any other method described herein comprises a computational process performed using any combination of hardware, firmware, and / or software. For example, various functions may be performed by a processor executing instructions stored in a memory. These methods may also be embodied as computer-usable instructions stored on a computer storage medium. These methods may be provided as a standalone application, a service, or a hosted service (standalone or in combination with another hosted service), or as a plug-in to another product, to name a few.
[0118] As shown in block 402, account data is accessed. For example, account data storage devices associated with an online trading platform (e.g., Figure 1 The account data storage device 110 associated with the online trading platform 104 accesses the account data. The account data includes data associated with each of a plurality of accounts on the online trading platform, including attribute values of account attributes of each account.
[0119] Average link probabilities for a plurality of linking strategies are generated based on the account data, as shown in block 404. Each linking strategy corresponds to a single account attribute or a combination of two or more account attributes of an account in the account data. In some aspects, multiple average link probabilities are generated for each linking strategy, where each average link probability for a linking strategy corresponds to the number of accounts that share the same linking strategy attribute value.
[0120] The average link probability is stored, as shown in block 406. The average link probability may be stored, for example, in a data storage device (e.g., Figure 1 To facilitate lookup, each average link probability may be stored along with a link policy identifier for the corresponding link policy and an indication of the number of accounts that share the same link policy attribute value.
[0121] Figure 5 The average link probability for the link generation strategy is shown in Figure 4 404 of the block 406). The method 500 may be performed by, for example Figure 1 The average link probability component 112 is executed. As shown in block 502, a link strategy is selected. The selected link strategy corresponds to one or more account attributes.
[0122] At block 504, a single attribute value is selected if the selected linking policy corresponds to a single account attribute, or a combination of attribute values is selected if the selected linking policy corresponds to multiple account attributes. For example, if the linking policy corresponds to an IP address, a specific IP address value is selected; and if the linking policy corresponds to an IP address and an email address, a combination of a specific IP address value and a specific email address value is selected.
[0123] The account data is accessed based on the selected linking strategy and the selected attribute values, as shown in block 506. Specifically, accounts having the selected attribute values for the account attributes of the selected linking strategy are identified. In addition, account clustering information for each of these identified accounts is accessed. The clustering information typically identifies the account cluster to which each of the identified accounts belongs by a cluster identifier. In some aspects, accessing the clustering information may include performing account clustering to generate the account clustering information, or otherwise accessing previously generated account clustering information (e.g., stored in a Figure 1 account data storage device 110).
[0124] As shown in block 508, the account data (including clustering information of accounts identified as having the selected attribute value) is used to generate a link probability for the selected linking strategy and the selected attribute value. In some aspects, the link probability is generated using equation (1) provided herein above.
[0125] At block 510, a determination is made as to whether another attribute value (or combination of attribute values) for the selected linking strategy is available. If so, the process of blocks 502 to 508 is repeated to generate a link probability for the next selected attribute value. Figure 5 While a configuration is shown in which link probabilities for attribute values of a selected linking strategy are generated serially, it should be understood that link probabilities can be generated in parallel. In some aspects, all attribute values corresponding to the selected linking strategy and available in the account data are processed. In other aspects, only a subset of attribute values (e.g., randomly selected) are selected for processing.
[0126] Once the link probabilities for the different attribute values have been generated, an average link probability for the selected link strategy is generated based on the link probabilities for the different attribute values, as shown in block 512. For example, the average link probability can be generated as an average of the link probabilities generated for the different attribute values for the selected link strategy. In some aspects, different average link probabilities for the link strategy are generated, where each average link probability corresponds to multiple accounts that share the same link strategy attribute value (or combination of attribute values).
[0127] Next turn Figure 6 , provides a flow chart illustrating a method 600 for performing fraud detection using the account linking probability of a pair of accounts. The method 600 may be performed by, for example Figure 1 The method 600 may be performed by the account link probability component 114 and the account link component 116 in the example embodiment of the present invention. The method 600 may be performed in the context of a variety of different fraud detection applications (e.g., determining whether to block a prospective transaction between two accounts, determining whether to report a past transaction between two accounts, determining whether to close one or both accounts, and a number of other applications).
[0128] One or more shared linking policies between two accounts are identified, as shown in block 602. Identification of the shared linking policies may be performed by comparing attribute values of the two accounts to determine linking policies where the two accounts share the same attribute value (in the case of a linking policy having a single attribute) or the same combination of attribute values (in the case of a linking policy having a combination of two or more attributes).
[0129] At block 604, the average link probability for each of the shared link strategies is accessed. This may include performing a lookup using a link strategy identifier for each shared link strategy (e.g., Figure 1 ) to retrieve the average link probability corresponding to each shared link strategy based on the total number of accounts sharing the same link strategy attribute value.
[0130] As shown in block 606, the average link probability for each shared linking strategy is used to generate an account linking probability for the two accounts. At block 608, the account linking probability is compared to a threshold, and at block 610, a determination is made as to whether the account linking probability meets the threshold. If the threshold is not met, the two accounts are not linked, and the process ends, as shown in block 612. Alternatively, if the threshold is met, the two accounts are linked, as shown in block 614. Additionally, an action is taken based on the two accounts being linked, as shown in block 616. Any of a variety of different fraud detection / prevention actions may be taken, such as generating and providing a notification regarding the linked accounts, blocking the intended transaction, reporting the previous transaction, or closing one or both accounts.
[0131] Figure 7 A flow chart illustrating a method 700 for generating an account linking probability for two accounts when the accounts share multiple linking policies is provided. The method 700 may be performed by Figure 1 The account link probability component 114 is configured to perform the following operations. As shown in block 702, a set of shared link policies for the two accounts is identified. The identification of shared link policies can be performed by comparing attribute values of the two accounts to determine link policies for which the two accounts share the same attribute value (in the case of a link policy with a single attribute) or the same attribute value combination (in the case of a link policy with a combination of two or more attributes).
[0132] A subset of shared linking policies for generating account linking probabilities for the two accounts is determined, as shown in block 704. The subset of shared linking policies can be determined by selecting the shared linking policies with the highest average linking probabilities (which eliminate any overlapping account attributes between the subset of linking policies). This can include sorting the shared linking policies by average linking probability and, starting with the highest average linking probability, iteratively selecting shared linking policies that do not include account attributes from the shared linking policies previously selected for the subset. In this manner, shared linking policies that include account attributes from the shared linking policies previously selected for the subset are excluded from the subset. In the event that the shared linking policy sets for the two accounts do not include any overlapping account attributes, the subset of shared linking policies can include all shared linking policies.
[0133] As shown in block 706, the account linking probability for the two accounts is generated using the average linking probability of the subset of shared linking strategies. If the subset includes only a single shared linking strategy, the account linking probability can be the average linking probability of that shared linking strategy. If the subset includes multiple shared linking strategies, a recursive function (e.g., using equation (2) discussed above) can be used to generate the account linking probability.
[0134] Next reference Figure 8, provides a flow chart illustrating a method 800 for using account linking probability to determine whether to prevent a transaction between two accounts on an online trading platform. The method 800 may be used by the fraud detection system 106 in conjunction with Figure 1 The transaction is performed on the online trading platform 104. As shown in block 802, an indication of a transaction request between two accounts on the online trading platform is received. This can be, for example, a purchase transaction on an e-commerce platform, a payment transaction on a payment platform, a remittance on a banking platform, or other transactions between two accounts on the online trading platform.
[0135] In response to the indication of the requested transaction, an account linking probability for the two accounts is generated, as shown in block 804. This may include identifying shared linking policies for the two accounts, retrieving an average linking probability for the shared linking policies based on a total number of accounts that share the same linking policy attribute value for each shared linking policy, and generating the account linking probability based on the average linking probability for the shared linking policies.
[0136] At block 806, the account linking probability is compared to a threshold, and at block 808, a determination is made as to whether the account linking probability satisfies the threshold. If the account linking probability does not satisfy the threshold, the transaction is allowed to proceed, as shown in block 810. Alternatively, if the account linking probability satisfies the threshold, the transaction is blocked by preventing the transaction from being processed, as shown in block 812. In some aspects, one or more other actions can be performed based on blocking the transaction, such as closing the account or providing a notification to an administrator about the blocked transaction.
[0137] Exemplary Operating Environment
[0138] Having described the embodiments of the present disclosure, the following describes an exemplary operating environment in which embodiments of the present technology can be implemented in order to provide a general context for various aspects of the present disclosure. Figure 9 , an exemplary operating environment for implementing embodiments of the present technology is shown and generally designated as computing device 900. Computing device 900 is only an example of a suitable computing environment and is not intended to suggest any limitation as to the scope of use or functionality of the present technology. Neither should computing device 900 be interpreted as having any dependency or requirement relating to any one or combination of components illustrated.
[0139] The present technology may be described in the general context of computer code or machine-usable instructions, including computer-executable instructions (such as program modules) executed by a computer or other machine (such as a personal data assistant or other handheld device). Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs specific tasks or implements specific abstract data types. The present technology can be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, and the like. The present technology may also be practiced in distributed computing environments, where tasks are performed by remote processing devices linked through a communications network.
[0140] refer to Figure 9 , computing device 900 includes a bus 910 that directly or indirectly couples the following devices: memory 912, one or more processors 914, one or more presentation components 916, input / output (I / O) ports 918, I / O components 920, and an illustrative power supply 922. Bus 910 represents one or more buses (such as an address bus, a data bus, or a combination thereof). Although for clarity, Figure 9 The various boxes of FIG are represented by lines, but in reality, it is not so clear to depict the various components, and metaphorically, the lines would more accurately be gray and fuzzy. For example, a presentation component such as a display device can be considered an I / O component. Furthermore, a processor has memory. The inventors recognize that this is the nature of the art and reiterate that Figure 9 The diagrams are merely illustrative of exemplary computing devices that may be used in conjunction with one or more embodiments of the present technology. No distinction is made between categories such as "workstation," "server," "laptop," "handheld device," etc., as all of these categories are in the Figure 9 are considered within the scope of and with reference to “computing devices”.
[0141] The computing device 900 typically includes a variety of computer-readable media. Computer-readable media can be any available media that can be accessed by the computing device 900 and includes both volatile and nonvolatile media, removable and non-removable media. By way of example and not limitation, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data.
[0142] Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage devices, magnetic cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and which can be accessed by the computing device 900. The terms "computer storage medium" and "computer storage medium" do not themselves include signals.
[0143] Communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal (e.g., a carrier wave or other transport mechanism), and includes any information transmission media. The term "modulated data signal" refers to a signal that is configured or changed in a manner that encodes the information in the signal, the signal having one or more characteristics. By way of example, and not limitation, communication media include wired media such as a wired network or direct wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media. Combinations of any of the above should also be included within the scope of computer-readable media.
[0144] Memory 912 includes computer storage media in the form of volatile and / or non-volatile memory. Memory can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, and the like. Computing device 900 includes one or more processors that read data from various entities such as memory 912 or I / O components 920. Presentation component 916 presents data indications to a user or other device. Exemplary presentation components include a display device, a speaker, a printing component, a vibration component, and the like.
[0145] I / O ports 918 allow computing device 900 to logically couple with other devices, including I / O components 920, some of which may be built-in. Illustrative components include a microphone, joystick, game controller, satellite dish, scanner, printer, wireless device, etc. I / O components 920 can provide a natural user interface (NUI) that processes in-air gestures, voice, or other physiological input generated by the user. In some cases, the input can be sent to appropriate network elements for further processing. The NUI can implement any combination of the following: voice recognition, touch and stylus recognition, facial recognition, biometric recognition, gesture recognition both on and near the screen, in-air gestures, head and eye tracking, and touch recognition associated with the display on computing device 900. Computing device 900 can be equipped with a depth camera (such as a stereo camera system, an infrared camera system, an RGB camera system, and combinations thereof) for gesture detection and recognition. In addition, computing device 900 can be equipped with an accelerometer or gyroscope capable of detecting motion.
[0146] The present technology has been described with reference to specific embodiments which are intended in all respects to be illustrative rather than restrictive. Alternative embodiments will become apparent to those skilled in the art to which the technology pertains without departing from the scope of the present technology.
[0147] Various components used herein have been identified, and it should be understood that any number of components and arrangements may be used to implement the desired functionality within the scope of this disclosure. For example, for conceptual clarity, the components in the embodiments depicted in the accompanying drawings are shown with lines. Other arrangements of these and other components may also be implemented. For example, although some components are depicted as single components, many elements described herein may be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and position. Some elements may be omitted entirely. In addition, the various functions performed by one or more entities described herein may be performed by hardware, firmware, and / or software as described below. For example, the various functions may be performed by a processor executing instructions stored in a memory. Therefore, other arrangements and elements (e.g., machines, interfaces, functions, commands, and functional groupings) may also be used in addition to or in place of the arrangements and elements shown.
[0148] The embodiments described herein may be combined with one or more of the alternatives specifically described. Specifically, the claimed embodiments may include reference to more than one other embodiment in the alternatives. The claimed embodiments may specify additional limitations of the claimed subject matter.
[0149] The subject matter of embodiments of the present technology is described herein with specificity to satisfy statutory requirements. However, this description itself is not intended to limit the scope of this patent. Rather, the inventors have contemplated that the claimed subject matter may also be embodied in other ways, in conjunction with other prior art or future technologies, to include different steps or combinations of steps similar to those described in this document. Furthermore, although the terms "step" and / or "box" may be used herein to refer to different elements of the method employed, these terms should not be interpreted as implying any particular order among or between the various steps disclosed herein unless the order of the various steps is explicitly described.
[0150] For purposes of this disclosure, the word "include" has the same broad meaning as the word "comprising," and the word "access" includes "receiving," "referencing," or "retrieving." Furthermore, the word "communicate" has the same broad meaning as the word "receiving" or "sending" facilitated by a software or hardware-based bus, receiver, or transmitter using the communication media described herein. Furthermore, unless otherwise indicated, words such as "a," "an," and the like include the plural as well as the singular. Thus, for example, the constraint of "a feature" is satisfied where there is one or more features. Furthermore, the term "or" includes conjunctions, disjunctions, and both (a or b therefore includes a or b, and a and b).
[0151] For the purposes of the detailed discussion above, embodiments of the present technology are described with reference to a distributed computing environment; however, the distributed computing environment described herein is exemplary only. Components may be configured to perform the novel embodiments described herein, where the term "configured to" may refer to being "programmed to" perform a specific task or implement a specific abstract data type using code. Furthermore, while embodiments of the present technology may generally refer to the technical solution environments and schematics described herein, it should be understood that the described technology can be extended to other implementation contexts.
[0152] From the foregoing it will be seen that the present technology is well adapted to attain all of the ends and purposes set forth above, together with other advantages which are obvious and inherent to the system and method. It will be understood that some features and subcombinations are of utility and may be employed without reference to other features and subcombinations. This is contemplated by and is within the scope of the claims.
Claims
1. One or more computer storage media storing computer-usable instructions that, when used by one or more computing devices, cause the one or more computing devices to perform operations comprising: generating an average link probability for each of a plurality of linking strategies based on account data of a plurality of accounts on an online trading platform to provide a plurality of average linking probabilities; storing the plurality of average link probabilities in a data storage device; determining one or more linking strategies shared by the two accounts; retrieving from the data storage device the average link probability for each of the one or more linking strategies shared by the two accounts; generating an account linking probability for the two accounts using the average linking probability of each of the one or more linking strategies shared by the two accounts; performing a comparison of the account linking probabilities of the two accounts with a threshold value to determine whether to link the two accounts; and In response to determining to link the two accounts based on the comparison, an action is performed.
2. One or more computer storage media according to claim 1, wherein: The plurality of average link probabilities includes a first set of average link probabilities for a first link strategy, each average link probability in the first set of average link probabilities for the first link strategy corresponding to a different number of accounts sharing a same attribute value of the first link strategy.
3. One or more computer storage media according to claim 2, wherein: The two accounts share the first linking strategy, and wherein retrieving an average linking probability for each of one or more linking strategies shared by the two accounts comprises: A first average link probability is retrieved from a first set of average link probabilities for the first link policy based on a total number of accounts that share the attribute value of the first link policy with the two accounts.
4. The one or more computer storage media of claim 1 , wherein: Generating an average link probability of a first link strategy among the multiple link strategies includes: determining a link probability for each of a plurality of attribute values for a first linking strategy in the account data to provide a plurality of link probabilities for the first linking strategy; and An average link probability of the first link strategy is generated using the plurality of link probabilities of the first link strategy.
5. One or more computer storage media according to claim 4, wherein: Determining the link probability of a first attribute value among the plurality of attribute values for the first link strategy includes: identifying a subset of accounts having the first attribute value for the first linking policy; accessing account clustering information for the subset of accounts; and The link probability of the first attribute value for the first link strategy is determined according to the number of accounts in each account cluster in the account clustering information and the total number of accounts in the account subset.
6. The one or more computer storage media of claim 1, wherein: The two accounts share two or more linking strategies, and wherein generating the account linking probability of the two accounts includes: selecting a subset of the two or more linking strategies having non-overlapping account attributes; and An account linking probability of the two accounts is generated according to a subset of the two or more linking strategies.
7. One or more computer storage media according to claim 6, wherein: Selecting a subset of the two or more linking strategies having non-overlapping account attributes includes: sorting the two or more linking strategies in descending order of average linking probability; and The two or more link strategies are iteratively evaluated in descending order of the average link probability to determine whether the link strategy at each iteration is included in the subset of the two or more link strategies, wherein if the link strategy does not correspond to the account attribute of the link strategy previously selected for the subset of the two or more link strategies, the link strategy evaluated at the iteration is included in the subset of the two or more link strategies.
8. The one or more computer storage media of claim 1, wherein: The performing of the action includes preventing an electronic transaction between the two accounts on the online transaction platform in response to determining that the two accounts are linked.
9. The one or more computer storage media of claim 1, wherein: Responsive to determining to link the two accounts, causing the action to be performed includes transmitting a user interface to the remote computing device over a network, the user interface presenting an indication that the two accounts have been linked.
10. A computer-implemented method comprising: An average link probability of each of a plurality of link strategies is generated using account data of a plurality of accounts on an online trading platform, wherein a first average link probability of a first link strategy is generated by: determining a link probability for each of a plurality of attribute values of an account attribute corresponding to the first linking strategy using account cluster information identifying account clusters, wherein the same number of accounts have each of the plurality of attribute values of the account attribute corresponding to the first linking strategy, generating a first average link probability of the first link policy according to link probabilities for a plurality of attribute values of the account attribute corresponding to the first link policy; and The average link probabilities for the plurality of linking strategies are stored in a data storage device, wherein each average linking probability is stored in association with a linking strategy identifier and an indication of a plurality of accounts that share a same linking strategy attribute value.
11. The computer-implemented method of claim 10, wherein: Determining a link probability for a first attribute value among a plurality of attribute values of the account attribute corresponding to the first link strategy includes: identifying a subset of accounts having the first attribute value of the account attribute corresponding to the first linking policy; determining an account cluster for each account in the account subset according to the account clustering information; and A link probability of a first attribute value of an account attribute corresponding to the first link strategy is determined according to the number of accounts in each account cluster and the total number of accounts in the account subset.
12. The computer-implemented method of claim 10, wherein: The method further comprises: Generate the account link probability of two accounts; linking the two accounts based on a comparison of an account linking probability of the two accounts with a threshold; and In response to linking the two accounts, an action is caused to be performed.
13. The computer-implemented method of claim 12, wherein: The probability of generating an account link between the two accounts includes: identifying one or more linking strategies shared by the two accounts and a total number of accounts sharing each of the one or more linking strategies; retrieving from the data storage device an average link probability for each of the one or more linking strategies based on a total number of accounts sharing each of the one or more linking strategies; and The account linking probability of the two accounts is generated according to the average linking probability of each of the one or more linking strategies.
14. The computer-implemented method of claim 13, wherein: The two accounts share two or more linking strategies, and wherein generating the account linking probabilities of the two accounts comprises: A subset of the two or more linking strategies having the highest average linking probability and no overlapping account attributes is selected.
15. The computer-implemented method of claim 13, wherein: The two accounts share two or more linking strategies, and wherein a recursive function is used to generate the account linking probabilities of the two accounts.
16. The computer-implemented method of claim 12, wherein: Such that performing the action includes preventing transactions between the two accounts.
17. A computer system comprising: one or more processors; as well as One or more computer storage media storing computer-usable instructions that, when used by the one or more processors, cause the computer system to perform operations comprising: The account linking probability of two accounts is generated in the following way: identifying one or more linking strategies shared by the two accounts, retrieving, from a data storage device, an average link probability corresponding to each of the one or more link strategies shared by the two accounts based on the total number of accounts that share each link strategy with the two accounts, the average link probability corresponding to each of the one or more link strategies shared by the two accounts, the data storage device storing the average link probability for each of a plurality of link strategies, wherein the average link probabilities for the plurality of link strategies are generated using account clustering information of a plurality of accounts on an online trading platform, and generating an account linking probability for the two accounts using an average linking probability corresponding to each of the one or more linking strategies shared by the two accounts; linking the two accounts based on the account linking probability; and One or more actions are caused to be performed in response to linking the two accounts.
18. The computer system of claim 17, wherein: The account linking probability is generated in response to a request for a transaction between the two accounts, and wherein the one or more actions include not allowing the transaction.
19. The computer system according to claim 17, wherein: The two accounts share two or more linking strategies, and wherein generating the account linking probabilities of the two accounts further comprises: selecting a subset of the two or more linking strategies having non-overlapping account attributes; and An account linking probability of the two accounts is generated according to a subset of the two or more linking strategies.
20. The computer system of claim 19, wherein: Selecting a subset of the two or more linking strategies having non-overlapping account attributes includes: sorting the two or more linking strategies in descending order of average linking probability; and The two or more link strategies are iteratively evaluated in descending order of average link probability to determine whether the link strategy at each iteration is included in the subset of the two or more link strategies, wherein if the link strategy does not correspond to the account attribute of the link strategy previously selected for the subset of the two or more link strategies, the link strategy evaluated at the iteration is included in the subset of the two or more link strategies.