Highly correlated group account identification method and system based on machine learning model
Through the high-relevant group account identification method based on machine learning model, the correlation density and connectivity graph rules are used to solve the problem that fraud group transaction accounts cannot be accurately identified in the existing technology, and the accurate identification and efficient analysis of fraud group is achieved.
Patent Information
- Application Number
- CN202410934822.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-12
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2044-07-12
AI Technical Summary
The existing transaction fraud group identification technology is difficult to accurately identify complex fraud group transaction accounts, and cannot distinguish between high-risk and risk-free accounts, and the community division method cannot accurately judge the fraudulent nature of new accounts and rented accounts.
A high-relevant group account identification method based on machine learning model is adopted. By calculating the correlation density between transaction accounts, filtering and iterating the target transaction account using the correlation density threshold, combining the connection diagram and group association density rules, the fraud group account is accurately identified.
It realizes accurate identification of fraudulent group transaction accounts, can more comprehensively analyze and judge the structure of group fraudulent accounts, and improves identification efficiency and accuracy.
Smart Images

Figure CN119005985B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of transaction fraud risk identification, and in particular to a method and system for identifying highly correlated group accounts based on a machine learning model. Background Art
[0002] The key to identifying fraudulent transaction accounts is identifying grouped transaction accounts, a crucial step in fraudulent transaction identification. Current techniques for identifying fraudulent transaction groups primarily rely on identifying capital chain communities. For example, association graph technology is used to mine relationships between different entities, constructing a network structure of entities such as customers, accounts, and transactions. Community identification, such as label propagation or the Louvain algorithm, is then used to segment communities and determine risky behaviors within these subcommunities.
[0003] However, current identification of trading accounts within fraudulent groups still faces several challenges and challenges that remain unresolved. Community segmentation cannot accurately identify group trading accounts within fraudulent groups, nor can it distinguish between risk-free trading accounts. First, the scale and organizational structure of fraudulent groups can be extremely complex, involving a large number of trading accounts and transactions. An effective feature system is required to accurately identify group-linked trading accounts. Second, not all trading accounts associated with a group exhibit obvious fraud risk characteristics. Within a group, both high-risk and risk-free trading accounts may coexist, making differentiation difficult based solely on community segmentation. The sources of fraudulent trading accounts include the opening of new accounts and the buyout or rental of existing ones. Therefore, a trading account's historical counterparties may be legitimate counterparties, and community segmentation alone cannot be used to determine that they are part of the same fraudulent group, simply because they reside within the same sub-community. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to address the deficiencies of the existing technology and specifically provide a method and system for identifying highly correlated group accounts based on a machine learning model, as follows:
[0005] 1) In the first aspect, the present invention provides a method for identifying highly correlated group accounts based on a machine learning model. The specific technical solution is as follows:
[0006] Using the trained machine learning model, calculate the degree of association between the suspected group fraud transaction account and each transaction account associated with the suspected group fraud transaction account. Using the association degree threshold, screen out target transaction accounts associated with the suspected group fraud transaction account. Iterate using the target transaction account as the suspected group fraud transaction account to obtain multiple target transaction accounts.
[0007] Determine whether there are any transaction accounts suspected of group fraud among all target accounts.
[0008] The beneficial effects of the highly correlated group account identification method based on a machine learning model provided by the present invention are as follows:
[0009] Through the trained machine learning model, the degree of correlation between the transaction account suspected of group fraud and each transaction account associated with the transaction account suspected of group fraud is obtained. Then, the target transaction accounts associated with the transaction account suspected of group fraud are screened out, and multiple iterations are performed to accurately identify multiple target transaction accounts. This can more comprehensively analyze and identify the structure of the transaction account of the complete fraud group, and more accurately determine whether there is a transaction account involved in group fraud.
[0010] Based on the above solution, the highly correlated group account identification method based on a machine learning model of the present invention can also be improved as follows.
[0011] Furthermore, it also includes:
[0012] A connectivity graph is constructed based on all target transaction accounts screened out each time;
[0013] Determine whether there are any suspected group fraud transaction accounts and all target accounts involved in group fraud, including:
[0014] Using the group association closeness rule, each connectivity graph is judged to determine whether there are transaction accounts suspected of group fraud and whether there are transaction accounts suspected of group fraud among all target accounts.
[0015] The beneficial effect of adopting the above further solution is that: by using the group association closeness rule, it is possible to accurately determine whether there are group fraud transaction accounts.
[0016] Furthermore, the process of obtaining transaction accounts suspected of group fraud includes:
[0017] Through high-risk group account screening rules, trading accounts suspected of group fraud are screened out from all trading accounts.
[0018] The beneficial effect of adopting the above further solution is: based on the high-risk group account screening rules, transaction accounts suspected of group fraud are screened out, the recognition range of the trained machine learning model is narrowed, and the recognition efficiency is improved.
[0019] Furthermore, the process of obtaining a trained machine learning model includes:
[0020] Using known group fraud accounts and risk-free accounts as samples, the machine learning model is trained to obtain a trained machine learning model.
[0021] The beneficial effect of adopting the above further solution is: making full use of known group fraud accounts and risk-free accounts, so that the machine learning model can better capture the characteristics of fraudulent behavior and improve the recognition accuracy of the trained machine learning model.
[0022] 2) In a second aspect, the present invention also provides a highly correlated group account identification system based on a machine learning model. The specific technical solution is as follows:
[0023] Including target transaction account acquisition module and judgment module;
[0024] The target transaction account acquisition module is used to: use the trained machine learning model to calculate the degree of association between the transaction account suspected of group fraud and each transaction account associated with the transaction account suspected of group fraud; use the association degree threshold to screen out target transaction accounts associated with the transaction account suspected of group fraud; and iterate using the target transaction account as the transaction account suspected of group fraud to obtain multiple target transaction accounts;
[0025] The judgment module is used to determine whether there are any transaction accounts suspected of group fraud among all target accounts.
[0026] Based on the above solution, the high-correlation group account identification system based on a machine learning model of the present invention can also be improved as follows.
[0027] Furthermore, a connectivity graph construction module is included, which is used to: construct a connectivity graph according to all target transaction accounts screened out each time;
[0028] The judgment module is specifically used to: use the group association closeness rule to judge each connected graph to determine whether there is a group fraud transaction account among the suspected group fraud transaction account and all target accounts.
[0029] Furthermore, it also includes a transaction account screening module, which is used to: use high-risk group account screening rules to screen out transaction accounts suspected of group fraud from all transaction accounts.
[0030] Furthermore, a model training module is included, which is used to:
[0031] Using known group fraud accounts and risk-free accounts as samples, the machine learning model is trained to obtain a trained machine learning model.
[0032] 3) In a third aspect, the present invention further provides an electronic device, comprising a processor coupled to a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor, so that the electronic device implements any of the above-mentioned high-correlation group account identification methods based on machine learning models.
[0033] 4) In a fourth aspect, the present invention further provides a computer-readable storage medium, in which at least one computer program is stored, and at least one computer program is loaded and executed by a processor to enable the computer to implement any of the above-mentioned high-correlation group account identification methods based on machine learning models.
[0034] It should be noted that the beneficial effects achieved by the technical solutions of the second to fourth aspects of the present invention and the corresponding possible implementation methods can be found in the above-mentioned technical effects of the first aspect and its corresponding possible implementation methods, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:
[0036] Figure 1 Schematic diagram of a process for identifying highly correlated group accounts based on a machine learning model according to an embodiment of the present invention;
[0037] Figure 2 This is a schematic diagram of the structure of a highly correlated group account identification system based on a machine learning model according to an embodiment of the present invention;
[0038] Figure 3 The figure is a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0039] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0040] like Figure 1 As shown, a method for identifying highly correlated group accounts based on a machine learning model according to an embodiment of the present invention includes the following steps:
[0041] S1. Using the trained machine learning model, calculate the degree of association between the transaction account suspected of group fraud and each transaction account associated with the transaction account suspected of group fraud. Using the association degree threshold, screen out target transaction accounts associated with the transaction account suspected of group fraud. Iterate using the target transaction account as the transaction account suspected of group fraud to obtain multiple target transaction accounts.
[0042] The process of screening out target transaction accounts associated with transaction accounts suspected of group fraud using the correlation closeness threshold can be achieved by any of the following methods:
[0043] 1) The first method: directly determine the transaction account with a correlation closeness greater than the correlation closeness threshold as the target transaction account.
[0044] 2) The second method: Calculate the degree of association between the transaction account suspected of group fraud and each transaction account associated with the transaction account suspected of group fraud, and sort all transaction accounts associated with the transaction account suspected of group fraud according to the degree of association to obtain a transaction account sequence. Filter out transaction accounts with a degree of association greater than a threshold value from the transaction account sequence, and identify the filtered transaction accounts with a degree of association greater than the threshold value as target transaction accounts.
[0045] The number of iterations can be set according to actual conditions. Taking two iterations as an example, the process of obtaining multiple target trading accounts is explained as follows:
[0046] S10: Filter out target transaction accounts associated with the transaction accounts suspected of group fraud using a correlation closeness threshold, and record the target transaction accounts obtained in this step as first target transaction accounts.
[0047] S11. First Iteration: Using the trained machine learning model, calculate the degree of association between each first target transaction account and each corresponding associated transaction account. Using the degree of association threshold, select the target transaction accounts associated with each first target transaction account. The target transaction accounts associated with the first target transaction account obtained in this step are recorded as second target transaction accounts.
[0048] S12. Second Iteration: Using the trained machine learning model, calculate the degree of association between each second target transaction account and each corresponding associated transaction account. Using the degree of association threshold, select the target transaction accounts associated with each second target transaction account. The target transaction accounts associated with the second target transaction account obtained in this step are recorded as third target transaction accounts.
[0049] After S10 to S12 , the obtained multiple target transaction accounts include all first target transaction accounts, all second target transaction accounts, and all third target transaction accounts.
[0050] The following describes the association closeness threshold used each time to obtain the target transaction account:
[0051] The association closeness threshold used for obtaining the target transaction account each time can be the same, or the association closeness threshold used for obtaining the target transaction account each time can be set according to actual conditions. The association closeness output by the trained machine learning model ranges from 0 to 1. When the association closeness is 0, it means there is no association relationship. The association closeness threshold can be set to 0.6, or it can be set according to actual conditions.
[0052] In another embodiment, by analyzing the score distribution output by the trained machine learning model and combining the accuracy rate and business interruption rate indicators, a suitable threshold is determined as the association closeness threshold to distinguish between group accounts (target transaction accounts associated with transaction accounts suspected of group fraud) and non-group accounts (target transaction accounts not associated with transaction accounts suspected of group fraud).
[0053] Precision is calculated as follows: Precision = TP / (TP + FP). TP (True Positives) is the number of true positives, or the number of positive samples correctly predicted by the model. FP (False Positives) is the number of false positives, or the number of negative samples incorrectly predicted by the model. Disturbrate is calculated as: Disturbrate = FP / (TP + FP).
[0054] The service disturbance rate Disturbrate is set to be less than or equal to 0.5. On this basis, when the precision of the trained machine learning model reaches the maximum, the corresponding output value (that is, the predicted value of the association density output by the trained machine learning model) is the association density threshold.
[0055] Transaction accounts with a correlation closeness greater than the correlation closeness threshold are determined as target transaction accounts (group accounts) associated with transaction accounts suspected of group fraud, and transaction accounts with a correlation closeness not greater than the correlation closeness threshold are determined as target transaction accounts (non-group accounts) not associated with transaction accounts suspected of group fraud.
[0056] S2. Determine whether there are any transaction accounts suspected of group fraud among the transaction accounts and all target accounts.
[0057] In a high-correlation group account identification method based on a machine learning model in an embodiment of the present invention, the correlation between the transaction account suspected of group fraud and each transaction account associated with the transaction account suspected of group fraud is obtained through a trained machine learning model, and then the target transaction accounts associated with the transaction account suspected of group fraud are screened out, and multiple iterations are performed to accurately identify multiple target transaction accounts. This can more comprehensively analyze and identify the structure of the transaction account of the complete fraud group, and can more accurately determine whether there is a transaction account of group fraud.
[0058] Optionally, in the above technical solution, the following is further included:
[0059] A connectivity graph is constructed based on all target transaction accounts screened out each time;
[0060] The specific process of constructing a connectivity graph is as follows:
[0061] 1) Based on all the target transaction accounts filtered out each time, perform intersection association on the node table and the associated edge table respectively, read the filtered node and edge data, and store them in the edge_data variable.
[0062] 2) Create an empty undirected graph G.
[0063] 3) Traverse the edge data and add each edge to the undirected graph G.
[0064] 4) Use the nx.connected_components method to obtain all connected components of the undirected graph G to obtain a connected graph.
[0065] In S2, the transaction accounts suspected of group fraud and whether there are any group fraud transaction accounts among all target accounts are determined, including:
[0066] S20. Using the group association closeness rule, each connectivity graph is judged to determine whether there is a group fraud transaction account among the suspected group fraud transaction account and all target accounts.
[0067] The group association closeness rule is:
[0068] For any connectivity graph, the number of target accounts within three degrees of association with each target account is calculated to determine whether the current connectivity graph represents a fraudulent group. Specific rules include: at least one of the following: the number of target accounts with a transaction amount of RMB 5,000 or more within the last seven days within the three degrees of association with each target account is greater than or equal to 5; and the number of target accounts with a transaction amount of RMB 10,000 or more within the last 30 days within the three degrees of association with each target account is greater than or equal to 5.
[0069] Using the group association density rule, each connected graph is judged. The specific process is as follows:
[0070] A connectivity graph is constructed based on the screened target accounts. For each suspected group fraud transaction account (group fraud transaction account) in the connectivity graph, the number of target accounts associated with it within three degrees is calculated according to different rules to determine whether it belongs to a fraud group.
[0071] The group association closeness rule includes at least one of the following: the number of group accounts with a transaction amount of RMB 5,000 or more within three consecutive connections within the past seven days is greater than or equal to five, and the number of target accounts with a transaction amount of RMB 10,000 or more within three consecutive connections within the past 30 days is greater than or equal to five. If either of these conditions is triggered, the group association closeness rule is considered met.
[0072] For the target accounts that have been screened and triggered the group association closeness rule, a connectivity graph is reconstructed, and the target accounts in the connectivity graph are fraud groups. In this embodiment, the group association closeness rule can be used to accurately determine whether there are transaction accounts that are group fraud.
[0073] Optionally, in the above technical solution, the process of obtaining transaction accounts suspected of group fraud includes: screening out transaction accounts suspected of group fraud from all transaction accounts through high-risk group account screening rules.
[0074] The specific rules for screening high-risk group accounts are as follows:
[0075] 1) Trading accounts that have financial transactions with known fraudulent accounts.
[0076] 2) Frequently changing counterparty accounts within a short period of time, for example, the number of counterparties with a single transaction amount greater than or equal to 5,000 in the past three days is greater than or equal to 10.
[0077] 3) Transaction accounts that share or aggregate information, for example, accounts that have used the same computer or mobile phone to transfer funds in the past seven days are greater than or equal to 5, and accounts that have transferred funds at the same GPS location in the past seven days are greater than or equal to 10.
[0078] In this embodiment, transaction accounts suspected of group fraud are screened out based on high-risk group account screening rules, narrowing the recognition scope of the trained machine learning model and improving recognition efficiency.
[0079] Optionally, in the above technical solution, the process of obtaining the trained machine learning model includes:
[0080] Using known group fraud accounts and risk-free accounts as samples, the machine learning model is trained to obtain a trained machine learning model.
[0081] Among them, known group fraud accounts are defined as negative samples, and risk-free accounts are defined as positive samples. The screening rules for negative samples and positive samples are as follows:
[0082] 1) Negative sample screening rules:
[0083] ① Historical fraud records: Accounts with historical fraud records are selected as negative samples.
[0084] ② Blacklist and graylist: Use accounts on internal or external blacklists and graylists as negative samples.
[0085] ③ Abnormal transaction patterns: Identify accounts that are significantly different from normal transaction patterns, such as frequent large-value transactions, transactions at abnormal times, etc.
[0086] ④ Social network analysis: Through social network analysis, identify accounts associated with members of known fraud groups.
[0087] ⑤ Fund flow analysis: Analyze fund flow paths and identify fund flow patterns associated with fraudulent activities.
[0088] ⑥ Expert rules: Filter out negative samples based on rules defined by business experts, such as frequent changes of counterparties in a short period of time.
[0089] ⑦ Regulatory and compliance data: Use fraud case data provided by regulatory agencies or accounts in compliance reports as negative samples.
[0090] 2) The screening rules for positive samples are as follows:
[0091] ① Long-term stable transactions: Accounts with long-term stable transaction history are selected as positive samples.
[0092] ② High credit score: Use the credit scoring model to select accounts with high credit scores as positive samples.
[0093] ③ Transaction diversity: Select accounts that have transaction records with multiple counterparties and no abnormal behavior.
[0094] ④Customer feedback and evaluation: accounts based on positive customer feedback and high satisfaction ratings.
[0095] ⑤ Compliance records: Select accounts with no violation records in compliance checks as positive samples.
[0096] ⑥Transaction frequency and amount: Select an account whose transaction frequency and amount are within the normal range.
[0097] ⑦ Customer relationship management: Identify customer accounts with high loyalty and no bad records through the customer relationship management system.
[0098] ⑧Community detection: Use community detection algorithms to identify normal transaction groups in social networks.
[0099] Through the above negative sample screening rules and positive sample screening rules, negative samples and positive samples that meet the corresponding screening rules can be effectively screened out from a large number of trading accounts. Based on the screened negative samples and positive samples, a feature system is first constructed. Specifically, a multi-dimensional feature system is constructed, including but not limited to transaction behavior features, counterparty features, capital flow features, time series features, etc.; then, feature selection is performed. Specifically, feature selection techniques such as mutual information, correlation coefficient, recursive feature elimination (RFE), etc. are applied to select the most influential features. For example, 1,000 features are initially constructed, and through analysis it is found that only 50 are valid features, so feature selection processing is required. These sample data (i.e., the selected features) will provide high-quality input for subsequent model training.
[0100] Machine learning models can include random forests, gradient boosted decision trees (GBDTs), support vector machines (SVMs), etc. These models can be set according to actual conditions, and during the training process, methods such as cross-validation and grid search can be used for model training and parameter tuning to achieve optimal performance.
[0101] In this embodiment, full use is made of known group fraud accounts and risk-free accounts, so that the machine learning model can better capture the characteristics of fraudulent behavior and improve the recognition accuracy of the trained machine learning model.
[0102] In another embodiment, comprising:
[0103] Phase 1: Sample screening and feature system construction:
[0104] A bank has a vast customer base and transaction data. The bank's data science team first screened the following samples from historical transaction records:
[0105] 1) Negative samples: Identify known group-related fraud accounts through internal fraud case databases and external reports.
[0106] 2) Positive samples: Customer accounts with no historical fraud records and stable trading behavior are selected as risk-free multi-counterparty accounts.
[0107] Using these samples, the team constructed a feature system, including but not limited to: transaction frequency, transaction amount fluctuations, counterparty change rate, capital inflow and outflow patterns, basic customer information (such as age and income level), customer activity patterns on social media, IP addresses, and device fingerprint information;
[0108] Phase 2: Model training and tuning:
[0109] Using these features, the team employed a variety of machine learning algorithms (such as random forests, gradient boosting trees, and neural networks) to train a group fraud detection model. Through cross-validation and parameter tuning, the team selected the best-performing model, LightGBM.
[0110] Phase 3: Prediction of suspected group fraud accounts:
[0111] A bank's risk control system can now use high-risk associated account rules to screen out suspected group fraud accounts. These rules include:
[0112] 1) Accounts that have financial transactions with known fraudulent accounts;
[0113] 2) Frequently changing counterparty accounts within a short period of time;
[0114] 3) Accounts whose transaction patterns are similar to known group fraud patterns;
[0115] The risk control system will build features for these accounts, then use the trained model to make predictions and output model scores.
[0116] Phase 4: Group account identification and associated subgraph analysis:
[0117] Based on the model's output, the bank sets a scoring threshold and labels accounts above or equal to this threshold as group accounts. Next, using social network analysis tools, the bank calculates the number of group accounts within each account's three-degree connections. If the number of group accounts within an account's three-degree connections exceeds a predetermined ratio, the subgraph is identified as a fraudulent group.
[0118] Implementation Results: By implementing this solution, a bank was able to identify and prevent potential group fraud, reduce false positives and missed reports, improve risk control identification efficiency, protect customer assets, enhance customer trust, meet regulatory requirements, and improve compliance.
[0119] In the above embodiments, although the steps are numbered S1, S2, etc., these are only specific embodiments given by the present invention. Those skilled in the art may adjust the execution order of S1, S2, etc. according to actual conditions, which is also within the scope of protection of the present invention. It can be understood that in some embodiments, some or all of the above embodiments may be included.
[0120] like Figure 2 As shown, a highly correlated group account identification system 200 based on a machine learning model according to an embodiment of the present invention includes a target transaction account acquisition module 201 and a judgment module 202;
[0121] The target transaction account acquisition module 201 is configured to use a trained machine learning model to calculate the degree of association between the transaction account suspected of group fraud and each transaction account associated with the transaction account suspected of group fraud, use a threshold for the degree of association, screen out target transaction accounts associated with the transaction account suspected of group fraud, and iterate using the target transaction account as the transaction account suspected of group fraud to obtain multiple target transaction accounts.
[0122] The judgment module 202 is used to determine whether there are any transaction accounts suspected of group fraud among the transaction accounts and all target accounts.
[0123] Optionally, the above technical solution further includes a connectivity graph construction module, which is used to: construct a connectivity graph according to all target transaction accounts screened out each time;
[0124] The judgment module 202 is specifically configured to judge each connectivity graph using a group association closeness rule to determine whether there is a group fraud transaction account among the suspected group fraud transaction accounts and all target accounts.
[0125] Optionally, the above technical solution further includes a transaction account screening module, which is used to: screen out transaction accounts suspected of group fraud from all transaction accounts through high-risk group account screening rules.
[0126] Optionally, the above technical solution further includes a model training module, which is used to:
[0127] Using known group fraud accounts and risk-free accounts as samples, the machine learning model is trained to obtain a trained machine learning model.
[0128] It should be noted that the beneficial effects of the high-correlation group account identification system 200 based on a machine learning model provided in the above embodiment are the same as the beneficial effects of the high-correlation group account identification method based on a machine learning model, which will not be repeated here. In addition, when implementing its functions, the system provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the system can be divided into different functional modules according to actual conditions to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0129] like Figure 3As shown, an electronic device 300 according to an embodiment of the present invention includes a processor 320 coupled to a memory 310. The memory 310 stores at least one computer program 330. The at least one computer program 330 is loaded and executed by the processor 320 to enable the electronic device 300 to implement any of the above-mentioned methods for identifying highly correlated group accounts based on a machine learning model. Specifically:
[0130] The electronic device 300 may have relatively large differences due to different configurations or performances, and may include one or more processors 320 (Central Processing Units, CPU) and one or more memories 310, wherein the one or more memories 310 store at least one computer program 330, and the at least one computer program 330 is loaded and executed by the one or more processors 320, so that the electronic device 300 implements any of the high-correlation group account identification methods based on machine learning models provided in the above embodiments. Of course, the electronic device 300 may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The electronic device 300 may also include other components for realizing device functions, which will not be elaborated here. The electronic device may specifically be a computer, etc.
[0131] A computer-readable storage medium according to an embodiment of the present invention stores at least one computer program, which is loaded and executed by a processor to enable a computer to implement any of the above-mentioned methods for identifying highly correlated group accounts based on machine learning models.
[0132] Alternatively, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, or the like.
[0133] In an exemplary embodiment, a computer program product or computer program is also provided, comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform any of the aforementioned methods for identifying highly associated group accounts based on a machine learning model.
[0134] It should be noted that the terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects and to define a specific order or precedence. Where appropriate, the order used for similar objects may be interchanged, such that the embodiments of the present application described herein can be implemented in an order other than the order shown or described.
[0135] Those skilled in the art will appreciate that the present invention may be implemented as a system, method, or computer program product. Therefore, the present invention may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, resident software, microcode, etc.), or in a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, the present invention may be implemented in the form of a computer program product embodied in one or more computer-readable media containing computer-readable program code.
[0136] Any combination of one or more computer-readable media can be used. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device.
[0137] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A method for identifying highly correlated group accounts based on a machine learning model, characterized in that: include: Using the trained machine learning model, calculate the degree of association between the transaction account suspected of group fraud and each transaction account associated with the transaction account suspected of group fraud. Using the association degree threshold, screen out target transaction accounts associated with the transaction account suspected of group fraud. Iterate using the target transaction account as the transaction account suspected of group fraud to obtain multiple target transaction accounts. When iterated twice, the following steps are included: S10. Using the association closeness threshold, screen out target transaction accounts associated with the transaction accounts suspected of group fraud, and record the target transaction accounts obtained in this step as first target transaction accounts; S11. First Iteration: Using the trained machine learning model, calculate the degree of association between each first target transaction account and each corresponding associated transaction account. Using the association degree threshold, select the target transaction accounts associated with each first target transaction account. The target transaction accounts associated with the first target transaction account obtained in this step are recorded as second target transaction accounts. S12, Second Iteration: Using the trained machine learning model, calculate the degree of association between each second target transaction account and each corresponding associated transaction account. Using the association degree threshold, select the target transaction accounts associated with each second target transaction account. The target transaction accounts associated with the second target transaction account obtained in this step are recorded as third target transaction accounts. After S10 to S12, the obtained multiple target transaction accounts include all first target transaction accounts, all second target transaction accounts, and all third target transaction accounts; The process of screening out target transaction accounts associated with suspected group fraud transaction accounts using the correlation closeness threshold is achieved through any of the following methods: 1) The first method: directly identify the transaction accounts with a correlation degree greater than the correlation degree threshold as target transaction accounts; 2) The second method: Calculate the degree of association between the transaction account suspected of group fraud and each transaction account associated with the transaction account suspected of group fraud, and sort all transaction accounts associated with the transaction account suspected of group fraud by degree of association to obtain a transaction account sequence. Filter out transaction accounts with a degree of association greater than a threshold value from the transaction account sequence, and identify the selected transaction accounts with a degree of association greater than the threshold value as target transaction accounts. Determining whether there are any group fraud transaction accounts among the suspected group fraud transaction accounts and all target accounts; Also includes: A connectivity graph is constructed based on all target transaction accounts screened out each time; The specific process of constructing a connectivity graph is as follows: Based on all the target trading accounts filtered out each time, perform intersection joins on the node table and the associated edge table, read the filtered node and edge data, and store them in the edge_data variable. Create an empty undirected graph G. Traverse the edge data and add each edge to the undirected graph G. Use the nx.connected_components method to obtain all connected components of the undirected graph G to obtain a connected graph. Determine whether there are any group fraud transaction accounts among the suspected group fraud transaction accounts and all target accounts, including: Using the group association closeness rule, each connectivity graph is judged to determine whether there is a group fraud transaction account among the suspected group fraud transaction account and all target accounts; The process of obtaining the transaction account suspected of group fraud includes: Using high-risk group account screening rules, we can identify suspected group fraud accounts from all trading accounts. By analyzing the distribution of scores output by the trained machine learning model and combining the precision and service disruption rate indicators, a threshold is determined as the association closeness threshold to distinguish target transaction accounts associated with transaction accounts suspected of gang fraud from target transaction accounts not associated with transaction accounts suspected of gang fraud. The precision rate is calculated as follows: Precision = TP / (TP+FP), where TP is the number of true positives, i.e., the number of actual positive samples correctly predicted as positive by the model, and FP is the number of false positives, i.e., the number of actual negative samples incorrectly predicted as positive by the model. The service disruption rate is: Disturbrate = FP / (TP+FP). Set the service disturbance rate (Disturbrate) to less than or equal to 0.
5. On this basis, when the precision of the trained machine learning model reaches the maximum, the corresponding output value is the association closeness threshold; Transaction accounts with a correlation degree greater than the correlation degree threshold are determined as target transaction accounts associated with the transaction accounts suspected of gang fraud, and transaction accounts with a correlation degree not greater than the correlation degree threshold are determined as target transaction accounts not associated with the transaction accounts suspected of gang fraud.
2. The method for identifying highly correlated group accounts based on a machine learning model according to claim 1, characterized in that: The process of obtaining the trained machine learning model includes: Using known group fraud accounts and risk-free accounts as samples, the machine learning model is trained to obtain a trained machine learning model.
3. A highly correlated group account identification system based on a machine learning model, characterized in that: Including target transaction account acquisition module and judgment module; The target transaction account acquisition module is configured to use a trained machine learning model to calculate the degree of association between a transaction account suspected of group fraud and each transaction account associated with the transaction account suspected of group fraud, and to screen target transaction accounts associated with the transaction account suspected of group fraud using a threshold value of the degree of association. The target transaction account is then used as the transaction account suspected of group fraud for iteration to obtain multiple target transaction accounts. When the iteration is repeated twice, the module includes the following steps: S10. Using the association closeness threshold, screen out target transaction accounts associated with the transaction accounts suspected of group fraud, and record the target transaction accounts obtained in this step as first target transaction accounts; S11. First Iteration: Using the trained machine learning model, calculate the degree of association between each first target transaction account and each corresponding associated transaction account. Using the association degree threshold, select the target transaction accounts associated with each first target transaction account. The target transaction accounts associated with the first target transaction account obtained in this step are recorded as second target transaction accounts. S12, Second Iteration: Using the trained machine learning model, calculate the degree of association between each second target transaction account and each corresponding associated transaction account. Using the association degree threshold, select the target transaction accounts associated with each second target transaction account. The target transaction accounts associated with the second target transaction account obtained in this step are recorded as third target transaction accounts. After S10 to S12, the obtained multiple target transaction accounts include all first target transaction accounts, all second target transaction accounts, and all third target transaction accounts; The process of screening out target transaction accounts associated with suspected group fraud transaction accounts using the correlation closeness threshold is achieved through any of the following methods: 1) The first method: directly identify the transaction accounts with a correlation degree greater than the correlation degree threshold as target transaction accounts; 2) The second method: Calculate the degree of association between the transaction account suspected of group fraud and each transaction account associated with the transaction account suspected of group fraud, and sort all transaction accounts associated with the transaction account suspected of group fraud by degree of association to obtain a transaction account sequence. Filter out transaction accounts with a degree of association greater than a threshold value from the transaction account sequence, and identify the selected transaction accounts with a degree of association greater than the threshold value as target transaction accounts. The judgment module is used to: determine whether there is a group fraud transaction account among the suspected group fraud transaction account and all target accounts; The system further includes a connectivity graph construction module, wherein the connectivity graph construction module is used to: construct a connectivity graph according to all target transaction accounts screened out each time; The specific process of constructing a connectivity graph is as follows: Based on all the target trading accounts filtered out each time, perform intersection joins on the node table and the associated edge table, read the filtered node and edge data, and store them in the edge_data variable. Create an empty undirected graph G. Traverse the edge data and add each edge to the undirected graph G. Use the nx.connected_components method to obtain all connected components of the undirected graph G to obtain a connected graph. The judgment module is specifically configured to: use a group association closeness rule to judge each connectivity graph to determine whether there is a group fraud transaction account among the suspected group fraud transaction account and all target accounts; The module also includes a transaction account screening module, which is used to screen out transaction accounts suspected of group fraud from all transaction accounts using high-risk group account screening rules; By analyzing the distribution of scores output by the trained machine learning model and combining the precision and service disruption rate indicators, a threshold is determined as the association closeness threshold to distinguish target transaction accounts associated with transaction accounts suspected of gang fraud from target transaction accounts not associated with transaction accounts suspected of gang fraud. The precision rate is calculated as follows: Precision = TP / (TP+FP), where TP is the number of true positives, i.e., the number of actual positive samples correctly predicted as positive by the model, and FP is the number of false positives, i.e., the number of actual negative samples incorrectly predicted as positive by the model. The service disruption rate is: Disturbrate = FP / (TP+FP). Set the service disturbance rate (Disturbrate) to less than or equal to 0.
5. On this basis, when the precision of the trained machine learning model reaches the maximum, the corresponding output value is the association closeness threshold; Transaction accounts with a correlation degree greater than the correlation degree threshold are determined as target transaction accounts associated with the transaction accounts suspected of gang fraud, and transaction accounts with a correlation degree not greater than the correlation degree threshold are determined as target transaction accounts not associated with the transaction accounts suspected of gang fraud.
4. The highly correlated group account identification system based on a machine learning model according to claim 3 is characterized in that: It also includes a model training module, which is used to: Using known group fraud accounts and risk-free accounts as samples, the machine learning model is trained to obtain a trained machine learning model.
5. An electronic device, characterized in that: The electronic device includes a processor, which is coupled to a memory. The memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor so that the electronic device implements a high-correlation group account identification method based on a machine learning model as described in any one of claims 1 to 2.
6. A computer-readable storage medium, characterized in that At least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by the processor to enable the computer to implement a high-correlation group account identification method based on a machine learning model as described in any one of claims 1 to 2.
Citation Information
Patent Citations
Fraud recognition method and fraud recognition device
CN107730262A
Method and device for determining abnormal interactive account
CN108295476A
Abnormal account determination method and device, and electronic equipment
CN113344621A