A clustering-based real-time fraud transaction detection method, device, equipment and medium

By performing offline clustering of large inflow and outflow transactions of historical accounts involved in cases, and combining Flink stream processing and transaction decision engine, the fraud risk score is calculated in real time. This solves the problem of high false positive and false negative rates in the detection of telecommunications fraud in bank transactions, and realizes real-time and accurate detection and risk assessment of fraudulent transactions.

CN121073483BActive Publication Date: 2026-05-15BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING UNIV OF POSTS & TELECOMM
Filing Date
2025-09-16
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing bank transaction fraud detection mechanisms have high false alarm or false negative rates in pre-event and in-event models, and lack sufficient analysis of common behavioral patterns among the groups of accounts involved, making it difficult to make accurate decisions in a very short time.

Method used

The DP-Means clustering algorithm is used to perform offline clustering of large inflow and outflow transactions of historical accounts involved in cases. Combined with the Flink stream processing framework and transaction decision engine, the fraud risk score of the transaction is calculated in real time. Risk decision is made by weighted fusion of pre-event and in-event risk scores.

Benefits of technology

It achieves real-time and accurate monitoring of telecom fraud transactions, reduces false alarm and false negative rates, provides better risk interpretation, and facilitates understanding and strategy adjustment by business personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121073483B_ABST
    Figure CN121073483B_ABST
Patent Text Reader

Abstract

The application discloses a real-time fraud transaction detection method and device based on clustering, equipment and medium, relates to the field of financial technology and anti-fraud technology, and the method comprises the following steps: obtaining a historical involved account list composed of accounts defined as fraud accounts; obtaining a large amount of incoming transaction risk cluster set and a large amount of outgoing transaction risk cluster set of the historical involved accounts based on large amount of incoming / outgoing transaction samples of the historical involved accounts; obtaining a first feature vector and a second feature vector of a target account; obtaining a pre-fraud risk score and an in-fraud risk score according to the first feature vector, the second feature vector, the large amount of incoming transaction risk cluster set and the large amount of outgoing transaction risk cluster set; comprehensively considering the pre-fraud risk score and the in-fraud risk score, calculating a total fraud risk probability of a current large amount of outgoing transaction, and making a risk decision according to the total fraud risk probability. The method has good real-time performance and better risk explanation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of financial technology and anti-fraud technology, and in particular to a cluster-based real-time telecom fraud transaction detection method, device, equipment and medium. Background Technology

[0002] In recent years, telecommunications fraud methods have been constantly evolving (such as identity forgery, AI simulation, and phishing links), with funds rapidly flowing to multi-level linked accounts. Banks need to use intelligent technologies to intercept fraudulent accounts and disrupt the funding chain. Current fraudulent accounts are characterized by high-frequency, small-amount transfers, rapid inflows and outflows, and abnormal cross-border transactions. They also disperse funds through interbank transfers and cryptocurrency transactions to evade detection. Therefore, banks need to build a multi-dimensional risk feature database, integrating account behavior, device fingerprints, and related network data. They should leverage streaming computing engines to achieve millisecond-level risk assessment, and rely on machine learning models (such as logistic regression and GBDT algorithms) and real-time monitoring systems to continuously optimize risk control models, strengthen in-process interception capabilities, and build a full-lifecycle intelligent risk control system to combat new types of fraud.

[0003] Existing telecom fraud detection mechanisms based on bank transactions are divided into three types: pre-event, during-event, and post-event. Among them, pre-event models rely on multi-dimensional data, but the data sources are scattered and may contain noise or be missing. At the same time, because decisions need to be made in a very short time, the model may have a high false positive rate or false negative rate due to unreasonable threshold settings or unbalanced sample distribution. Moreover, existing rule-based strategies, machine learning models, and graph association models are insufficient in mining the commonalities of pre-event and during-event transaction behaviors of the account groups involved in the case. Summary of the Invention

[0004] This application provides a clustering-based real-time telecom fraud transaction detection method, device, equipment, and medium. It can calculate and fuse the similarity of risk clusters from pre-event clustering and in-event clustering results for online real-time large-amount transfer transactions, and then output the telecom fraud risk score of the current large-amount transfer transaction and make telecom fraud risk decisions. This makes the telecom fraud transaction monitoring more real-time, reduces the false alarm rate or false negative rate, and has better risk interpretability.

[0005] To achieve the above objectives, this application adopts the following technical solution:

[0006] Firstly, this application provides a clustering-based real-time telecom fraud transaction detection method, which includes:

[0007] Obtain a list of historical accounts involved in fraud cases, defined as such. Based on large-amount inflow transactions of these historical accounts, perform offline pre-event behavioral clustering using the DP-Means clustering algorithm to obtain a risk cluster set for large-amount inflow transactions of these historical accounts. Based on large-amount outflow transactions of these historical accounts, perform offline in-event behavioral clustering using the DP-Means clustering algorithm to obtain a risk cluster set for large-amount outflow transactions of these historical accounts. Obtain the first feature vector of a large-amount inflow transaction associated with a large-amount outflow transaction of the target account, calculate the maximum similarity between the first feature vector and the risk cluster set for large-amount inflow transactions, and obtain a pre-event fraud risk score. Obtain the second feature vector of a large-amount outflow transaction of the target account, calculate the maximum similarity between the second feature vector and the risk cluster set for large-amount outflow transactions, and obtain an in-event fraud risk score. Combine the pre-event fraud risk score and the in-event fraud risk score to calculate the total fraud risk probability of a large-amount outflow transaction of the target account, and make a risk decision based on the total fraud risk probability.

[0008] In some possible implementations, obtaining the first feature vector of a large outflow transaction associated with a large inflow transaction of the target account, calculating the maximum similarity between the first feature vector and the risk cluster set of large inflow transactions, and obtaining a pre-emptive telecom fraud risk score includes:

[0009] Obtain the first feature vector of a large outflow transaction associated with a large inflow transaction of the target account, calculate the maximum similarity between the first feature vector and the risk cluster set of large inflow transactions, determine whether the large outflow transaction of the target account corresponds to multiple large inflow transactions, and obtain a first judgment result; if the first judgment result indicates that the large outflow transaction of the target account corresponds to multiple large inflow transactions, then determine the maximum similarity between the first feature vector and the risk cluster set of large inflow transactions as the pre-existing telecom fraud risk score of the large outflow transaction of the target account.

[0010] In some possible implementations, the method further includes:

[0011] If the first judgment result indicates that there is no corresponding large-amount transfer-in transaction for the target account during the transaction, then the pre-transfer fraud risk score of the target account for the large-amount transfer-out transaction during the transaction is determined to be zero.

[0012] In some possible implementations, the method further includes:

[0013] A risk decision probability threshold is set. When the total probability of telecom fraud is greater than or equal to the risk decision probability threshold, a large outflow transaction of the target account is determined to be at risk of telecom fraud and interception is triggered.

[0014] In some possible implementations, the method further includes:

[0015] The pre-fraud risk score and the in-process fraud risk score are weighted according to a preset weight ratio to obtain the total fraud risk probability; wherein the sum of the weight ratios is 1.

[0016] In some possible implementations, the method further includes:

[0017] The construction of the first feature vector includes at least one of the following features: number of days since account opening, limit query operation, limit increase behavior, small trial transaction, transaction channel, whether the transaction amount is an integer multiple, number of counterparty accounts, number of transactions, transaction amount surge multiple, transactions during sensitive time periods, and fund transition indicators.

[0018] In some possible implementations, the DP-Means clustering algorithm includes:

[0019] Randomly select a transaction data point as the centroid of the first cluster; traverse each data point to be processed, calculate the distance between the data point to be processed and the centroids of all existing clusters, if there is a cluster with a distance less than or equal to a preset dynamic threshold, then assign the data point to be processed to the nearest cluster, otherwise generate a new cluster and use the data point as the centroid of the new cluster; recalculate the centroids of all data points within all clusters, and use them as the new centroids of each cluster; repeat the assignment until all data points have been processed.

[0020] Secondly, this application provides a clustering-based real-time telecom fraud transaction detection device, which includes:

[0021] The acquisition module is used to obtain a list of historical accounts involved in cases that have been defined as fraudulent accounts;

[0022] The clustering module is used to perform offline pre-event behavioral clustering based on large-amount inflow transactions of historical accounts involved in cases, using the DP-Means clustering algorithm to obtain a set of risk clusters for large-amount inflow transactions of historical accounts involved in cases; and to perform offline in-event behavioral clustering based on large-amount outflow transactions of historical accounts involved in cases, using the DP-Means clustering algorithm to obtain a set of risk clusters for large-amount outflow transactions of historical accounts involved in cases.

[0023] The calculation module is used to obtain the first feature vector of a large transfer-in transaction associated with a large transfer-out transaction of the target account, calculate the maximum similarity between the first feature vector and the risk cluster set of large transfer-in transactions, and obtain a pre-emptive telecom fraud risk score; and to obtain the second feature vector of a large transfer-out transaction of the target account, calculate the maximum similarity between the second feature vector and the risk cluster set of large transfer-out transactions, and obtain a real-time telecom fraud risk score.

[0024] The decision-making module is used to combine the pre-fake fraud risk score and the in-fake fraud risk score to calculate the total fraud risk probability of a large transfer transaction of the target account, and to make a risk decision based on the total fraud risk probability.

[0025] Thirdly, this application provides a computing device, including a memory and a processor;

[0026] The memory stores one or more computer programs, the one or more computer programs including instructions; when the instructions are executed by the processor, the computing device performs the method as described in any one of the first aspects.

[0027] Fourthly, this application provides a computer-readable storage medium for storing a computer program for performing the method as described in any one of the first aspects.

[0028] Fifthly, this application provides a computer program product comprising one or more computer instructions, wherein when the computer instructions are executed by a computer, the computer performs the method as described in any one of the first aspects.

[0029] As can be seen from the above technical solution, this application has at least the following beneficial effects:

[0030] In this application, a list of historical accounts involved in fraud cases, defined as fraudulent accounts, is obtained. Based on large-amount inflow transaction examples of these historical accounts, offline pre-event behavioral clustering is performed using the DP-Means clustering algorithm to obtain a risk cluster set of large-amount inflow transactions of these historical accounts. Based on large-amount outflow transaction examples of these historical accounts, offline in-event behavioral clustering is performed using the DP-Means clustering algorithm to obtain a risk cluster set of large-amount outflow transactions of these historical accounts. Then, the first feature vector of a large-amount inflow transaction associated with a large-amount outflow transaction of the target account is obtained, and the maximum similarity between the first feature vector and the risk cluster set of large-amount inflow transactions is calculated to obtain a pre-event fraud risk score. The second feature vector of a large-amount outflow transaction of the target account is obtained, and the maximum similarity between the second feature vector and the risk cluster set of large-amount outflow transactions is calculated to obtain an in-event fraud risk score. Combining the pre-event fraud risk score and the in-event fraud risk score, the total fraud risk probability of a large-amount outflow transaction of the target account is calculated, and a risk decision is made based on the total fraud risk probability.

[0031] Existing technologies utilize traditional clustering algorithms (such as K-Means and DBSCAN) for unsupervised analysis of bank transaction data to assist in identifying telecom fraud risks. However, K-Means requires a preset number of clusters, and DBSCAN is sensitive to parameters. Furthermore, pre-detection relies on multi-dimensional and scattered data, which is prone to high false positive and false negative rates due to thresholds, while in-process models are insufficient in mining the commonalities of the behavior of involved accounts. Therefore, this application performs offline clustering on large-amount inflow (pre-detection) and large-amount outflow (in-process) transactions of historical involved accounts, and integrates the similarity between the current transaction and the two risk clusters during real-time detection, covering the entire process of telecom fraud funds from "inflow to outflow" and avoiding the limitations of single-dimensional detection. It also uses the Flink stream processing framework or transaction decision engine to achieve millisecond-level risk assessment to meet the needs of real-time interception. Risk scores are generated by cluster similarity matching. Compared with black-box models, the risk assessment logic is more transparent, making it easier for business personnel to understand and adjust strategies.

[0032] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description

[0033] Figure 1 A flowchart illustrating a clustering-based real-time telecom fraud transaction detection method provided in this application embodiment;

[0034] Figure 2 A schematic diagram of a cluster-based real-time telecom fraud transaction detection device provided in an embodiment of this application;

[0035] Figure 3 This is a schematic diagram of a computing device provided in an embodiment of this application. Detailed Implementation

[0036] The terms "first," "second," and "third," etc., used in this application specification and accompanying drawings are used to distinguish different objects, not to limit a specific order.

[0037] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0038] To ensure clarity and conciseness in the description of the following embodiments, a brief introduction to the related technologies is given first:

[0039] Clustering refers to the process of dividing a set of physical or abstract objects into multiple clusters composed of similar objects. It belongs to unsupervised learning methods. Its core is to maximize the similarity of objects within the same cluster and minimize the similarity between different clusters, without the need to predefine classification criteria.

[0040] Telecommunications fraud refers to criminal acts committed with the intent to illegally possess property, using non-contact remote means such as telephone, text message, and the internet to fabricate false information or forge legitimate identities (such as public security, procuratorate, court, bank, e-commerce customer service, etc.) to induce victims to transfer money or disclose sensitive information. Technically, it utilizes modern communication technologies and online platforms to carry out the fraud. Common forms include: impersonating public security, procuratorate, and court officials; fraudulent online shopping scams offering rebates; fake investment and wealth management products; and online loan fraud.

[0041] Flink is an open-source distributed stream processing framework that achieves precise real-time feature processing through state management and event-time processing. For real-time trading scenarios, Flink can process tens of thousands of transactions per second in the following ways: statistically analyze transaction frequency or amount distribution using time windows (such as a 5-second sliding window); track continuous user trading behavior (such as multiple abnormal operations within 10 minutes) and trigger risk control rules.

[0042] The Transaction Decision Engine is a real-time automated decision-making system designed specifically for financial trading scenarios. By integrating a rule engine with real-time computing capabilities, it performs instant feature processing and strategy execution on massive amounts of transaction data. Its core functions include: dynamically calculating features such as transaction frequency and amount distribution based on time windows (e.g., sliding windows, session windows), and integrating batch and streaming data to achieve millisecond-level response; supporting rule sets, decision trees, and other structures, and triggering anti-fraud rules by combining user profiles, device fingerprints, and other data.

[0043] Currently, traditional solutions use conventional clustering algorithms to perform unsupervised analysis on bank transaction data, mining the inherent distribution structure of the data and classifying similar transaction behaviors into potential categories to help identify the risk of telecom fraud. However, traditional clustering algorithms have significant drawbacks: the K-Means algorithm requires a pre-set number of clusters, making it difficult to adapt to the dynamic distribution of real data; the DBSCAN algorithm is sensitive to parameters, making parameter tuning difficult and affecting the stability of clustering results.

[0044] Furthermore, existing pre-emptive telecom fraud detection relies on multi-dimensional data (such as transaction amounts, user behavior, and device fingerprints), but the data sources are scattered and may contain noise or be incomplete. Real-time pre-emptive detection needs to make decisions within a very short time, but the model may suffer from high false positive or false negative rates due to unreasonable threshold settings or uneven sample distribution. In addition, current in-process telecom fraud detection models, which employ existing rule-based strategies, machine learning models, and graph association models, lack sufficient ability to identify commonalities in the pre-emptive and in-process transaction behaviors of the affected account groups.

[0045] In view of this, embodiments of this application provide a clustering-based real-time telecom fraud transaction detection method. This method involves obtaining a list of historical accounts involved in fraudulent transactions, defined as such; performing offline pre-event behavioral clustering using the DP-Means clustering algorithm based on large-amount inflow transaction examples from these historical accounts to obtain a set of risk clusters for large-amount inflow transactions; and performing offline in-event behavioral clustering using the DP-Means clustering algorithm based on large-amount outflow transaction examples from these historical accounts to obtain a set of risk clusters for large-amount outflow transactions. Then… Obtain the first feature vector of a large outflow transaction associated with a large inflow transaction in the target account. Calculate the maximum similarity between the first feature vector and the risk cluster set of large inflow transactions to obtain a pre-emptive fraud risk score. Obtain the second feature vector of a large outflow transaction in the target account. Calculate the maximum similarity between the second feature vector and the risk cluster set of large outflow transactions to obtain an in-process fraud risk score. Combine the pre-emptive and in-process fraud risk scores to calculate the total fraud risk probability of a large outflow transaction in the target account. Make a risk decision based on the total fraud risk probability. Therefore, this application performs offline clustering of large-amount inflow (pre-event) and large-amount outflow (in-event) transactions of historical accounts involved in cases. During real-time detection, it integrates the similarity between the current transaction and the two risk clusters, covering the entire process of "transfer in to out" of funds involved in telecom fraud, avoiding the limitations of single-dimensional detection. It also uses the Flink stream processing framework or transaction decision engine to achieve millisecond-level risk assessment, meeting the needs of real-time interception. By generating risk scores through cluster similarity matching, the risk assessment logic is more transparent than the black-box model, making it easier for business personnel to understand and adjust strategies.

[0046] To make the technical solution of this application clearer and easier to understand, a clustering-based real-time telecom fraud transaction detection method provided by an embodiment of this application is described below with reference to the accompanying drawings. This method is applied to processing devices, such as… Figure 1 As shown, this figure is a flowchart of a cluster-based real-time telecom fraud transaction detection method provided in an embodiment of this application. The cluster-based real-time telecom fraud transaction detection method includes:

[0047] S101. Obtain a list of historical accounts involved in a case that have been defined as fraudulent accounts.

[0048] Accounts involved in cases refer to bank accounts used by telecommunications fraudsters to receive and transfer funds. These accounts are investigated and included in the financial institutions' collaborative reporting mechanism for control. The historical list of accounts involved in cases is reported after investigation and falls under the control of this mechanism, possessing legal validity and authority. These accounts are explicitly used for receiving and transferring funds related to telecommunications fraud and form the core sample basis for detecting such transactions. Typically, historical accounts involved in cases within the past N days (e.g., N=180 days) are obtained to ensure the sample covers recent characteristics of telecommunications fraud activities. Key fields for accounts involved in cases include the issuance date, card number, and account status, stored in an account-related card table. The bank system connects to the system via a dedicated line or secure interface, periodically (e.g., daily at midnight) synchronizing the latest list of accounts involved in cases to ensure the timeliness of data used for offline clustering and real-time detection. For example, the latest issued account-related file is automatically retrieved after midnight each day, overwriting historical records with updated status.

[0049] S102. Based on examples of large-amount inflow transactions of historical accounts involved in cases, offline pre-event behavior clustering is performed using the DP-Means clustering algorithm to obtain a set of risk clusters for large-amount inflow transactions of historical accounts involved in cases.

[0050] In this embodiment, a risk sample is constructed for each large, unfamiliar inflow transaction (defined as 5000 yuan or more) within 30 days of the issuance date for the account in question. The set of counterparties for each account's inflow and outflow transactions within the previous N days (N=3) and M days (M=365) is compiled (i.e., all accounts that have had financial transactions). If the counterparty for the current inflow transaction is not in this set, the current inflow transaction is defined as an unfamiliar inflow. For each risk sample, a feature vector is calculated based on the feature design described in the next section. This risk sample and feature vector are used for unsupervised DP-Means clustering. Examples of feature vector design are shown in Table 1.

[0051] Table 1: Example of feature vector design for risk samples constructed based on large-amount transfer transactions.

[0052]

[0053] A current large inflow transaction refers to a large inflow transaction into the target account; a current large outflow transaction refers to a large outflow transaction from the target account; and a day transaction refers to a transaction that occurs on the day a large inflow transaction into the target account occurs.

[0054] Offline clustering feature processing relies on Spark or Hadoop big data platforms and uses HiveQL for distributed processing.

[0055] DP-Means clustering is an improved clustering algorithm based on nonparametric Bayesian principles. Its core objective is to overcome the limitation of the traditional K-Means algorithm, which requires pre-specifying the number of clusters, by dynamically adjusting the number of clusters (K value). Core principle and algorithm idea: Dynamic cluster number adjustment: DP-Means introduces a nonparametric model of the Dirichlet Process, treating the number of clusters as a random variable. When the distance between a data point and the existing cluster center exceeds a threshold λ (λ=0.95 in this invention), a new cluster is automatically generated without manually pre-setting the K value. The role of the threshold parameter λ: A smaller λ leads to more small clusters, suitable for fine-grained clustering; a larger λ tends to merge similar clusters, reducing complexity.

[0056] The DP-Means clustering algorithm includes the following steps: randomly selecting a data point as the centroid of the first cluster; traversing each data point to be processed, calculating the distance between the data point to be processed and the centroids of all existing clusters; if there is a cluster with a distance less than or equal to a preset dynamic threshold, then the data point to be processed is assigned to the nearest cluster; otherwise, a new cluster is generated and the data point is used as the centroid of the new cluster; recalculating the centroids of all data points within each cluster, and using them as the new centroids of each cluster; repeating the assignment until all data points have been processed.

[0057] Based on the list of historical accounts involved in the case obtained from S101, large-amount transfer transaction examples corresponding to the historical accounts involved in the case are obtained from the information of the historical accounts involved in the case. Offline pre-event behavior clustering is performed by the DP-Means clustering algorithm to obtain a set of risk clusters of large-amount transfer transactions of historical accounts involved in the case. The data format of each risk cluster is: {risk cluster ID, number of risk members, {feature 01 value, feature 02 value, ..., feature K value}}.

[0058] In this step, on the one hand, the DP-Means clustering algorithm eliminates the need for pre-setting the number of clusters, dynamically adjusting based on the data distribution. This solves the problem of manually determining the number of clusters required by the traditional K-Means algorithm, making it more suitable for the complex and variable nature of telecom fraud transaction data and improving the adaptability and accuracy of clustering. On the other hand, clustering large-amount inflow transactions of historically involved accounts can reveal common behavioral patterns of telecom fraud accounts during the fund inflow stage, such as specific transaction channels and amount characteristics, providing strong feature support for pre-emptive risk assessment. Moreover, the highly similar transaction behaviors within each risk cluster represent typical telecom fraud inflow risk patterns, facilitating rapid matching and identification of similar risky transactions during subsequent real-time detection, thus improving detection efficiency and accuracy.

[0059] S103. Based on examples of large-amount outflow transactions of historical accounts involved in cases, offline in-process behavior clustering is performed using the DP-Means clustering algorithm to obtain a risk cluster set of large-amount outflow transactions of historical accounts involved in cases.

[0060] Large-amount outflow transactions include withdrawals via quick payment, transfer, QR code, digital RMB, and cash withdrawal. In this embodiment, for each large-amount outflow transaction (defined as ≥2000 RMB, <2000 RMB, and requiring the counterparty to be different) after a large-amount inflow to the account in question within 30 days of the issuance date, a risk sample is constructed. For each risk sample, a feature vector is obtained, and examples of the feature vector design are shown in Table 2.

[0061] Table 2: Example of feature vector design for risk samples constructed based on large-amount transfer transactions.

[0062]

[0063] Based on the list of historical accounts involved in the case obtained from S101, large-amount transfer transaction examples corresponding to the historical accounts involved in the case are obtained from the information of the historical accounts involved in the case. Offline in-process behavior clustering is performed by the DP-Means clustering algorithm to obtain a set of risk clusters of large-amount transfer transactions of historical accounts involved in the case. The data format of each risk cluster is: {risk cluster ID, number of risk members, {feature 01 value, feature 02 value, ..., feature K value}}.

[0064] This step, combined with pre-emptive clustering, provides comprehensive coverage of the entire "transfer-in" and "transfer-out" process of fraudulent funds, overcoming the shortcomings of existing technologies in identifying commonalities in real-time transaction behavior and making risk detection more comprehensive. By clustering large-amount outflow transactions, typical behaviors of fraudulent accounts during fund transfers can be identified, such as high-frequency small-amount transfers, rapid in-and-out transactions, and transactions during sensitive time periods, providing clear risk patterns for real-time interception. Leveraging the advantages of the DP-Means algorithm, it eliminates the need for pre-setting the number of clusters, adapting to the constantly changing fund transfer methods of fraudsters and ensuring that the clustering results always reflect the latest risk characteristics.

[0065] S104. Obtain the first feature vector of a large transfer-in transaction associated with a large transfer-out transaction of the target account, calculate the maximum similarity between the first feature vector and the risk cluster set of large transfer-in transactions, and obtain the pre-emptive telecom fraud risk score.

[0066] Using the Flink stream processing framework or transaction decision engine, calculate the first feature vector of a large transfer-in transaction associated with a large transfer-out transaction of the target account, calculate the maximum similarity between the first feature vector and the risk cluster set of large transfer-in transactions, determine whether the large transfer-out transaction of the target account corresponds to multiple large transfer-in transactions, and obtain the first judgment result.

[0067] The formula for calculating the similarity between the first feature vector and the risk cluster set of large-amount inflow transactions is as follows:

[0068]

[0069] in, Let A be the similarity between the first feature vector and the set of risk clusters for large-amount transfer transactions; A is the first feature vector, B is a risk cluster for large-amount transfer transactions in the set of risk clusters for large-amount transfer transactions, and n is the dimension of the first feature vector. .

[0070] If the first judgment result indicates that the target account's large outflow transaction corresponds to multiple large inflow transactions, then the maximum similarity between the first feature vector and the risk cluster set of large inflow transactions is determined as the pre-emptive telecom fraud risk score of the target account's large outflow transaction.

[0071] If the first judgment result indicates that there is no corresponding large-amount transfer-in transaction for the target account during the transaction, then the pre-transfer fraud risk score for the target account's large-amount transfer-out transaction during the transaction is determined to be zero.

[0072] The maximum similarity is calculated by taking the pre-transfer feature vector of a large outflow transaction of the target account, each large inflow transaction in the past 24 hours, and each risk cluster in the pre-clustering result cluster set loaded into memory, and then taking the maximum similarity value.

[0073] The construction of the first feature vector includes at least one of the following features: number of days since account opening, limit query operation, limit increase behavior, small trial transaction, transaction channel, whether the transaction amount is an integer multiple, number of counterparty accounts, number of transactions, transaction amount surge multiple, transactions during sensitive time periods, and capital transition indicators. The above features are shown in Table 1.

[0074] By analyzing the characteristics of inbound transactions associated with large outflows from target accounts, risks can be assessed at the source of fund inflows. If an inbound transaction is highly similar to a historical cluster of inbound transactions related to telecom fraud, it indicates that the outflow transaction may involve telecom fraud, improving the foresight of risk assessment. When a target outflow transaction corresponds to multiple inbound transactions, the highest similarity is used as the pre-emptive risk score. This allows for a comprehensive consideration of the risk situation of multiple inbound transactions, avoiding misjudgments based on a single inbound transaction, and making risk assessment more comprehensive and accurate. If there are no associated large inbound transactions, the pre-emptive risk score is set to 0, avoiding unfounded risk misjudgments, improving the reliability of detection, and reducing the false positive rate.

[0075] S105. Obtain the second feature vector of a large outflow transaction of the target account, calculate the maximum similarity between the second feature vector and the risk cluster set of large outflow transactions, and obtain the in-process telecom fraud risk score.

[0076] Using the Flink stream processing framework or transaction decision engine, the second feature vector of large outflow transactions of the target account is calculated. The maximum similarity between the second feature vector and the risk cluster set of large outflow transactions is calculated. The maximum similarity between the second feature vector and the risk cluster set of large outflow transactions is determined as the real-time fraud risk score of the target account's large outflow transactions.

[0077] The maximum similarity is calculated by combining the feature vector of a large outflow transaction of the target account with each risk cluster in the in-process clustering result set that has been loaded into memory, resulting in multiple similarities. The maximum similarity value is then selected based on these multiple similarities.

[0078] Analyzing the characteristics of the target outgoing transaction itself and matching it with historical real-time risk clusters allows for direct identification of whether the transaction matches the outgoing behavior pattern of a fraudulent account, providing direct evidence for real-time interception. The second feature vector encompasses various risk characteristics of the outgoing transaction, such as transaction channel, type, amount, and fund transition. Through similarity calculation with risk clusters, it integrates risk information from multiple dimensions, improving the accuracy and comprehensiveness of risk assessment. Furthermore, leveraging the Flink stream processing framework or transaction decision engine, it can calculate the real-time risk score for each large outgoing transaction, meeting the time requirements for real-time detection and interception of fraudulent transactions, ensuring risk detection and intervention at the first moment of fund transfer.

[0079] S106. Combine the pre-event fraud risk score and the in-event fraud risk score to calculate the total fraud risk probability of a large outflow transaction from the target account, and make risk decisions based on the total fraud risk probability.

[0080] The pre-fraud risk score and the in-flight fraud risk score are weighted according to preset weight ratios to obtain the total fraud risk probability, where the sum of the weight ratios is 1. The calculation formula is as follows:

[0081] The expression A + B = 1.0 is satisfied. Wherein, This represents the total probability of telecommunications fraud. As a pre-emptive score for the risk of telecom fraud, The system assigns a risk score to the transaction in question related to telecom fraud. A risk decision probability threshold is set; when the total telecom fraud risk probability is greater than or equal to the threshold, a large outflow transaction from the target account is deemed to pose a telecom fraud risk and is blocked.

[0082] Specifically, in practice, two risk importance parameters are selected based on business needs, generally A=0.6 and B=0.4. If the total probability of telecom fraud risk is greater than or equal to the risk decision probability threshold (e.g., 0.8), then a large outflow transaction from the target account is considered to pose a telecom fraud risk.

[0083] Firstly, it weights and synthesizes risk scores from both the pre- and during-event stages, fully considering the risks associated with both fund inflows and outflows, avoiding the limitations of single-dimensional detection and making risk assessment more comprehensive and accurate. Secondly, by pre-setting weight ratios (e.g., A=0.6, B=0.4), the importance of pre- and during-event risks can be flexibly adjusted according to business needs, allowing the detection model to better adapt to different business scenarios and risk preferences. Thirdly, it sets a risk decision probability threshold; when the total probability of telecom fraud exceeds the threshold, it triggers interception, accurately identifying and promptly blocking high-risk transactions, effectively preventing the transfer of telecom fraud funds and protecting user funds. Furthermore, compared to traditional black-box models, generating risk scores through cluster similarity matching makes the risk assessment logic more transparent, facilitating business personnel's understanding and analysis of risk causes, and providing strong support for subsequent strategy adjustments and model optimization.

[0084] Based on the above, this application performs offline clustering of large-amount inflow (pre-event) and large-amount outflow (in-event) transactions of historical accounts involved in cases. During real-time detection, it integrates the similarity between the current transaction and the two risk clusters, covering the entire process of "transfer in to out" of funds involved in telecom fraud, avoiding the limitations of single-dimensional detection. It also uses the Flink stream processing framework or transaction decision engine to achieve millisecond-level risk assessment, meeting the needs of real-time interception. By generating risk scores through cluster similarity matching, the risk assessment logic is more transparent than the black-box model, making it easier for business personnel to understand and adjust strategies.

[0085] The above text combined Figure 1 The clustering-based real-time telecom fraud transaction detection method provided in this application embodiment has been described in detail. The apparatus and equipment provided in this application embodiment will be described below with reference to the accompanying drawings.

[0086] This application also provides a clustering-based real-time telecom fraud transaction detection device, such as... Figure 2 As shown in the figure, this figure is a schematic diagram of a cluster-based real-time telecom fraud transaction detection device provided in an embodiment of this application. The device includes: an acquisition module 201, a clustering module 202, a calculation module 203, and a decision module 204.

[0087] Module 201 is used to obtain a list of historical accounts involved in a case that have been defined as fraudulent accounts;

[0088] Clustering module 202 is used to perform offline pre-event behavioral clustering based on large-amount inflow transaction samples of historical accounts involved in cases, using the DP-Means clustering algorithm to obtain a set of risk clusters for large-amount inflow transactions of historical accounts involved in cases; and to perform offline in-event behavioral clustering based on large-amount outflow transaction samples of historical accounts involved in cases, using the DP-Means clustering algorithm to obtain a set of risk clusters for large-amount outflow transactions of historical accounts involved in cases.

[0089] The calculation module 203 is used to obtain the first feature vector of a large transfer-in transaction associated with a large transfer-out transaction of the target account, calculate the maximum similarity between the first feature vector and the risk cluster set of large transfer-in transactions, and obtain a pre-emptive telecom fraud risk score; obtain the second feature vector of a large transfer-out transaction of the target account, calculate the maximum similarity between the second feature vector and the risk cluster set of large transfer-out transactions, and obtain a real-time telecom fraud risk score.

[0090] The decision module 204 is used to combine the pre-fake fraud risk score and the in-fake fraud risk score to calculate the total fraud risk probability of a large transfer transaction of the target account, and to make a risk decision based on the total fraud risk probability.

[0091] In some possible implementations, the calculation module 203 is specifically used to obtain the first feature vector of a large transfer-in transaction associated with a large transfer-out transaction of the target account, calculate the maximum similarity between the first feature vector and the risk cluster set of large transfer-in transactions, determine whether the large transfer-out transaction of the target account corresponds to multiple large transfer-in transactions, and obtain a first judgment result; if the first judgment result indicates that the large transfer-out transaction of the target account corresponds to multiple large transfer-in transactions, then the maximum similarity between the first feature vector and the risk cluster set of large transfer-in transactions is determined as the pre-emptive fraud risk score of the large transfer-out transaction of the target account.

[0092] In some possible implementations, the calculation module 203 is further configured to determine that the pre-fake fraud risk score of the large-amount transfer-out transaction of the target account is zero if the first judgment result indicates that there is no corresponding large-amount transfer-in transaction for the large-amount transfer-out transaction of the target account.

[0093] In some possible implementations, the cluster-based real-time telecom fraud transaction detection device further includes an interception module, specifically used to set a risk decision probability threshold. When the total telecom fraud risk probability is greater than or equal to the risk decision probability threshold, the module determines that a large outflow transaction of the target account has a telecom fraud risk and triggers interception.

[0094] In some possible implementations, the decision module 204 is specifically used to perform a weighted calculation on the pre-fake fraud risk score and the in-fake fraud risk score according to a preset weight ratio to obtain the total fraud risk probability; wherein the sum of the weight ratios is 1.

[0095] In some possible implementations, the construction of the first feature vector includes at least one of the following features: number of days since account opening, limit query operation, limit increase behavior, small trial transaction, transaction channel, whether the transaction amount is an integer multiple, number of counterparty accounts, number of transactions, transaction amount surge multiple, transactions during sensitive time periods, and fund transition indicators.

[0096] In some possible implementations, the DP-Means clustering algorithm includes: randomly selecting a transaction data point as the centroid of the first cluster; traversing each data point to be processed, calculating the distance between the data point to be processed and the centroids of all existing clusters, and if there is a cluster with a distance less than or equal to a preset dynamic threshold, then assigning the data point to be processed to the nearest cluster; otherwise, generating a new cluster and using the data point as the centroid of the new cluster; recalculating the centroids of all data points within all clusters as the new centroids of each cluster; repeating the assignment until all data points have been processed.

[0097] The cluster-based real-time telecom fraud transaction detection device according to the embodiments of this application can correspondingly execute the method described in the embodiments of this application, and the other operations and / or functions of each module / unit of the cluster-based real-time telecom fraud transaction detection device are respectively for implementing Figure 1 For the sake of brevity, the corresponding processes of each method in the illustrated embodiments will not be described in detail here.

[0098] This application also provides a computing device. For example... Figure 3 As shown in the figure, this is a schematic diagram of a computing device provided in an embodiment of this application. The computing device 400 includes a bus 401, a processor 402, a communication interface 403, and a memory 404. The processor 402, the memory 404, and the communication interface 403 communicate with each other via the bus 401.

[0099] Bus 401 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 3 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0100] Processor 402 can be any one or more of the following processors: central processing unit (CPU), graphics processing unit (GPU), microprocessor (MP), or digital signal processor (DSP).

[0101] Communication interface 403 is used for communication with external devices.

[0102] Memory 404 may include volatile memory, such as random access memory (RAM). Memory 404 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0103] The memory 404 stores executable code, and the processor 402 executes the executable code to perform the aforementioned cluster-based real-time telecom fraud transaction detection method.

[0104] Specifically, in achieving Figure 2 In the case of the illustrated embodiment, and Figure 2 When the modules or units of the cluster-based real-time telecom fraud transaction detection device described in the embodiments are implemented in software, the execution... Figure 2 The software or program code required for the functions of each module / unit can be partially or wholly stored in memory 404. Processor 402 executes the program code corresponding to each unit stored in memory 404, and executes the aforementioned cluster-based real-time telecom fraud transaction detection method.

[0105] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-described cluster-based real-time telecom fraud transaction detection method.

[0106] This application also provides a computer program product comprising one or more computer instructions. When the computer instructions are loaded and executed on a computing device, all or part of the processes or functions described in this application are generated.

[0107] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center via wired (e.g., coaxial cable, fiber optic) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0108] When the computer program product is executed by a computer, the computer executes any of the methods described in the aforementioned cluster-based real-time telecom fraud transaction detection method. The computer program product can be a software installation package; when any of the aforementioned cluster-based real-time telecom fraud transaction detection methods needs to be used, the computer program product can be downloaded and executed on the computer.

[0109] The descriptions of the processes or structures corresponding to the above figures each have their own emphasis. For parts of a process or structure that are not described in detail, please refer to the relevant descriptions of other processes or structures.

[0110] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be covered within the scope of protection of this application.

Claims

1. A cluster-based real-time telecom fraud transaction detection method, characterized in that, The method includes: Obtain a list of historical accounts involved in cases that have been defined as fraudulent accounts; Based on examples of large-amount inflow transactions of historical accounts involved in cases, offline pre-event behavioral clustering is performed using the DP-Means clustering algorithm to obtain a set of risk clusters for large-amount inflow transactions of historical accounts involved in cases. Based on examples of large outflow transactions from historical accounts involved in cases, offline in-process behavior clustering is performed using the DP-Means clustering algorithm to obtain a set of risk clusters for large outflow transactions from historical accounts involved in cases. Obtain the first feature vector of a large outflow transaction associated with a large inflow transaction of the target account, calculate the maximum similarity between the first feature vector and the risk cluster set of large inflow transactions, and obtain the pre-emptive telecom fraud risk score. Obtain the second feature vector of a large outflow transaction of the target account, calculate the maximum similarity between the second feature vector and the risk cluster set of the large outflow transaction, and obtain the in-process telecom fraud risk score; By combining the pre-fake fraud risk score and the in-fake fraud risk score, the total fraud risk probability of a large outflow transaction from the target account is calculated, and a risk decision is made based on the total fraud risk probability. The process involves obtaining a first feature vector of a large outflow transaction associated with a large inflow transaction in the target account, calculating the maximum similarity between the first feature vector and the risk cluster set of large inflow transactions, and obtaining a pre-emptive telecom fraud risk score, including: Obtain the first feature vector of a large transfer-out transaction associated with a large transfer-in transaction in the target account, calculate the maximum similarity between the first feature vector and the risk cluster set of large transfer-in transactions, determine whether the large transfer-out transaction in the target account corresponds to multiple large transfer-in transactions, and obtain the first judgment result. If the first judgment result indicates that the large outflow transaction of the target account corresponds to multiple large inflow transactions, then the maximum similarity between the first feature vector and the risk cluster set of large inflow transactions is determined as the pre-existing telecom fraud risk score of the large outflow transaction of the target account. If the first judgment result indicates that there is no corresponding large-amount transfer-in transaction for the target account during the transaction, then the pre-transfer fraud risk score of the target account for the large-amount transfer-out transaction during the transaction is determined to be zero. The formula for calculating the similarity between the first feature vector and the risk cluster set of large-amount inflow transactions is as follows: in, Let A be the similarity between the first feature vector and the set of risk clusters for large-amount transfer transactions; A is the first feature vector, B is a risk cluster for large-amount transfer transactions in the set of risk clusters for large-amount transfer transactions, and n is the dimension of the first feature vector. ; The maximum similarity between the first feature vector and the risk cluster set of large transfer-in transactions is calculated by taking the feature vector of each large transfer-in transaction of the target account in the past 24 hours and the pre-transfer behavior feature vector of the transfer-in transaction, and each risk cluster in the pre-clustering result cluster set loaded into memory, and obtaining multiple similarities. The maximum similarity value is taken based on the multiple similarities. The construction of the first feature vector includes at least one of the following features: number of days since account opening, limit query operation, limit increase behavior, small trial transaction, transaction channel, whether the transaction amount is an integer multiple, number of counterparty accounts, number of transactions, transaction amount surge multiple, transactions during sensitive time periods, and fund transition indicators; The second feature vector should include at least the transaction channel, type, amount characteristics, and fund transition of the transfer transaction.

2. The method according to claim 1, characterized in that, The method further includes: A risk decision probability threshold is set. When the total probability of telecom fraud is greater than or equal to the risk decision probability threshold, a large outflow transaction of the target account is determined to be at risk of telecom fraud and interception is triggered.

3. The method according to claim 1, characterized in that, The method further includes: The pre-fraud risk score and the in-process fraud risk score are weighted according to a preset weight ratio to obtain the total fraud risk probability; wherein the sum of the weight ratios is 1.

4. The method according to claim 1, characterized in that, The DP-Means clustering algorithm includes: Randomly select a transaction data point as the centroid of the first cluster; Iterate through each data point to be processed, calculate the distance between the data point to be processed and all existing cluster centroids. If there is a cluster with a distance less than or equal to a preset dynamic threshold, assign the data point to be processed to the nearest cluster. Otherwise, generate a new cluster and use the data point as the centroid of the new cluster. Recalculate the centroids of all data points within each cluster, and use them as the new centroids for each cluster. Repeat the assignment until all data points have been processed.

5. A cluster-based real-time telecom fraud transaction detection device, characterized in that, The device includes: The acquisition module is used to obtain a list of historical accounts involved in cases that have been defined as fraudulent accounts; The clustering module is used to perform offline pre-event behavioral clustering based on large-amount inflow transactions of historical accounts involved in cases, using the DP-Means clustering algorithm to obtain a set of risk clusters for large-amount inflow transactions of historical accounts involved in cases; and to perform offline in-event behavioral clustering based on large-amount outflow transactions of historical accounts involved in cases, using the DP-Means clustering algorithm to obtain a set of risk clusters for large-amount outflow transactions of historical accounts involved in cases. The calculation module is used to obtain a first feature vector of a large transfer-in transaction associated with a large outflow transaction of a target account, calculate the maximum similarity between the first feature vector and the risk cluster set of large transfer-in transactions, and obtain a pre-emptive telecom fraud risk score; obtain a second feature vector of a large outflow transaction of a target account, calculate the maximum similarity between the second feature vector and the risk cluster set of large outflow transactions, and obtain a real-time telecom fraud risk score; the step of obtaining the first feature vector of a large transfer-in transaction associated with a large outflow transaction of a target account, calculating the maximum similarity between the first feature vector and the risk cluster set of large transfer-in transactions, and obtaining a pre-emptive telecom fraud risk score includes: Obtain the first feature vector of a large transfer-out transaction associated with a large transfer-in transaction in the target account, calculate the maximum similarity between the first feature vector and the risk cluster set of large transfer-in transactions, determine whether the large transfer-out transaction in the target account corresponds to multiple large transfer-in transactions, and obtain the first judgment result. If the first judgment result indicates that the target account's large outflow transaction corresponds to multiple large inflow transactions, then the maximum similarity between the first feature vector and the risk cluster set of large inflow transactions is determined as the pre-emptive fraud risk score of the target account's large outflow transaction; if the first judgment result indicates that the target account's large outflow transaction does not correspond to a large inflow transaction, then the pre-emptive fraud risk score of the target account's large outflow transaction is determined to be zero. The formula for calculating the similarity between the first feature vector and the risk cluster set of large-amount inflow transactions is as follows: in, Let A be the similarity between the first feature vector and the set of risk clusters for large-amount transfer transactions; A is the first feature vector, B is a risk cluster for large-amount transfer transactions in the set of risk clusters for large-amount transfer transactions, and n is the dimension of the first feature vector. ; The maximum similarity between the first feature vector and the risk cluster set of large-amount transfer-in transactions is calculated by taking the feature vector of each large-amount transfer-in transaction in the past 24 hours of a large-amount transfer-out transaction of the target account and the pre-transfer-in transaction's behavioral feature vector, and each risk cluster in the pre-clustering result cluster set loaded into memory, to obtain multiple similarities. The maximum similarity value is taken from the multiple similarities. The construction of the first feature vector includes at least one of the following features: number of days since account opening, limit query operation, limit increase behavior, small-amount trial transaction, transaction channel, whether the transaction amount is an integer multiple, number of counterparty accounts, number of transactions, transaction amount surge multiple, transactions in sensitive time periods, and fund transition indicators. The second feature vector includes at least the transaction channel, type, amount characteristics, and fund transition of the transfer-out transaction. The decision-making module is used to combine the pre-fake fraud risk score and the in-fake fraud risk score to calculate the total fraud risk probability of a large transfer transaction of the target account, and to make a risk decision based on the total fraud risk probability.

6. A computing device, characterized in that, Including memory and processor; The memory stores one or more computer programs, the one or more computer programs including instructions; when the instructions are executed by the processor, the computing device performs the method as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program for performing the method as described in any one of claims 1 to 4.