Abnormal transaction detection method and device, storage medium and program product
By performing clustering and outlier metric analysis on transaction data, structured intermediate results are generated, solving the accuracy and real-time issues of abnormal transaction detection in existing technologies, and enabling efficient identification and tracing of new or collaborative abnormal transactions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INDUSTRIAL AND COMMERCIAL BANK OF CHINA
- Filing Date
- 2026-02-11
- Publication Date
- 2026-05-15
AI Technical Summary
Existing abnormal transaction detection solutions struggle to balance accuracy and real-time performance in large-scale financial transaction scenarios, and are unable to effectively identify new or collaborative abnormal transactions that circumvent rules.
By acquiring the transaction behavior feature vectors from transaction data, clustering is performed to determine outlier metrics and identification information, generating structured intermediate results, and identifying abnormal transactions in the case of discrete transaction data.
It enables rapid screening of potential abnormal transactions from massive transaction data, improving detection accuracy and traceability, and effectively identifying new or collaborative abnormal transactions that circumvent rules.
Smart Images

Figure CN122048367A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of financial technology and intelligent risk control technology, and in particular to a method, device, storage medium and program product for detecting abnormal transactions. Background Technology
[0002] With the development of financial technology and the diversification of payment channels, the daily transaction data volume of financial institutions has increased dramatically, covering multiple channels such as counter services, mobile banking, cross-border settlement and third-party payment. Its high-dimensionality, strong temporal sequence and heterogeneous characteristics pose a severe challenge to the detection of abnormal transactions.
[0003] At present, abnormal transactions are mainly identified through set rules. That is, risk control personnel set fixed thresholds or logical conditions (e.g., exceeding the limit for daily transaction frequency or large cross-regional transfers), and the system matches the transaction flow in real time and triggers alarms or interceptions.
[0004] However, this method relies on human experience and has rigid rules, making it difficult to detect new or collaborative abnormal transactions that are not explicitly defined. Furthermore, it suffers from a tradeoff between efficiency and accuracy when dealing with large-scale transaction data, resulting in limited real-time detection effectiveness. Summary of the Invention
[0005] This invention provides a method, device, storage medium, and program product for detecting abnormal transactions, which solves the technical problems of existing abnormal transaction detection schemes in large-scale financial transaction scenarios, which are unable to balance detection accuracy and real-time performance, and cannot effectively identify new or collaborative abnormal transactions that circumvent rules. It can quickly screen potential abnormal transactions in massive transaction data, and can also perform in-depth identification of suspicious transactions, thereby improving detection accuracy and traceability.
[0006] According to one aspect of the present invention, a method for detecting abnormal transactions is provided, the method comprising: Obtain transaction data for each transaction generated by the target transaction system within a preset time period, and determine the transaction behavior feature vector that matches each of the transaction data. Clustering is performed on the feature vectors of each transaction behavior, and outlier metric values are determined for each transaction data based on the clustering results. Identification information for each transaction data is then determined based on the outlier metric values. Based on the target transaction data, the outlier metric of the target transaction data, and the identification information of the target transaction data, a structured intermediate result of the target transaction data is generated; If the target transaction data is determined to be discrete transaction data based on the structured intermediate results of each transaction data to be analyzed, the target transaction corresponding to the target transaction data is identified as an abnormal transaction.
[0007] According to another aspect of the present invention, an abnormal transaction detection device is provided, the device comprising: The acquisition module is used to acquire transaction data of each transaction generated by the target transaction system within a preset time period, and to determine the transaction behavior feature vector that matches each transaction data. The first clustering module is used to perform clustering processing on the feature vectors of each transaction behavior, determine the outlier metric of each transaction data according to the clustering results, and determine the identification information of each transaction data according to each outlier metric. The structured intermediate result generation module is used to generate structured intermediate results of the target transaction data based on the target transaction data, the outlier metric of the target transaction data, and the identification information of the target transaction data. The second clustering module is used to identify the target transaction corresponding to the target transaction data as an abnormal transaction when the target transaction data is determined to be discrete transaction data based on the structured intermediate results of each transaction data to be analyzed.
[0008] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the abnormal transaction detection method according to any embodiment of the present invention.
[0009] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the abnormal transaction detection method according to any embodiment of the present invention.
[0010] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the abnormal transaction detection method described in any embodiment of the present invention.
[0011] The technical solution of this invention involves acquiring transaction data of each transaction generated by a target trading system within a preset time period, and determining transaction behavior feature vectors that match each transaction data. The transaction behavior feature vectors are then clustered, and outlier metrics are determined based on the clustering results. Identification information for each transaction data is also determined based on these outlier metrics. A structured intermediate result for the target transaction data is generated based on the target transaction data, its outlier metrics, and its identification information. If the target transaction data is determined to be discrete based on the structured intermediate result of each transaction data to be analyzed, the target transaction corresponding to the target transaction data is identified as an abnormal transaction. This solution addresses the technical problems of existing abnormal transaction detection schemes, which struggle to balance detection accuracy and real-time performance in large-scale financial transaction scenarios and cannot effectively identify novel or collaborative abnormal transactions that circumvent rules. It enables rapid screening of potential abnormal transactions from massive amounts of transaction data and deep identification of suspicious transactions, improving detection accuracy and traceability.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a flowchart of an abnormal transaction detection method provided by an embodiment of the present invention; Figure 2 This is a flowchart of another abnormal transaction detection method provided by an embodiment of the present invention; Figure 3 This is a flowchart of another abnormal transaction detection method provided by an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of an abnormal transaction detection device provided according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of an electronic device that implements the abnormal transaction detection method of this invention. Detailed Implementation
[0015] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0016] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0017] Figure 1 This is a flowchart of an abnormal transaction detection method according to an embodiment of the present invention. This embodiment is applicable to the accurate detection of abnormal transactions in a transaction system. The method can be executed by an abnormal transaction detection device, which can be implemented in hardware and / or software and can be configured in electronic devices such as computers, servers, or tablet computers. Figure 1 As shown, the method includes: Step 110: Obtain transaction data for each transaction generated by the target trading system within a preset time period, and determine the transaction behavior feature vector that matches each transaction data.
[0018] The target transaction system can be any financial transaction processing system, such as a bank's core system or a third-party payment platform. The preset time can be a sliding or fixed time window, such as the most recent 5 minutes, 1 hour, or 1 day, etc., and is not limited in this embodiment.
[0019] Transaction data can be the raw record of each transaction, and may include fields such as transaction timestamp, transaction amount, payer identifier, payee identifier, IP (Internet Protocol) address, device fingerprint, geographic location, and merchant category. Transaction behavior feature vectors can be generated by transforming multiple heterogeneous fields in the transaction data into numerical vectors; for example, encoding geographic location as latitude and longitude, mapping device fingerprints to hash values, and converting transaction time periods into normalized hourly values.
[0020] Optionally, in this embodiment, multiple original transaction records generated within a preset time window can be obtained from the data interface of the target transaction system. Each original transaction record may contain multiple structured fields, such as transaction timestamp, transaction amount, payer identifier, payee identifier, network address information, device unique identifier, and geographic location information.
[0021] Subsequently, for each original transaction record, feature vectorization processing is performed to generate a corresponding transaction behavior feature vector. Specifically, time period features are extracted from the transaction timestamp and numerically encoded; the transaction amount is normalized or logarithmically transformed; the payer identifier, payee identifier, and device unique identifier are converted into fixed-dimensional numerical representations through hash mapping or embedding tables, respectively; the network address information is parsed to determine its location or its original byte sequence is preserved in numerical form; the geographic location information is converted into latitude and longitude coordinates, and optionally, the Euclidean distance or spherical distance between it and the user's historically frequently used locations is calculated. After standardization, the above-mentioned features are concatenated in a predefined order to form a single numerical vector, which serves as the transaction behavior feature vector for that transaction.
[0022] For example, in a payment platform, the preset time is set to the most recent 10 minutes. A total of 12,843 transaction records are captured within this window. The raw data of one transaction includes: transaction time 2026-01-27T03:42:18Z, amount 15,200, payer ID ACC_8891, payee ID MCH_3345, device description d8a3b1c7, IP address 123.51.100.22, and geographic location acd. After feature engineering, this transaction is converted into a 64-dimensional floating-point vector, where the 0th dimension is the time normalized value 0.154 (3.7 hours / 24), the 1st dimension is log(15200)≈9.63, the 2nd–3rd dimensions are the latitude and longitude [x, y], and the remaining dimensions are generated by hash embedding. This vector is used as input for subsequent unsupervised clustering calculations.
[0023] It should be noted that in this embodiment, the transaction data is obtained only after user authorization, and the method of acquisition is reasonable and legal.
[0024] Step 120: Perform clustering processing on the feature vectors of each transaction behavior, determine the outlier metric of each transaction data based on the clustering results, and determine the identification information of each transaction data based on each outlier metric.
[0025] Clustering refers to the process of dividing transaction behavior feature vectors in a high-dimensional feature space into several groups (i.e., clusters) based on their similarity to each other. The clustering result can refer to the objective result output by the algorithm after execution. For example, it can include: sample cluster labels, i.e., the cluster number assigned to each transaction; cluster center coordinates, the position of the center point of each cluster in the feature space; and the distance between the sample point and the cluster center, which can be used to measure the degree of deviation of a single transaction from the typical pattern of its group.
[0026] Outlier metric values are scalar values calculated based on the distance between a sample point and the cluster center, used to quantitatively characterize the degree of anomaly in a transaction. They reflect the extent to which the transaction deviates from the norm of its group. For example, the distance can be used directly as the outlier metric, or it can be normalized or standardized to fall within a uniform numerical range (e.g., between 0 and 1). Identification information is an auxiliary label generated based on the outlier metric values for subsequent processing. For example, if the outlier metric value is greater than a dynamic threshold (e.g., 0.7 or 0.8, which is not limited in this embodiment), it is identified as high-risk; otherwise, it is low-risk. In this embodiment, the identification information can be a Boolean value, a category label, or an encoded string.
[0027] In an optional implementation of this embodiment, after the transaction behavior feature vectors are matched with each transaction data, the transaction behavior feature vectors can be further clustered. Based on the clustering results, the outlier metric of each transaction data can be determined, and the identification information of each transaction data can be determined based on each outlier metric. Optionally, all transaction behavior feature vectors can be processed based on an unsupervised clustering model to determine the k-nearest neighbor set for each feature vector, where k is a preset positive integer, typically ranging from 10 to 50. Further, the local reachability distance of each feature vector is calculated, defined as the larger of the distance from the target feature vector to any of its k-nearest neighbor feature vectors and the maximum distance between the k-nearest neighbors of its corresponding nearest neighbor. Then, the reciprocal form of the local reachability density of each feature vector (i.e., the local reachability metric) is calculated, and an outlier metric is calculated based on this metric. This outlier metric is equal to the ratio of the average local reachability metric of the target feature vector's k-nearest neighbors to the local reachability metric of the target feature vector itself. After calculating all outlier metrics, a dynamic risk threshold for the current risk control scenario is obtained. This dynamic risk threshold is dynamically adjusted based on the proportion of abnormal transactions verified within the most recent sliding time window and corrected in conjunction with the characteristics of the business period. Furthermore, the outlier metric corresponding to each transaction is compared with the dynamic risk threshold. If the outlier metric is greater than or equal to the threshold, the identifier of the transaction data is set to the first preset code 1; otherwise, it is set to the second preset code 0.
[0028] Step 130: Based on the target transaction data, the outlier metric of the target transaction data, and the identification information of the target transaction data, generate a structured intermediate result for the target transaction data.
[0029] The target transaction data can be any of the transaction data obtained above, and this embodiment does not limit it. The structured intermediate result is a composite data object formed by organizing the multi-dimensional and multi-type information generated in the anomaly detection pipeline for each target transaction data according to a predefined logic and format. It may include a feature vector field (carrying the original transaction behavior feature vector or its derived representation), a risk indicator field (carrying the calculated outlier metric value or its risk score after further processing), and a control parameter field (carrying the process control instructions directly mapped or derived from the identification information).
[0030] Optionally, in this embodiment, after obtaining the outlier metric and identification information of the target transaction data based on the above steps, the transaction behavior feature vector, outlier metric, and identification information corresponding to each target transaction data can be obtained from the processing cache. Following a predefined data structure template, the transaction behavior feature vector is written to the feature vector field, the outlier metric is written to the risk indicator field and retained to three decimal places, the identification information is written to the control parameter field, and the original target transaction data is written to the fourth field to retain the business context. Further, the data structure is serialized into a string in a preset format, and a timestamp and a unique message identifier are appended as metadata. Finally, the string is published to a specified topic in the message queue.
[0031] In an optional implementation of this embodiment, for each transaction, a structured record containing multiple fields is generated based on its affiliation and degree of anomalousness in the initial clustering: one field identifies the cluster or anomalous category to which the transaction belongs; another field is a binary label indicating whether the transaction needs to enter the fine-tuning stage; when the label is of the first value, it indicates that the transaction is initially judged as suspicious and requires further analysis; when it is of the second value, it indicates that the transaction belongs to the normal pattern and can be directly released; in addition, a numerical field is included to reflect the degree to which the transaction deviates from the normal behavior pattern, serving as a risk reference input for the subsequent fine-tuning model. The above set of structured records together constitutes the output of the initial selection stage, providing a standardized data interface for the two-stage detection architecture.
[0032] In one example of this embodiment, after completing the initial clustering of a batch of transactions, a structured record can be generated for one of the transactions identified as being far from the main cluster: the record includes its anomaly category identifier, a binary label set to 1, and a risk score of 2.15; while for transactions located within dense clusters, their binary label is set to 0 and they do not participate in subsequent processing; the structured records corresponding to all transactions are uniformly organized into an intermediate dataset and passed to the anomaly detection module in the next stage for identifying potential risks.
[0033] Step 140: If the target transaction data is determined to be discrete transaction data based on the structured intermediate results of each transaction data to be analyzed, the target transaction corresponding to the target transaction data is identified as an abnormal transaction.
[0034] The transaction data to be analyzed consists of the set of transactions that have completed the initial screening stage and generated structured intermediate results, serving as input for the fine screening stage. Discrete transaction data refers to isolated points identified in the local neighborhood analysis of the fine screening stage as not forming density-connected or correlated subgraphs with any other suspicious transactions.
[0035] Optionally, in this embodiment, after obtaining the structured intermediate results of each transaction data, the structured intermediate results of all transaction data can be read from the storage unit, and the transaction dataset to be analyzed with the identification information of the first preset code can be selected; based on the relational feature vector of each transaction in the transaction dataset to be analyzed, the correlation strength between any two transactions is calculated. When the correlation strength exceeds a preset threshold, a connection edge is established between the two to construct a transaction correlation graph; connected component detection is performed on the transaction correlation graph to obtain several connected subgraphs; for each suspicious transaction, the number of nodes in its connected subgraph is determined. If the number of nodes is equal to 1, the transaction data is determined to be discrete transaction data; in response to determining that the target transaction data is discrete transaction data, the status of the target transaction corresponding to the target transaction data is updated to abnormal transaction and recorded in the abnormal transaction log.
[0036] For example, 50 suspicious transactions identified as high-risk are received during the initial screening stage. By analyzing their device fingerprints and payee accounts, it is found that 47 of these transactions can be divided into three related subgraphs (containing 20, 18, and 9 transactions respectively), while the remaining three transactions are independent and do not share devices with any other suspicious transactions. These three transactions each constitute a connected component of size 1 and are identified as discrete transaction data. Furthermore, their corresponding target transactions can be directly marked as abnormal transactions, triggering subsequent review processes.
[0037] In another optional implementation of this embodiment, after filtering out the transaction dataset to be analyzed whose identification information is a first preset code, the corresponding transaction behavior feature vector is obtained from the structured intermediate results of the transaction dataset to be analyzed, forming a fine screening feature matrix; local density clustering processing is performed on the fine screening feature matrix, and a neighborhood-based unsupervised clustering algorithm is used to identify several high-density regions as normal suspicious pattern clusters, and transaction vectors not assigned to any cluster are marked as noise points; for each suspicious transaction, if its transaction behavior feature vector is determined to be a noise point in the local density clustering processing, the transaction data is determined to be discrete transaction data; in response to determining that the target transaction data is discrete transaction data, the status of the target transaction corresponding to the target transaction data is updated to abnormal transaction and recorded in the abnormal transaction log.
[0038] The technical solution of this embodiment acquires transaction data of each transaction generated by the target transaction system within a preset time period and determines the transaction behavior feature vector matching each transaction data. It then performs clustering processing on each transaction behavior feature vector, determines the outlier metric of each transaction data based on the clustering results, and determines the identification information of each transaction data based on each outlier metric. Based on the target transaction data, the outlier metric of the target transaction data, and the identification information of the target transaction data, it generates a structured intermediate result for the target transaction data. If the target transaction data is determined to be discrete transaction data based on the structured intermediate result of each transaction data to be analyzed, the target transaction corresponding to the target transaction data is identified as an abnormal transaction. This solves the technical problem that existing abnormal transaction detection schemes struggle to balance detection accuracy and real-time performance in large-scale financial transaction scenarios, and cannot effectively identify new or collaborative abnormal transactions that circumvent rules. It can quickly screen potential abnormal transactions from massive transaction data and perform in-depth identification of suspicious transactions, improving detection accuracy and traceability.
[0039] Figure 2 This is a flowchart of another abnormal transaction detection method provided by an embodiment of the present invention. This embodiment is a further refinement of the above technical solution, and the technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 2 As shown, the method includes: Step 210: Obtain transaction data for each transaction generated by the target transaction system within a preset time period, and determine the transaction behavior feature vector that matches each transaction data.
[0040] Optionally, in this embodiment, acquiring transaction data of each transaction generated by the target transaction system within a preset time period and determining transaction behavior feature vectors that match each transaction data can include: preprocessing each transaction data to obtain standard transaction data; determining transaction behavior features of each transaction based on the standard transaction data, and concatenating the transaction behavior features to obtain transaction behavior feature vectors that match each transaction data.
[0041] Among them, the characteristics of transaction behavior include at least one of the following: statistical characteristics of amount, characteristics of time frequency, characteristics of transaction relationship, and characteristics of regional heterogeneity.
[0042] Optionally, in this embodiment, raw transaction data can be read from the transaction log interface, records with the source identifier "test environment" can be removed, and missing geographic location fields can be filled in using the IP address geolocation database to form standard transaction data. For each standard transaction data, its amount statistical characteristics are calculated, including the current transaction amount and the cumulative transaction amount of the payer's account in the past 24 hours; time frequency characteristics are calculated, including the hour value of the transaction occurrence time and the transaction frequency of the payer in the last 60 minutes; transaction relationship characteristics are calculated, including the number of transactions between the payer and the current payee in the past 30 days and the number of shared device identifiers; and regional heterogeneity characteristics are calculated, including whether the transaction currency is not the benchmark currency, whether the region where the payer's registration is located is different from the region where the payee's registration is located, and whether the region where the transaction initiating IP address is located is inconsistent with the region where the payer's historically frequently used location is located. The above features are arranged in the order of amount statistical characteristics, time frequency characteristics, transaction relationship characteristics, and regional heterogeneity characteristics, and all values are normalized. Finally, a fixed-dimensional floating-point vector is generated by concatenating the values, which serves as the transaction behavior feature vector corresponding to the transaction.
[0043] In one example of this embodiment, an original transaction record is obtained, with an amount of 8000 units in the first currency. The transaction timestamp corresponds to 2:00 AM in a certain time zone. The payer's registered location is in region A, the payee's registered location is in region B, and the transaction initiating IP address belongs to region C. After preprocessing, the transaction amount is converted to 57600 units in the base currency according to the real-time exchange rate, the timestamp is uniformly converted to 2:00 PM in the target time zone, and the geographic location field is completed to region C. Subsequently, features are extracted: the amount statistics feature is [57600, 120000] (current transaction amount and the payer's cumulative amount in the past 24 hours), the time frequency feature is [14, 3] (hour of transaction occurrence and transaction frequency in the past 60 minutes), the transaction relationship feature is [0, 0] (no interaction record with the payee in the past 30 days, no shared device identifier), and the regional heterogeneity feature is [1, 1, 0]. [1] (The transaction currency is not the benchmark currency, the payer and payee are registered in different regions, and the IP address region is inconsistent with the historically frequently used region of the payer); Finally, the above features are concatenated into a 9-dimensional floating-point vector, which serves as the transaction behavior feature vector corresponding to the transaction.
[0044] The solution in this embodiment divides transaction data into four categories of behavioral features: amount statistics, time frequency, transaction relationship, and cross-border. These features are then extracted and spliced together in a targeted manner. This invention achieves a systematic representation of high-dimensional heterogeneous transaction information and effectively integrates multi-dimensional signals such as single transaction attributes, user behavior habits, inter-subject correlations, and regional risks. It avoids feature shifts caused by differences in data sources and provides a stable, reliable, and high signal-to-noise ratio input foundation for subsequent anomaly detection.
[0045] Step 220: Cluster the feature vectors of each transaction behavior, determine the outlier metric of each transaction data based on the clustering results, and determine the identification information of each transaction data based on each outlier metric.
[0046] Optionally, in this embodiment, clustering is performed on each transaction behavior feature vector, and outlier values for each transaction data are determined based on the clustering results. Identification information for each transaction data is then determined based on each outlier value. This process may include: performing unsupervised clustering on the transaction behavior feature vectors to obtain multiple transaction feature clusters; determining the cluster center of the cluster to which the target transaction data belongs, and determining the target distance between the transaction behavior feature vector of the target transaction data and the cluster center; determining the target distance as the outlier value of the target transaction data; if the outlier value of the target transaction data is greater than a first preset threshold, the target transaction data is identified as transaction data to be analyzed and its identification information is 1; if the outlier value of the target transaction data is less than or equal to the first preset threshold, the target transaction data is identified as non-transaction data to be analyzed and its identification information is 0.
[0047] Unsupervised clustering refers to the algorithmic process of grouping transaction behavior feature vectors according to similarity without relying on labels, such as K-Means and Mini-Batch K-Means. Each transaction feature cluster in the clustering result represents a set of transactions with similar behavioral patterns. The cluster center is the central point of each cluster, which can be the mean vector of all vectors within that cluster. The target distance is the distance between the feature vector of the target transaction and its cluster center; the outlier metric is the target distance, and the larger the value, the more the transaction deviates from its normal pattern cluster. The identifier is 1 or 0, where 1 indicates high risk (transaction to be analyzed) and 0 indicates low risk (transaction not to be analyzed).
[0048] In one optional implementation of this embodiment, after determining the transaction behavior feature vectors of each transaction data, an unsupervised clustering algorithm can be used to process all transaction behavior feature vectors, initialize several (e.g., 50 or 60) cluster centers, and iteratively update the positions of each cluster center until a preset convergence condition is met, completing the clustering and assigning a cluster to each transaction; for each target transaction data, obtain the cluster center vector corresponding to its cluster, calculate the Euclidean distance between the transaction behavior feature vector of the target transaction data and the cluster center vector, and use it as the target distance; use the target distance as the outlier metric of the target transaction data. Read a first preset threshold from the system configuration parameters, which is set according to the distance distribution statistical characteristics of historical normal transaction samples; compare the outlier metric with the first preset threshold. If the outlier metric is greater than the first preset threshold, mark the state of the target transaction data as transaction data to be analyzed and set the identification information to 1; if the outlier metric is less than or equal to the first preset threshold, mark the state of the target transaction data as transaction data not to be analyzed and set the identification information to 0.
[0049] Optionally, in this embodiment, K-Means can be used to quickly coarsely screen the entire dataset; its objective function is: ;in For the sample Cluster center. The degree of anomaly is measured by calculating the distance between the sample and the cluster center: Set threshold If the distance between a sample and the cluster center satisfies the following formula: If the transaction is deemed suspicious, it will be classified as a suspicious sample and proceed to the next stage.
[0050] For example, K-Means clustering is performed on the feature vectors of a batch of transactions to form 40 clusters; one transaction is assigned to the 12th cluster, and the Euclidean distance between its feature vector and the cluster center is 4.2; the first preset threshold read is 3.5; since 4.2 > 3.5, the transaction is marked as transaction data to be analyzed, and the identification information is set to 1; another transaction is 2.8 away from its cluster center, which is less than the threshold, so the identification information is set to 0 and it does not enter the subsequent processing flow.
[0051] The solution in this embodiment quantifies the outlier degree of transactions by using an unsupervised clustering method based on cluster center distance. This invention achieves adaptive modeling of normal transaction patterns and avoids the limitations of relying on manually set fixed thresholds. Using the distance from the target transaction to its cluster center as an outlier metric can effectively reflect its deviation from the local normal behavior distribution and improve the sensitivity of the initial screening stage to novel abnormal transactions.
[0052] Step 230: Based on the target transaction data, the outlier metric of the target transaction data, and the identification information of the target transaction data, generate a structured intermediate result for the target transaction data.
[0053] Optionally, in this embodiment, generating a structured intermediate result of the target transaction data based on the target transaction data, the outlier metric of the target transaction data, and the identification information of the target transaction data may include: using the target transaction data, the outlier metric of the target transaction data, and the identification information of the target transaction data as input parameters of a first expression; mapping the target transaction data, the outlier metric of the target transaction data, and the identification information of the target transaction data through the first expression to obtain the structured intermediate result of the target transaction data.
[0054] The first expression can refer to a predefined data transformation rule or function used to organize input parameters into structured data objects. In this embodiment, it can be a constructor function, serialization template, mapping script, or configurable data assembly logic in the code; it is not a mathematical formula, but rather a data encapsulation operation in the technical process.
[0055] Optionally, in this embodiment, the target transaction data (not the target transaction data itself, but a transaction behavior feature vector matching the target transaction data), the outlier metric of the target transaction data, and the identification information of the target transaction data can be used as input parameters of the first expression; the target transaction data, the outlier metric of the target transaction data, and the identification information of the target transaction data are mapped through the first expression to obtain a structured intermediate result of the target transaction data, which may specifically include the following process: First, after completing the outlier metric calculation and identifier generation, three related data items for each target transaction can be obtained: the original transaction record (transaction behavior feature vector or unique transaction identifier), the corresponding outlier metric, and the identifier information. Then, the preset data assembly logic is called. This logic takes these three data items as input parameters and writes each parameter into the corresponding field position according to the pre-configured field order and data type requirements.
[0056] Furthermore, the padded data structure is serialized into a byte stream or string in a unified format that supports cross-module parsing, such as a lightweight data exchange format or a binary serialization protocol. Finally, the generated serialized object is the structured intermediate result of the target transaction data and is stored in a message queue or shared memory buffer for consumption in the subsequent anomaly detection phase.
[0057] The solution in this embodiment uses predefined data encapsulation logic to uniformly map target transaction data, outlier metrics, and identification information. This invention achieves standardization of output results and interface normalization in the initial screening stage, effectively avoiding parsing errors or information loss caused by inconsistent data formats during multi-stage processing. The structured intermediate results explicitly retain the original transaction context and quantitative risk indicators, providing a complete and traceable input basis for subsequent fine-grained analysis.
[0058] In an optional implementation of this embodiment, the first expression may include a feature field, a feature field, and a scoring field. Accordingly, mapping the target transaction data, the outlier metric of the target transaction data, and the identification information of the target transaction data through the first expression to obtain a structured intermediate result of the target transaction data may include: filling the transaction behavior feature vector of the target transaction data into the feature field of the first expression; filling the outlier metric of the target transaction data into the scoring field of the first expression; and filling the identification information into the admission control field of the first expression to obtain a structured intermediate result.
[0059] Optionally, in this embodiment, after completing the cluster analysis and identification determination in the initial screening stage, three core data items corresponding to each target transaction can be obtained: transaction behavior feature vector, outlier metric, and identification information; at the same time, a predefined first expression template is loaded, which contains three preset fields: feature field, scoring field, and admission control field, each field having a clear data type and position index.
[0060] Furthermore, the field filling operation is performed: the transaction behavior feature vector is written into the feature field of the first expression, maintaining the original dimension and numerical precision; the outlier metric is written into the scoring field, and the specified number of decimal places are retained according to the system configuration; the identification information is written into the admission control field, ensuring that its value conforms to the preset enumeration range (e.g., 0 or 1); after filling is completed, the entire first expression instance is serialized to generate a standardized data object, which is the structured intermediate result of the target transaction data.
[0061] In this embodiment, the first expression can be a mapping function. : ; ;in, This indicates entry into the fine-tuning model; This indicates a normal transaction. This serves as an additional risk indicator input into subsequent models. : The preprocessed and expanded transaction feature matrix, i.e., the feature set of all samples; : A mapping function that maps the output of the initial model (e.g., the distance from each sample to the cluster center and the cluster assignment) to pairs of labels and risk values; The output set of the mapping function records whether each sample has entered the fine selection stage and its risk indicators, which are used for subsequent processing.
[0062] For example, a transaction's 128-dimensional transaction behavior feature vector, outlier metric 2.1547, and identifier 1; according to a preset first expression template, the 128-dimensional vector is filled into... Field 2.155 to fill in Field 1, fill in The field; after serialization, it forms a structured record, which is then pushed to the downstream processing queue.
[0063] In this embodiment, by filling the transaction behavior feature vector, outlier metric, and identification information into the feature field, scoring field, and admission control field of the first expression, the present invention achieves refined organization and semantic clarity of the initial screening output data. This allows subsequent processing modules to directly parse and reuse high-dimensional features and risk scores, avoiding computational redundancy caused by repeated feature extraction. The introduction of the admission control field provides an explicit flow control mechanism for the two-stage architecture, ensuring that only suspicious transactions are passed to the resource-intensive fine screening module, significantly improving the overall system throughput.
[0064] Step 240: Based on the structured intermediate results of each transaction data to be analyzed, determine that the target transaction data is discrete transaction data.
[0065] Optionally, in this embodiment, determining that the target transaction data is discrete transaction data based on the structured intermediate results of each transaction data to be analyzed may include: extracting the transaction behavior feature vectors of each transaction data to be analyzed from the structured intermediate results of each transaction data to be analyzed; determining the Euclidean distance between the transaction behavior feature vector of the target transaction data and the transaction behavior feature vectors of each transaction data to be analyzed; determining the number of transaction data to be analyzed whose Euclidean distance is less than or equal to a preset neighborhood radius; if the number is less than a preset minimum neighborhood point threshold, then the target transaction data is determined to be discrete transaction data.
[0066] In an optional implementation of this embodiment, after obtaining the structured intermediate results of each transaction data and determining the transaction data to be analyzed, the structured intermediate results of all the transaction data to be analyzed can be further loaded from the message queue or cache, and the transaction behavior feature vectors contained in each record can be parsed to form a set of feature vectors to be analyzed; at the same time, the transaction behavior feature vectors of the current target transaction data can be obtained. Further, each vector in the set of feature vectors to be analyzed is traversed, and the Euclidean distance between it and the transaction behavior feature vector of the target transaction data is calculated to obtain a set of distance values.
[0067] Furthermore, the distance value can be compared one by one with the preset neighborhood radius, and the number of distance values less than or equal to the radius can be counted. This count represents the number of points in the neighborhood of the target transaction in the feature space. The preset minimum neighborhood point threshold can be read from the system configuration. If the number of points in the neighborhood is less than the threshold, it is determined that the target transaction data lacks sufficient similar suspicious transactions to support it in the local area, and it is identified as discrete transaction data. Otherwise, it is considered to be in a dense suspicious area and no discrete processing is performed.
[0068] In this embodiment, density clustering analysis can be performed on the suspicious transaction sample set screened in the initial screening stage. A neighborhood density-based clustering algorithm is used to further identify abnormal transactions in complex patterns. Using preset neighborhood radius and minimum neighborhood point threshold as parameters, the local density attribute of each sample in the feature space is determined: if the number of samples contained in the neighborhood of a sample is greater than or equal to the minimum neighborhood point threshold, then the sample is determined to be a core point; if a sample does not meet the core point condition but is located in the neighborhood of at least one core point, then the sample is determined to be a boundary point; if a sample neither meets the core point condition nor is adjacent to any core point, then the sample is determined to be an isolated point. The isolated point is the identified abnormal transaction, and finally an abnormal label set is generated, where each label corresponds to the abnormal state of the corresponding transaction, with a value of 0 or 1, representing normal or abnormal, respectively.
[0069] The solution in this embodiment determines whether a target transaction is discrete transaction data based on Euclidean distance and neighborhood density during the fine screening stage. This effectively identifies isolated abnormal behaviors that lack group collaboration characteristics and avoids misjudging such high-risk transactions as normal patterns. In addition, this strategy complements the outlier measurement in the initial screening stage, constructing a multi-dimensional anomaly identification mechanism from global deviation to local isolation, thereby improving the overall coverage of low-frequency and hidden transactions.
[0070] Step 250: If the target transaction data is determined to be discrete transaction data based on the structured intermediate results of each transaction data to be analyzed, the target transaction corresponding to the target transaction data is identified as an abnormal transaction.
[0071] Step 260: If the transaction amount of the target transaction is greater than or equal to a preset amount threshold, the target transaction is intercepted; if the transaction amount of the target transaction is less than the preset amount threshold and the transaction frequency is greater than or equal to a preset frequency threshold, a monitoring instruction for the target transaction is generated and executed; if the target transaction has regional heterogeneity characteristics, the target transaction is reported; if the transaction amount of the target transaction is less than the preset amount threshold, the transaction frequency is less than the preset frequency threshold, and there are no regional heterogeneity characteristics, the target transaction is reviewed.
[0072] The preset amount threshold is a pre-configured threshold value, such as 200,000, 500,000, or 1,000,000, which is not limited in this embodiment. Transaction frequency is the number of transactions initiated by the same account per unit time (e.g., 1 hour). The preset frequency threshold can be 10 transactions / hour or 20 transactions / hour, etc., to identify high-frequency probing behavior.
[0073] Optionally, in this embodiment, after identifying the target transaction corresponding to the target transaction data as an abnormal transaction, the transaction amount can be obtained from the standard transaction data of the target transaction, and the number of transactions of the payer's account in the most recent hour can be read as the transaction frequency. At the same time, the judgment result of the regional heterogeneity feature can be obtained. The transaction amount is compared with a preset amount threshold. If the transaction amount is greater than or equal to the preset amount threshold, the interception interface of the transaction gateway is called to prevent the target transaction from completing the payment. If the transaction amount is less than the preset amount threshold, the transaction frequency is further compared with a preset frequency threshold. If the transaction frequency is greater than or equal to the preset frequency threshold, a monitoring instruction containing transaction identifier and monitoring level fields is generated and sent to the real-time behavior monitoring module for continuous tracking. If the transaction frequency is less than the preset frequency threshold, it is further determined whether the regional heterogeneity feature is true. If it is true, the complete transaction record of the target transaction is packaged into structured reporting data and pushed to the reporting interface of the regulatory information platform through a secure channel. If the regional heterogeneity feature is false, the target transaction is marked as pending review and added to the manual review task queue for subsequent review by risk control personnel.
[0074] For example, after a transaction is determined to be an abnormal transaction, if the amount is 80,000 yuan, which exceeds the preset threshold of 50,000 yuan, it will be directly blocked; if the amount is 30,000 yuan, which does not exceed the threshold, but is the 12th transaction of the account within 1 hour (frequency threshold is 10), a monitoring instruction will be generated and real-time tracking will be initiated; if the amount is 20,000 yuan and the frequency is 2, but the payer's registered location and the payee's registered location are in different administrative regions, it will be reported to the compliance platform; if the amount is 5,000 yuan, the frequency is 1, and there is no regional heterogeneity, it will be sent to the manual review queue.
[0075] This embodiment of the solution constructs a multi-dimensional, tiered handling mechanism based on the heterogeneous characteristics of amount, frequency, and region. This allows for precise matching of response strategies to different types of abnormal transactions: high-amount transactions are immediately intercepted to effectively prevent significant financial losses; high-frequency, low-amount transactions are dynamically monitored, balancing user experience and risk coverage; transactions with inconsistent regional characteristics trigger compliance reporting; and transactions with a combination of low-risk characteristics are subject to manual review to avoid automated misjudgments. This tiered handling logic significantly improves the decision-making precision and resource allocation efficiency of the risk control system, while enhancing its adaptability and compliance in complex business scenarios.
[0076] In one example of this embodiment, during an e-commerce promotional event, a transaction is initiated by a user account to purchase multiple items. The transaction amount is high but does not reach the large-amount interception threshold. The transaction time is during the late night of the overnight promotion phase. The payee is a merchant account. The transaction device has a unique identifier, and the IP address belongs to a certain region. The geolocation field is missing in the original record.
[0077] First, the original transaction records are preprocessed: data with the source identifier of test environment is removed, the transaction amount is converted into the base currency according to the real-time exchange rate, the timestamp is converted to the standard time zone, and the geographical location is supplemented by the IP address location database to form standard transaction data with complete fields and uniform format.
[0078] Furthermore, multi-dimensional transaction behavior features are extracted based on the standard transaction data, including: the current transaction amount and the user's cumulative transaction amount over the past 24 hours; the specific hour in which the transaction occurred and the frequency of transactions initiated in the past hour; the number of interactions between the user and the current receiving merchant over the past 30 days and the number of shared devices; and whether the payer's registered location, the receiving party's registered location, and the IP address's region are consistent. Analysis showed that these three are different, indicating the existence of regional heterogeneity. After normalization, these features are concatenated in a preset order to generate a fixed-dimensional transaction behavior feature vector.
[0079] Further, in the initial screening stage, unsupervised clustering is performed on the feature vectors of all transactions on the day to obtain multiple transaction feature clusters. The transaction is assigned to one of the clusters, and if the Euclidean distance from its feature vector to the center of its cluster exceeds a first preset threshold, it is identified as a suspicious transaction and its identification information is set to 1. Subsequently, the transaction behavior feature vector, outlier metric, and identification information of the transaction are filled into a predefined data template to generate structured intermediate results, which are then pushed to a message queue for consumption by the fine screening module.
[0080] Furthermore, in the fine screening stage, transaction behavior feature vectors are extracted from the structured intermediate results of all transactions to be analyzed that are marked as 1, and a suspicious transaction feature set is constructed; the Euclidean distance between the feature vector of the target transaction and the remaining vectors in the set is calculated, and the number of transactions whose distance is less than or equal to the preset neighborhood radius is counted; since this number is lower than the preset minimum neighborhood point threshold, the target transaction is determined to be discrete transaction data.
[0081] Ultimately, based on its real-time transaction attributes, a tiered approach was taken: because the transaction amount was less than the preset amount threshold, no interception was triggered; however, the transaction frequency was higher than the preset frequency threshold and there were regional heterogeneous characteristics; according to the preset rule priority, the reporting process was carried out, the complete information of the transaction was encapsulated and pushed to the compliance and supervision interface, and a monitoring instruction was generated to track the subsequent behavior of the user account in real time in order to identify potential batch probing fraudulent behavior.
[0082] Based on the above technical solution, the target transaction system may involve multiple transaction participants. Before clustering the feature vectors of each transaction behavior, the method may further include: acquiring at least two transaction participants related to the target transaction data; determining at least one interaction indicator based on the interaction behavior of at least two transaction participants within a preset time window; the interaction indicator includes: the transaction frequency between the two parties, the average transaction amount per transaction, the number of times shared device identifiers are used, or the number of jointly associated third-party accounts; combining the interaction indicators into a relationship feature vector, and concatenating the relationship feature vector to the transaction behavior feature vector.
[0083] The participants in a transaction refer to the entities involved in the transaction, which may include the payer and the payee, and may also extend to intermediaries, guarantors, etc.
[0084] In an optional implementation of this embodiment, before clustering each transaction behavior feature vector, at least two transaction participants related to the target transaction data can be obtained; based on the interaction behavior of at least two transaction participants within a preset time window, at least one interaction indicator is determined; the interaction indicator includes: the transaction frequency between the two parties, the average transaction amount per transaction, the number of times shared device identifiers are used, or the number of jointly associated third-party accounts; combining each interaction indicator into a relationship feature vector and concatenating the relationship feature vector to the transaction behavior feature vector may include the following process: First, the payer's account identifier and the payee's account identifier are parsed from the target transaction data, identifying them as the two core transaction participants. Using the current transaction time as a baseline, a preset time window (e.g., 30 days) is traced back, and all interaction records between these two parties are retrieved from historical transaction logs, including transaction details, device usage logs, and account association graphs. Subsequently, several interaction metrics are calculated based on these records: the total number of transactions is counted as the transaction frequency; the total transaction amount is divided by the number of transactions to obtain the average transaction amount per transaction; the sets of device identifiers used by both parties are compared, and the number of overlapping elements is counted as the frequency of shared device identifiers; the historical counterparty sets of both parties are queried, and the size of their intersection is calculated as the number of jointly associated third-party accounts.
[0085] Furthermore, the aforementioned interaction indicators are organized into a numerical array in a fixed order to form a relational feature vector; simultaneously, the transaction behavior feature vector (containing basic features such as amount, time, and region) already generated for the target transaction is obtained; all elements of the relational feature vector are appended to the end of the transaction behavior feature vector to generate an enhanced feature vector with expanded dimensions; this enhanced feature vector will serve as the input for subsequent unsupervised clustering processing to more comprehensively characterize the risk attributes of the transaction.
[0086] For example, user A initiates a payment to merchant B. The transaction itself is of moderate amount and occurs at a normal time, but judging from a single transaction alone, it's difficult to determine if it's abnormal. To more comprehensively assess the risk, after generating the transaction behavior feature vector and before clustering, the historical interaction between A and B is further analyzed. Using the current transaction time as a benchmark, data from the past 30 days is reviewed, revealing that A and B have had 5 transactions totaling 100,000 yuan, averaging 20,000 yuan per transaction. Furthermore, A and B used the same mobile phone 3 times to complete their respective operations (e.g., A making a payment, B logging into the backend). The account relationship graph also reveals that A and B each have multiple counterparties, including 2 accounts where both parties have had financial transactions.
[0087] Based on this information, four interaction metrics were calculated: transaction frequency of 5, average transaction amount of 20,000, number of times devices were shared (3), and number of jointly associated third-party accounts (2). These four values were arranged in a fixed order to form a relational feature vector [5, 20000, 3, 2], which was then appended to the end of the original transaction behavior feature vector (such as a vector composed of features like amount, time, and geographical heterogeneity) to form a higher-dimensional enhanced feature vector. This enhanced vector was then used for unsupervised clustering. Because it contained strong correlation signals such as high frequency, stable amount, device sharing, and social overlap, the clustering algorithm categorized it into a normal transaction cluster representing acquaintances or long-term partners, thus avoiding misjudging a transaction as high-risk due to seemingly abnormal behavior in a single instance.
[0088] The solution in this embodiment, by introducing a relational feature vector based on the historical interaction behavior of transaction participants before clustering, effectively integrates long-term association information between transaction entities with the instantaneous behavioral characteristics of a single transaction, significantly improving the semantic richness of feature representation and risk discrimination capability.
[0089] To better understand the detection of abnormal transactions involved in the embodiments of the present invention, Figure 3 This is a flowchart of another abnormal transaction detection method provided by an embodiment of the present invention, referred to [reference]. Figure 3The process includes the following steps: First, raw transaction data is acquired through a data acquisition module; next, the raw transaction data is input into a feature engineering module to extract multi-dimensional transaction behavior features and generate transaction behavior feature vectors; next, the feature vectors of all transactions are input into a preliminary clustering model M1, and the K-Means algorithm is used for preliminary clustering to select a set D of suspicious transaction samples with high outlier rates; finally, the set D of suspicious transaction samples is input into an expression connection module to execute a mapping function. The process involves generating structured intermediate results, inputting these results into a refined clustering model M2, and further analyzing them using the DBSCAN algorithm. Isolated points are identified based on neighborhood density to obtain an anomaly label set. Finally, anomaly detection and labeling are performed based on transaction information within the anomaly label set.
[0090] After anomaly labeling is completed, each marked transaction is handled in a tiered manner: if the transaction amount is greater than or equal to a preset amount threshold, it is identified as a high-amount suspicious transaction and real-time interception is triggered; if the transaction amount is less than the preset amount threshold, it is identified as a low-amount suspicious transaction, and its transaction frequency is further assessed to determine if it is greater than or equal to a preset frequency threshold; if the transaction frequency meets the criteria, it is identified as a high-frequency small-amount transaction and added to the key monitoring list; if the transaction frequency does not meet the criteria, it is assessed to determine if it has specific risk characteristics such as cross-regional transactions; if so, it is identified as a transaction with significant cross-border characteristics and reported to the compliance audit department; if not, and without other obvious characteristics, it is identified as another suspicious transaction and enters the manual review process.
[0091] The solution of this invention, by constructing a collaborative architecture of dual unsupervised models, efficiently identifies complex anomaly patterns such as high-frequency small-amount splitting and account relay transfers without the need for manual annotation, significantly improving detection coverage and adaptability to unknown fraud. It utilizes an expression connection mechanism to achieve complete and traceable data transfer between the two stages, ensuring interpretability of results to meet regulatory audit requirements. Simultaneously, based on multi-branch processing strategies with heterogeneous characteristics of amount, frequency, and region, it intercepts high-risk transactions in real time, monitors or reviews medium- and low-risk transactions in stages, and guides cross-border transactions towards compliance. This approach balances risk control with user experience and regulatory compliance. The overall solution combines high efficiency, high accuracy, good real-time performance, and scalability, making it suitable for intelligent risk control in large-scale financial transaction scenarios.
[0092] Figure 4 This is a schematic diagram of the structure of an abnormal transaction detection device provided according to an embodiment of the present invention. Figure 4 As shown, the device includes: an acquisition module 410, a first clustering module 420, a structured intermediate result generation module 430, and a second clustering module 440.
[0093] The acquisition module 410 is used to acquire transaction data of each transaction generated by the target transaction system within a preset time period, and to determine the transaction behavior feature vector that matches each transaction data. The first clustering module 420 is used to perform clustering processing on the feature vectors of each transaction behavior, determine the outlier metric of each transaction data according to the clustering results, and determine the identification information of each transaction data according to the outlier metric. The structured intermediate result generation module 430 is used to generate structured intermediate results of the target transaction data based on the target transaction data, the outlier metric of the target transaction data, and the identification information of the target transaction data. The second clustering module 440 is used to identify the target transaction corresponding to the target transaction data as an abnormal transaction when the target transaction data is determined to be discrete transaction data based on the structured intermediate results of each transaction data to be analyzed.
[0094] In an optional implementation of this embodiment, the abnormal transaction detection device further includes: an anomaly handling module, used for: If the transaction amount of the target transaction is determined to be greater than or equal to a preset amount threshold, the target transaction will be intercepted. If it is determined that the transaction amount of the target transaction is less than a preset amount threshold and the transaction frequency is greater than or equal to a preset frequency threshold, a monitoring instruction for the target transaction is generated and executed. If it is determined that the target transaction has regional heterogeneity, the target transaction will be reported. If the transaction amount of the target transaction is less than a preset amount threshold, the transaction frequency is less than a preset frequency threshold, and there are no regional heterogeneous characteristics, the target transaction will be reviewed.
[0095] In an optional implementation of this embodiment, the acquisition module 410 is specifically used to preprocess each of the transaction data to obtain standard transaction data; Based on the standard transaction data, the transaction behavior characteristics of each transaction are determined, and the transaction behavior characteristics are concatenated to obtain a transaction behavior feature vector that matches each transaction data. The transaction behavior characteristics include at least one of the following: amount statistics characteristics, time frequency characteristics, transaction relationship characteristics, and regional heterogeneity characteristics.
[0096] In an optional implementation of this embodiment, the first clustering module 420 is specifically used to perform unsupervised clustering on the transaction behavior feature vector to obtain multiple transaction feature clusters; Determine the cluster center of the cluster to which the target transaction data belongs, and determine the target distance between the transaction behavior feature vector of the target transaction data and the cluster center; The target distance is determined as the outlier metric of the target transaction data; If the outlier metric of the target transaction data is determined to be greater than a first preset threshold, the target transaction data is identified as transaction data to be analyzed and its identification information is set to 1. If the outlier metric of the target transaction data is determined to be less than or equal to a first preset threshold, the target transaction data is determined to be non-analyzable transaction data and its identification information is 0.
[0097] In an optional implementation of this embodiment, the structured intermediate result generation module 430 is specifically used to take the target transaction data, the outlier metric of the target transaction data, and the identification information of the target transaction data as input parameters of the first expression; The first expression is used to map the target transaction data, the outlier metric of the target transaction data, and the identification information of the target transaction data to obtain a structured intermediate result of the target transaction data.
[0098] In an optional implementation of this embodiment, the first expression includes a feature field, a feature field, and a rating field; The structured intermediate result generation module 430 is also specifically used to fill the transaction behavior feature vector of the target transaction data into the feature field of the first expression; Fill the outlier metric of the target transaction data into the scoring field of the first expression; The identification information is filled into the admission control field of the first expression to obtain the structured intermediate result.
[0099] In an optional implementation of this embodiment, the second clustering module 440 is specifically used to extract the transaction behavior feature vector of each transaction data to be analyzed from the structured intermediate results of each transaction data to be analyzed; Determine the Euclidean distance between the transaction behavior feature vector of the target transaction data and the transaction behavior feature vector of each of the transaction data to be analyzed; Determine the number of transaction data to be analyzed whose Euclidean distance is less than or equal to a preset neighborhood radius; If the number is less than the preset minimum neighborhood point threshold, then the target transaction data is determined to be discrete transaction data.
[0100] In one optional implementation of this embodiment, the target transaction system involves multiple transaction participants; Before clustering the feature vectors of each transaction behavior, the method further includes: Obtain at least two transaction participants related to the target transaction data; Based on the interaction behavior of at least two of the transaction participants within a preset time window, at least one interaction indicator is determined; the interaction indicator includes: the frequency of transactions between the two parties, the average transaction amount per transaction, the number of times shared device identifiers are used, or the number of jointly associated third-party accounts; The interaction indicators are combined into a relationship feature vector, and the relationship feature vector is concatenated to the transaction behavior feature vector.
[0101] The abnormal transaction detection device provided in this embodiment of the invention can execute the abnormal transaction detection method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0102] The collection, storage, use, processing, transmission, provision, and disclosure of transaction data involved in the technical solutions of this invention comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0103] Figure 5 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0104] like Figure 5As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded into the RAM 13 from the storage unit 18. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0105] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0106] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods described above, such as methods for detecting anomalous transactions.
[0107] In some embodiments, the abnormal transaction detection method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the abnormal transaction detection method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the abnormal transaction detection method by any other suitable means (e.g., by means of firmware).
[0108] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0109] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0110] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, erasable programmable read-only memory (EPROM), optical fibers, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0111] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0112] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0113] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system. It addresses the shortcomings of traditional physical hosts and Virtual Private Servers (VPS) in terms of management difficulty and weak business scalability.
[0114] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0115] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
[0116] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements a database detection method as provided in any embodiment of this application.
[0117] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including LANs or WANs—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0118] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, they do not mean that the solution has been or necessarily used.
[0119] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A method for detecting abnormal transactions, characterized in that, The method includes: Obtain transaction data for each transaction generated by the target transaction system within a preset time period, and determine the transaction behavior feature vector that matches each of the transaction data. Clustering is performed on the feature vectors of each transaction behavior, and outlier metric values are determined for each transaction data based on the clustering results. Identification information for each transaction data is then determined based on the outlier metric values. Based on the target transaction data, the outlier metric of the target transaction data, and the identification information of the target transaction data, a structured intermediate result of the target transaction data is generated; If the target transaction data is determined to be discrete transaction data based on the structured intermediate results of each transaction data to be analyzed, the target transaction corresponding to the target transaction data is identified as an abnormal transaction.
2. The method for detecting abnormal transactions according to claim 1, characterized in that, After identifying the target transaction corresponding to the target transaction data as an abnormal transaction, the method further includes: If the transaction amount of the target transaction is determined to be greater than or equal to a preset amount threshold, the target transaction will be intercepted. If it is determined that the transaction amount of the target transaction is less than a preset amount threshold and the transaction frequency is greater than or equal to a preset frequency threshold, a monitoring instruction for the target transaction is generated and executed. If it is determined that the target transaction has regional heterogeneity, the target transaction will be reported. If the transaction amount of the target transaction is less than a preset amount threshold, the transaction frequency is less than a preset frequency threshold, and there are no regional heterogeneous characteristics, the target transaction will be reviewed.
3. The method for detecting abnormal transactions according to claim 1, characterized in that, The step of acquiring transaction data for each transaction generated by the target transaction system within a preset time period, and determining a transaction behavior feature vector matching each of the transaction data, includes: The transaction data is preprocessed to obtain standard transaction data; Based on the standard transaction data, the transaction behavior characteristics of each transaction are determined, and the transaction behavior characteristics are concatenated to obtain a transaction behavior feature vector that matches each transaction data. The transaction behavior characteristics include at least one of the following: amount statistics characteristics, time frequency characteristics, transaction relationship characteristics, and regional heterogeneity characteristics.
4. The method for detecting abnormal transactions according to claim 1, characterized in that, The process of clustering the feature vectors of each transaction behavior, determining the outlier metric for each transaction data based on the clustering results, and determining the identification information for each transaction data based on the outlier metric, includes: Unsupervised clustering is performed on the transaction behavior feature vectors to obtain multiple transaction feature clusters; Determine the cluster center of the cluster to which the target transaction data belongs, and determine the target distance between the transaction behavior feature vector of the target transaction data and the cluster center; The target distance is determined as the outlier metric of the target transaction data; If the outlier metric of the target transaction data is determined to be greater than a first preset threshold, the target transaction data is identified as transaction data to be analyzed and its identification information is set to 1. If the outlier metric of the target transaction data is determined to be less than or equal to a first preset threshold, the target transaction data is determined to be non-analyzable transaction data and its identification information is 0.
5. The method for detecting abnormal transactions according to claim 1, characterized in that, The process of generating structured intermediate results for the target transaction data based on the target transaction data, the outlier metric of the target transaction data, and the identification information of the target transaction data includes: The target transaction data, the outlier metric of the target transaction data, and the identification information of the target transaction data are used as input parameters for the first expression; The first expression is used to map the target transaction data, the outlier metric of the target transaction data, and the identification information of the target transaction data to obtain a structured intermediate result of the target transaction data.
6. The method for detecting abnormal transactions according to claim 5, characterized in that, The first expression includes feature fields, feature fields, and a rating field; By mapping the target transaction data, the outlier metric of the target transaction data, and the identification information of the target transaction data using the first expression, a structured intermediate result of the target transaction data is obtained, including: Fill the feature field of the first expression with the transaction behavior feature vector of the target transaction data; Fill the outlier metric of the target transaction data into the scoring field of the first expression; The identification information is filled into the admission control field of the first expression to obtain the structured intermediate result.
7. The method for detecting abnormal transactions according to claim 1, characterized in that, Based on the structured intermediate results of each transaction data to be analyzed, the target transaction data is determined to be discrete transaction data, including: Extract the transaction behavior feature vector of each transaction data to be analyzed from the structured intermediate results of each transaction data to be analyzed; Determine the Euclidean distance between the transaction behavior feature vector of the target transaction data and the transaction behavior feature vector of each of the transaction data to be analyzed; Determine the number of transaction data to be analyzed whose Euclidean distance is less than or equal to a preset neighborhood radius; If the number is less than the preset minimum neighborhood point threshold, then the target transaction data is determined to be discrete transaction data.
8. The method for detecting abnormal transactions according to any one of claims 1-7, characterized in that, The target trading system involves multiple trading participants; Before clustering the feature vectors of each transaction behavior, the method further includes: Obtain at least two transaction participants related to the target transaction data; Based on the interaction behavior of at least two of the transaction participants within a preset time window, at least one interaction indicator is determined; the interaction indicator includes: the frequency of transactions between the two parties, the average transaction amount per transaction, the number of times shared device identifiers are used, or the number of jointly associated third-party accounts; The interaction indicators are combined into a relationship feature vector, and the relationship feature vector is concatenated to the transaction behavior feature vector.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the abnormal transaction detection method according to any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the abnormal transaction detection method according to any one of claims 1-8.
11. A computer program product comprising a computer program that, when executed by a processor, implements the method for detecting abnormal transactions according to any one of claims 1-8.