Financial transaction system anomaly identification method based on isolated forest algorithm

Through the isolated forest algorithm, financial transaction data is processed, combined with multi-dimensional data processing and dynamic model updates, the problems of low accuracy, poor adaptability and insufficient real-time identification of abnormal transactions in the existing technology are solved, and efficient and accurate abnormal transaction identification and risk management are achieved.

CN120338955AInactive Publication Date: 2025-07-18赵奕涵
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510479155.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing abnormal transaction identification methods are not very accurate, poor adaptability, insufficient real-time, low interpretability, and difficult to effectively process multi-dimensional and heterogeneous financial transaction data.

Method used

The isolated forest algorithm is used to process financial transaction data, combine multi-dimensional data processing, dynamic model update and manual auditing to build an abnormal transaction identification system, calculate the abnormality score through data cleaning, feature engineering and the isolated forest model, and generate an abnormal transaction report.

Benefits of technology

It significantly improves the accuracy and real-time nature of abnormal transaction identification, reduces the false positive rate, improves the adaptability and interpretability of the system, and provides a powerful risk management tool.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338955A_ABST
    Figure CN120338955A_ABST
Patent Text Reader

Abstract

The invention relates to the field of financial data analysis, in particular to a financial transaction system anomaly identification method based on an isolated forest algorithm, and relates to big data analysis and machine learning. The method comprises the steps of obtaining multi-dimensional transaction data, executing data cleaning and feature engineering, constructing an isolated forest model, and calculating an anomaly score. The method has the remarkable advantages in the aspects of accuracy, real-time performance, adaptability and interpretability, abnormal transactions can be effectively recognized, financial environment changes can be rapidly adapted, and a powerful and flexible risk management tool is provided for financial institutions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of financial data analysis, and particularly to a method for identifying anomalies in a financial trading system based on the Isolation Forest algorithm. Background Art

[0002] With the rapid development of financial technology, electronic payment and online trading have become important components of the modern financial system. However, along with the continuous increase in the scale and complexity of transactions, financial fraud and abnormal transactions have shown an increasingly complex and diverse trend. Against this background, how to effectively identify and prevent abnormal transactions and ensure the security and stability of the financial system has become a major challenge faced by financial institutions.

[0003] Traditional methods for identifying abnormal transactions mainly rely on rule-based systems. This method determines whether a transaction is abnormal by setting a series of fixed rules and thresholds. For example, setting an upper limit for the amount of a single transaction, or stipulating the maximum number of transactions within a certain period of time, etc. Although this method is simple to implement and has a fast response speed, it has many limitations. First, fixed rules are difficult to cope with constantly changing fraud means and new abnormal patterns. Second, the formulation and maintenance of rules require a large amount of human input and professional knowledge. Moreover, overly strict rules may lead to a large number of false alarms, affecting the normal operation of the business; while overly loose rules may allow real abnormal transactions to escape detection.

[0004] With the development of machine learning technology, some methods based on statistical learning have been introduced into the field of abnormal transaction identification. For example, algorithms such as Support Vector Machine (SVM), decision tree, and neural network are used to build abnormal detection models. This type of method has stronger adaptability and generalization ability compared to rule-based systems. However, in practical applications, it still faces some challenges. First, these methods usually require a large amount of labeled data for training, and in the financial field, labeled data for abnormal transactions is often scarce and costly to obtain. Second, in the face of high-dimensional and large-scale transaction data, the computational efficiency and real-time performance of traditional machine learning methods are difficult to meet the requirements. Moreover, most of these methods belong to "black box" models, and their decision-making processes lack interpretability, which has become an important issue today when financial supervision is becoming increasingly strict.

[0005] In addition, most of the existing abnormal transaction identification systems are static and lack dynamic update and adaptive mechanisms. Financial fraud means and abnormal transaction patterns are constantly evolving, while the update cycle of the model often lags behind this change, resulting in a decline in detection effectiveness over time. At the same time, existing systems also face challenges in processing multi-dimensional and heterogeneous data. Financial transactions involve numerous data dimensions, including transaction amount, time, location, device information, etc. How to effectively integrate this information to improve the detection accuracy is a problem that has not been fully solved by the existing technology. Summary of the Invention

[0006] The present invention aims to solve the problems existing in the existing abnormal transaction recognition methods, such as low accuracy, poor adaptability, insufficient real-time performance, and low interpretability. By innovatively applying the Isolation Forest algorithm to the abnormal recognition of financial transactions and combining technologies such as multi-dimensional data processing and dynamic model updating, the present invention provides an efficient, accurate, and interpretable abnormal transaction recognition method.

[0007] The present invention proposes an abnormal recognition method for a financial trading system based on the Isolation Forest algorithm, including:

[0008] An acquisition step, including:

[0009] Acquire multi-dimensional transaction data in the financial trading system, where the multi-dimensional transaction data includes transaction objects, transaction amounts, and transaction time information;

[0010] A processing step, including:

[0011] Based on the multi-dimensional transaction data, perform data cleaning and feature engineering to obtain preprocessed transaction data;

[0012] According to the preprocessed transaction data, construct an Isolation Forest model;

[0013] Based on the Isolation Forest model, calculate the abnormality score of the transaction data;

[0014] An output step, including:

[0015] According to the abnormality score, determine abnormal transactions and generate an abnormal transaction report.

[0016] Preferably, the acquisition step specifically includes:

[0017] Extract transaction records from the database of the financial trading system;

[0018] Perform a preliminary screening on the transaction records to remove records containing obvious errors or missing values;

[0019] Convert the screened transaction records into a standardized data format.

[0020] Preferably, the data cleaning specifically includes:

[0021] Detect and process outliers, and replace values outside the normal range with statistical means or medians;

[0022] Fill in missing values, and estimate missing data items using interpolation or the nearest neighbor method;

[0023] Remove duplicate records and retain the latest or most complete transaction information.

[0024] "Outliers" processed during data cleaning: mainly refer to values that significantly deviate from physical laws or business logics due to data entry errors, transmission errors, system failures, etc., such as negative transaction amounts or extreme values exceeding the system limit, like a transaction amount being negative or exceeding the system limit.

[0025] "Abnormal transactions" to be ultimately identified: refer to actual transaction behavior patterns that may involve fraud, money laundering, or other risks, which are the core objectives of this invention.

[0026] Preferably, the data cleaning step in the method of this invention can be further refined into:

[0027] Handling of distinguishable outliers: Classify outliers into two categories:

[0028] Outliers of obvious error type: Values that are physically or logically impossible, such as negative transaction amounts, transactions exceeding the system maximum limit (e.g., 10 billion);

[0029] Suspicious outliers: Values that are numerically abnormal but physically possible, such as transactions exceeding 3 standard deviations of the user's historical transaction mean but within the system limit;

[0030] Differentiated processing strategy:

[0031] For outliers of obvious error type, they can be directly replaced or corrected;

[0032] For suspicious outliers, the original information should be retained, a marked field should be created, and the corrected value should be added for model training;

[0033] Preferably, the feature engineering specifically includes:

[0034] Constructing time-related features, including periodic indicators such as the hour, week, month of the transaction;

[0035] Calculating transaction frequency features, such as the number of transactions and the total transaction amount within a unit time;

[0036] Generating historical behavior features of the transaction object, such as the statistical transaction pattern over a certain period in the past.

[0037] Preferably, the steps of constructing the isolation forest model specifically include:

[0038] Randomly selecting subsets of training samples;

[0039] Constructing decision trees for each subsample, where each node randomly selects a feature and a splitting point;

[0040] Repeating the construction of multiple decision trees to form an isolation forest.

[0041] Preferably, during the process of constructing the decision tree:

[0042] Set the maximum tree depth and the minimum sample number threshold;

[0043] When the maximum tree depth is reached or the number of samples in a node is less than the minimum threshold, set the current node as a leaf node;

[0044] Calculate the anomaly score for each leaf node.

[0045] Preferably, the step of calculating the anomaly score of the transaction data specifically includes:

[0046] Input the transaction data to be detected into the isolation forest model;

[0047] Calculate the path length for the data sample to reach the leaf node in each decision tree;

[0048] Based on the path lengths of all decision trees, calculate the average path length;

[0049] According to the average path length, use a predefined anomaly calculation formula to obtain the final anomaly score.

[0050] Preferably, it further includes a model update step:

[0051] Periodically or when a new anomaly pattern is detected, retrain the isolation forest model;

[0052] Add the newly identified abnormal transaction data to the training set to improve the detection ability of the model.

[0053] Preferably, it further includes a manual review step:

[0054] Conduct a manual review of transactions with an anomaly score exceeding the preset threshold;

[0055] According to the results of the manual review, update the abnormal transaction database and the normal transaction database;

[0056] Based on the updated transaction database, adjust the threshold of the anomaly score.

[0057] Preferably, it further includes an alarm and handling step:

[0058] When an abnormal transaction is detected, immediately send an alarm to the system administrator;

[0059] According to the type and severity of the abnormal transaction, automatically execute preset risk control measures, such as suspending the transaction, restricting the transaction amount, or requiring additional identity verification.

[0060] The beneficial effects of the present invention are mainly reflected in the following aspects:

[0061] First, the present invention significantly improves the recognition accuracy of abnormal transactions. Through multi-dimensional data processing and feature engineering, combined with the advantages of the Isolation Forest algorithm, this method can capture more subtle and complex abnormal patterns. Experimental results show that the detection accuracy and recall rate of this method both exceed 90%, and the F1 score reaches 94.5%, far exceeding traditional methods and other machine learning methods.

[0062] Second, the present invention significantly reduces the false alarm rate. In financial trading systems, a low false alarm rate is crucial. The false alarm rate of this method is only 0.7%, significantly lower than traditional methods. This not only improves the reliability of the system but also reduces the unnecessary manual review workload and enhances the overall operational efficiency.

[0063] Furthermore, the present invention has excellent real-time performance. Despite using complex algorithms, this method still maintains a fast detection speed (15 milliseconds per transaction on average), capable of meeting the requirements of large-scale real-time trading systems. This performance advantage enables this method to provide a high level of security protection without affecting the user experience.

[0064] In addition, the dynamic model update mechanism of the present invention significantly enhances the adaptability of the system. Through regular or trigger-based model updates, this method can quickly adapt to new trading patterns and fraud means, maintaining the timeliness of the model. Compared with other methods, the model update time of the present invention is shorter (2.5 hours vs 6 hours), which means the system can respond more quickly to market changes.

[0065] Finally, the present invention also makes a breakthrough in interpretability. The working principle of the Isolation Forest algorithm is relatively intuitive and easy to explain to non-technical personnel. At the same time, this method can provide specific explanations for each abnormal determination by analyzing the path of abnormal transactions in the decision tree, which is of great significance for financial supervision and customer communication.

[0066] In summary, the method for identifying abnormal financial transactions based on the Isolation Forest algorithm provided by the present invention demonstrates significant advantages in multiple aspects such as accuracy, real-time performance, adaptability, and interpretability. This method can not only effectively identify various types of abnormal transactions but also quickly adapt to the changing financial environment, providing a powerful and flexible risk management tool for financial institutions. In the face of increasingly complex financial fraud challenges, the application of this method will bring significant security improvements and potential economic benefits to financial institutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 is the overall flowchart of the present invention.

[0068] Figure 2 is the flowchart of data cleaning and feature engineering of the present invention.

[0069] Figure 3 It is the flowchart for the construction of the forest model and the calculation of anomaly scores of the present invention.

[0070] Figure 4 It is the flowchart for the model update and manual review of the present invention. Specific Embodiments

[0071] Please refer to the attached Figures 1-4 , the present invention provides a method for anomaly recognition in a financial trading system based on the isolation forest algorithm. This method aims to effectively identify abnormal transactions in the financial trading system and improve trading security and system reliability. The technical solution of the present invention will be described in detail below.

[0072] The method of the present invention includes an acquisition step, a processing step, and an output step. In the acquisition step, this method acquires multi-dimensional transaction data in the financial trading system, and this data includes transaction object, transaction amount, and transaction time information. These multi-dimensional data provide a comprehensive information basis for subsequent anomaly recognition.

[0073] In the processing step, this method first performs data cleaning and feature engineering based on the acquired multi-dimensional transaction data to obtain preprocessed transaction data. This step aims to improve data quality, remove noise and redundant information, and at the same time construct a more effective feature representation. Next, this method constructs an isolation forest model according to the preprocessed transaction data. The isolation forest algorithm is an efficient anomaly detection algorithm, especially suitable for processing high-dimensional data. Finally, this method calculates the anomaly score of the transaction data based on the constructed isolation forest model.

[0074] In the output step, this method determines abnormal transactions according to the calculated anomaly scores and generates corresponding abnormal transaction reports. This step presents the anomaly detection results in an easy-to-understand and operable form, facilitating subsequent risk management and decision-making.

[0075] Furthermore, the acquisition step of the present invention specifically includes three sub-steps. First, extract transaction records from the database of the financial trading system. Here, the database can be a relational database such as MySQL or Oracle, or a non-relational database such as MongoDB. Second, perform a preliminary screening on the extracted transaction records to remove records containing obvious errors or missing values. This step helps to improve the efficiency and accuracy of subsequent processing. Finally, convert the screened transaction records into a standardized data format. The standardized data format can be CSV, JSON, or other formats suitable for machine learning processing.

[0076] Preferably, in one embodiment of the present invention, the extraction of transaction records can be performed using SQL queries or NoSQL query statements, and filtered according to conditions such as time range, transaction type, etc. For example, all transaction records within the most recent 30 days can be extracted, or only transaction records of a specific type (such as large transfers) can be extracted.

[0077] In the preliminary screening stage, some simple rules can be set to identify obvious errors or outliers. For example, records with negative transaction amounts or exceeding a certain reasonable upper limit (such as 100 million yuan) can be marked as abnormal and removed from the dataset. For missing values, if the missing proportion is below a certain threshold (such as 5%), these records can be directly deleted; if the missing proportion is relatively high, more complex processing methods may be required, such as mean filling or advanced interpolation techniques.

[0078] The data cleaning steps of the present invention specifically include three sub-steps. First is to detect and process outliers, replacing values outside the normal range with the statistical mean or median. The "normal range" here can be determined by statistical methods, such as using the three-sigma rule or the quartile rule. Second is to fill in missing values, estimating the missing data items using interpolation methods or the nearest neighbor method. Finally is to remove duplicate records, retaining the latest or most complete transaction information.

[0079] Preferably, in one embodiment of the present invention, the following method can be used to process outliers: for numerical features (such as transaction amounts), the mean μ and standard deviation σ can be calculated, and then values outside the range [μ - 3σ, μ + 3σ] are regarded as outliers. The outliers can be replaced with the boundary values of this range, or replaced with the median. This method can be expressed as:

[0080]

[0081] where x is the original value, and x cleaned is the value after cleaning.

[0082] For filling in missing values, appropriate methods can be selected according to the characteristics of the data. For example, for time series data, linear interpolation can be used; for categorical data, the mode can be used for filling; for continuous data, the K-nearest neighbor (KNN) algorithm can be used for filling. The basic idea of KNN filling is to find the K samples most similar to the sample where the missing value is located, and then use the mean or median of these K samples to fill in the missing value.

[0083] In a financial trading system, data cleaning and abnormal transaction detection are two key but different processes. Correctly handling the relationship between the two is crucial for improving the accuracy and reliability of the system. The following is a specific description of how to coordinate these two processes:

[0084] 1. "Outliers" in data cleaning: Such outliers generally refer to situations where data values do not conform to actual physical laws or business logics due to various reasons (such as input errors, transmission problems, or system failures). For example, negative transaction amounts or large transactions beyond a reasonable range.

[0085] 2. "Abnormal transactions" to be identified: This refers to those actual transaction patterns that may involve fraud or other risks. This is the core concern of the present invention, aiming to identify potential risky behaviors by analyzing transaction patterns.

[0086] To better coordinate these two processes and ensure that no real abnormal transactions are missed due to data cleaning, the following measures can be taken:

[0087] Differentiated outlier handling: Classify outliers into obvious error outliers and suspicious outliers.

[0088] Obvious error outliers can be directly corrected or replaced.

[0089] Suspicious outliers need to retain the original information and create a marker field, while adding the corrected value for subsequent analysis.

[0090] Establish an outlier tracking mechanism that not only records the corrected value but also saves the information of the original outliers for in-depth analysis.

[0091] During the application of anomaly detection algorithms such as the Isolation Forest model, introduce a cross-validation step to compare the difference in anomaly scores calculated using the original value and the corrected value. If the difference exceeds a preset threshold, mark the transaction as "requiring special review".

[0092] In the manual review process, give priority to those transactions with a large difference in anomaly scores before and after outlier correction to ensure that these potentially overlooked abnormal transactions are promptly reviewed.

[0093] Adjust the anomaly score strategy and use the method of `max(s(x_original), s(x_corrected))` to determine the final anomaly score, which can effectively avoid the occurrence of missed anomaly detections caused by data correction.

[0094] Through the above improvement measures, while ensuring data quality, the impact of data cleaning on abnormal transaction detection can be reduced, thereby achieving more accurate and reliable identification of abnormal transactions. This method not only guarantees the consistency and accuracy of data but also improves the risk resistance of the entire system, helping financial institutions to more effectively prevent various transaction risks.

[0095] When removing duplicate records, it can be judged based on the unique identifier of the transaction (such as the transaction ID). If duplicate records are found, the most recent record or the record with the most complete information is preferably retained. This can be achieved by comparing the timestamps of the records and the number of non-empty fields.

[0096] Through the above detailed data cleaning steps, the present invention can significantly improve the data quality, laying a solid foundation for subsequent feature engineering and model construction. High-quality data can not only improve the accuracy of anomaly detection, but also reduce the complexity of model training and improve the efficiency of the overall system.

[0097] The feature engineering steps of the present invention specifically include three sub-steps. First is to construct time-related features, including periodic indicators such as the hour, week, month of the transaction. Second is to calculate the transaction frequency features, such as the number of transactions and the total transaction amount within a unit time. Finally is to generate the historical behavior features of the transaction object, such as the statistical pattern of transactions in a certain past period.

[0098] In a preferred embodiment of the present invention, the construction of time-related features can be more refined. For example, a day can be divided into multiple time periods, such as early morning (0 - 6 o'clock), morning (6 - 12 o'clock), afternoon (12 - 18 o'clock) and evening (18 - 24 o'clock), and a corresponding time period label is assigned to each transaction. In addition, the difference between weekdays and weekends, as well as the impact of holidays, can also be considered. These time features can capture the periodic patterns of transaction behaviors and help identify abnormal transactions that do not conform to the normal transaction habits of users.

[0099] For the transaction frequency features, the present method can use the sliding window technique to calculate. For example, three time windows of 1 hour, 24 hours and 7 days can be set, and the number of transactions and the total transaction amount within these windows are calculated respectively. Such multi-scale frequency features can help the system identify short-term and long-term abnormal behavior patterns. Specifically, the frequency features can be calculated using the following formula:

[0100]

[0101] where t is the current time, w is the size of the time window, I(transaction i ) is the indicator function, which is 1 when there is a transaction at time i and 0 otherwise, and amount i is the transaction amount at time i.

[0102] When generating the historical behavior features of a trading object, the method of the present invention can consider multiple dimensions. For example, statistics such as the average transaction amount within the past 30 days, the standard deviation of the number of transactions, and the maximum single transaction amount can be calculated. In addition, a behavior portrait of the trading object can be constructed, such as the common trading time, common trading location, common trading type, etc. These historical behavior features can establish the normal behavior pattern of the trading object, making it easier to identify abnormal transactions that deviate from this pattern.

[0103] The steps of constructing the isolation forest model of the present invention specifically include three sub-steps. First is to randomly select a subset of training samples. Second is to construct a decision tree for each subsample, where each node randomly selects a feature and a splitting point. Finally, repeat the construction of multiple decision trees to form an isolation forest.

[0104] In an embodiment of the present invention, the method of randomly selecting a subset of training samples can adopt the sampling without replacement method. Assuming the total number of samples is N, the size of each subsample can be set to the square root of N, that is This can reduce the computational complexity while ensuring the representativeness of the subsamples.

[0105] For the construction of the decision tree, this method adopts the strategy of random feature selection and random splitting point selection. Specifically, at each node, a feature f is randomly selected, and then a splitting point p is randomly selected within the value range of this feature. The sample set S is divided into two subsets S l and S r :

[0106] S l ={x∈S|x f <p},

[0107] S r ={x∈S|x f ≥p},

[0108] where x f represents the value of sample x on feature f.

[0109] This random selection strategy makes the isolation forest model particularly sensitive to outliers. Outliers usually have rare feature values, so they are more likely to be "isolated", that is, to reach the leaf node earlier in the decision tree.

[0110] Preferably, in the method of the present invention, the number of decision trees can be set between 100 and 1000. The more the number of trees, the better the stability of the model, but the higher the computational complexity. In practical applications, methods such as cross-validation can be used to determine the optimal number of trees.

[0111] In the process of constructing the decision tree of the present invention, there are also three specific steps. First, set the maximum tree depth and the minimum sample number threshold. Second, when the maximum tree depth is reached or the number of samples in a node is less than the minimum threshold, set the current node as a leaf node. Finally, calculate the anomaly score of each leaf node.

[0112] In a preferred embodiment of the present invention, the maximum tree depth can be set to This is because in an ideal situation, the depth of a balanced binary tree should be at the logarithmic level of the number of samples. The minimum sample number threshold can be set to 1 or 2 to ensure that each sample can be "isolated" sufficiently.

[0113] The anomaly score of a leaf node can be calculated by the following formula:

[0114]

[0115] where h(x) is the path length for the sample x to reach the leaf node, E(h(x)) is the expected value of h(x), and c(n) is the average path length given n samples, which can be approximated by the following formula:

[0116] c(n) = 2H(n - 1) - (2(n - 1) / n),

[0117] where H(i) is the i-th harmonic number, which can be approximated as ln(i) + 0.5772156649 (Euler's constant).

[0118] The steps for the present invention to calculate the anomaly score of transaction data specifically include four sub-steps. First, input the transaction data to be detected into the isolation forest model. Second, calculate the path length for the data sample to reach the leaf node in each decision tree. Third, calculate the average path length based on the path lengths of all decision trees. Finally, use the predefined anomaly calculation formula according to the average path length to obtain the final anomaly score.

[0119] In the specific implementation of this method, for each transaction data x to be detected, we need to pass it through each decision tree in the isolation forest. In each tree, we record the path length h i (x) from the root node to the leaf node. Then, we can calculate the average path length:

[0120]

[0121] where t is the total number of decision trees.

[0122] Finally, we use the anomaly calculation formula mentioned above to obtain the final anomaly score. The range of the anomaly score is [0, 1], where a value close to 1 indicates a higher likelihood of being an anomaly.

[0123] Preferably, the method of the present invention can set an anomaly threshold, such as 0.6 or 0.7. When the anomaly score of a certain transaction exceeds this threshold, the system will mark it as a potential abnormal transaction and trigger a further review process.

[0124] Through the above detailed steps, the method of the present invention can effectively identify abnormal transactions in the financial trading system. This method based on the isolation forest algorithm has the advantages of high computational efficiency, strong applicability to high-dimensional data, and does not require prior knowledge of abnormal patterns, and is particularly suitable for real-time anomaly detection in large-scale financial trading systems.

[0125] The method of the present invention also includes a model update step. This step specifically includes two aspects: retraining the isolation forest model regularly or when new abnormal patterns are detected; and adding the newly identified abnormal transaction data to the training set to improve the detection ability of the model.

[0126] In a preferred embodiment of the present invention, the model update can be carried out in a sliding window manner. For example, a fixed-size time window can be set, such as the transaction data in the most recent 90 days. Every certain period (such as weekly or monthly), the system will retrain the model using the latest data within this sliding window. This way can enable the model to adapt to changes in trading patterns in a timely manner and improve the accuracy of anomaly detection.

[0127] In addition, the present method also adopts the idea of incremental learning. When the system detects new abnormal transactions and after manual confirmation, these abnormal samples will be added to the training set. In this way, the model can learn new abnormal patterns and continuously improve its detection ability. Specifically, an abnormal sample library can be maintained, and the samples in this library are regularly mixed with normal transaction samples for retraining of the model.

[0128] Preferably, the method of the present invention can use the following formula to determine whether to perform model update:

[0129]

[0130] where N new is the number of newly added transactions, N total is the total number of transactions, F false is the number of false alarms, F total is the total number of alarms, T is the time since the last update (in days). α, β, and γ are weight parameters that can be adjusted according to specific business requirements. When U exceeds a preset threshold (such as 0.5), model update is triggered.

[0131] The method of the present invention further includes a manual review step. This step includes three specific operations: manually reviewing transactions with an anomaly score exceeding a preset threshold; updating the abnormal transaction library and the normal transaction library according to the manual review results; and adjusting the threshold of the anomaly score based on the updated transaction library.

[0132] In an embodiment of the present invention, the manual review process can be designed as a multi-level review mechanism. For example, when the anomaly score of a transaction is between 0.7 and 0.8, it is reviewed by a junior reviewer; when the score is between 0.8 and 0.9, it is reviewed by a senior reviewer; when the score exceeds 0.9, multiple people need to review it together. This grading mechanism can balance the efficiency and accuracy of the review.

[0133] The feedback of the review results is crucial for improving the system performance. This method adds the abnormal transactions confirmed by the review to the abnormal transaction library, and adds the misreported transactions to the normal transaction library. These two libraries are not only used for the retraining of the model, but also for dynamically adjusting the threshold of the anomaly score.

[0134] Preferably, the method of the present invention can use the following formula to dynamically adjust the threshold of the anomaly score:

[0135]

[0136] Where TP is the number of true positives (correctly identified abnormal transactions), FP is the number of false positives (false alarms), target precision is the target precision rate (such as 0.95), and λ is the learning rate (such as 0.01). This formula can automatically adjust the threshold according to the actual performance of the system to achieve the expected precision rate.

[0137] The method of the present invention further includes an alarm and handling step. This step specifically includes two aspects: when an abnormal transaction is detected, an alarm is immediately sent to the system administrator; and according to the type and severity of the abnormal transaction, preset risk control measures are automatically executed, such as suspending the transaction, restricting the transaction amount, or requiring additional identity verification.

[0138] In a preferred embodiment of the present invention, the alarm mechanism can be designed to be multi-channel and multi-level. For example, for minor anomalies, the system can send an email notification; for moderate anomalies, the system will send a text message and an email; for severe anomalies, the system will trigger a phone alarm and push it to the mobile App. The alarm information should include key information of the transaction, such as the transaction amount, time, location, involved accounts, etc., as well as the anomaly score and a preliminary analysis of the anomaly cause.

[0139] The automatic risk control measures of this method are an important innovation point. The system can automatically take corresponding risk control measures according to pre-set rules. For example, when detecting a possible account theft, the system can temporarily freeze the account; when detecting an abnormal large transfer, the system can require additional identity verification or manual review; when an abnormal transaction frequently occurs for a certain merchant, the system can reduce the transaction limit of that merchant.

[0140] Preferably, the method of the present invention can use a decision tree or a rule engine to implement the selection of automatic risk control measures. The decision-making process can consider multiple factors, such as the anomaly score, transaction amount, account history, transaction type, etc. For example:

[0141]

[0142]

[0143] In this way, the method of the present invention can quickly respond when detecting an abnormal transaction, minimizing potential financial risks. At the same time, the automated risk control measures can also reduce the pressure of manual review and improve the operating efficiency of the entire system.

[0144] Generally speaking, the method for identifying anomalies in a financial trading system based on the isolation forest algorithm provided by the present invention constructs a comprehensive, efficient, and reliable anomaly transaction identification system through multi-dimensional data processing, an efficient anomaly detection algorithm, a dynamic model update mechanism, manual review feedback, and intelligent risk control measures. This method can not only accurately identify various types of abnormal transactions, but also continuously learn and adapt to new trading patterns and fraud means, providing a powerful risk management tool for financial institutions.

[0145] To verify the effectiveness and superiority of the method for identifying anomalies in a financial trading system based on the isolation forest algorithm of the present invention, we designed a simulation experiment to simulate a credit card trading system of a large commercial bank. This system processes approximately 1 million transactions per day and involves 500,000 users. We selected the transaction data for 30 consecutive days as the experimental data set, which includes normal transactions and artificially implanted abnormal transactions.

[0146] Example 1 adopts the method of the present invention, including multi-dimensional data processing, the isolation forest algorithm, dynamic model update, manual review feedback, and intelligent risk control measures.

[0147] Comparative Example 1 adopts a traditional rule-based anomaly detection method, using fixed thresholds and rules to identify abnormal transactions.

[0148] Comparative Example 2 adopts an anomaly detection method based on support vector machines (SVM), using historical data to train an SVM model to identify abnormal transactions.

[0149] We selected the following metrics to evaluate the performance of various methods:

[0150] 1. Detection Precision: The proportion of correctly identified abnormal transactions among all transactions identified as abnormal.

[0151] 2. Detection Recall: The proportion of correctly identified abnormal transactions among all actual abnormal transactions.

[0152] 3. F1 Score: The harmonic mean of precision and recall, comprehensively reflecting the performance of the model.

[0153] 4. False Positive Rate: The proportion of normal transactions misjudged as abnormal among all normal transactions.

[0154] 5. Average Detection Time: The average abnormal detection time for a single transaction.

[0155] 6. Model Update Time: The time required for a complete model update.

[0156] Detection Method: We used the cross-validation method to divide the 30-day data into a training set (the first 20 days) and a test set (the last 10 days). For Example 1 and Comparative Example 2, we used the training set to initialize and train the model, and then evaluated the performance on the test set. For Comparative Example 1, we set the rules and thresholds according to the experience of domain experts.

[0157] The following are the performance comparison results of each method on the test set:

[0158] Index Example 1 Comparative Example 1 (rule-based) Comparative Example 2 (SVM) Detection accuracy 95.8% 82.3% 89.1% Detection recall 93.2% 76.5% 85.7% F1 score 94.5% 79.3% 87.3% False alarm rate 0.7% 3.2% 1.9% Average detection time 15 ms 8 ms 25 ms Model update time 2.5 hours N / A 6 hours

[0159] As can be seen from the above results, the method of the present invention (Example 1) is significantly superior to the traditional method and other machine learning methods in most metrics. The specific analysis is as follows:

[0160] 1. Detection Precision and Recall: The method of the present invention reaches more than 90% in these two key metrics, far exceeding the other two methods. This means that this method can not only accurately identify abnormal transactions, but also capture most of the abnormal situations, greatly reducing financial risks.

[0161] 2. F1 Score: The F1 score of the method of the present invention reaches 94.5%, which is 15.2 and 7.2 percentage points higher than that of Comparative Example 1 and Comparative Example 2 respectively. This shows that this method achieves a good balance between precision and recall, and has the optimal overall performance.

[0162] 3. False alarm rate: The false alarm rate of the method of the present invention is only 0.7%, far lower than that of the other two methods. In a financial trading system, a low false alarm rate is crucial because false alarms may cause normal transactions to be blocked, affecting user experience and banking operations.

[0163] 4. Average detection time: Although the method of the present invention (15 ms) is slightly slower than the rule-based method (8 ms), it is still much faster than the SVM method (25 ms). Considering the high accuracy and low false alarm rate of this method, this slight time trade-off is completely acceptable.

[0164] 5. Model update time: The method of the present invention can complete model update within 2.5 hours, while the SVM method requires 6 hours. This means that this method can adapt to new trading patterns and fraud means faster, improving the real-time performance and adaptability of the system.

[0165] The superiority of the method of the present invention is mainly reflected in the following aspects:

[0166] 1. Multi-dimensional data processing: Through comprehensive feature engineering, this method can capture more subtle abnormal patterns, which is difficult to achieve by traditional rule-based methods.

[0167] 2. Advantages of the Isolation Forest algorithm: Compared with SVM, the Isolation Forest algorithm has more advantages in dealing with high-dimensional data and large-scale data sets, which is reflected in higher detection accuracy and shorter detection time.

[0168] 3. Dynamic model update: This method can incorporate new trading patterns in a timely manner, keeping the model in the best state all the time. This is particularly important when dealing with rapidly changing financial fraud means.

[0169] 4. Low false alarm rate: While maintaining a high detection rate, this method significantly reduces the false alarm rate. This not only improves the reliability of the system but also reduces the unnecessary workload of manual review.

[0170] 5. Real-time performance: Despite using a more complex algorithm, this method still maintains a fast detection speed and can meet the requirements of a real-time trading system.

[0171] In summary, the method for identifying anomalies in a financial trading system based on the Isolation Forest algorithm provided by the present invention performs excellently in terms of accuracy, comprehensiveness, real-time performance, and scalability. This method can effectively identify various types of abnormal transactions, while maintaining a low false alarm rate and a high processing speed, providing a powerful and flexible risk management tool for financial institutions. In the face of increasingly complex financial fraud means, the adaptive ability and high efficiency of this method can bring significant security improvements and potential economic benefits to financial institutions.

[0172] It should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An abnormal identification method for a financial trading system based on the isolation forest algorithm, characterized in that, It includes: The obtaining step includes: Obtain multi-dimensional transaction data in the financial trading system, where the multi-dimensional transaction data includes transaction objects, transaction amounts, and transaction time information; The processing step includes: Based on the multi-dimensional transaction data, perform data cleaning and feature engineering to obtain preprocessed transaction data; According to the preprocessed transaction data, construct an isolation forest model; Based on the isolation forest model, calculate the anomaly score of the transaction data; The output step includes: According to the anomaly score, determine abnormal transactions and generate an abnormal transaction report.

2. The method according to claim 1, wherein The obtaining step specifically includes: Extract transaction records from the database of the financial trading system; Perform a preliminary screening on the transaction records to remove records containing obvious errors or missing values; Convert the screened transaction records into a standardized data format.

3. The method according to claim 1, wherein The data cleaning specifically includes: Detect and process outliers, and replace values outside the normal range with the statistical mean or median; Fill in missing values, and estimate missing data items using interpolation or the nearest neighbor method; Remove duplicate records and retain the latest or most complete transaction information.

4. The method according to claim 1, wherein The feature engineering specifically includes: Construct time-related features, including periodic indicators such as the hour, week, and month of the transaction; Calculate transaction frequency features, such as the number of transactions and the total transaction amount within a unit of time; Generate historical behavior features of the transaction object, such as the statistical pattern of transactions in a certain past period.

5. The method according to claim 1, wherein The steps of constructing the isolation forest model specifically include: Randomly select a subset of training samples; Construct a decision tree for each subsample, where each node randomly selects a feature and a splitting point; Repeat the construction of multiple decision trees to form an isolation forest.

6. The method according to claim 5, wherein During the process of constructing the decision tree: Set the maximum tree depth and the minimum sample number threshold; When the maximum tree depth is reached or the number of samples at a node is less than the minimum threshold, set the current node as a leaf node; Calculate the anomaly score of each leaf node.

7. The method according to claim 1, characterized in that The steps of calculating the anomaly score of the transaction data specifically include: Input the transaction data to be detected into the isolation forest model; Calculate the path length of the data sample reaching the leaf node in each decision tree; Based on the path lengths of all decision trees, calculate the average path length; According to the average path length, use a predefined anomaly score calculation formula to obtain the final anomaly score.

8. The method according to claim 1, characterized in that, It also includes a model update step: Periodically or when a new abnormal pattern is detected, retrain the isolation forest model; Add the newly identified abnormal transaction data to the training set to improve the detection ability of the model.

9. The method according to claim 1, wherein It also includes an artificial review step: Conduct an artificial review on transactions with an anomaly score exceeding the preset threshold; According to the results of the artificial review, update the abnormal transaction database and the normal transaction database; Based on the updated transaction database, adjust the threshold of the anomaly score.

10. The method according to claim 1, characterized in that, It also includes an alarm and handling step: When an abnormal transaction is detected, immediately send an alarm to the system administrator; According to the type and severity of the abnormal transaction, automatically execute preset risk control measures, such as suspending the transaction, restricting the transaction amount, or requiring additional identity verification.

Citation Information

Cited By

  • Big data intelligent cleaning method and system based on machine learning

    CN122019990A

  • Machine learning-based big data intelligent cleaning method and system

    CN122019990B