Financial statement fraud automatic detection method based on improved isolated forest algorithm
By improving the isolation forest algorithm combined with incremental learning and graph convolutional networks, the inefficiency and misjudgment problems of traditional financial auditing methods in large-scale and complex financial data are solved, and efficient and accurate automatic detection of financial fraud is achieved, adapting to the dynamic changes of financial data and improving the accuracy and real-time performance of detection.
Patent Information
- Application Number
- CN202510869959.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional financial auditing methods are inefficient and prone to missed detections when faced with large-scale and complex financial data. The isolation forest algorithm has a high misjudgment rate when processing complex financial data, making it difficult to capture fraud hidden in financial statements. Existing technologies also face challenges in feature selection, data preprocessing, and dynamic threshold adjustment.
Combining the improved isolation forest algorithm, incremental learning mechanism and graph convolutional network, through real-time monitoring and anomaly detection of financial statement data, the graph convolutional network is used to extract the dependencies between financial data, and combined with the dynamic threshold adjustment mechanism, an improved isolation forest model is constructed to achieve efficient and accurate identification of financial data.
It improves the accuracy and real-time performance of financial fraud detection, enables automated financial monitoring in large-scale enterprises and listed companies, reduces false positives and missed reports, and enhances the adaptability and intelligence of the system.
Smart Images

Figure CN120807182A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of financial audit and fraud detection, and particularly relates to a financial statement fraud automatic detection method based on an improved isolation forest algorithm. BACKGROUND
[0002] With the expansion of enterprise scale and the increase of market complexity, traditional financial audit methods have gradually been unable to effectively cope with the increasingly complex financial fraud detection tasks. The traditional method mainly relies on manual inspection and manual experience to analyze each item of the abnormal points in the financial statements. Although this method can be applied to the financial statements of small-scale enterprises, it is inefficient and prone to missed detection when facing large-scale enterprises and complex financial structures. Financial fraud means is also evolving, and conventional financial audit methods often fail to detect hidden fraud, so there is an urgent need for more efficient and intelligent financial fraud detection methods.
[0003] In recent years, machine learning-based anomaly detection methods have gradually become a research hotspot in financial fraud detection. Isolation Forest algorithm, as an efficient anomaly detection algorithm, is widely used in the financial field due to its efficient computing performance and strong anomaly detection ability. However, the traditional Isolation Forest algorithm has some problems when dealing with complex financial data. Isolation Forest algorithm identifies abnormal points by building multiple decision trees, but it is prone to misjudgment when facing complex financial statement data, especially in the case of high data dimension and complex features. In addition, the splitting mechanism of Isolation Forest is based on random selection of features, which lacks deep modeling of the potential dependencies between features, making it underperform in capturing hidden fraud behaviors in financial statements.
[0004] To make up for these shortcomings, recent research has focused on improving the Isolation Forest algorithm and combining it with Graph Convolutional Network (GCN) to extract dependencies between financial data. The ability of the graph convolutional network can better capture the internal relationships between various data in the financial statements. In addition, the introduction of incremental learning mechanism enables the model to update itself when new financial data arrives, constantly adapting to changing financial data, thereby improving the real-time detection capability of the model. However, existing technologies still face challenges in feature selection, data preprocessing, and efficient dynamic adjustment of threshold values. For example, how to reasonably adjust the threshold value of the abnormal score based on historical data and real-time feedback information to ensure the accuracy and sensitivity of the detection is still a problem to be solved. In addition, financial data often contains missing values and noise, and how to handle these problems in the graph convolutional network is still a technical difficulty. Although existing methods have improved the accuracy of anomaly detection by enhancing the functionality of the Isolation Forest algorithm, further optimization is still needed to cope with the complexity and dynamic changes of large-scale enterprise and listed company financial statements.
[0005] The financial statement fraud automatic detection method based on the improved isolation forest algorithm combines incremental learning, graph convolution network and dynamic threshold adjustment technology, aiming to make up for the shortcomings in the prior art. The method can more accurately identify potential fraud behaviors in financial statements, and can adapt to data changes in real time, improving financial monitoring and auditing efficiency. SUMMARY
[0006] One object of the present application is to provide a financial statement fraud automatic detection method based on an improved isolation forest algorithm. The present application makes full use of machine learning, anomaly detection and graph convolution network technology, combines an incremental learning mechanism, and automatically identifies potential financial fraud behaviors through real-time monitoring and anomaly detection of financial statement data. The method improves the adaptability of the model to complex financial data by improving the isolation forest algorithm, extracts the dependency relationship between financial data using the graph convolution network, and improves the accuracy and real-time performance of the detection by combining the dynamic threshold adjustment mechanism. The method has the advantages of high efficiency, accuracy and real-time updating, and can realize automatic financial monitoring in large-scale enterprises and listed companies, improving the accuracy and efficiency of financial fraud detection.
[0007] The financial statement fraud automatic detection method based on the improved isolation forest algorithm according to the embodiment of the present application comprises the following steps:
[0008] S1, collecting financial statement data and preprocessing the financial statement data;
[0009] S2, performing feature engineering and weighted feature selection on the preprocessed financial statement data, and using a graph convolution network to extract the dependency relationship between financial data;
[0010] S3, combining the dependency relationship between financial data, updating the isolation forest model using an incremental learning mechanism, and adjusting the anomaly score threshold value through a dynamic threshold adjustment mechanism to build an improved isolation forest model;
[0011] S4, training the improved isolation forest model using a historical financial statement data set, and evaluating the performance of the improved isolation forest model using cross-validation;
[0012] S5, using the trained improved isolation forest model to detect outliers of new financial statement data, and generating a risk assessment report;
[0013] S6, integrating the improved isolation forest algorithm into a real-time financial data monitoring system to monitor the financial statements of listed companies and enterprises in real time.
[0014] Optionally, the financial statement data includes an income statement, a balance sheet and a cash flow statement, and the financial statement data is preprocessed, and the preprocessing steps include data cleaning, missing value filling, standardization and normalization of specific financial indicators.
[0015] Optionally, the S2 specifically includes:
[0016] S21. Performing feature engineering on the preprocessed financial statement data, wherein the feature engineering step includes extracting original financial features from the financial statements, transforming and deriving the original financial features, calculating financial ratios closely related to financial health status, and generating new financial features;
[0017] S22. Perform weighted feature selection on the new financial features. The weighted feature selection strategy is based on the financial importance of the features and the influence of historical fraud behavior patterns. Specifically, the weighted value W of each financial feature is calculated. i To assess the susceptibility of these financial characteristics to financial fraud;
[0018] S23. Use a graph convolutional network to extract the dependencies between financial data. The graph convolutional network uses each financial item in the financial statement as a node. The dependencies between the nodes constitute the edges of the graph. The feature representation of each financial item is transmitted and updated through the graph convolutional network. The adjacency matrix is used to represent the dependencies between the financial items. The embedded feature representation H of each node learned by the graph convolutional network (l+1) ;
[0019] S24. Combining the financial data dependency graph obtained by the graph convolutional network with the weighted feature selection results, further screening all financial features to generate a final feature representation. The feature representation includes the dependency relationships between various features in the financial statements.
[0020] Optionally, the improved isolation forest model includes the following: first, the pre-processed financial statement data is combined with weighted feature selection and graph convolutional networks to extract the dependencies between financial data; in the process of building the tree structure, the features and splitting points are determined by weighted feature selection and the dependency graph of the financial data; the improved isolation forest model adopts an incremental learning mechanism, and each time new financial data arrives, the model updates the existing parameters without retraining the entire model. In addition, the improved isolation forest model introduces a dynamic threshold adjustment mechanism, which automatically adjusts the threshold of the anomaly score based on the historical financial statement data set and real-time feedback. In addition, the improved isolation forest model combines local and global anomaly detection strategies, and comprehensively considers the abnormality of each feature in the financial statement through the financial data dependency learned by the graph convolutional network.
[0021] Optionally, the S3 specifically includes:
[0022] S31, update the improved Isolation Forest model using the incremental learning mechanism, each decision tree in the improved Isolation Forest model can be fine-tuned according to the new financial statement data. In the process of building each decision tree, the new financial statement data is used to update the split point and depth of the existing decision tree;
[0023] S32, under the incremental learning mechanism, the improved Isolation Forest model updates the weight matrix in the Isolation Forest by calculating the impact of the new financial statement data on the features;
[0024] S33, through the dynamic threshold adjustment mechanism, automatically adjust the threshold of the anomaly score according to the historical financial statement data set and real-time feedback information; the historical financial statement data set includes past financial statement data and known abnormal behavior patterns, which is used to provide the model with normal and abnormal distribution characteristics of financial data; real-time feedback information includes the results of anomaly detection, false positive rate and false negative rate of the model, and results from manual review and user feedback;
[0025] S34, when performing anomaly detection, each financial statement data point will calculate an anomaly score, if the anomaly score exceeds the current threshold, the data point is determined to be abnormal, wherein the threshold is dynamically adjusted according to the historical financial statement data set and real-time feedback information.
[0026] S35, based on the labeled abnormal data points and normal data points in the historical financial statement data set, calculate the score distribution of the abnormal data points; then, according to the newly added abnormal score and normal score in the real-time data, calculate the threshold adjustment amount ΔT;
[0027] S36, when the real-time feedback result shows that the false positive rate and the false negative rate are too high, adjust the threshold to make the anomaly detection more strict, when the false positive and false negative are reduced, adjust the threshold to relax the standard.
[0028] Optionally, the dynamic threshold adjustment mechanism takes into account the sensitivity of the model to the anomaly score, and the updated weight matrix enables the improved Isolation Forest model to adjust the anomaly score according to the new feature distribution.
[0029] Optionally, the S4 specifically comprises:
[0030] S41, use the historical financial statement data set to train the improved Isolation Forest model, the historical financial statement data set includes a plurality of labeled financial data points, including normal data points and abnormal data points, wherein the anomaly score S i of each data point is calculated as the input data of the training.
[0031] S42, the historical financial statement dataset is divided into K subsets, and in the cross-validation process, K-1 subsets are selected as the training set each time, and the remaining one subset is selected as the validation set, the process is repeated K times to ensure that each subset participates in model evaluation as a validation set.
[0032] S43, in each training process, the improved isolation forest model optimizes model parameters based on the training set, the improved isolation forest model minimizes errors and optimizes anomaly detection capability by adjusting the depth of the decision tree, the split point and the feature selection weight; after each training, the improved isolation forest model performs performance evaluation on the validation set, and the anomaly score S i is calculated, and compared with the set reference threshold value, to evaluate whether the data point is correctly marked as an anomaly.
[0033] S44, after the cross-validation ends, the average performance indicators of all K validations are calculated to evaluate the overall performance of the improved isolation forest model, the performance indicators include accuracy Acc, recall Rec, precision Pre, F1 score F1,
[0034] S45, based on the average performance indicators of cross-validation, the hyperparameters of the model are adjusted, including the number of trees, the depth of the tree and the feature selection weight.
[0035] Optionally, the S5 specifically comprises:
[0036] S51, using the trained improved isolation forest model to detect anomalies in new financial statement data, the financial statement data includes a plurality of financial data points, and the anomaly score S i of each data point is calculated as the basis for anomaly identification;
[0037] S52, the anomaly score S i of each data point is calculated, and compared with the set threshold value T, if S i >T, the data point is determined to be an anomaly. The threshold value T is dynamically adjusted according to the historical financial statement dataset and real-time feedback information;
[0038] S53, through a multi-dimensional anomaly score system, a comprehensive anomaly score S c is provided for each anomaly data point to represent the severity of the anomaly data point, and the anomaly data is classified into multiple levels, including slight anomaly, moderate anomaly and severe anomaly; the comprehensive anomaly score S c is used to determine the degree of anomaly of the data point, and each data point is assigned to the appropriate anomaly level according to the comprehensive anomaly score S c value;
[0039] S55, based on the detected abnormal data points, a risk assessment is performed using a risk level assessment mechanism based on model score, for each abnormal data point, according to the comprehensive abnormal score S c and the abnormal level, the risk level of the data point in the overall financial statements is determined; the risk assessment report includes the following contents: the abnormal score S i and the comprehensive abnormal score S c of each data point; the abnormal type and abnormal level of each data point;
[0040] Risk analysis of potential financial fraud behavior;
[0041] S56, a detailed risk assessment report is generated, which lists the abnormal values of each financial indicator in the financial statements, abnormal patterns and potential fraud risk analysis; the risk assessment report provides specific financial data point analysis and risk assessment suggestions according to the detected abnormalities and severity.
[0042] Optionally, the S6 specifically includes:
[0043] S61, the improved isolation forest algorithm is integrated into a real-time financial data monitoring system, which can receive and process real-time financial statement data from listed companies and enterprises;
[0044] S62, the real-time monitoring system uses an incremental learning mechanism, and the improved isolation forest model is automatically updated whenever new financial statement data arrives;
[0045] S63, in the real-time data processing process, the improved isolation forest model compares the abnormal score S i of each data point with the set dynamic threshold T, if S i >T, the data point is determined to be abnormal, and further determined to be a potential financial fraud behavior;
[0046] S64, when the monitoring system detects abnormal data, the real-time feedback mechanism automatically triggers the review mechanism of the securities regulatory authority, and provides detailed abnormal analysis and risk assessment report, helping the regulatory personnel to take review and investigation.
[0047] The beneficial effects of the present application are:
[0048] The improved isolation forest algorithm combined with the incremental learning mechanism and the graph convolution network provides a new automatic financial statement fraud detection method. First, the improved isolation forest algorithm can automatically update the model when new financial data arrives by using the incremental learning mechanism, ensuring that the financial data detection system can always adapt to the latest data changes. This method significantly improves the real-time adaptability of the system, avoiding the slow response of traditional static models to new changes in data, making fraud detection more sensitive and efficient.
[0049] The graph convolution network is used to extract the dependency between financial data, further enhancing the anomaly detection capability of the model. The data in the financial statements often depend on each other, and the graph convolution network can accurately capture these dependencies, optimizing the feature selection and splitting process. Unlike the traditional isolation forest algorithm which relies on random selection of features, the present invention can more accurately identify and classify abnormal patterns in financial data, improving the accuracy of anomaly detection.
[0050] The dynamic threshold adjustment mechanism introduced in the present invention allows the anomaly score to be flexibly adjusted according to historical data and real-time feedback, thereby improving the accuracy and robustness of the detection. Through multidimensional analysis and real-time adjustment of the anomaly score, the system can optimize the detection standard according to the characteristics of different enterprises and financial statements, reducing the false positive rate and the false negative rate. This mechanism makes the financial fraud detection system more adaptable and intelligent, capable of handling different types of financial statements.
[0051] In summary, the automatic financial statement fraud detection method provided by the present invention, through deep learning and real-time monitoring of financial data, not only improves the accuracy and efficiency of anomaly detection, but also enhances the adaptability and real-time response capability of the system, with significant technical advantages and wide application prospects. BRIEF DESCRIPTION OF DRAWINGS
[0052] The accompanying drawings are included to provide a further understanding of the present invention, and constitute a part of the specification, together with the embodiments of the present invention, to explain the present invention, and do not constitute a limitation on the present invention. In the drawings:
[0053] Figure 1 The overall flowchart of the automatic financial statement fraud detection method based on the improved isolation forest algorithm proposed by the present invention;
[0054] Figure 2 The structure diagram of the improved isolation forest algorithm model proposed by the present invention. DETAILED DESCRIPTION
[0055] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams that only schematically illustrate the basic structure of the present invention, and therefore only show the components related to the present invention.
[0056] Reference Figure 1 and Figure 2 The automatic detection method for financial statement fraud based on the improved isolation forest algorithm comprises the following steps:
[0057] S1, collecting financial statement data and preprocessing the financial statement data;
[0058] S2, performing feature engineering and weighted feature selection on the preprocessed financial statement data, and using a graph convolution network to extract the dependency relationship between financial data;
[0059] S3, combining the dependency relationship between the financial data, updating the isolation forest model using an incremental learning mechanism, and adjusting the abnormal score threshold value through a dynamic threshold adjustment mechanism to build an improved isolation forest model;
[0060] S4, training the improved isolation forest model using a historical financial statement data set, and evaluating the performance of the improved isolation forest model using cross-validation;
[0061] S5, using the trained improved isolation forest model to detect abnormal values in new financial statement data and generating a risk assessment report;
[0062] S6, integrating the improved isolation forest algorithm into a real-time financial data monitoring system to monitor the financial statements of listed companies and enterprises in real time.
[0063] In this embodiment, the automatic detection method for financial statement fraud based on the improved isolation forest algorithm is characterized in that the financial statement data includes a profit statement, a balance sheet, and a cash flow statement, and the financial statement data is preprocessed, the preprocessing step including data cleaning, missing value filling, standardization processing, and normalization processing of specific financial indicators.
[0064] In this embodiment, S2 specifically comprises:
[0065] S21, performing feature engineering on the preprocessed financial statement data, the feature engineering step comprising extracting original financial features from the financial statement and transforming and deriving them, calculating financial ratios closely related to financial health, such as asset-liability ratio, accounts receivable turnover ratio, and liquidity ratio, to generate new financial features. These new financial features are used for subsequent analysis and provide input for the graph convolution network.
[0066] S22, performing weighted feature selection on the extracted financial features, the weighting strategy being based on the financial importance of the features and the influence of historical fraud behavior patterns, specifically comprising calculating the weighted value of each financial feature to assess its sensitivity to financial fraud. The weighted value Wi The formula is as follows:
[0067]
[0068] wherein, is the standardized i-th financial feature value, and a i is the financial importance coefficient of the feature, which is learned from historical data.
[0069] S23, use a graph convolution network (GCN) to extract the dependency relationship between the financial data, which takes each financial item in the financial statement as a node, and the dependency relationship between the nodes constitutes the edge of the graph. The feature representation of each financial item is updated through the graph convolution network. The adjacency matrix The embedding feature representation H of each node learned by the graph convolution network (l+1) is updated by the following formula:
[0070]
[0071] wherein, H (l) is the feature representation after the l-th layer of graph convolution, is the normalized adjacency matrix of the graph, and W (l) is the weight matrix of the l-th layer, and σ is the activation function, usually ReLU function. Through the graph convolution operation, the network can learn the complex dependency relationship between the financial items, thereby generating a higher quality feature representation of each financial item.
[0072] S24, combine the financial data dependency graph obtained by the graph convolution network with the weighted feature selection result, further filter all financial features, and generate the final feature representation. The feature representation contains the dependency relationship between each feature in the financial statement, and preferentially retains the financial features that best reflect the fraud behavior. These final feature representations are used for subsequent isolation forest model training, thereby improving the accuracy and sensitivity of financial fraud detection.
[0073] The financial statement data is processed through a series of steps to improve the accuracy and sensitivity of financial fraud detection. First, the original financial features are extracted from the financial statements, and financial ratios related to financial health, such as asset-liability ratio, accounts receivable turnover ratio, etc., are calculated to generate new financial features. Next, a weighted feature selection method is used to assign weights to the features based on their importance and sensitivity to fraudulent behavior, thereby optimizing the feature selection process. Then, a graph convolution network (GCN) is used to extract the dependency relationships between financial data, treating each financial item in the financial statement as a node, updating the embedded feature representation of the node through an adjacency matrix and a weight matrix, and learning the complex dependency relationships between the financial items. Finally, the dependency graph extracted by the graph convolution network and the weighted feature selection results are combined to further filter the financial features, generating the final feature representation, which prioritizes the financial features that best reflect fraudulent behavior, for subsequent isolation forest model training to improve the accuracy and sensitivity of financial fraud detection.
[0074] In this embodiment, S3 specifically includes:
[0075] S31, extract the dependency relationships between the financial items in the financial data through a graph convolution network (GCN), which treats each financial item as a node in the graph, and the dependency relationships between the nodes are updated through an adjacency matrix The graph convolution operation transmits the information of each node through the adjacency matrix. The update formula of each node is:
[0076]
[0077] where H (l) is the node feature representation after the lth layer of graph convolution, is the normalized adjacency matrix of the graph, W (l) is the weight matrix of the lth layer, and σ is the activation function (such as ReLU). This process generates the embedded feature representation of each financial item, capturing the relationships between the financial items and used for subsequent model updates.
[0078] S32, update the improved isolation forest model using an incremental learning mechanism, which enables the model to locally update the existing model based on new data each time new financial data arrives, without the need to retrain the entire model. The incremental learning mechanism adjusts the parameters and structure of the model step by step to adapt to new data and ensure the efficiency and accuracy of the model. Specifically, the incremental learning mechanism includes the following:
[0079] Incremental tree updates: Each decision tree in the improved isolation forest model is fine-tuned based on the newly added data, rather than being completely rebuilt. During the construction of each decision tree, the new data is used to update the split point and depth of the existing tree. The incremental tree update formula is as follows:
[0080] T new =T old +ΔT;
[0081] Among them, T new is the updated tree structure, T old is the original tree structure, ΔT is the tree structure adjustment calculated based on the newly added data, which means that the split points and nodes of the tree are adjusted by the incremental data.
[0082] Weight matrix update: Under the incremental learning mechanism, the model updates the weight matrix W in the isolation forest by calculating the impact of new data on features. The specific formula for incremental update is:
[0083] W new =W old +ΔW;
[0084] Among them, W new is the updated model parameter, W old is the old model parameter, ΔW is the parameter increment calculated based on the new data, which indicates the importance of financial features in the new data.
[0085] S33. Automatically adjust the threshold for anomaly scoring based on historical data and real-time feedback through a dynamic threshold adjustment mechanism. Specifically, in each round of detection, the model calculates the score of anomaly data points and dynamically adjusts the threshold based on real-time feedback and anomaly manifestations in historical data. The purpose of this adjustment is to optimize the sensitivity and accuracy of detection and prevent excessive false positives or false negatives. The adjustment process is as follows:
[0086] Anomaly score and threshold: When performing anomaly detection, an anomaly score S is calculated for each financial statement data point. i If S i >T, the data point is judged as an anomaly, where T is the anomaly scoring threshold.
[0087] Threshold adjustment: Based on historical data analysis and combined with real-time data feedback, the threshold of anomaly scoring is dynamically adjusted. The adjustment formula is:
[0088] T new =T old +ΔT;
[0089] Among them, T new is the updated anomaly score threshold, T oldis the original threshold value, and AT is the adjustment amount. The threshold adjustment amount AT is calculated based on statistical analysis of historical abnormal data and real-time feedback information, which can make the model more adaptable to the current data environment.
[0090] Feedback mechanism: When the real-time feedback result shows that the false positive rate or false negative rate is too high, AT will increase, making the threshold adjustment more stringent; on the contrary, when the false positives and false negatives are reduced, AT will decrease, reducing the stringency of the threshold, thereby optimizing the precision of anomaly detection.
[0091] The application uses graph convolution network (GCN) to extract the dependency relationship between each financial item in the financial data, represents the dependency between nodes through adjacency matrix, and updates the feature representation of each node through graph convolution operation. This process generates embedded feature representation for each financial item to capture the relationship between financial items and is used for subsequent model updating. Through the incremental learning mechanism, the improved isolation forest model can locally update the model when new financial data arrives, avoiding retraining the entire model, ensuring efficiency and accuracy. Each decision tree adjusts the splitting point and depth of the tree according to the new data, while updating the weight matrix to reflect the influence of new data on features. In addition, the dynamic threshold adjustment mechanism combines historical data and real-time feedback to automatically adjust the threshold of abnormal score to optimize the sensitivity and accuracy of detection, preventing false positives and false negatives. When real-time feedback shows that false positives or false negatives are too high, the threshold will be adjusted to be more stringent, and vice versa, further improving the adaptability and accuracy of the model.
[0092] In this embodiment, the S4 specifically includes:
[0093] S41, using a historical financial statement dataset to train the improved isolation forest model, the historical financial statement dataset including a plurality of labeled financial statement data, including normal data and abnormal data. The abnormal score S of each financial statement data point will be calculated as the basis for training. i
[0094] S42, during the training process, a cross-validation method is used to evaluate the model, which divides the historical financial statement dataset into multiple subsets as training set and validation set. Specifically, the dataset is divided into K subsets, and K-1 subsets are selected as training set and the remaining one subset as validation set. This process is repeated K times, and different validation sets are selected each time to ensure that each subset can participate in evaluation as a validation set. After each training, the abnormal detection performance indicators of the model on the validation set are calculated, such as accuracy, recall rate, F1 score, etc.
[0095] S43. In each round of cross-validation, the improved isolation forest model uses the training set to adjust the model parameters and evaluates the performance of the model on the validation set. Specifically, for each abnormal point in the validation set, the model calculates its abnormal score S i and compare it with the baseline threshold to determine whether it is correctly identified as an abnormality.
[0096] S44. After the cross-validation process is completed, the average performance indicators of all rounds are calculated to evaluate the overall performance of the model. Specific evaluation indicators include but are not limited to accuracy Acc, recall Rec, precision Pre, F1 score F1, etc. Each performance indicator is calculated using the following formula:
[0097]
[0098] Among them, TP is the number of true positive examples (detected anomalies), TN is the number of true negative examples (normal data points), FP is the number of false positive examples (normal data points mistakenly identified as anomalies), and FN is the number of false negative examples (abnormal data points mistakenly identified as normal).
[0099] S45. Based on the results of cross-validation, select the best model parameters, such as the number of trees, tree depth, feature selection weights, etc., and perform final optimization on the model to ensure its performance stability in practical applications.
[0100] The present invention trains an improved isolation forest model by using a historical financial statement dataset, wherein the dataset contains financial statement data points marked as normal and abnormal, and the anomaly score of each data point is calculated and used as the basis for training. During the training process, a cross-validation method is used to evaluate the performance of the model. By dividing the dataset into multiple subsets, training and validation are repeated to ensure that each subset can participate in the evaluation as a validation set. After each round of training, the model calculates the anomaly detection performance indicators on the validation set, such as accuracy, recall rate, F1 score, etc., and evaluates the anomaly score of each data point. After the cross-validation is completed, the average performance indicators of all rounds are calculated to comprehensively evaluate the effect of the model, and the model parameters are optimized based on the results to ensure the stability and efficiency of the model in practical applications.
[0101] In this embodiment, the S5 specifically includes:
[0102] S51, using the trained improved isolation forest model to perform outlier detection on new financial statement data, wherein the financial statement data includes multiple financial data points, and the anomaly score S of each data point is i is calculated as the basis for anomaly identification.
[0103] S52. Calculate the abnormal score S for each data point iand compared with a set threshold T, if S i >T, the data point is determined to be abnormal. The threshold T is dynamically adjusted according to historical financial statement data sets and real-time feedback information, ensuring that the detection standard adapts to changing financial data.
[0104] S53, through a multi-dimensional abnormal score system, a comprehensive abnormal score S c is provided for each abnormal data point to represent the severity of its abnormality, and the abnormal data is classified into multiple levels (minor abnormality, moderate abnormality, severe abnormality). The calculation formula of the comprehensive abnormal score S c is as follows:
[0105] S c =αS i +βW;
[0106] Wherein, α and β are weighting coefficients, S i is the abnormal score, W is the feature weight, and the comprehensive abnormal score S c is used to determine the degree of abnormality of the data point.
[0107] S54, based on the comprehensive abnormal score S c , each data point is divided into different levels of abnormality, including: minor abnormality, moderate abnormality, severe abnormality. Each data point is assigned to the appropriate abnormality level according to its S c value, helping regulators more clearly identify and handle potential risks.
[0108] S55, based on the detected abnormal data points, risk assessment is carried out, and a risk level assessment mechanism based on model score is adopted. For each abnormal data point, according to its comprehensive abnormal score S c and abnormal level, the risk level of the data point in the overall financial statement is determined. The risk assessment report includes the following contents:
[0109] The abnormal score S i and the comprehensive abnormal score S c of each data point;
[0110] The abnormal type and abnormal level (minor, moderate, severe) of each data point;
[0111] Risk analysis of potential financial fraud behavior.
[0112] S56, generate a detailed risk assessment report, which lists the abnormal values of each financial indicator in the financial statement, abnormal patterns and potential fraud risk analysis. The report provides specific financial data point analysis and further risk assessment suggestions according to the detected abnormalities and their severity.
[0113] The application utilizes the trained improved isolation forest model to detect outliers on new financial statement data, and calculates the anomaly score of each data point, which is compared with the dynamically adjusted threshold to determine whether the data point is abnormal. Through the multi-dimensional anomaly scoring system, a comprehensive anomaly score is calculated for each abnormal data point to represent the severity of the anomaly, and the data points are divided into multiple levels such as slight anomaly, moderate anomaly and severe anomaly. Based on the comprehensive anomaly score and the anomaly level, a risk assessment report is generated, which includes the anomaly score, comprehensive anomaly score, anomaly type and anomaly level of each data point, and risk analysis of potential financial fraud behavior. This method helps regulators to clearly identify and handle potential risks, and provides strong support for further review and decision-making.
[0114] In this embodiment, S6 specifically includes:
[0115] S61, integrate the improved isolation forest algorithm into a real-time financial data monitoring system, which can receive and process real-time financial statement data from listed companies and enterprises;
[0116] S62, the real-time monitoring system adopts an incremental learning mechanism, and the improved isolation forest model is automatically updated whenever new financial statement data arrives;
[0117] S63, in the real-time data processing process, the improved isolation forest model calculates the anomaly score S i and compares it with the set dynamic threshold T, if S i >T, the data point is determined to be abnormal, and further determined to be a potential financial fraud behavior;
[0118] S64, when the monitoring system detects abnormal data, the real-time feedback mechanism automatically triggers the review mechanism of the securities regulatory agency, and provides detailed anomaly analysis and risk assessment report, helping the regulators to take review and investigation.
[0119] The application integrates the improved isolation forest algorithm into a real-time financial data monitoring system, which can receive and process real-time financial statement data from listed companies and enterprises. Through the incremental learning mechanism, the system automatically updates the improved isolation forest model every time new financial data arrives, and calculates the anomaly score of each data point in real time. The data point is compared with the dynamic threshold, if it exceeds the threshold, it is determined to be abnormal, and further identifies potential financial fraud behavior. When detecting abnormality, the system automatically triggers the review mechanism of the securities regulatory agency through the real-time feedback mechanism, and provides detailed anomaly analysis and risk assessment report, helping the regulators to take review and investigation measures in time.
[0120] Example 1:
[0121] To verify the feasibility of the application in practice, the application is applied to the fraud detection of the financial statements of a certain listed company in this paper. The company is a multinational manufacturing enterprise with a market value of billions of dollars, and its annual financial statements involve tens of thousands of financial data. The challenge faced by the company is that the amount of financial data is huge and complex, and traditional manual auditing methods have been unable to efficiently and accurately detect potential fraud in them. Especially between multiple indicators of financial statements, there may be small and difficult to detect fraud, and the manual auditing has a high rate of missed detection. Therefore, how to use modern intelligent detection technology to automatically and accurately identify fraud in financial statements is a problem that the company urgently needs to solve.
[0122] In this scenario, we applied the financial statement fraud automatic detection method based on the improved isolation forest algorithm. Specifically, we first input the historical financial statement data set of the company into the improved isolation forest model for training. The model is continuously optimized through incremental learning mechanism, and uses graph convolution network to extract the dependency between financial data, in order to capture the complex interaction between financial items. The real-time feedback mechanism of the system automatically adjusts the model according to new financial data and performs anomaly detection to automatically determine whether there is a risk of financial fraud.
[0123] The specific operation steps are as follows: we input the financial data of the past three years into the system, including the detailed data of balance sheet, income statement and cash flow statement. The anomaly score of each financial statement data point is compared with the threshold value adjusted dynamically according to the model output. When the anomaly score exceeds the threshold value, the system will mark the data point as a potential fraud. Through this method, we can effectively identify financial abnormal data that may be missed by traditional manual auditing.
[0124] We used this method to conduct a financial fraud detection, covering the company's financial statement data for the past five years. By comparing with the results of manual auditing, we found that the model showed significant advantages in detection accuracy, efficiency and adaptability. In terms of data processing, traditional manual auditing takes weeks to analyze each financial data, while through the automated improved isolation forest model, the whole process is completed within a few days, and many hidden financial anomalies are accurately identified.
[0125] For example, in the 2024 annual financial statements, the traditional auditing method failed to find a financial data about the abnormal turnover rate of accounts receivable. However, after applying the technology of the application, the system successfully detected this abnormal data and further revealed the fraud of the company by artificially increasing the amount of accounts receivable to fake profits. If this is not discovered in time, it will cause great losses to investors and shareholders. The model not only improves the accuracy in detecting such abnormal behaviors, but also can monitor new data in real time and make dynamic adjustments.
[0126] To further verify the effectiveness of the present application, the following is a data table generated based on the company's actual financial data. The table shows the results of financial anomaly detection after applying the present application technology, including anomaly score, anomaly level and corresponding risk assessment.
[0127] Table 1 2024 financial statement anomaly detection results
[0128]
[0129] Table 1 2024 financial statement anomaly detection results shows the anomaly score of financial items, comprehensive anomaly score, anomaly level and corresponding risk assessment. Through the application of this model, regulatory personnel can timely discover abnormal financial behavior and further analyze the underlying financial fraud risk. Compared with traditional manual audit, the system can more accurately and quickly identify potential fraud, providing reliable decision support for regulatory agencies.
[0130] Through this method, the improved isolation forest algorithm can effectively improve the efficiency and accuracy of financial statement fraud detection, especially when faced with large-scale and complex financial data, it has significant advantages compared with traditional manual audit.
[0131] The above is only the preferred embodiment of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can make equivalent substitutions or changes within the technical scope disclosed by the present application and according to the technical solutions and inventive concepts of the present application, which should be covered within the protection scope of the present application.
Claims
1. An automatic detection method for financial statement fraud based on an improved isolation forest algorithm, characterized by: The steps include: S1. Collect financial statement data and pre-process the financial statement data; S2. Perform feature engineering and weighted feature selection on the preprocessed financial statement data, and use a graph convolutional network to extract the dependencies between financial data. S3. Based on the dependencies between financial data, an incremental learning mechanism is used to update the isolation forest model. The anomaly score threshold is adjusted through a dynamic threshold adjustment mechanism to construct an improved isolation forest model. S4. Using the historical financial statement dataset to train the improved isolation forest model, and using cross-validation to evaluate the performance of the improved isolation forest model; S5. Use the trained improved isolation forest model to detect outliers in new financial statement data and generate a risk assessment report; S6. Integrate the improved isolation forest algorithm into the real-time financial data monitoring system to conduct real-time monitoring of the financial statements of listed companies and enterprises.
2. The automatic detection method for financial statement fraud based on the improved isolation forest algorithm according to claim 1 is characterized in that: The financial statement data includes an income statement, a balance sheet and a cash flow statement, and the financial statement data is preprocessed. The preprocessing steps include data cleaning, missing value filling, standardization and normalization of specific financial indicators.
3. The automatic detection method for financial statement fraud based on the improved isolation forest algorithm according to claim 1 is characterized in that: The S2 specifically includes: S21. Performing feature engineering on the preprocessed financial statement data, wherein the feature engineering step includes extracting original financial features from the financial statements, transforming and deriving the original financial features, calculating financial ratios closely related to financial health status, and generating new financial features; S22. Perform weighted feature selection on the new financial features. The weighted feature selection strategy is based on the financial importance of the features and the influence of historical fraud behavior patterns. Specifically, the weighted value W of each financial feature is calculated. i To assess the susceptibility of these financial characteristics to financial fraud; S23. Use a graph convolutional network to extract the dependencies between financial data. The graph convolutional network uses each financial item in the financial statement as a node. The dependencies between the nodes constitute the edges of the graph. The feature representation of each financial item is transmitted and updated through the graph convolutional network. The adjacency matrix is used to represent the dependencies between the financial items. The embedded feature representation H of each node learned by the graph convolutional network (l+1) ; S24. Combining the financial data dependency graph obtained by the graph convolutional network with the weighted feature selection results, further screening all financial features to generate a final feature representation. The feature representation includes the dependency relationships between various features in the financial statements.
4. The automatic detection method for financial statement fraud based on the isolation forest algorithm according to claim 1 is characterized in that: The improved isolation forest model includes the following contents: In the process of building the tree structure, features and splitting points are determined through weighted feature selection and the dependency graph of financial data. A dynamic threshold adjustment mechanism is introduced to automatically adjust the threshold of anomaly scores based on historical financial statement datasets and real-time feedback. It combines local and global anomaly detection strategies, and comprehensively considers the abnormalities of various features in the financial statements through the financial data dependency learned through the graph convolutional network.
5. The automatic detection method for financial statement fraud based on the isolation forest algorithm according to claim 1 is characterized in that: The S3 specifically includes: S31. Using an incremental learning mechanism to update the improved isolation forest model, each decision tree in the improved isolation forest model is fine-tuned based on the new financial statement data; during the construction of each decision tree, the new financial statement data is used to update the split point and depth of the existing decision tree; S32. Under the incremental learning mechanism, the improved isolation forest model updates the weight matrix in the isolation forest by calculating the impact of new financial statement data on features; S33. Automatically adjust the anomaly scoring threshold using a dynamic threshold adjustment mechanism based on historical financial statement datasets and real-time feedback information; the historical financial statement datasets include past financial statement data and known abnormal behavior patterns, which are used to provide the model with normal and abnormal distribution characteristics of financial data; the real-time feedback information includes the model's anomaly detection results, false positive and false negative rates, as well as results from manual review and user feedback; S34. During anomaly detection, an anomaly score is calculated for each financial statement data point. If the anomaly score exceeds the current threshold, the data point is considered an anomaly. The threshold is dynamically adjusted based on the historical financial statement data set and real-time feedback information. S35. Calculate the score distribution of the abnormal data points based on the marked abnormal data points and normal data points in the historical financial statement data set; then calculate the threshold adjustment ΔT based on the newly added abnormal scores and normal scores in the real-time data; S36. When the real-time feedback results show that the false alarm rate and the missed alarm rate are too high, the threshold is adjusted to make the anomaly detection more stringent. When the false alarm and missed alarm rates decrease, the threshold is adjusted to relax the standards.
6. The automatic detection method for financial statement fraud based on the isolation forest algorithm according to claim 5 is characterized in that: The dynamic threshold adjustment mechanism takes into account the model's sensitivity to anomaly scores, and the updated weight matrix enables the improved isolation forest model to adjust the anomaly scores based on the new feature distribution.
7. The automatic detection method for financial statement fraud based on the isolation forest algorithm according to claim 1 is characterized in that: The S4 specifically includes: S41. Use a historical financial statement dataset to train the improved isolation forest model. The historical financial statement dataset includes a plurality of labeled financial data points, including normal data points and abnormal data points, wherein the abnormality score S of each data point is i is calculated as input data for training; S42. The historical financial statement dataset is divided into K subsets. During the cross-validation process, K-1 subsets are selected each time as training sets, and the remaining subset is selected as a validation set. The process is repeated K times to ensure that each subset is used as a validation set for model evaluation. S43. During each training process, the improved isolation forest model optimizes the model parameters based on the training set. The improved isolation forest model minimizes the error and optimizes the anomaly detection capability by adjusting the depth, splitting point and feature selection weight of the decision tree. After each training, the improved isolation forest model is evaluated on the validation set by calculating the anomaly score S. i And compare it with the set benchmark threshold to evaluate whether the data point is correctly marked as an anomaly; S44. After the cross-validation is completed, the average performance index of all K validations is calculated to evaluate the overall performance of the improved isolation forest model, wherein the performance index includes accuracy Acc, recall Rec, precision Pre, and F1 score F1; S45. Based on the average performance index of cross-validation, adjust the model hyperparameters, including the number of trees, tree depth, and feature selection weights.
8. The automatic detection method for financial statement fraud based on the isolation forest algorithm according to claim 1 is characterized in that: The S5 specifically includes: S51, using the trained improved isolation forest model to perform outlier detection on new financial statement data, wherein the financial statement data includes multiple financial data points, and the anomaly score S of each data point is i is calculated as the basis for anomaly identification; S52. Calculate the abnormal score S for each data point i , and compare it with the set threshold T. If S i >T, the data point is determined to be abnormal; the threshold T is dynamically adjusted based on the historical financial statement data set and real-time feedback information; S53, through the multi-dimensional anomaly scoring system, provide a comprehensive anomaly score S for each abnormal data point c , to indicate the severity of the abnormal data point anomaly, and to perform multi-level classification of abnormal data, including slight abnormality, moderate abnormality and severe abnormality; comprehensive abnormality score S c Used to determine the degree of abnormality of the data point, each data point is scored according to the comprehensive abnormality score S c Values are assigned to appropriate exception levels; S55, based on the detected abnormal data points, risk assessment is performed, and a risk level assessment mechanism based on model scoring is adopted. For each abnormal data point, the risk level is assessed based on the comprehensive abnormal score S c and abnormality level, determine the risk level of the data point in the overall financial statement; the risk assessment report includes the following: abnormality score S for each data point i and comprehensive abnormality score S c ; Anomaly type and anomaly level for each data point; Risk analysis of potential financial fraud; S56. Generate a detailed risk assessment report that lists abnormal values, abnormal patterns, and potential fraud risk analysis for each financial indicator in the financial statements; the risk assessment report provides analysis of specific financial data points and provides risk assessment recommendations based on the detected anomalies and their severity.
9. The automatic detection method for financial statement fraud based on the isolation forest algorithm according to claim 1 is characterized in that: The S6 specifically includes: S61. Integrating the improved isolation forest algorithm into a real-time financial data monitoring system, wherein the real-time financial data monitoring system is capable of receiving and processing real-time financial statement data from listed companies and enterprises; S62. The real-time monitoring system uses an incremental learning mechanism to automatically update and improve the isolation forest model whenever new financial statement data arrives. S63. In the real-time data processing process, the improved isolation forest model is used to score the abnormality of each data point S. i Compare with the set dynamic threshold T, if S i >T, the data point is judged as abnormal and further as a potential financial fraud; S64. When the monitoring system detects abnormal data, the real-time feedback mechanism automatically triggers the review mechanism of the securities regulatory agency and provides detailed abnormality analysis and risk assessment reports to help regulators conduct reviews and investigations.
Citation Information
Cited By
Sandbox terminal file protection method based on improved isolated forest
CN121327822A
A sandbox terminal file protection method based on improved isolated forest
CN121327822B