Financial counterfeiting risk detection method and device
By using a method of combining multi-dimensional feature extraction and optimization algorithms in financial fraud risk detection, three-dimensional statistical features of opportunities, motivation and traces are designed, and detection models are constructed through machine learning algorithms, the problems of strong rule dependence and insufficient feature mining in the existing technology are solved, and higher detection accuracy and recall rate are achieved.
Patent Information
- Application Number
- CN202411941069.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-06-10
AI Technical Summary
The existing financial fraud risk detection methods have problems such as strong dependence on rules, insufficient feature mining and limited detection effects, and it is difficult to effectively capture potential features under the complex and changeable financial fraud model.
A method combining multi-dimensional feature extraction and optimization algorithm is adopted to design statistical features covering three-dimensional opportunities, motivations and traces, and a financial fraud risk detection model is constructed through machine learning algorithms to improve the accuracy and recall of detection.
It significantly improves the accuracy and recall rate of financial fraud risk detection, preferably up to 10%-15%, and shows stronger robustness when processing high-dimensional data and unbalanced samples.
Smart Images

Figure CN120125359A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of big data and artificial intelligence, and particularly to a method and device for detecting financial fraud risks. Background Art
[0002] Financial fraud refers to the act of enterprise management or financial personnel violating national laws, regulations and related systems, and using means such as false statements, forgery or tampering of accounting information to cover up the true financial situation, operating results and cash flows of the enterprise. Such acts not only seriously damage the transparency and credibility of financial information, but may also mislead investors into making wrong decisions, resulting in significant economic losses.
[0003] Current financial fraud risk detection methods are mainly divided into two categories: The first category of methods relies on expert rules and conducts detection by artificially setting rules. Due to the high dependence on expert experience, this method has strong subjectivity and is difficult to effectively capture potential features outside the rules. Especially when dealing with complex and changing financial fraud models, it shows obvious limitations. The second category of methods is based on artificial intelligence algorithms. Usually, a single model is difficult to achieve a high detection accuracy. Therefore, the industry generally adopts the method of multi-model ensemble learning to improve the effect. However, these methods have insufficient in-depth mining and understanding of key features. In particular, feature selection overly relies on financial data, while the mining of non-financial features is significantly insufficient, resulting in severely limited diversity and discrimination ability of features. Financial fraud often uses carefully designed means (such as fictitious transactions, related-party transactions and complex accounting techniques) to make the data in the financial statements seem reasonable. It is difficult to comprehensively reveal fraud behaviors only relying on financial data features. This status quo of feature homogenization limits the discriminant ability of the model. Even if more models are used, it is difficult to fundamentally improve the detection accuracy and recall rate. Therefore, existing methods show obvious deficiencies when facing the actual production environment, and there is an urgent need for an innovative method that can make up for the defects of current technologies, start from the diversity of features and the in-depth mining of key features, and achieve more accurate detection and effective early warning of financial fraud behaviors. Summary of the Invention
[0004] In view of the defects of strong rule dependence, insufficient feature mining and limited detection effect in existing financial fraud risk detection methods, the present invention provides a method combining multi-dimensional feature extraction and optimization algorithms. Through innovative feature extraction and optimization algorithms, the accuracy and recall rate of financial fraud risk detection are significantly improved. Preferably, the improvement range can reach 10%-15%. Compared with existing methods, the method of the present invention shows stronger robustness when dealing with high-dimensional data and unbalanced samples.
[0005] A method for detecting financial fraud risks, the method comprising:
[0006] Step S1: Sample Screening. Construct a representative positive and negative sample set through authoritative data sources. The specific method is to screen out the listed companies that are clearly identified as having financial fraud and the years to which the financial reports containing fraud data belong as positive samples according to authoritative information such as CSRC announcements. At the same time, credible negative samples are comprehensively selected through multiple factors such as market value scale, industry distribution, and regulatory intensity.
[0007] Step S2: Feature Design. Design statistical features covering three dimensions: opportunity, motivation, and trace. Specific features may include corporate governance structure, capital requirements, and financial anomaly indicators. Preferably, key features are mined by combining financial data and non-financial data to ensure the discriminative ability and representativeness of the features.
[0008] The opportunity-related features may include, but are not limited to, the following indicators: relevant data reflecting the power distribution of the enterprise's management, external audit intensity, and governance structure, such as the total number of directors, supervisors, and senior management (referred to as "DSMEs"), the proportion of non-independent directors among DSMEs, the turnover rate of DSMEs, the concentration of ownership, whether the audit opinion on the financial statements is "unqualified opinion", and the number of independent audit institutions providing audit services to the enterprise in the past five years.
[0009] The motivation-related features may include, but are not limited to, the following indicators: representing the driving force of financial fraud, data related to the enterprise's capital requirements, such as the number of equity pledges, the proportion of short selling in market value, the number of seasoned equity offerings within a year, and the proportion of restricted stock unlock.
[0010] The trace-related features may include, but are not limited to, the following indicators: statistical features used to reveal abnormal behaviors in financial statements, such as accounts receivable index, asset quality index, depreciation rate index, accrual coefficient, gross profit margin index, operating revenue index, selling and administrative expense index, and financial leverage index.
[0011] Step S3: Dataset Construction. Based on the screened samples and designed features, construct a multi-dimensional dataset suitable for analysis. The specific method is: by collecting the annual financial data and relevant non-financial data of the sample enterprises, combined with the results of the three-dimensional feature design, form a complete basic dataset for analysis. The dataset not only covers financial statement data, but may also include corporate governance structure information, industry indicators, and other external public data to ensure the comprehensiveness and representativeness of the dataset.
[0012] Step S4: Data preprocessing to improve data quality and consistency and address the problem of sample imbalance. Specific methods include handling missing values and outliers, normalization operations, and sample balance adjustment. Missing value handling can be based on interpolation, filling, or deletion strategies; outlier handling is performed by marking or correcting through statistical methods; normalization processing is used to eliminate the influence of different feature dimensions. To address the problem of scarce positive samples, preferably, SMOTE or other oversampling algorithms suitable for sample balance can be adopted to enhance the balance and representativeness of the dataset by increasing or adjusting the distribution of sample points.
[0013] Step S5: Feature calculation. Based on the three-dimensional feature combination and the preprocessed dataset, statistical feature values are calculated.
[0014] Step S6: Model construction. For the dataset after data processing, machine learning algorithms are used for model construction. Preferably, the algorithms can adopt random forest, gradient boosting decision tree (GBDT), or other machine learning algorithms suitable for high-dimensional data processing.
[0015] Step S7: Risk detection. Use the constructed final model to detect the financial fraud risk of any enterprise's financial report. The detection results generate classification labels, which can provide data support for internal audit and external supervision of enterprises and are significantly superior to the existing technologies in terms of accuracy and recall rate.
[0016] A financial fraud risk detection device, the device includes: a sample screening module for screening positive and negative samples according to authoritative data to construct a representative dataset; a feature design module for designing three-dimensional features of opportunity, motivation, and trace and extracting key feature indicators; a dataset construction module for collecting financial and non-financial data to construct a multi-dimensional analysis dataset; a data preprocessing module for handling missing values and outliers and improving data quality through normalization and oversampling; a feature calculation module for calculating the feature values of samples according to the three-dimensional feature combination; a model construction module for training a financial fraud risk detection model based on machine learning algorithms; a risk detection module for generating classification labels or scores of financial fraud risks using the trained model.
[0017] A computer device includes a memory and a processor, and a computer program is stored in the memory. When the processor executes the computer program, it can implement the steps of the financial fraud risk monitoring method provided in any embodiment of the present application.
[0018] A computer-readable storage medium has a computer program stored thereon. When the program is executed by a processor, it can implement the steps of the financial fraud risk monitoring method provided in any embodiment of the present application.
[0019] The method and device of the present invention are not only applicable to the detection of the risk of corporate financial fraud, but also can be extended to other anomaly detection scenarios, such as:
[0020] Financial field: used to identify loan fraud, investment fraud and market manipulation behaviors;
[0021] Tax review: used to detect corporate tax violations and false declarations;
[0022] Corporate compliance detection: evaluate whether there are anomalies in the corporate governance structure or non-compliance in information disclosure;
[0023] Public security field: applied to anti-money laundering, monitoring of fund flows and other business scenarios. Description of the Drawings
[0024] Figure 1 It is a schematic flow chart of the financial fraud risk detection method in an embodiment.
[0025] Figure 2 It is a structural block diagram of the financial risk detection device in an embodiment. Detailed Description of the Invention
[0026] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0027] Embodiment 1, as Figure 1 shown, provides a method for detecting the risk of financial fraud, including the following steps.
[0028] Step S1: Sample screening. In this step, according to authoritative data sources (such as CSRC announcements, etc.), representative positive and negative sample sets are screened.
[0029] Positive samples: Screen listed companies that are clearly identified as having financial fraud and the years to which their financial reports belong.
[0030] Negative samples: Through factors such as market value scale, industry distribution and regulatory intensity, comprehensively select credible negative samples.
[0031] This step ensures that the samples are representative and comparable, laying a foundation for subsequent feature extraction and model training.
[0032] Step S2: Feature design. In this step, statistical features covering three dimensions of opportunity, motivation and trace are designed to comprehensively reflect possible financial fraud behaviors of enterprises.
[0033] (1) Opportunity - type features, which characterize the possibility of a company's management engaging in financial fraud and involve indicators such as the distribution of management power and the intensity of external supervision. Include but are not limited to: the total number of directors, supervisors, and senior management (referred to as "the board, supervisors, and senior management"), the proportion of non - independent directors among the board, supervisors, and senior management, the turnover rate of the board, supervisors, and senior management, the concentration of ownership, the type of audit opinion on the financial statements, and the number of independent audit institutions providing audit services to the company within a certain time range.
[0034] (2) Motivation - type features, which characterize the driving force for a company to engage in financial fraud and are usually related to the company's capital needs. Include but are not limited to: the number of pledged shares, the proportion of short - selling to market value, the number of seasoned equity offerings within a year, and the proportion of restricted stock unlock.
[0035] (3) Trace - type features, which reveal the statistical features of abnormal behaviors in financial statements. Include but are not limited to: accounts receivable index, asset quality index, depreciation rate index, accrual coefficient, gross profit margin index, operating revenue index, selling and administrative expense index, and financial leverage index.
[0036] Step S3: Dataset construction. In this step, based on the selected samples and designed features, a multi - dimensional analysis dataset is constructed. Annual financial data and relevant non - financial data of sample companies are collected. The dataset covers but is not limited to: financial statement data, corporate governance structure information, industry indicators, and other publicly available external data. Ensure that the dataset meets the requirements of subsequent analysis in terms of coverage, data quality, and dimensional richness.
[0037] Step S4: Data pre - processing. In this step, the data is pre - processed to improve data quality and consistency and solve the problem of sample imbalance. It includes the following steps:
[0038] (1) Missing value handling: Interpolation, filling, or deletion strategies are used to handle missing values.
[0039] (2) Outlier handling: Statistical methods are used for marking or correction.
[0040] (3) Normalization: Feature values with different dimensions are mapped to a unified range (such as [0, 1] or [-1, 1]).
[0041] (4) Sample balance adjustment: Oversampling algorithms (such as SMOTE, etc.) are used to balance the proportion of positive and negative samples, improving the representativeness of the dataset and the generalization ability of the model. Since the number of positive samples of financial fraud released by authoritative institutions is very limited, significantly less than the negative samples, an oversampling algorithm is used to moderately increase the number of positive samples.
[0042] Step S5: Feature calculation. In this step, the feature values are calculated based on the three-dimensional feature combination and the preprocessed data set. The current period refers to the year in which the financial report data being tested is counted, and the previous period refers to the previous financial report statistical year. For example:
[0043] Opportunity characteristics
[0044] The ratio of non-independent directors to directors, supervisors and senior managers = the number of non-independent directors in this period / the total number of directors, supervisors and senior managers in this period.
[0045] Director, Supervisor and Senior Manager Resignation Rate = Number of Directors, Supervisors and Senior Managers Resigned in this Period / Total Number of Directors, Supervisors and Senior Managers in this Period. Exceptions include changes of term or retirement, which may not be counted in the number of resignations.
[0046] The Herfindahl-Hirschman Index (HHI) can be used to calculate equity concentration, which is to first calculate the shareholding ratio of each shareholder and then take the sum of the square values.
[0047] The types of audit opinions on financial statements and the number of independent audit firms that provided audit services to the company within a certain period of time can be counted. The number of independent audit firms in the past three years, five years or longer time spans can be counted.
[0048] Motivational characteristics
[0049] Number of pledged shares = number of shares still pledged by all shareholders in this period - number of shares released from pledge in this period.
[0050] The proportion of short selling to market value = the total amount of new short selling in this period / the ending balance of the total market value of the enterprise in this period.
[0051] The proportion of restricted shares released = the number of restricted shares released in this period / the ending balance of the total number of shares of the enterprise in this period.
[0052] Trace characteristics
[0053] Accounts receivable index = the proportion of accounts receivable in current period to operating income / the proportion of accounts receivable in previous period to operating income.
[0054] The asset quality index detects whether a company is likely to control profits by manipulating non-physical asset items. The calculation formula is: asset quality index = current period non-physical asset ratio / previous period non-physical asset ratio. The non-physical assets can be calculated by deducting fixed assets from total assets, or by adding up the identified non-physical assets, such as intangible assets, goodwill, and long-term deferred expenses.
[0055] The depreciation rate index characterizes the possibility that an enterprise is upwardly adjusting the useful asset life assumption or adopting a new method that is favorable to revenue. Its calculation formula is: Depreciation rate index = Depreciation rate in the previous statistical year / Depreciation rate in the current year.
[0056] The accrual coefficient evaluates the degree of a company's manipulation of net profit through accruals. Its calculation formula is: Accruals = (Net profit after deducting non-recurring gains and losses - Net cash flow from operating activities) / Total assets.
[0057] Gross profit margin index = Gross profit margin in the previous period / Gross profit margin in the current period.
[0058] Operating revenue index = Operating revenue in the current period / Operating revenue in the previous period.
[0059] Sales and administrative expense index = Ratio of current period's sales and administrative expenses to operating revenue / Ratio of previous period's sales and administrative expenses to operating revenue.
[0060] Financial leverage index = Asset-liability ratio in the current period / Asset-liability ratio in the previous period.
[0061] The above feature calculations can be adjusted according to the actual data characteristics and scenario requirements to better adapt to different financial fraud scenarios. For example, for enterprises with a relatively small floating stock, replacing the total market value with the floating market value will make the features more sensitive.
[0062] Step S6: Model construction. In this step, based on the preprocessed dataset, a machine learning algorithm is used to construct a financial fraud risk detection model. It includes the following steps:
[0063] (1) Algorithm selection: Random forest, Bayesian model, decision tree, etc.
[0064] (2) Model training: Use the training set to train the model to ensure that the model can effectively process high-dimensional data and unbalanced samples and reduce the risk of overfitting.
[0065] (3) Parameter tuning: Optimize the parameters through techniques such as grid search and nested cross-validation to improve the accuracy and generalization ability of the model.
[0066] Preferably, any one machine learning algorithm can be used for grid search first, then nested cross-validation can be implemented using the best parameter combination, and finally the algorithm with the best performance can be selected for secondary parameter tuning. The purpose of this step is to further verify the generalization ability of the model to ensure that the model not only performs well during hyperparameter tuning but also has stable performance on the entire dataset.
[0067] Step S7: Risk detection. In this step, use the constructed final model to detect the financial fraud risk of the enterprise's financial reports.
[0068] Input: Enterprise financial data and non-financial data.
[0069] Output: Generate classification labels or risk scores for internal audit and external supervision of enterprises for reference.
[0070] Technical effect: In this embodiment, nested cross-validation was performed on the random forest, SVM model, Bayesian model, logistic regression model, decision tree model, and K-nearest neighbor model using optimized parameters. The random forest model performed the best in the nested cross-validation, with an average accuracy of 0.8688 + / - 0.0642. After performing grid search tuning on the random forest model again, the recognition accuracy for positive samples reached 93%, the recall rate was 86%, and the f1-score was 0.86. Its stability and accuracy are significantly better than the prior art. Using the same model training and tuning method, compared with only using the trace features of financial data, the accuracy and recall rate of the model trained with three-dimensional features can be improved by 10% - 15%.
[0071] Embodiment 2: Device implementation. Refer to Figure 2 The present invention also provides a financial fraud risk detection device, including:
[0072] Sample screening module: Used to construct positive and negative sample sets.
[0073] Feature design module: Used to design three-dimensional features of opportunity, motivation, and trace.
[0074] Dataset construction module: Used to collect multi-dimensional data and construct a dataset.
[0075] Data preprocessing module: Used to process missing values, outliers, perform normalization, and sample balance adjustment.
[0076] Feature calculation module: Used to calculate three-dimensional feature values.
[0077] Model construction module: Used to perform model training using machine learning algorithms.
[0078] Risk detection module: Used to generate classification labels or scores for financial fraud risks.
[0079] This device can be deployed based on hardware devices, cloud platforms, or distributed architectures to meet the requirements of different application scenarios.
[0080] The above embodiments only represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for detecting financial fraud risk, the method comprising: Step S1: Sample screening, constructing representative positive and negative sample sets based on authoritative data sources; Step S2 :Feature design, which designs statistical features covering three dimensions: opportunity, motivation and trace. The opportunity-type features are used to characterize the possibility of financial fraud by the management of an enterprise, which is related to the power distribution of the management and the intensity of external supervision. The motivation-type features are used to characterize the driving force of financial fraud, which is related to the capital demand of the enterprise. The trace-type features are used to reveal the statistical characteristics of abnormal behavior in financial statements. Step S3: Dataset construction: Based on the selected samples and designed features, annual financial data and related non-financial data of sample enterprises are collected to construct a multi-dimensional dataset for analysis; Step S4: Data preprocessing, processing missing values and outliers, normalization, and sample balance adjustment; Step S5: feature calculation, statistical feature values according to the three-dimensional feature combination and the preprocessed data set; Step S6: Model construction, training a financial fraud risk detection model based on a machine learning algorithm; Step S7: Risk detection: Use the trained model to detect the risk of financial fraud in corporate financial statements and generate classification labels or risk scores.
2. The method for detecting financial fraud risk according to claim 1, wherein: The opportunity characteristics include relevant data reflecting the power distribution of the enterprise management, the intensity of external audit and the governance structure. The data include but are not limited to: the total number of directors, supervisors and senior management, the proportion of non-independent directors to directors, supervisors and senior management, the turnover rate of directors, supervisors and senior management, equity concentration, the type of financial statement audit opinion, and the number of independent audit firms providing audit services to the enterprise within a specific time period. The time period may include the past several years, a specific financial reporting period or other reasonable time interval.
3. The method for detecting financial fraud risk according to claim 1, wherein: The motivational characteristics include the driving force of financial fraud and data related to the capital demand of the enterprise, including but not limited to: the number of pledged equity, the proportion of securities lending to market value, the number of additional issuances within the year, and the proportion of restricted shares released.
4. The method for detecting financial fraud risk according to claim 1, wherein: The trace features include statistical features used to reveal abnormal behaviors in financial statements, and the data include but are not limited to: accounts receivable index, asset quality index, depreciation rate index, accrual coefficient, gross profit margin index, operating income index, sales and administrative expenses index, and financial leverage index.
5. The method for detecting financial fraud risk according to claim 1, wherein: In the data preprocessing step, SMOTE or other oversampling algorithms suitable for sample balance are preferably used.
6. A financial fraud risk detection device, comprising: The sample screening module is used to construct representative positive and negative sample sets based on authoritative data sources; Feature design module, used to design statistical features covering three dimensions of opportunity, motivation and trace, and extract key feature indicators; The data set construction module is used to collect financial and non-financial data and construct a multi-dimensional analysis data set; Data preprocessing module, used to handle missing values, outliers, and improve data quality through normalization and oversampling; A feature calculation module is used to calculate the feature value of the sample according to the three-dimensional feature combination; Model building module, used to train financial fraud risk detection models based on machine learning algorithms; The risk detection module is used to generate classification labels or scores for financial fraud risks using the trained model.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the method for detecting the risk of financial fraud described in any one of claims 1 to 5 can be implemented.
8. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the steps of the method for detecting the risk of financial fraud as described in any one of claims 1 to 5.