A financial report risk identification method and device, electronic equipment and storage medium

By constructing labeled and unlabeled sample sets of financial statement datasets, and using supervised learning models to identify and analyze financial manipulation behavior, this solves the problem of difficulty in identifying financial statement manipulation in existing technologies, and achieves efficient financial risk identification and rule updates.

CN116307712BActive Publication Date: 2026-01-06CHINA BOHAI BANK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310243791.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-14
Publication Date
2026-01-06
Estimated Expiration
2043-03-14

AI Technical Summary

Technical Problem

The lack of effective means to identify financial statement manipulation by existing technologies leads to increased financial risks, financial losses, and disruption of economic order.

Method used

By acquiring financial statement datasets, preprocessing them, constructing labeled and unlabeled sample sets, building feature matrices, using supervised learning models for comparative learning, filtering out suspected financial statement manipulation data, performing attribution analysis, and updating the financial statement manipulation rule base.

Benefits of technology

It improves the detection efficiency of financial statement manipulation, effectively identifies unknown financial manipulation situations, and reduces financial risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116307712B_ABST
    Figure CN116307712B_ABST
Patent Text Reader

Abstract

The application provides a financial report risk identification method and device, electronic equipment and a storage medium. The method comprises the following steps: obtaining a financial report data set, preprocessing and data cleaning the financial report data set, constructing a feature matrix, comparing and learning the financial report data of different enterprises in the corresponding time dimension, obtaining the implicit feature matrix of each enterprise in different time dimensions, then constructing a supervised learning model, screening out target samples that cannot be detected by existing rules from the unannotated sample set, then screening out suspected financial report data from the target samples, performing attribution analysis on the suspected financial report data, obtaining the financial report and the corresponding financial report behavior of the financial report, and finally updating the financial report financial report rule library based on the financial report and the financial report behavior. The application can identify the financial report financial report situation to a high degree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, electronic device and storage medium for identifying financial statement risks. Background Technology

[0002] Financial statements must accurately and comprehensively reflect a company's financial condition and operating results, meeting the information needs of stakeholders and ensuring the accuracy and reliability of all data provided to users. However, the practice of financial statement manipulation is currently severe and quite sophisticated. These inaccurate accounting statements convey incorrect information, misleading intended users and leading to flawed decision-making, increased financial risk, and financial losses. They also disrupt economic order, resulting in tax evasion and losses to government and banking funds.

[0003] Therefore, it is essential to carefully analyze the reasons for financial statement manipulation and identify such manipulation to the greatest extent possible. However, existing technologies lack corresponding solutions. Summary of the Invention

[0004] In view of this, embodiments of this application provide a financial statement risk identification method, apparatus, electronic device, and storage medium, which can identify financial manipulation to a high degree.

[0005] The technical solution of this application embodiment is implemented as follows:

[0006] In a first aspect, embodiments of this application provide a method for identifying financial statement risks, comprising the following steps:

[0007] Obtain the financial statement dataset and preprocess it to obtain a labeled sample set and an unlabeled sample set. The labeled sample set represents financial statement data that has been determined to be manipulated, while the unlabeled sample set represents financial statement data that cannot be determined to be manipulated.

[0008] The unlabeled sample set is cleaned, and a feature matrix is ​​constructed based on the cleaned unlabeled sample set, wherein the feature matrix is ​​used to represent the financial report data of each enterprise at different time dimensions;

[0009] Based on the feature matrix, the financial report data of different companies in the corresponding time dimension are compared and learned to obtain the latent feature matrix of each company in different time dimensions. The latent feature matrix includes discriminative features to describe whether there is embellishment.

[0010] A supervised learning model is constructed based on the labeled sample set and the latent feature matrix of each enterprise at different time dimensions, and the supervised learning model is used to filter out target samples that cannot be detected by existing rules from the unlabeled sample set.

[0011] Suspected financial statement manipulation data are selected from the target sample. Attribution analysis is performed on the suspected financial statement manipulation data to obtain the probability that each suspected financial statement manipulation data is manipulated. Based on the probability, the manipulated financial statement and the corresponding manipulation behavior are obtained.

[0012] The financial statement manipulation rule base is updated based on the stated manipulation of financial statements and the stated manipulation behavior.

[0013] In one possible implementation, the labeled sample set includes a positive sample set and a negative sample set, and the preprocessing of the financial statement dataset to obtain the labeled sample set and the unlabeled sample set includes:

[0014] The financial statement dataset is composed of at least one financial statement data point collected from the target data source.

[0015] The financial statement data that has been determined to be unembellished is designated as the positive sample set, the financial statement data that has been determined to be embellished is designated as the negative sample set, and the financial statement data whose embellishment status cannot be determined is designated as the unlabeled sample set. In one possible implementation, the method further includes:

[0016] In one possible implementation, constructing the feature matrix based on the cleaned unlabeled sample set includes:

[0017] A first feature matrix is ​​constructed for each enterprise based on its enterprise information and industry information. The enterprise information includes basic enterprise information and enterprise operating information, and the industry information includes basic industry information and industry operating information.

[0018] Based on the time dimension of the enterprise operating information and the industry operating information, the first feature matrix corresponding to each enterprise is processed to obtain the second feature matrix corresponding to each enterprise, and the second feature matrix corresponding to each enterprise is used as the feature matrix.

[0019] In one possible implementation, the step of performing comparative learning processing on the financial statement data of different companies at corresponding time dimensions based on the feature matrix to obtain the latent feature matrix of each company at different time dimensions includes:

[0020] Using enterprises as the unit, financial data from any quarter of any enterprise is selected as the anchor sample, financial data from enterprises that are not anchor samples are used as negative samples, and financial data from other quarters of the selected enterprise are used as positive samples.

[0021] Based on the anchor sample, the negative sample, and the positive sample selected each time, the positive sample selected in this time is input into a deep neural network for comparative learning processing, and the discriminative features of each enterprise in different time dimensions are determined by the comparative learning loss function.

[0022] The discriminative features are added to the feature matrix to obtain the latent feature matrix.

[0023] In one possible implementation, the step of filtering target samples from the unlabeled sample set using the supervised learning model that cannot be detected by existing rules includes:

[0024] A multilayer perceptron is used as the supervised learning model, wherein the activation function of the neurons in the multilayer perceptron is a noise linear rectified function;

[0025] The unlabeled sample set is used as input parameters to the supervised learning model to obtain the probability of embellishment behavior.

[0026] Financial statement data with a probability of embellishment exceeding a preset first threshold are identified as the target samples.

[0027] In one possible implementation, the step of filtering out suspected window-dressing financial statement data from the target sample includes:

[0028] Determine the current industry in which the target sample belongs;

[0029] Calculate the proportion of known financial statement data with embellishment practices within the current industry to the total sample of the industry.

[0030] A target number of financial statement data are taken from the target sample as the suspected window dressing financial statement data, wherein the target number is the product of the total sample and the proportion of the total sample.

[0031] In one possible implementation, the attribution analysis of the suspected window dressing financial statement data to obtain the probability that each suspected window dressing financial statement data is window dressing, and the window dressing financial statement and the corresponding window dressing behavior based on the probability, includes:

[0032] The Shapley value of the suspected window dressing financial statements is predicted using a gradient boosting tree, wherein the Shapley value represents the probability that the suspected window dressing financial statements are indeed window dressing.

[0033] The suspected financial statement data with a probability greater than a preset second threshold are processed to obtain the financial statement and the corresponding manipulation behavior.

[0034] Secondly, embodiments of this application also provide a financial statement risk identification device, the device comprising:

[0035] The acquisition module is used to acquire a financial statement dataset and preprocess the financial statement dataset to obtain a labeled sample set and an unlabeled sample set. The labeled sample set represents financial statement data that has been determined to be embellished, while the unlabeled sample set represents financial statement data that cannot be determined to be embellished.

[0036] A construction module is used to clean the unlabeled sample set and construct a feature matrix based on the cleaned unlabeled sample set, wherein the feature matrix is ​​used to represent the financial report data of each enterprise at different time dimensions;

[0037] The comparison module is used to perform comparative learning processing on the financial statement data of different companies in the corresponding time dimension based on the feature matrix, so as to obtain the latent feature matrix of each company in different time dimensions, wherein the latent feature matrix includes a discriminative feature for describing whether there is embellishment;

[0038] The filtering module is used to construct a supervised learning model based on the labeled sample set and the latent feature matrix of each enterprise at different time dimensions, and to filter out target samples that cannot be detected by existing rules from the unlabeled sample set through the supervised learning model.

[0039] The analysis module is used to filter out suspected window dressing financial statement data from the target sample, perform attribution analysis on the suspected window dressing financial statement data, obtain the probability that each suspected window dressing financial statement data has been window dressing, and obtain the window dressing financial statement and the window dressing behavior corresponding to the window dressing financial statement based on the probability.

[0040] The update module is used to update the financial statement manipulation rule base based on the manipulated financial statements and the manipulation behavior.

[0041] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the financial statement risk identification method described in any of the first aspects.

[0042] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the financial statement risk identification method described in any of the first aspects.

[0043] The embodiments of this application have the following beneficial effects:

[0044] 1. With minimal financial statement manipulation, self-supervised learning on a large number of unlabeled samples can effectively learn the hidden representation (latent feature matrix) of the samples. This representation has high similarity to similar classes and high distinguishability between different classes.

[0045] 2. Using this hidden layer representation, an effective classification model (i.e., a supervised learning model) can be learned with a small number of labeled samples, and the classification model can filter out target samples that existing rules cannot detect from the unlabeled sample set. 3. Attribution analysis can more effectively pinpoint unknown financial statement manipulation behaviors, improving the detection efficiency of such behaviors. Attached Figure Description

[0046] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 This is a flowchart illustrating steps S101-S106 provided in the embodiments of this application;

[0048] Figure 2 This is a flowchart illustrating steps S1011-S1012 provided in the embodiments of this application;

[0049] Figure 3 This is a flowchart illustrating steps S1021-S1022 provided in the embodiments of this application;

[0050] Figure 4 This is a flowchart illustrating steps S1031-S1033 provided in the embodiments of this application;

[0051] Figure 5 This is a flowchart illustrating steps S1041-S1043 provided in the embodiments of this application;

[0052] Figure 6 This is a flowchart illustrating steps S1051-S1053 provided in the embodiments of this application;

[0053] Figure 7 This is a flowchart illustrating steps A1051-A1052 provided in the embodiments of this application;

[0054] Figure 8 This is a schematic diagram provided in an embodiment of this application;

[0055] Figure 9 This is an information diagram of the first feature matrix construction provided in the embodiments of this application;

[0056] Figure 10 This is the second feature matrix diagram provided in the embodiments of this application;

[0057] Figure 11 This is a diagram of the contrastive learning network structure provided in the embodiments of this application;

[0058] Figure 12 This is a schematic diagram of the attribution analysis principle provided in the embodiments of this application;

[0059] Figure 13 This is a schematic diagram of the financial statement risk identification device provided in the embodiments of this application;

[0060] Figure 14 This is a schematic diagram of the composition structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0061] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0062] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0063] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0064] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0065] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.

[0066] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application and is not intended to limit this application.

[0067] During the implementation of the embodiments of this application, the applicant discovered the following problems:

[0068] (1) The definition of financial embellishment is difficult to determine, and the lack of objective benchmarks makes it relatively difficult to classify bad samples for modeling.

[0069] (2) Most of the fraudulent financial statements used by enterprises are publicly reported, and the amount of data is relatively small.

[0070] (3) Traditional identification methods use linear regression models, which detect changes in the proportion of a single subject or indicator. The overall trend of related subjects is difficult to explain, so the interpretability of the model results is not strong.

[0071] (4) Since a single indicator has a weak impact on the final result given by the traditional model, the difference between the capture rate and the false positive rate in each interval is not obvious, and the verification and optimization of the traditional model is difficult.

[0072] (5) Relying on expert rules and embellishing the sample database to summarize rules, the data processing of accumulated expert experience rules is cumbersome and rule changes are delayed.

[0073] See Figure 1 , Figure 1 This is a flowchart illustrating steps S101-S106 of the financial statement risk identification method provided in this application embodiment, which will be combined with... Figure 1 Steps S101-S106 shown will be explained.

[0074] Step S101: Obtain the financial statement dataset and preprocess the financial statement dataset to obtain a labeled sample set and an unlabeled sample set. The labeled sample set represents financial statement data that has been determined to be embellished, and the unlabeled sample set represents financial statement data that cannot be determined to be embellished.

[0075] Step S102: Clean the unlabeled sample set and construct a feature matrix based on the cleaned unlabeled sample set, wherein the feature matrix is ​​used to represent the financial report data of each enterprise at different time dimensions;

[0076] Step S103: Based on the feature matrix, perform comparative learning processing on the financial statement data of different companies in the corresponding time dimension to obtain the latent feature matrix of each company in different time dimensions, wherein the latent feature matrix includes a discriminative feature used to describe whether there is embellishment.

[0077] Step S104: Construct a supervised learning model based on the labeled sample set and the latent feature matrix of each enterprise at different time dimensions, and use the supervised learning model to filter out target samples that cannot be detected by existing rules from the unlabeled sample set;

[0078] Step S105: Select suspected financial statement data from the target sample, perform attribution analysis on the suspected financial statement data to obtain the probability that each suspected financial statement data is manipulated, and obtain the manipulated financial statement and the corresponding manipulation behavior based on the probability.

[0079] Step S106: Update the financial statement manipulation rule base based on the manipulated financial statements and the manipulation behavior.

[0080] The above-mentioned financial statement risk identification method has the following beneficial effects:

[0081] 1. With minimal financial statement manipulation, self-supervised learning on a large number of unlabeled samples can effectively learn the hidden representation (latent feature matrix) of the samples. This representation has high similarity to similar classes and high distinguishability between different classes.

[0082] 2. Using this hidden layer representation, an effective classification model (i.e., a supervised learning model) can be learned with a small number of labeled samples, and the classification model can filter out target samples that existing rules cannot detect from the unlabeled sample set. 3. Attribution analysis can more effectively pinpoint unknown financial statement manipulation behaviors, improving the detection efficiency of such behaviors.

[0083] The exemplary steps described above in the embodiments of this application will be explained below.

[0084] In step S101, a financial statement dataset is obtained and preprocessed to obtain a labeled sample set and an unlabeled sample set. The labeled sample set represents financial statement data that has been determined to be embellished, while the unlabeled sample set represents financial statement data that cannot be determined to be embellished.

[0085] In some embodiments, see Figure 8 , Figure 8 This is a schematic diagram provided in the embodiments of this application, such as... Figure 8 As shown, by preprocessing the financial statement dataset, labeled sample sets (the dataset with and without financial statement manipulation shown in the figure) and unlabeled sample sets (the unlabeled dataset) are obtained.

[0086] In some embodiments, see Figure 2 , Figure 2 This is a flowchart illustrating steps S1011-S1012 provided in the embodiments of this application. Figure 1 The step S101 shown can be implemented through steps S1011-S1012, which will be explained in conjunction with each step.

[0087] In step S1011, at least one financial report data is collected from the target data source to form the financial report dataset.

[0088] Here, the amount of financial report data collected should be as large as possible. Specific data sources include: financial report disclosures of listed companies, historical financial report data accumulated within the industry, and financial report data from publicly reported cases of financial fraud and embellishment. The financial report data can be monthly, quarterly, or annual. Considering the continuity, density, and stability of the data, quarterly financial report data is selected as the dataset for subsequent modeling.

[0089] In step S1012, the financial statement data that has been determined to be unembellished is identified as the positive sample set, the financial statement data that has been determined to be embellished is identified as the negative sample set, and the financial statement data that cannot be determined to be embellished is identified as the unlabeled sample set.

[0090] Here, we construct the following from the collected financial statement data:

[0091] a. Positive sample set: Financial statements and companies without fraud or embellishment; generally, financial statements of companies with many years of stable development, good corporate reputation, and industry leadership are selected.

[0092] b. Negative sample set: Financial statements and companies found to have engaged in window dressing practices, based on publicly reported data and industry experience.

[0093] c. Unlabeled sample set: It is impossible to clearly distinguish whether there is fraudulent financial data and companies.

[0094] In step S102, the unlabeled sample set is cleaned, and a feature matrix is ​​constructed based on the cleaned unlabeled sample set, wherein the feature matrix is ​​used to represent the financial statement data of each enterprise at different time dimensions.

[0095] In some embodiments, see continue to see Figure 8 ,like Figure 8As shown, data cleaning can be performed on the unlabeled sample set (unlabeled dataset) to re-examine and verify the data. The purpose is to remove duplicate information, correct existing errors, and provide data consistency. Then, a feature matrix is ​​constructed based on the cleaned unlabeled sample set.

[0096] In some embodiments, see Figure 3 , Figure 3 This is a flowchart illustrating steps S1021-S1022 provided in the embodiments of this application. The construction of the feature matrix based on the cleaned unlabeled sample set can be achieved through steps S1021-S1022. Each step of the set will be explained.

[0097] In step S1021, a first feature matrix corresponding to each enterprise is constructed based on the enterprise information and industry information of each enterprise. The enterprise information includes basic enterprise information and enterprise operating information, and the industry information includes basic industry information and industry operating information.

[0098] For example, see Figure 9 , Figure 9 This is an information graph of the first feature matrix construction provided in the embodiments of this application, such as... Figure 9 As shown, the construction of the feature matrix (first feature matrix) focuses on the following four aspects: basic enterprise information, enterprise operation information, basic industry information, and industry operation information. Specifically, enterprise information includes basic enterprise information and enterprise operation information, and industry information includes basic industry information and industry operation information.

[0099] Basic company information includes, but is not limited to, the following: industry, establishment time, location, and years, shareholder information, senior management disclosures, and size: number of salespersons and personnel.

[0100] Business operation information includes, but is not limited to, the following: corporate financial data, business transaction data, energy consumption, etc.

[0101] Basic industry information includes, but is not limited to, the following: the industry itself, and the number of times the industry has historically engaged in financial fraud.

[0102] Industry operating information includes, but is not limited to, the following: industry prosperity and macroeconomic information facing the industry.

[0103] In step S1022, based on the time dimension of the enterprise operating information and the industry operating information, the first feature matrix corresponding to each enterprise is subjected to derivation processing to obtain the second feature matrix corresponding to each enterprise, and the second feature matrix corresponding to each enterprise is used as the feature matrix.

[0104] For example, see Figure 10 , Figure 10This is the second feature matrix diagram provided in the embodiments of this application, such as... Figure 10 As shown, basic information mainly consists of factual indicators, which are relatively stable; therefore, the current indicator values ​​are primarily used to construct the feature matrix. In contrast, operational information indicators are closely related to a company's repayment ability and exhibit significant fluctuations over time. Therefore, different scales are applied to the operational information indicators along the time dimension. Let M represent the set of firms, where M is the number of firms. For enterprises A collection of quarterly financial report release dates. For enterprises The number of financial reports released during the statistical window period. For enterprises Financial reports were released at time t.

[0105] A company's financial statements will show different data at different points in time, therefore and It is a one-to-many relationship. and The resulting matrix corresponds row-wise with the feature matrix on the right; in the feature matrix derived from the above features, the corresponding feature vector for each enterprise can have approximately 300 dimensions.

[0106] In step S103, based on the feature matrix, the financial statement data of different companies in the corresponding time dimension are compared and learned to obtain the latent feature matrix of each company in different time dimensions. The latent feature matrix includes discriminative features to describe whether there is embellishment.

[0107] In some embodiments, see continue to see Figure 8 ,like Figure 8 As shown, after obtaining the feature matrix (second feature matrix), comparative learning can be performed on the feature matrix to learn the common features between similar instances, distinguish the differences between dissimilar instances, and add the differences as discriminative features to the feature matrix to form a latent feature matrix.

[0108] In some embodiments, see Figure 4 , Figure 4 This is a flowchart illustrating steps S1031-S1033 provided in the embodiments of this application. Figure 1 The step S103 shown can be implemented through steps S1031-S1033, and will be explained in conjunction with each step.

[0109] In step S1031, taking an enterprise as a unit, the financial report data of any enterprise in any quarter is selected as the anchor sample, the financial report data of the enterprise that is not the anchor sample is used as the negative sample, and the financial report data of the selected enterprise in other quarters is used as the positive sample.

[0110] Before conducting comparative learning, positive and negative samples must first be constructed on the unlabeled sample set.

[0111] a1. Construct positive and negative samples, using enterprises as the unit.

[0112] a2. When the model is trained in batches, any sample is fixed and called the anchor sample, while other samples (the companies where the non-anchor samples are located) are negative samples;

[0113] a3. Determine the company corresponding to the anchor sample, and use the characteristics corresponding to other quarterly data of that company as positive samples.

[0114] In step S1032, based on the anchor sample, the negative sample, and the positive sample selected each time, the selected positive sample is input into a deep neural network for contrastive learning processing, and the discriminative features of each enterprise in different time dimensions are determined by the contrastive learning loss function.

[0115] For example, see Figure 11 , Figure 11 This is a diagram of the contrastive learning network structure provided in the embodiments of this application, such as... Figure 11 As shown, the anchor sample can be "Feature Vector of Company A-2021-Q3", the negative samples are "Feature Vector of Company B-2021-Q3", "Feature Vector of Company C-2021-Q2", and "Feature Vector of Company D-2020-Q3", and the positive samples are "Feature Vector of Company A-2021-Q2" and "Feature Vector of Company A-2021-Q1". For the selected anchor sample, negative sample, and positive sample, the "Feature Vector of Company A-2021-Q2" and "Feature Vector of Company A-2021-Q1" in the positive sample can be compared. After processing by the encoder multilayer fully connected network and the projector multilayer fully connected network, the corresponding feature vectors used in the next task can be obtained respectively. Then, the comparison loss of the two feature vectors is obtained through the loss function, and thus the discriminative feature is obtained.

[0116] It should be noted that the encoder is a multi-layer fully connected network, consisting of one input layer and multiple hidden layers. The activation function can be, for example, a linear rectification function (RCF), also known as a rectified linear unit, which can be expressed by the formula: The number of hidden layers depends on the amount of data, and is generally 2-5 layers. The projector also corresponds to a multi-layer fully connected network, including one input layer and multiple hidden layers. The activation function can be, for example, a linear rectified function. The number of neurons in the hidden layer can be the same as that in the input layer. The number of hidden layers depends on the amount of data, and is generally 2-5 layers.

[0117] The goal of the contrastive learning loss function is to maximize the similarity between positive sample pairs and minimize the similarity between negative sample pairs. For example, the InfoNCE loss can be used, which can be expressed by the following formula:

[0118]

[0119]

[0120] in, Z represents the inner product of the corresponding vectors of two positive examples. i and Z j Let represent two feature vectors to be compared, and represent the parameters. It is a constant, typically 0.1 or 0.2.

[0121] In step S1033, the discriminative features are added to the feature matrix to obtain the hidden feature matrix.

[0122] Here, after obtaining the discriminative features, these discriminative features can be added to the original feature matrix to obtain the latent feature matrix.

[0123] In step S104, the supervised learning model is used to filter out target samples from the unlabeled sample set that cannot be detected by existing rules.

[0124] In some embodiments, see Figure 8 ,like Figure 8 As shown, a supervised learning model can be built based on the latent feature matrix and labeled sample sets (datasets with and without financial statements), thereby filtering out target samples that cannot be detected by existing rules in the latent feature matrix.

[0125] In some embodiments, see Figure 5 , Figure 5 This is a flowchart illustrating steps S1041-S1043 provided in the embodiments of this application. The supervised learning model filters out target samples that cannot be detected by existing rules from the unlabeled sample set. This can be achieved through steps S1041-S1043, which will be explained in conjunction with each step.

[0126] In step S1041, a multilayer perceptron is used as the supervised learning model, wherein the activation function of the neurons in the multilayer perceptron is a noise linear rectified function.

[0127] Here, the supervised learning model is essentially a classification model that ultimately outputs the probability of embellishment behavior. The activation function of the neurons in this classification model is a noisy linear rectified function. The number of neurons is the same as the dimension of the input feature vector; the number of hidden layers is a hyperparameter, and its specific value is related to the actual amount of data, generally ranging from 2 to 5.

[0128] In step S1042, the unannotated sample set is used as an input parameter to the supervised learning model to obtain the probability of embellishment behavior.

[0129] For example, the input to a supervised learning model (i.e., a classification model) is... , To apply the feature matrix, also known as the latent feature matrix, of the enterprise at different time points obtained by the contrastive learning model on unlabeled samples, S unknown This is an unlabeled sample set. The network output is the probability value of each sample having engaged in embellishment behavior; the higher the probability value, the more likely embellishment behavior is to occur.

[0130] In step S1043, financial statement data with a probability of embellishment greater than a preset first threshold are identified as the target sample.

[0131] Here, the model trained using this model and the dataset is denoted as... , To pass the model Predicted suspected embellishment sample set, target sample , .

[0132] In step S105, suspected financial statement manipulation data are screened from the target sample, attribution analysis is performed on the suspected financial statement manipulation data to obtain the probability that each suspected financial statement manipulation data is manipulated, and the manipulated financial statement and the corresponding manipulation behavior are obtained based on the probability.

[0133] In some embodiments, see continue to see Figure 8 ,like Figure 8As shown, after filtering out samples (target samples) that cannot be detected by existing rules, it is necessary to further filter out suspected financial statement manipulation data from the target samples and perform attribution analysis on the suspected financial statement manipulation data. When performing attribution analysis on the suspected financial statement manipulation data, it is necessary to first input the suspected financial statement manipulation data into the classification model (i.e., supervised learning model) for identification. The suspected financial statement manipulation data with a recognition probability greater than or equal to the preset recognition probability are used as input samples for attribution analysis. Then, attribution analysis is performed on the input samples to determine the financial statement manipulation and the corresponding manipulation behavior. The classification model (i.e., the supervised learning model) here can be a decision tree model, such as Gradient Boosting Decision Tree (GBDT) or Extreme Gradient Boosting (XGB). Compared to GBDT, XGB performs a second-order Taylor expansion. GBDT fits along the direction of the negative gradient and only uses the first-order gradient information, while XGB directly performs a second-order Taylor expansion on the loss function. Compared to GBDT, it has a more accurate fitting direction and is faster.

[0134] In some embodiments, see Figure 6 , Figure 6 This is a flowchart illustrating steps S1051-S1053 provided in the embodiments of this application. The step of filtering out suspected embellished financial statement data from the target sample can be achieved through steps S1051-S1053, which will be explained in conjunction with each step.

[0135] In step S1051, the current industry of the target sample is determined.

[0136] In step S1052, the proportion of financial statement data known to have engaged in embellishment within the current industry is calculated to represent the overall sample size of the total sample within that industry.

[0137] In step S1053, a target number of financial statement data is taken from the target sample as the suspected window dressing financial statement data, wherein the target number is the product of the total sample and the proportion of the total sample.

[0138] For example, a sample set of suspected embellishments is taken from different industries, based on industry... For example, let's illustrate how to select a sample set suspected of being embellished:

[0139] First, the computing industry The percentage of companies in the industry that have consistently engaged in financial embellishment is [not specified in the original text]. ; For the industry The total number of samples in the entire sample set, through industry The number of known financial statements that have been manipulated and The ratio is determined .

[0140] Then apply the model. right Predict all samples and output the probability value of each sample having engaged in embellishment.

[0141] Finally take The sample set with the highest probability value was selected as the suspected data for manipulated financial statements and added to the set. .

[0142] In some embodiments, see Figure 7 , Figure 7 This is a flowchart illustrating steps A1051-A1052 provided in the embodiments of this application. The step of performing attribution analysis on the suspected embellished financial statement data to obtain the probability that each suspected embellished financial statement data is embellished, and obtaining the embellished financial statement and the embellished behavior corresponding to the embellished financial statement based on the probability, can be achieved through steps A1051-A1052, which will be explained in conjunction with specific steps.

[0143] In step A1051, the Shapley value of the suspected embellished financial statement data is predicted using a gradient boosting tree, wherein the Shapley value represents the probability that the suspected embellished financial statement data is embellished.

[0144] In step A1052, the suspected window dressing financial statement data with a probability greater than a preset second threshold are processed to obtain the window dressing financial statement and the window dressing behavior corresponding to the window dressing financial statement.

[0145] For example, see Figure 12 , Figure 12 This is a schematic diagram of the attribution analysis principle of an embodiment of this application, as shown below. Figure 12 As shown, the Shapley value method is used to analyze the sample set. Attribution analysis was performed on each sample to obtain the following results: Figure 12 The data results after analysis are shown.

[0146] It's important to note that the Shapley value attribute is visualized as a "force," where each feature value represents a force that either increases or decreases the prediction. Predictions begin with a baseline, which is the average of all predictions. Each Shapley value is an arrow that either increases (positive) or decreases (negative) the prediction.

[0147] Specifically, in the sample set Train a classification model, such as a limit gradient boosting tree, and obtain the model after training. , here A collection of known, unembellished financial statements, a binary tuple. This indicates that the financial report released by company i at time t was not embellished.

[0148] The model output value in the figure yes The logarithmic probability transformation can be expressed by the following formula:

[0149]

[0150] Both have the same monotonicity, as shown above. The probability of the corresponding financial statement being manipulated is 0.7068. As can be easily seen from the graph above, the longer the progress bar to the left of f(x), the more likely the corresponding feature dimension is to be manipulated. Using shapely values, several key features can be selected as suspected fraudulent activities, and further judgment can be made by combining expert knowledge (or an expert knowledge base) to obtain the manipulated financial statement and the corresponding manipulation behavior.

[0151] In step S106, the financial statement manipulation rule base is updated based on the manipulated financial statements and the manipulation behavior.

[0152] Here, the financial statement manipulation rule base is updated based on the obtained financial statements and manipulation behaviors. Subsequently, financial statement risk identification can be carried out based on the financial statement manipulation rule base. During the identification process, the financial statement manipulation rule base can be continuously updated and iterated to make the results more accurate.

[0153] In summary, the embodiments of this application have the following beneficial effects:

[0154] 1. With minimal financial statement manipulation, self-supervised learning on a large number of unlabeled samples can effectively learn the hidden representation (latent feature matrix) of the samples. This representation has high similarity to similar classes and high distinguishability between different classes.

[0155] 2. Utilizing this hidden layer representation, an effective classification model (supervised learning model) can be learned using a small number of labeled samples. Furthermore, the classification model can filter out target samples from the unlabeled sample set that cannot be detected by existing rules. 3. Attribution analysis can more effectively pinpoint unknown financial statement manipulation behaviors, improving the detection efficiency of such behaviors.

[0156] Based on the same inventive concept, this application also provides a financial statement risk identification device corresponding to the financial statement risk identification method in the first embodiment. Since the principle of the device in this application is similar to the above-mentioned financial statement risk identification method, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0157] like Figure 13 As shown, Figure 13 This is a schematic diagram of the structure of the financial statement risk identification device 1300 provided in this application embodiment. The financial statement risk identification device 1300 includes:

[0158] The acquisition module 1301 is used to acquire a financial statement dataset and preprocess the financial statement dataset to obtain a labeled sample set and an unlabeled sample set, wherein the labeled sample set represents financial statement data that has been determined to be embellished, and the unlabeled sample set represents financial statement data that cannot be determined to be embellished.

[0159] The construction module 1302 is used to clean the unlabeled sample set and construct a feature matrix based on the cleaned unlabeled sample set, wherein the feature matrix is ​​used to represent the financial report data of each enterprise in different time dimensions.

[0160] The comparison module 1303 is used to perform comparative learning processing on the financial statement data of different companies in the corresponding time dimension based on the feature matrix, and obtain the latent feature matrix of each company in different time dimensions, wherein the latent feature matrix includes a discriminative feature for describing whether there is embellishment.

[0161] The filtering module 1304 is used to construct a supervised learning model based on the labeled sample set and the latent feature matrix of each enterprise in different time dimensions, and to filter out target samples that cannot be detected by existing rules from the unlabeled sample set through the supervised learning model.

[0162] The analysis module 1305 is used to filter out suspected window dressing financial statement data from the target sample, perform attribution analysis on the suspected window dressing financial statement data, obtain the probability that each suspected window dressing financial statement data has been window dressing, and obtain the window dressing financial statement and the window dressing behavior corresponding to the window dressing financial statement based on the probability.

[0163] Update module 1306 is used to update the financial statement manipulation rule base based on the manipulated financial statements and the manipulation behavior.

[0164] Those skilled in the art should understand that Figure 13 The functions of each unit in the financial statement risk identification device 1300 shown can be understood by referring to the relevant description of the aforementioned financial statement risk identification method. Figure 13 The functions of each unit in the financial statement risk identification device 1300 shown can be implemented by a program running on a processor or by specific logic circuits.

[0165] In one possible implementation, the labeled sample set includes a positive sample set and a negative sample set. The acquisition module 1301 preprocesses the financial statement dataset to obtain a labeled sample set and an unlabeled sample set, including:

[0166] The financial statement dataset is composed of at least one financial statement data point collected from the target data source.

[0167] The financial statement data that has been determined to be unembellished is identified as the positive sample set, the financial statement data that has been determined to be embellished is identified as the negative sample set, and the financial statement data that cannot be determined to be embellished is identified as the unlabeled sample set.

[0168] In one possible implementation, the construction module 1302 constructs a feature matrix based on the cleaned unlabeled sample set, including:

[0169] A first feature matrix is ​​constructed for each enterprise based on its enterprise information and industry information. The enterprise information includes basic enterprise information and enterprise operating information, and the industry information includes basic industry information and industry operating information.

[0170] Based on the time dimension of the enterprise operating information and the industry operating information, the first feature matrix corresponding to each enterprise is processed to obtain the second feature matrix corresponding to each enterprise, and the second feature matrix corresponding to each enterprise is used as the feature matrix.

[0171] In one possible implementation, the comparison module 1303 performs comparative learning processing on the financial statement data of different companies in the corresponding time dimension based on the feature matrix, to obtain the implicit feature matrix of each company in different time dimensions, including:

[0172] Using enterprises as the unit, financial data from any quarter of any enterprise is selected as the anchor sample, financial data from enterprises that are not anchor samples are used as negative samples, and financial data from other quarters of the selected enterprise are used as positive samples.

[0173] Based on the anchor sample, the negative sample, and the positive sample selected each time, the positive sample selected in this time is input into a deep neural network for comparative learning processing, and the discriminative features of each enterprise in different time dimensions are determined by the comparative learning loss function.

[0174] The discriminative features are added to the feature matrix to obtain the latent feature matrix.

[0175] In one possible implementation, the screening module 1304 uses the supervised learning model to screen target samples from the unlabeled sample set that cannot be detected by existing rules, including:

[0176] A multilayer perceptron is used as the supervised learning model, wherein the activation function of the neurons in the multilayer perceptron is a noise linear rectified function;

[0177] The unlabeled sample set is used as input parameters to the classification model to obtain the probability of embellishment behavior.

[0178] Financial statement data with a probability of embellishment exceeding a preset first threshold are identified as the target samples.

[0179] In one possible implementation, the analysis module 1305 filters out suspected window dressing financial statement data from the target sample, including:

[0180] Determine the current industry in which the target sample belongs;

[0181] Calculate the proportion of known financial statement data with embellishment practices within the current industry to the total sample of the industry.

[0182] A target number of financial statement data are taken from the target sample as the suspected window dressing financial statement data, wherein the target number is the product of the total sample and the proportion of the total sample.

[0183] In one possible implementation, the analysis module 1305 performs attribution analysis on the suspected window dressing financial statement data to obtain the probability that each suspected window dressing financial statement data is window dressing, and obtains the window dressing financial statement and the corresponding window dressing behavior based on the probability, including:

[0184] The Shapley value of the suspected window dressing financial statements is predicted using a gradient boosting tree, wherein the Shapley value represents the probability that the suspected window dressing financial statements are indeed window dressing.

[0185] The suspected financial statement data with a probability greater than a preset second threshold are processed to obtain the financial statement and the corresponding manipulation behavior.

[0186] The aforementioned financial statement risk identification device has the following beneficial effects:

[0187] 1. With minimal financial statement manipulation, self-supervised learning on a large number of unlabeled samples can effectively learn the hidden representation (latent feature matrix) of the samples. This representation has high similarity to similar classes and high distinguishability between different classes.

[0188] 2. Utilizing this hidden layer representation, an effective classification model (supervised learning model) can be learned using a small number of labeled samples. Furthermore, the classification model can filter out target samples from the unlabeled sample set that cannot be detected by existing rules. 3. Attribution analysis can more effectively pinpoint unknown financial statement manipulation behaviors, improving the detection efficiency of such behaviors.

[0189] like Figure 14 As shown, Figure 14 This is a schematic diagram of the composition structure of the electronic device 1400 provided in the embodiments of this application. The electronic device 1400 includes:

[0190] The device 1400 includes a processor 1401, a storage medium 1402, and a bus 1403. The storage medium 1402 stores machine-readable instructions executable by the processor 1401. When the electronic device 1400 is running, the processor 1401 communicates with the storage medium 1402 via the bus 1403. The processor 1401 executes the machine-readable instructions to perform the steps of the financial statement risk identification method described in the embodiments of this application.

[0191] In practical applications, the various components in the electronic device 1400 are coupled together via bus 1403. It is understood that bus 1403 is used to achieve communication between these components. In addition to a data bus, bus 1403 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 14 The general designated all buses as Bus 1403.

[0192] The above-mentioned electronic devices have the following beneficial effects:

[0193] 1. With minimal financial statement manipulation, self-supervised learning on a large number of unlabeled samples can effectively learn the hidden representation (latent feature matrix) of the samples. This representation has high similarity to similar classes and high distinguishability between different classes.

[0194] 2. Utilizing this hidden layer representation, an effective classification model (supervised learning model) can be learned using a small number of labeled samples. Furthermore, the classification model can filter out target samples from the unlabeled sample set that cannot be detected by existing rules. 3. Attribution analysis can more effectively pinpoint unknown financial statement manipulation behaviors, improving the detection efficiency of such behaviors.

[0195] This application also provides a computer-readable storage medium storing executable instructions, which, when executed by at least one processor 801, implement the financial statement risk identification method described in this application.

[0196] In some embodiments, the storage medium may be a magnetic random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM), etc.; or it may be a device that includes one or any combination of the above-mentioned memories.

[0197] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0198] As an example, executable instructions may, but do not necessarily, correspond to files in the file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborating files (e.g., a file that stores one or more modules, subroutines, or code sections).

[0199] As an example, executable instructions can be deployed to execute on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.

[0200] The aforementioned computer-readable storage media have the following beneficial effects:

[0201] 1. With minimal financial statement manipulation, self-supervised learning on a large number of unlabeled samples can effectively learn the hidden representation (latent feature matrix) of the samples. This representation has high similarity to similar classes and high distinguishability between different classes.

[0202] 2. Utilizing this hidden layer representation, an effective classification model (supervised learning model) can be learned using a small number of labeled samples. Furthermore, the classification model (i.e., supervised learning model) can filter out target samples from the unlabeled sample set that cannot be detected by existing rules. 3. Attribution analysis can more effectively pinpoint unknown financial statement manipulation behaviors, improving the detection efficiency of such behaviors.

[0203] In the several embodiments provided in this application, it should be understood that the disclosed methods and electronic devices can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components may be combined, or integrated into another system, or some features may be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0204] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0205] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0206] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a platform server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0207] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A financial report risk identification method, characterized by, The method comprises the following steps: obtaining a financial report dataset and preprocessing the financial report dataset to obtain a labeled sample set and an unlabeled sample set, wherein the labeled sample set represents financial report data whose polishing has been determined, and the unlabeled sample set represents financial report data whose polishing cannot be determined; performing data cleaning on the unlabeled sample set, and constructing a feature matrix based on the cleaned unlabeled sample set, wherein the feature matrix is used to represent the financial report data of each enterprise in different time dimensions; based on the feature matrix, performing comparative learning processing on the financial report data of different enterprises in the corresponding time dimensions to obtain an implicit feature matrix of each enterprise in different time dimensions, wherein the implicit feature matrix includes a discrimination feature for describing whether there is polishing; constructing a supervised learning model based on the labeled sample set and the implicit feature matrix of each enterprise in different time dimensions, and screening out target samples that cannot be detected by existing rules from the unlabeled sample set through the supervised learning model; screening out suspected polished financial report data from the target samples, performing attribution analysis on the suspected polished financial report data to obtain the probability of existence of polishing for each suspected polished financial report data, and obtaining polished financial reports and polishing behaviors corresponding to the polished financial reports according to the probability; updating a financial report polishing rule library based on the polished financial reports and the polishing behaviors; wherein, based on the feature matrix, comparative learning processing is performed on the financial report data of different enterprises in the corresponding time dimensions to obtain an implicit feature matrix of each enterprise in different time dimensions, comprising: selecting the financial report data of any enterprise in any quarter as an anchor sample, selecting the financial report data of the enterprise of the non-anchor sample as a negative sample, and selecting the financial report data of other quarters of the enterprise as a positive sample; based on the anchor sample, the negative sample and the positive sample selected each time, inputting the positive sample selected this time into a deep neural network for comparative learning processing, and determining the discrimination feature of each enterprise in different time dimensions through a comparative learning loss function; adding the discrimination feature to the feature matrix to obtain the implicit feature matrix.

2. The method of claim 1, wherein, The labeled sample set includes a positive sample set and a negative sample set, and the preprocessing of the financial report dataset to obtain a labeled sample set and an unlabeled sample set comprises: collecting at least one financial report from a target data source to form the financial report dataset; determining the financial report data that has been determined to have no polishing as the positive sample set, determining the financial report data that has been determined to have polishing as the negative sample set, and determining the financial report data that cannot be determined whether it has polishing as the unlabeled sample set.

3. The method of claim 1, wherein, The construction of the feature matrix based on the cleaned unlabeled sample set comprises: constructing a first feature matrix corresponding to each enterprise according to the enterprise information and industry information of each enterprise, wherein the enterprise information includes enterprise basic information and enterprise operating information, and the industry information includes industry basic information and industry operating information; Derive the first feature matrix corresponding to each enterprise to obtain a second feature matrix corresponding to each enterprise based on a time dimension of the enterprise operation information and the industry operation information, and take the second feature matrix corresponding to each enterprise as the feature matrix.

4. The method of claim 1, wherein, The target sample that cannot be detected by the existing rules is screened out from the unlabeled sample set by the supervised learning model, and the target sample that cannot be detected by the existing rules is screened out from the unlabeled sample set by the supervised learning model. The multilayer perceptron is taken as the supervised learning model, and the activation function of a neuron in the multilayer perceptron is a noise linear rectifier function. The unlabeled sample set is taken as an input parameter to input the supervised learning model to obtain a falsification behavior probability. The financial report data with the falsification behavior probability greater than a preset first threshold value is determined as the target sample.

5. The method of claim 1, wherein, The target sample that cannot be detected by the existing rules is screened out from the unlabeled sample set by the supervised learning model, and the target sample that cannot be detected by the existing rules is screened out from the unlabeled sample set by the supervised learning model. The current industry in which the target sample is located is determined. The overall sample proportion of the financial report data with the falsification behavior in the total sample in the current industry is calculated. A target number of financial report data are taken from the target sample as the suspected falsification financial report data, and the target number is a product of the total sample and the overall sample proportion.

6. The method of claim 1, wherein, The suspected falsification financial report data are subjected to attribution analysis to obtain a probability that each suspected falsification financial report data has falsification, and a falsification financial report and a falsification behavior corresponding to the falsification financial report are obtained according to the probability, and the suspected falsification financial report data are subjected to attribution analysis to obtain a probability that each suspected falsification financial report data has falsification, and a falsification financial report and a falsification behavior corresponding to the falsification financial report are obtained according to the probability. The Shapley value of the suspected falsification financial report data is predicted by the gradient boosting tree, and the Shapley value represents the probability that the suspected falsification financial report data has falsification. The suspected falsification financial report data with the probability greater than a preset second threshold value are subjected to discrimination processing to obtain a falsification financial report and a falsification behavior corresponding to the falsification financial report.

7. A financial report risk identification device characterized by comprising: The device comprises: An acquisition module is configured to acquire a financial report data set, and pre-process the financial report data set to obtain a labeled sample set and an unlabeled sample set. The labeled sample set represents financial report data that has been determined to have falsification or not, and the unlabeled sample set represents financial report data that cannot be determined to have falsification or not. A construction module is configured to perform data cleaning on the unlabeled sample set, and construct a feature matrix based on the cleaned unlabeled sample set. The feature matrix is used to represent the financial report data of each enterprise in different time dimensions. A comparison module is configured to perform comparative learning processing on the financial report data of different enterprises in corresponding time dimensions based on the feature matrix to obtain an implicit feature matrix of each enterprise in different time dimensions. The implicit feature matrix includes a distinguishing feature for describing whether to have falsification. A screening module is configured to construct a supervised learning model based on the labeled sample set and the implicit feature matrix of each enterprise in different time dimensions, and screen out a target sample that cannot be detected by the existing rules from the unlabeled sample set by the supervised learning model. The analysis module is configured to screen out suspected financial report polishing data from the target sample, perform attribution analysis on the suspected financial report polishing data, obtain a probability of existence of polishing for each suspected financial report polishing data, and obtain polished financial reports and polishing behaviors corresponding to the polished financial reports according to the probability. The updating module is configured to update a financial report polishing rule library based on the polished financial reports and the polishing behaviors. The comparison module is configured to perform comparative learning processing on the financial report data of different enterprises in the corresponding time dimension based on the feature matrix, to obtain an implicit feature matrix of each enterprise in different time dimensions, including: selecting, in units of enterprises, financial report data of an arbitrary enterprise in an arbitrary quarter as anchor samples, financial report data of enterprises in which non-anchor samples are located as negative samples, and financial report data of other quarters of the selected enterprise as positive samples; performing comparative learning processing on the positive samples selected this time by inputting the anchor samples, the negative samples and the positive samples selected each time into a deep neural network, and determining the degree of distinction features of each enterprise in different time dimensions through a comparative learning loss function; adding the degree of distinction features to the feature matrix to obtain the implicit feature matrix.

8. An electronic device, comprising: The processor, the storage medium and the bus, the storage medium stores the machine readable instructions executable by the processor, when the electronic equipment runs, the processor and the storage medium are communicated through the bus, the processor executes the machine readable instructions, to execute the financial report risk identification method as any one of claims 1 to 6. The computer readable storage medium stores a computer program, and the computer program is executed by the processor to perform the financial report risk identification method as any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, ​

Citation Information

Patent Citations

  • Model training method and device, abnormal data detection method and device and electronic equipment

    CN111428757A