Risk information pushing method and device, electronic equipment and storage medium
By acquiring indicator values from financial reports and related information, and using binning intervals and weights to generate risk probability representation values, the problem of low accuracy of financial report fraud risk information in existing technologies has been solved, enabling more accurate risk information delivery and enhancing the reliability of user judgment.
Patent Information
- Application Number
- CN202511016923.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-11-21
AI Technical Summary
In existing technologies, the accuracy of financial report fraud risk information generated based on industry expert comparison rules is low, making it difficult for users to accurately judge the authenticity of financial reports.
By acquiring textual data, related information, and public opinion data from financial reports, the indicator values of various financial, corporate attributes, and rule indicators are determined. Indicator characteristics are identified within binning intervals, and risk probability representation values and risk values are generated using weights and pre-built parameters. Abnormal indicators are identified, and more accurate risk information is generated.
It improves the accuracy of financial reporting fraud risk information, enabling more accurate delivery of risk information to users, reducing the impact of abnormal indicators, providing overall risk values and risk values for abnormal indicators, and enhancing the reliability of users' judgments.
Smart Images

Figure CN120994832A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a risk information push method, apparatus, electronic device, and storage medium. Background Technology
[0002] To protect their investment interests, investors often focus on the financial reports of listed companies. Financial reports disclose information such as a company's financial condition, operating results, and cash flows; therefore, they are not only information-rich but also highly technical. Furthermore, listed companies, for their own benefit, may falsify their financial reports. Faced with such information-rich and technical financial reports, ordinary users often struggle to fully understand the disclosed information, making it even more difficult to detect whether the reports are falsified.
[0003] In light of the above, the trading platform provides users with risk information regarding the potential for falsified financial reports, helping them determine whether the financial reports are fraudulent, thereby encouraging users to trade investment products on the platform.
[0004] Currently, financial reports are typically compared based on comparison rules provided by industry experts. Then, risk information indicating the risk of financial report fraud is generated based on the comparison results. However, the accuracy of the generated risk information is low, and the accuracy of the risk information pushed to users is also low. Summary of the Invention
[0005] The purpose of this invention is to provide a risk information push method, apparatus, electronic device, and storage medium to generate more accurate risk information characterizing the risk of fraud in the financial reports of a target company, thereby improving the accuracy of risk information pushed to users. The specific technical solution is as follows:
[0006] According to one aspect of the present invention, a risk information push method is provided, the method comprising:
[0007] Based on the textual data of the target company's financial reports, the values of each financial indicator are obtained.
[0008] Based on the text data of the target company's associated information, the index values of each company attribute indicator and each rule indicator of the target company are obtained.
[0009] Within the binning intervals of each target indicator, the binning interval to which the indicator value of each target indicator belongs is determined. Based on the determined binning intervals, the indicator characteristics of the target indicators are obtained. The target indicators include: each financial indicator, each company attribute indicator, and each rule indicator.
[0010] Based on the weight and characteristics of each target indicator, a risk probability representation value is generated. The risk probability representation value represents the ratio between the probability that the financial report has a risk of fraud and the probability that it does not have a risk of fraud.
[0011] A first risk value for the financial report to be fraudulent is generated based on the risk probability characterization value, the pre-constructed first parameter and the second parameter; and a second risk value for the financial report to be fraudulent in each target indicator is generated based on the second parameter, the weight of each target indicator and the indicator characteristics.
[0012] Based on the obtained second risk value, identify the abnormal indicators in the target indicators;
[0013] Based on the first risk value and the abnormal indicators, risk information is generated that indicates the risk of fraud in the financial report.
[0014] The risk information is pushed to the target user's device.
[0015] In one embodiment of the present invention, generating a risk probability characterization value based on the weight and characteristics of each target indicator includes: for each target indicator, determining a risk probability benchmark value corresponding to the indicator characteristics of the target indicator based on the indicator characteristics of the target indicator, calculating the product of the determined risk probability benchmark value and the weight of the target indicator as the probability score corresponding to the target indicator, and using the sum of the probability scores of each target indicator as the risk probability characterization value.
[0016] and / or
[0017] The step of generating a first risk value for the financial report to have a risk of fraud based on the risk probability characterization value, a pre-constructed first parameter, and a second parameter includes: using the product of the pre-constructed second parameter and the risk probability characterization value as the risk increase value of the financial report, using the pre-constructed first parameter as a benchmark risk value, and determining the first risk value for the financial report to have a risk of fraud based on the difference between the benchmark risk value and the risk increase value.
[0018] and / or
[0019] The step of generating a second risk value for the financial report in each target indicator based on the second parameter, the weight of each target indicator, and the indicator characteristics includes: for each target indicator, determining the risk probability benchmark value corresponding to the target indicator based on the indicator characteristics of the target indicator, and using the product of the determined risk probability benchmark value, the weight of the target indicator, and the second parameter as the second risk value characterizing the financial report in each target indicator.
[0020] In one embodiment of the present invention, the method further includes:
[0021] Obtain a third-party audit opinion on the aforementioned financial reports;
[0022] Determine the risk level described in the third-party audit opinion;
[0023] The risk information generated from the first risk value and anomaly indicators to produce the financial report includes:
[0024] The risk information in the financial report is generated based on the risk level, the first risk value, and the anomaly indicator.
[0025] In one embodiment of the present invention, the target indicator further includes: public opinion indicators;
[0026] The index values of the aforementioned public opinion indicators are obtained in the following manner:
[0027] Obtain public opinion data about the target company;
[0028] Determine the sentiment direction of the obtained public opinion data, wherein the sentiment direction is positive or negative;
[0029] Using time units as the unit, determine the first quantity of positive public opinion data and the second quantity of negative public opinion data in each time unit covered by the obtained public opinion data;
[0030] The index values of public opinion indicators are generated based on the first quantity, the second quantity, and the time difference between each time unit and the current time.
[0031] In one embodiment of the present invention, the first parameter and the second parameter are constructed in the following manner:
[0032] Sample indicator values for each target indicator were obtained for the sample companies;
[0033] Based on the sample index values, a multinomial fitting is performed with each target index as the independent variable to obtain the weight of each target index.
[0034] Within the binning intervals of each target indicator, determine the binning interval to which the sample indicator value of each target indicator belongs;
[0035] For each target indicator, the sample distribution difference value for each bin interval under that target indicator is determined, wherein the sample distribution difference value represents the difference between the number of positive samples and the number of negative samples in the sample indicator value belonging to the bin interval.
[0036] Based on the sample distribution differences of each bin interval under each target indicator and the weight of each target indicator, the standard risk probability characterization value is determined.
[0037] The second parameter is constructed based on the preset risk increase value and the standard risk probability characterization value;
[0038] The first parameter is constructed based on the preset standard risk value, the standard risk probability representation value, and the second parameter.
[0039] In one embodiment of the present invention, the sample indicator value is labeled;
[0040] The sample index value is determined to be a negative sample based on any of the following events identified by the label:
[0041] Financial restatement events of the sample companies;
[0042] Financial inquiry events at the sample companies;
[0043] Tax violations by the sample companies;
[0044] Financial fraud cases involving sample companies.
[0045] In one embodiment of the present invention, the characteristic indicators used for risk assessment are determined from among the candidate indicators in the following manner:
[0046] For each candidate indicator, the sample indicator values belonging to positive samples and the sample indicator values belonging to negative samples are obtained for the sample companies. The candidate indicators include: candidate financial indicators, candidate company attribute indicators and candidate rule indicators.
[0047] Based on the sample index values corresponding to each candidate index, determine the impact value of each candidate index on the risk probability characterization value;
[0048] Based on the determined impact values of each candidate indicator, the target indicator is selected from the candidate indicators.
[0049] According to another aspect of the present invention, a risk information push device is provided, the device comprising:
[0050] The indicator value acquisition module is used to obtain the indicator values of each financial indicator based on the text data of the target company's financial report; and to obtain the indicator values of each company attribute indicator and each rule indicator of the target company based on the text data of the target company's related information.
[0051] The indicator feature acquisition module is used to determine the binning interval to which the indicator value of each target indicator belongs in the binning interval of each target indicator, and to obtain the indicator features of the target indicator based on the determined binning interval. The target indicators include: each financial indicator, each company attribute indicator and each rule indicator.
[0052] The representation value generation module is used to generate a risk probability representation value based on the weight and characteristics of each target indicator. The risk probability representation value represents the ratio between the probability that the financial report has a risk of fraud and the probability that it does not have a risk of fraud.
[0053] The risk value generation module is used to generate a first risk value for the financial report to have a risk of fraud based on the risk probability characterization value, a pre-constructed first parameter and a second parameter, and to generate a second risk value for the financial report to have a risk of fraud in each target indicator based on the second parameter, the weight of each target indicator and the indicator characteristics.
[0054] The abnormal indicator determination module is used to determine the abnormal indicators in the target indicators based on the obtained second risk value.
[0055] The risk information generation module is used to generate risk information that characterizes the financial report as having a risk of being falsified, based on the first risk value and the abnormal indicators.
[0056] The risk information push module is used to push the risk information to the target user terminal.
[0057] In one embodiment of the present invention, the characterization value generation module is specifically used to determine the risk probability benchmark value corresponding to the indicator characteristics of each target indicator based on the indicator characteristics of the target indicator, calculate the product of the determined risk probability benchmark value and the weight of the target indicator as the probability score corresponding to the target indicator, and use the sum of the probability scores of each target indicator as the risk probability characterization value.
[0058] and / or
[0059] The risk value generation module is specifically used to take the product of the pre-constructed second parameter and the risk probability characterization value as the risk increase value of the financial report, take the pre-constructed first parameter as the benchmark risk value, and determine the first risk value of the financial report as having a risk of fraud based on the difference between the benchmark risk value and the risk increase value.
[0060] and / or
[0061] The risk value generation module is specifically used to determine the risk probability benchmark value corresponding to each target indicator based on the indicator characteristics of the target indicator, and to use the product of the determined risk probability benchmark value, the weight of the target indicator and the second parameter as the second risk value characterizing the financial report as having a risk of fraud in each target indicator.
[0062] In one embodiment of the present invention, the apparatus further includes:
[0063] Audit Opinion Obtaining Module. Used to obtain a third-party audit opinion on the financial statements.
[0064] The risk level determination module is used to determine the risk level described in the third-party audit opinion;
[0065] The risk information generation module is specifically used to generate risk information for the financial report based on the risk level, the first risk value, and the anomaly indicator.
[0066] In one embodiment of the present invention, the target indicator further includes: public opinion indicators;
[0067] The index values of the aforementioned public opinion indicators are obtained in the following manner:
[0068] Obtain public opinion data about the target company;
[0069] Determine the sentiment direction of the obtained public opinion data, wherein the sentiment direction is positive or negative;
[0070] Using time units as the unit, determine the first quantity of positive public opinion data and the second quantity of negative public opinion data in each time unit covered by the obtained public opinion data;
[0071] The index values of public opinion indicators are generated based on the first quantity, the second quantity, and the time difference between each time unit and the current time.
[0072] In one embodiment of the present invention, the first parameter and the second parameter are constructed in the following manner:
[0073] Sample indicator values for each target indicator were obtained for the sample companies;
[0074] Based on the sample index values, a multinomial fitting is performed with each target index as the independent variable to obtain the weight of each target index.
[0075] Within the binning intervals of each target indicator, determine the binning interval to which the sample indicator value of each target indicator belongs;
[0076] For each target indicator, the sample distribution difference value for each bin interval under that target indicator is determined, wherein the sample distribution difference value represents the difference between the number of positive samples and the number of negative samples in the sample indicator value belonging to the bin interval.
[0077] Based on the sample distribution differences of each bin interval under each target indicator and the weight of each target indicator, the standard risk probability characterization value is determined.
[0078] The second parameter is constructed based on the preset risk increase value and the standard risk probability characterization value;
[0079] The first parameter is constructed based on the preset standard risk value, the standard risk probability representation value, and the second parameter.
[0080] In one embodiment of the present invention, the sample indicator value is labeled;
[0081] The sample index value is determined to be a negative sample based on any of the following events identified by the label:
[0082] Financial restatement events of the sample companies;
[0083] Financial inquiry events at the sample companies;
[0084] Tax violations by the sample companies;
[0085] Financial fraud cases involving sample companies.
[0086] In one embodiment of the present invention, the characteristic indicators used for risk assessment are determined from among the candidate indicators in the following manner:
[0087] For each candidate indicator, the sample indicator values belonging to positive samples and the sample indicator values belonging to negative samples are obtained for the sample companies. The candidate indicators include: candidate financial indicators, candidate company attribute indicators and candidate rule indicators.
[0088] Based on the sample index values corresponding to each candidate index, determine the impact value of each candidate index on the risk probability characterization value;
[0089] Based on the determined impact values of each candidate indicator, the target indicator is selected from the candidate indicators.
[0090] According to another aspect of the present invention, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0091] Memory, used to store computer programs;
[0092] The processor, when executing a program stored in memory, implements any of the aforementioned risk information push methods.
[0093] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein a computer program is stored therein, and the computer program, when executed by a processor, implements any of the above-described risk information push methods.
[0094] According to another aspect of the present invention, a computer program product containing instructions is provided, which, when run on a computer, causes the computer to execute any of the risk information push methods described above.
[0095] Beneficial effects of the embodiments of the present invention:
[0096] In the risk information push method provided by this invention, the indicator values of target indicators, including various financial indicators, company attribute indicators, and rule indicators, are obtained. By determining the indicator characteristics corresponding to the binning interval to which the indicator value of the target indicator belongs, the impact of abnormal indicator values of the target indicators on the generated risk information can be reduced. Furthermore, by using indicators of multiple dimensions, risk probability representation values can be generated more accurately based on the sum and characteristics of each target indicator. Based on the accurate risk probability representation value and the pre-constructed first and second parameters, a first risk value that can represent the probability of fraud in financial reports can be generated. Also, for each target indicator, according to the first... The second risk value of the target indicator is generated by two parameters, the weight of the target indicator, and the indicator characteristics. The second risk value can accurately represent the probability of fraud risk in financial reports under different target indicators. Based on such a second risk value, abnormal indicators in the target indicators can be identified. Then, the risk information of the financial report generated based on the first risk value and abnormal indicators includes the risk value of fraud risk in the financial report as a whole, as well as the abnormal indicators with existing fraud risk and the risk value of fraud risk in the financial report under abnormal indicators. Such risk information can more accurately represent the fraud risk of the target company's financial report and can push more accurate risk information to users.
[0097] Of course, implementing any product or method of the present invention does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description
[0098] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings.
[0099] Figure 1 This is a flowchart illustrating a risk information push method provided in an embodiment of the present invention;
[0100] Figure 2 A flowchart illustrating another risk information push method provided in an embodiment of the present invention;
[0101] Figure 3 A flowchart illustrating a method for obtaining the index value of a public opinion indicator provided in an embodiment of the present invention;
[0102] Figure 4 A flowchart illustrating a parameter construction method provided in an embodiment of the present invention;
[0103] Figure 5This is a schematic diagram of the structure of a risk information push device provided in an embodiment of the present invention;
[0104] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0105] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art based on the present invention are within the scope of protection of the present invention.
[0106] The execution subject of the embodiments of the present invention will be described below.
[0107] The solutions provided in the embodiments of the present invention can be applied to electronic devices such as desktop computers, laptops, tablets, and servers, and can also be applied to information push platforms for risk information push. For ease of description, the executing entity of the risk information push method provided in the embodiments of the present invention is collectively referred to as the information push platform.
[0108] The following is an explanation of each target indicator.
[0109] I. Financial Indicators
[0110] Financial metrics are indicators that reflect the financial information of a target company, determined based on data from its financial reports. Examples of financial metrics include: the ratio of operating cash flow to net profit, the ratio of accounts receivable to total assets, accounts receivable index, asset quality index, depreciation rate index, total asset growth rate, sales growth rate, inventory growth rate, total asset turnover, debt-to-equity ratio, etc.
[0111] For example, financial indicators such as total asset growth rate and sales growth rate can include the values of total asset growth rate and sales growth rate.
[0112] II. Company Attribute Indicators
[0113] Company attribute indicators are metrics that reflect the basic attribute information of the target company. For example, company attribute indicators may include: company type indicators, company legal attribute indicators, company equity attribute indicators, company industry indicators, company brand influence indicators, public opinion indicators, etc. The method for obtaining the value of the public opinion indicator will be explained in the examples below and will not be detailed here.
[0114] For company attribute indicators, such as company type and industry, which do not have specific numerical values, the company type and industry can be converted into corresponding representation values. For example, if the company attribute indicator includes a company type indicator, the indicator value can be the representation value corresponding to the target company's company type. For instance, if the company types include state-owned enterprises, private enterprises, foreign-invested enterprises, and public enterprises, the company type representation value for state-owned enterprises can be pre-set as 1, for private enterprises as 2, for foreign-invested enterprises as 3, and for public enterprises as 4. If the target company is a private enterprise, then the indicator value for the company type indicator could be 2.
[0115] III. Rules and Indicators
[0116] The rule indicators are indicators obtained by judging the company's related party information according to preset judgment rules. For each rule indicator, the judgment rules can be preset. For example, the judgment rules for a rule indicator may include: whether the company has had a financial restatement event, whether the company has received an inquiry letter, whether the company has had tax violations, whether the company has committed financial fraud, whether the company has engaged in contract fabrication, whether the company has signed but not delivered orders, whether the company has delayed recognizing expenses, etc. For example, the indicator value of a rule indicator may include a confirmation value or a denial value, where the confirmation value can be 1 and the denial value can be 0. Therefore, a rule indicator value of 1 indicates that the judgment result of the company's related party information under this rule indicator is yes, and a rule indicator value of 0 indicates that the judgment result of the company's related party information under this rule indicator is no.
[0117] The risk information push method provided in the embodiments of the present invention will be described in detail below.
[0118] In one embodiment of the present invention, see Figure 1 A flowchart illustrating a risk information push method is provided, comprising the following steps S101-S107:
[0119] Step S101: Based on the text data of the target company's financial report, obtain the indicator values of each financial indicator; based on the text data of the target company's related information, obtain the indicator values of each company attribute indicator and each rule indicator of the target company.
[0120] The method for obtaining indicator values is explained below.
[0121] I. Financial Indicators
[0122] Financial indicators can include those recorded in financial reports, or those calculated based on the textual data of financial reports. When financial indicators include those recorded in financial reports, the information push platform can identify the corresponding indicators from the publicly available textual data of the target company's financial reports, and use the value of the corresponding indicator in the financial report as the indicator value of the financial indicator. When financial indicators include those calculated based on the data of indicators recorded in financial reports, input indicators can be pre-set. The financial indicator can then be calculated based on these input indicators. The information push platform can then identify the corresponding indicators from the textual data of the financial reports, use the value of the corresponding indicator in the financial report as the value of the input indicator, and obtain the indicator value of the financial indicator based on the value of the input indicator.
[0123] II. Company Attribute Indicators
[0124] The information push platform can obtain text data of the target company's related information, and then determine the target company's indicator value under each company attribute indicator based on the obtained text data of related information. For example, for each company attribute indicator, it identifies the text data related to the company's attribute indicator in the text data of the related information, performs semantic parsing on the text data related to the company's attribute indicator, and obtains the indicator value of the company's attribute indicator. For a specific example of generating indicator values, please refer to the above embodiment, which will not be detailed here.
[0125] III. Rules and Indicators
[0126] The information push platform can obtain the associated information of the target company, and then judge the associated information of the company according to the preset judgment rules corresponding to the rule indicator for each rule indicator, and determine the indicator value of the rule indicator.
[0127] Information push platforms can obtain text data related to target companies through the following methods: calling preset data acquisition interfaces to obtain text data related to target companies, downloading text data of publicly available information about target companies from preset information disclosure platforms, and using web scraping tools to obtain text data related to target companies from the internet. For example, web scraping tools can be used to obtain text data related to target companies from internet platforms such as social media platforms, news platforms, and forum discussion platforms.
[0128] Step S102: Determine the binning interval to which the indicator value of each target indicator belongs in the binning interval of each target indicator, and obtain the indicator characteristics of the target indicator based on the determined binning interval.
[0129] The target indicators include: various financial indicators, various company attribute indicators, and various rule indicators. The binning ranges for each target indicator are predetermined.
[0130] A target indicator includes multiple sub-bins, and each sub-bin has corresponding indicator features. Taking the target indicator as the total asset growth rate as an example, if the numerical range of the total asset growth rate indicator is (-100%, 200%), then the corresponding bin intervals for this target indicator include: (-100%, -50%), (-50%, -25%), (-25%, 0%), (0%, 5%), (5%, 10%), (10%, 25%), (25%, 50%), (50%, 100%), (100%, 150%), and (150%, 200%). The indicator characteristics corresponding to the bin intervals are as follows: the indicator characteristic corresponding to the bin interval (-100%, -50%) is 1; the indicator characteristic corresponding to the bin interval (-50%, -25%) is 2; the indicator characteristic corresponding to the bin interval (-25%, 0%) is 3; and the indicator characteristic corresponding to the bin interval (0%, 5%) is... The corresponding indicator characteristics are: 4. The indicator characteristics corresponding to the binning interval (5%, 10%) are: 5. The indicator characteristics corresponding to the binning interval (10%, 25%) are: 6. The indicator characteristics corresponding to the binning interval (25%, 50%) are: 7. The indicator characteristics corresponding to the binning interval (50%, 100%) are: 8. The indicator characteristics corresponding to the binning interval (100%, 150%) are: 9. The indicator characteristics corresponding to the binning interval (150%, 200%) are: 10. If the obtained total asset growth rate indicator value is 13%, then from the binning interval corresponding to the predetermined total asset growth rate target indicator, the binning interval to which the total asset growth rate indicator value belongs is determined to be (10%, 25%). Therefore, the indicator characteristic corresponding to (10%, 25%) is determined to be 6. Thus, the indicator characteristic of the obtained total asset growth rate indicator value is 6.
[0131] The following explains how to determine the binning intervals for each target indicator.
[0132] In one implementation, the information push platform can obtain sample indicator values for each target indicator from the sample companies, and then perform binning on the obtained sample indicator values for each target indicator to obtain the binning intervals under each target indicator.
[0133] For example, information push platforms can employ at least one of the following binning methods: equal-width binning, equal-frequency binning, chi-square binning, quantile binning, clustering binning, and supervised binning using decision trees. For instance, supervised binning using decision trees can avoid the subjective influence of manually setting binning boundary values. Further discretizing the values of each target indicator can effectively improve the robustness of risk information generation.
[0134] Step S103: Generate risk probability representation values based on the weights and characteristics of each target indicator.
[0135] The risk probability representation value represents the ratio between the probability of financial reporting being fraudulent and the probability of it not being fraudulent.
[0136] In one implementation, for each target indicator, the information push platform can determine the risk probability benchmark value corresponding to the indicator characteristics of the target indicator based on the indicator characteristics of the target indicator and the sample distribution characteristics of the sample companies representing positive and negative samples in each target indicator. The platform can then calculate the product of the determined risk probability benchmark value and the weight of the target indicator as the probability score corresponding to the target indicator. Finally, the sum of the probability scores of each target indicator is used as the risk probability representation value.
[0137] Specifically, for each target metric, the information push platform can pre-record the correspondence between the target metric and its weight. The above sample distribution characteristics are determined based on the sample metric values of each target metric for the sample companies.
[0138] The following explains the correspondence between target indicators and weights, and the correspondence between indicator characteristics and risk probability benchmark values.
[0139] Each target indicator is assigned a corresponding weight. For example, if there are n target indicators used to generate risk information, then there are also n weights.
[0140] Each risk probability benchmark value corresponds one-to-one with the indicator characteristics of the target indicator. As mentioned above, a target indicator includes multiple bin intervals, and each bin interval has corresponding indicator characteristics. If the first target indicator includes k1 bin intervals, then the first target indicator also includes k1 indicator characteristics. Similarly, the first target indicator also includes k1 risk probability benchmark values.
[0141] The following example illustrates how to generate risk probability representation values.
[0142] In one approach, after obtaining the indicator characteristics of the target indicator from the information push platform, the risk probability benchmark value corresponding to the indicator characteristics can be determined based on the sample distribution characteristics of sample companies representing positive and negative samples for each target indicator. The weight of the target indicator is determined according to the correspondence between the target indicator and its weight. Then, the product of the risk probability benchmark value and the weight of the target indicator is calculated, and the sum of the probability scores of each target indicator is calculated to obtain the risk probability representation value.
[0143] Suppose there are n target metrics, and the i-th target metric includes k. i Given a set of indicator characteristics, the i-th target indicator includes k... iGiven a risk probability benchmark, if the indicator feature of the i-th target indicator obtained by the information push platform is k... i If the j-th indicator feature is selected from the 10 indicator features, then the risk probability benchmark value of the j-th indicator feature under the 1-th target indicator can be determined. If the weight corresponding to the i-th target indicator is β i The calculated probability score corresponding to the i-th target indicator is then: Therefore, the risk probability representation value is:
[0144] In this way, the risk probability benchmark value corresponding to the indicator characteristics of the target indicator can be adjusted by using the weights corresponding to different target indicators. This allows for a more accurate determination of the probability scores corresponding to different target indicators, resulting in a more accurate calculated risk probability representation value.
[0145] Furthermore, the magnitude of the indicator characteristics of the target indicator has a two-way impact on the probability of financial reporting fraud. That is, the larger the value of the indicator characteristics of the target indicator, the greater the probability of financial reporting fraud, and vice versa. Generating a risk probability representation value in the above manner can correct the two-way impact of the magnitude of different indicator characteristics on the probability of financial reporting fraud, and obtain a risk probability representation value that can accurately represent the probability of financial reporting fraud relative to the probability of no fraud.
[0146] Step S104: Generate a first risk value for the financial report to have a risk of fraud based on the risk probability characterization value, the pre-constructed first parameter and second parameter, and generate a second risk value for the financial report to have a risk of fraud in each target indicator based on the second parameter, the weight of each target indicator and the indicator characteristics.
[0147] The method for generating the first risk value is explained below.
[0148] In one implementation, the product of a pre-constructed second parameter and a risk probability representation value is used as the risk increase value of the financial report, and the pre-constructed first parameter is used as a benchmark risk value. Based on the difference between the benchmark risk value and the risk increase value, a first risk value indicating the risk of fraud in the financial report is determined.
[0149] The difference between the benchmark risk value and the risk augmentation value can be the difference between the benchmark risk value and the risk augmentation value. In this case, the higher the resulting first risk value, the higher the probability of fraud in the financial report. Alternatively, the difference between the benchmark risk value and the risk augmentation value can also be the difference between the risk augmentation value and the benchmark risk value. In this case, the higher the resulting first risk value, the lower the probability of fraud in the financial report.
[0150] In this way, the risk increase value of financial reporting can be accurately calculated, and the difference between the benchmark risk value and the risk increase value can be used to generate the first risk value for the risk of financial reporting fraud.
[0151] The method for generating the second risk value is explained below.
[0152] In one implementation, for each target indicator, a risk probability benchmark value corresponding to the target indicator is determined based on the indicator characteristics of the target indicator and the sample distribution characteristics of the sample companies representing positive and negative samples in each target indicator. The product of the determined risk probability benchmark value, the weight of the target indicator, and the second parameter is used as the second risk value representing the risk of fraud in the financial report in each target indicator.
[0153] The second risk value is illustrated below with an example.
[0154] Suppose there are n target metrics, and the i-th target metric includes k. i Given a set of indicator characteristics, the i-th target indicator includes k... i Given a risk probability benchmark, if the indicator feature of the i-th target indicator obtained by the information push platform is k... i If the j-th indicator feature is selected from the 10 indicator features, then the risk probability benchmark value of the j-th indicator feature under the 1-th target indicator can be determined. If the second parameter is The weight corresponding to the i-th target indicator is: β i The generated second risk value can be: The generated second risk value can also be:
[0155] In this way, a corresponding second risk value can be generated for different target indicators. When the first risk value indicates a high probability of financial reporting fraud, attribution analysis can be performed based on the value of the second risk value of different target indicators. This analysis can identify the indicators that indicate a high probability of financial reporting fraud and identify abnormal indicators with high risk.
[0156] Step S105: Determine the abnormal indicators in the target indicators based on the obtained second risk value.
[0157] In one implementation, for each target indicator, the information push platform can determine whether the target indicator is an abnormal indicator by judging whether the second risk value corresponding to the target indicator is less than a preset abnormal threshold. Assume that the second risk value of the target indicator is: like Then the target indicator to which the second risk value belongs is determined to be an abnormal indicator.
[0158] In another implementation, the information push platform can determine that the target indicator to which the second risk value within the preset abnormal indicator range belongs is an abnormal indicator.
[0159] The following explains how to determine whether the second risk value falls within the preset abnormality index range.
[0160] In one scenario, a lower second risk value indicates a higher probability of financial reporting fraud under the target indicators. In this case, the second risk value can be compared to a preset upper limit to determine if it falls within the preset abnormal indicator range. If the second risk value is less than or equal to the preset upper limit, it is determined to be within the preset abnormal indicator range; if it is greater than the preset upper limit, it is determined to be outside the preset abnormal indicator range. For example, the preset abnormal indicator range can be set as [Risk...]. min Risk up ], among which, Risk min Risk is the minimum value within the range of values for the second risk value. up This is the preset upper limit of the risk value.
[0161] In another scenario, a higher second risk value indicates a higher probability of financial reporting fraud under the target indicators. In this case, the second risk value can be compared to a preset lower limit to determine if it falls within the preset abnormal indicator range. If the second risk value is greater than or equal to the preset lower limit, it is determined to be within the preset abnormal indicator range; if it is less than the preset lower limit, it is determined to be outside the preset abnormal indicator range. For example, the preset abnormal indicator range can be set as [Risk...]. down Risk max ], among which, Risk max Risk is the maximum value within the range of the second risk value. down This is the preset lower limit of the risk value.
[0162] Step S106: Based on the first risk value and anomaly indicators, generate risk information that characterizes the risk of fraud in financial reports.
[0163] In one implementation, the information push platform can also display the value of the first risk value, the abnormal indicator, and the second risk value corresponding to the abnormal indicator as risk information representing the risk of fraud in the financial report. Furthermore, the information push platform can determine the risk level of the financial report fraud risk based on the first risk value, and generate a visual chart based on the abnormal indicator and the second risk value corresponding to the abnormal indicator, displaying the risk information including the risk level and the visual chart. The risk level can include: low risk level, medium risk level, and high risk level, and can also include: risk level 1, risk level 2, risk level 3, risk level 4, etc., for example, a higher risk level indicates a higher probability of financial report fraud.
[0164] Step S107: Push risk information to the target user's terminal.
[0165] In one implementation, the information push platform can respond to a risk information request sent by a user's client, designate the client that sent the risk information request as the target client, and push the risk information to the target client.
[0166] The risk information request can be a request sent by a user through a client application. This request may include a company identifier. The information push platform can identify the company corresponding to that identifier as the target company and display its corresponding risk information to the user. The information push platform can periodically generate risk information for each target company according to a preset frequency, such as once daily or once every three days. The platform can then store this generated risk information and respond to any risk information request based on the latest stored risk information for the target company.
[0167] Information push platforms can use chart generation tools to create visual charts displaying risk information, including the primary risk value and anomaly indicators. Based on these generated charts, they can respond to risk information requests and display risk information to users. For example, chart generation tools could include ECharts (a data visualization tool) and EasyExcel (a chart processing tool). For instance, the risk level corresponding to the primary risk value can be determined based on the primary risk value, and the corresponding risk level can be displayed using an ECharts dashboard. For each anomaly indicator and its corresponding secondary risk value, EasyExcel can be used to export visual tables for easy viewing by users.
[0168] In another implementation, users can subscribe to the risk information push service of the target company through their client. The information push platform can use the client that has subscribed to the risk information push service of the target company as the target client. The information push platform can push the latest risk information to the target client each time the risk information of the target company is generated.
[0169] In the risk information push method provided by this invention, the indicator values of target indicators, including various financial indicators, company attribute indicators, and rule indicators, are obtained. By determining the indicator characteristics corresponding to the binning interval to which the indicator value of the target indicator belongs, the impact of abnormal indicator values of the target indicators on the generated risk information can be reduced. Furthermore, by using indicators of multiple dimensions, risk probability representation values can be generated more accurately based on the sum and characteristics of each target indicator. Based on the accurate risk probability representation value and the pre-constructed first and second parameters, a first risk value that can represent the probability of fraud in financial reports can be generated. Also, for each target indicator, according to the first... The second risk value of the target indicator is generated by two parameters, the weight of the target indicator, and the indicator characteristics. The second risk value can accurately represent the probability of fraud risk in financial reports under different target indicators. Based on such a second risk value, abnormal indicators in the target indicators can be identified. Then, the risk information of the financial report generated based on the first risk value and abnormal indicators includes the risk value of fraud risk in the financial report as a whole, as well as the abnormal indicators with existing fraud risk and the risk value of fraud risk in the financial report under abnormal indicators. Such risk information can more accurately represent the fraud risk of the target company's financial report and can push more accurate risk information to users.
[0170] This allows for the delivery of more accurate risk information to users, enabling them to make more accurate assessments of target companies and improving the user experience.
[0171] In addition, the solution provided in this embodiment of the invention uses a variety of target indicators, including various financial indicators, various company attribute indicators, and various rule indicators. Such target indicators are updated frequently, which can cope with the high-frequency changes in information of the target company and increase the frequency of generating effective risk information of the target company.
[0172] The following describes another way to generate risk information for financial reports.
[0173] In one embodiment of the present invention, see Figure 2 A flowchart of another risk information push method is provided. The method includes steps S201-S209. Step S106 can also be implemented through step S208.
[0174] Step S201: Based on the text data of the target company's financial report, obtain the indicator values of each financial indicator; based on the text data of the target company's related information, obtain the indicator values of each company attribute indicator and each rule indicator of the target company.
[0175] Step S202: Determine the binning interval to which the indicator value of each target indicator belongs in the binning interval of each target indicator, and obtain the indicator characteristics of the target indicator based on the determined binning interval.
[0176] Step S203: Generate risk probability representation values based on the weights and characteristics of each target indicator.
[0177] Step S204: Generate a first risk value for the financial report to have a risk of fraud based on the risk probability characterization value, the pre-constructed first parameter and second parameter, and generate a second risk value for the financial report to have a risk of fraud in each target indicator based on the second parameter, the weight of each target indicator and the indicator characteristics.
[0178] Step S205: Determine the abnormal indicators in the target indicators based on the obtained second risk value.
[0179] The steps S201-S205 above are the same as those S101-S105 above, and will not be described in detail here.
[0180] Step S206: Obtain a third-party audit opinion on the financial statements.
[0181] The information push platform can call preset data acquisition interfaces to obtain third-party audit opinions on the target company's financial reports, download third-party audit opinions on the target company's financial reports from preset information disclosure platforms, and use web information crawling tools to obtain third-party audit opinions on the target company's financial reports from the network.
[0182] Step S207: Determine the risk level described in the third-party audit opinion.
[0183] In one implementation, the information push platform can determine the risk level of a third-party audit opinion that is a standard unqualified opinion as: Level 1 risk. The information push platform can also determine the risk level of a third-party audit opinion that is a qualified opinion or a disclaimer of opinion as: Level 2 risk. The probability of financial reporting fraud represented by Level 2 risk is higher than the probability of financial reporting fraud represented by Level 1 risk.
[0184] Step S208: Generate risk information for the financial report based on the risk level, the first risk value, and the anomaly indicators.
[0185] In one scenario, if the risk level is Level 1, the information push platform will generate risk information for the financial report based on the Level 1 risk value and abnormal indicators.
[0186] In another scenario, if the risk level is Level 2, the information push platform generates risk information in the financial report, including the third-party audit opinion, the Level 2 risk level, and anomaly indicators. The information push platform can also adjust the Level 1 risk value based on the risk level, and generate risk information in the financial report based on the adjusted Level 1 risk value and anomaly indicators. For example, the Level 1 risk value can be adjusted based on a preset adjustment threshold.
[0187] Step S209: Push risk information to the target user's terminal.
[0188] Step S209 is the same as step S107 above, and will not be described in detail here.
[0189] As can be seen from the above, using the risk level described in the third-party audit opinion to generate risk information for financial reports can further refine the probability of fraud in the financial reports based on the first risk value and anomaly indicators, thereby improving the accuracy of the generated risk information. Furthermore, when both the third-party audit opinion and the first risk value indicate a high probability of fraud in the financial reports, or both indicate a low probability, the third-party audit opinion can serve as factual support for the generated risk information, enhancing its credibility.
[0190] In one embodiment of the present invention, the target indicator further includes: public opinion indicators. See also Figure 3 A flowchart illustrating a method for obtaining the value of a public opinion indicator is provided. The method includes the following steps S301-S304.
[0191] Step S301: Obtain public opinion data of the target company.
[0192] In one implementation, public opinion data of the target company can be obtained from a preset public opinion data source.
[0193] For example, using web scraping tools or data acquisition interfaces. Public opinion data sources can include: social media platforms, news platforms, and forum discussion platforms, etc.
[0194] Step S302: Determine the sentiment direction of the obtained public opinion data.
[0195] The emotional orientation is either positive or negative.
[0196] In one implementation, a pre-trained sentiment classification model can be used to classify the obtained public opinion data to obtain the sentiment direction of the public opinion data.
[0197] Specifically, the obtained public opinion data can be preprocessed to obtain the public opinion characteristics of the target company. These characteristics are then input into a pre-trained sentiment classification model to determine the sentiment direction of the public opinion data. Data preprocessing can include data cleaning and natural language preprocessing. For example, data cleaning can include data deduplication and irrelevant data removal, while natural language preprocessing can include text segmentation, stop word removal, and part-of-speech tagging. The sentiment classification model can be a sentiment analysis model, a topic model, etc. For instance, the sentiment classification model can be the BERT model (a pre-trained language representation model based on the Transformer architecture, a deep learning model architecture based on attention mechanisms). The sentiment classification model can include multiple stacked Transformer encoders, each containing a self-attention layer and a feedforward layer.
[0198] Step S303: Using time units as the unit, determine the first quantity of positive public opinion data and the second quantity of negative public opinion data in each time unit covered by the obtained public opinion data.
[0199] For example, the time unit can be one day. For each time unit, the number of positive public opinion data in the t-th time unit can be expressed as: Num 1,t The number of negative public opinion data in the t-th time unit can be represented as: Num 2,t .
[0200] Step S304: Generate the index value of the public opinion index based on the first quantity, the second quantity, and the time difference between each time unit and the current time.
[0201] Specifically, based on the time difference between each time unit and the current time, the time decay coefficient representing the importance of public opinion data in each time unit can be determined. Then, based on the difference between the first and second quantities of each time unit and the time decay coefficient, the index value of the public opinion indicator can be generated.
[0202] For example, the difference between the first and second quantities in the t-th time unit can be expressed as: S t =Num 1,t -Num 2,t The time decay coefficient of the t-th time unit can be expressed as: Where N is the total number of time units and T is the current time, the index value of the public opinion indicator can be expressed as:
[0203] In this way, by combining representative and diverse public opinion data, the risk of fraud in the target company's financial reports can be analyzed from different perspectives, generating more accurate and comprehensive risk information about the financial reports.
[0204] In one embodiment of the present invention, see Figure 4 A flowchart illustrating a parameter construction method is provided, comprising the following steps S401-S407. The first and second parameters can be constructed according to steps S401-S407.
[0205] Step S401: Obtain sample indicator values for each target indicator for the sample companies.
[0206] The method of obtaining the sample indicator value in step S401 is similar to the method of obtaining the indicator value in step S101. The difference is that the names of the sample company and the target company are different, and the names of the sample indicator value and the indicator value are different. These will not be described in detail here.
[0207] Step S402: Based on the sample index values, perform multinomial fitting with each target index as the independent variable to obtain the weight of each target index.
[0208] One implementation method is to use a multinomial fitting approach for regression equations, with the target indicators as independent variables to perform multinomial fitting and obtain the weights of each target indicator.
[0209] For example, if the number of target indicators is n, the resulting polynomial can be expressed as: Where, β i Let x be the weight corresponding to the i-th target indicator. i Let be the independent variable corresponding to the i-th target indicator, and β0 be the intercept term of the regression equation. The weights of each target indicator are obtained as β. i , i∈[1,n].
[0210] Step S403: Determine the binning interval to which the sample indicator value of each target indicator belongs within the binning interval of each target indicator.
[0211] The method for determining the binning interval in step S403 is similar to that in step S102, except that the sample index values and the names of the index values are different, which will not be described in detail here.
[0212] Step S404: For each target indicator, determine the sample distribution difference value for each bin interval under that target indicator.
[0213] Among them, the sample distribution difference value represents the difference between the number of positive samples and the number of negative samples in the sample index values belonging to the binning interval.
[0214] Specifically, for each target indicator, a first ratio of the number of positive samples in each binning interval to the number of positive samples in all binning intervals can be calculated, and a second ratio of the number of negative samples in each binning interval to the number of negative samples in all binning intervals can be calculated. Based on the first and second ratios, the sample distribution difference value of each binning interval under the target indicator can be determined.
[0215] For example, if the number of target indicators is n, and the i-th target indicator includes k i The first ratio in the j-th bin interval of the i-th target indicator is: in, Good represents the number of positive samples in the j-th bin interval of the i-th target indicator. sum_i Let be the number of positive samples in all bin intervals of the i-th target metric. The second ratio in the j-th bin interval of the i-th target metric is: in, Bad represents the number of negative samples in the j-th bin interval of the i-th target indicator. sum_i denoted as the number of negative samples in all bin intervals of the i-th target indicator.
[0216] Furthermore, the difference between the natural logarithm of the first ratio and the natural logarithm of the second ratio can be used as the sample distribution difference value, or the quotient between the natural logarithm of the first ratio and the natural logarithm of the second ratio can be used as the sample distribution difference value. For example, the sample distribution difference value of the j-th bin interval of the i-th target indicator can be: The sample distribution difference value of the j-th bin interval of the i-th target indicator can also be:
[0217] Step S405: Based on the sample distribution differences of each bin interval under each target indicator and the weight of each target indicator, determine the standard risk probability characterization value.
[0218] For example, if the resulting polynomial is expressed as: Therefore, the sum of the sample distribution differences of each bin interval of the i-th target indicator can be used as x. i Calculate the value of . The obtained standard risk probability characterization value is: Among them, E i For the i-th target indicator, k is included. i The sum of the sample distribution differences for each bin interval:
[0219] In addition, in step S103, based on the indicator characteristics of the target indicator and the sample distribution characteristics of the sample companies representing positive and negative samples in each target indicator, the risk probability benchmark value corresponding to the target indicator is determined. The sample distribution characteristics of the sample companies in each target indicator can be obtained based on the sample distribution difference value of the expression of the above standard risk probability characterization value. The sample distribution difference value is obtained based on the sample indicator value of each target indicator of the sample companies.
[0220] For example, combining the expression for the standard risk probability representation value: And the expression for the sum of the sample distribution differences: An expression for the standard risk probability representation value can be obtained.
[0221] For each target indicator, the expression for the standard risk probability representation value includes multiple sample distribution difference values corresponding to that target indicator. These multiple sample distribution difference values correspond one-to-one with each bin interval. For example, the i-th target indicator includes k values under the i-th target indicator. i The sample distribution difference values of each bin interval: (Assuming the number of bin intervals is greater than or equal to 3, that is, k) i ≥3), Corresponding to the first compartment, Corresponding to the second compartment, and so on, With the kth i Each binning interval corresponds to a specific target indicator. Based on the expression of the standard risk probability representation value, the correspondence between the binning interval and the sample distribution difference value under each target indicator can be determined. The correspondence between the binning interval and the sample distribution difference value under each target indicator is used as the sample distribution characteristics of the sample company under each target indicator.
[0222] Therefore, when determining the risk probability benchmark value corresponding to the target indicator, the information push platform can determine the sample distribution difference value corresponding to the binning interval of the target indicator's indicator characteristics from the sample distribution characteristics, and use this as the risk probability benchmark value corresponding to the target indicator's indicator characteristics. For example, under the i-th target indicator, the obtained indicator characteristics are set in the second binning interval, and the information push platform can determine the risk probability benchmark value corresponding to the k binning intervals included under the i-th target indicator based on the indicator characteristics of the i-th target indicator and the sample distribution characteristics. i The sample distribution difference values of each bin interval: Determine the sample distribution difference value corresponding to the second bin interval. The risk probability benchmark value corresponding to the indicator characteristics of the i-th target indicator.
[0223] Step S406: Construct the second parameter based on the preset risk increase value and the standard risk probability characterization value.
[0224] Specifically, the quotient of the natural logarithm of the preset risk increase value and the standard risk probability representation value under the preset multiple can be calculated as the second parameter.
[0225] Step S407: Construct the first parameter based on the preset standard risk value, standard risk probability representation value, and second parameter.
[0226] Specifically, the product of the natural logarithm of the standard risk probability representation value and the second parameter can be calculated, and the sum of the above product and the preset standard risk value can be used as the first parameter.
[0227] This allows for more accurate construction of the first and second parameters, improving the accuracy of generating the first risk value for financial reporting fraud risks using the first and second parameters, and improving the accuracy of generating the second risk value for financial reporting fraud risks in each target indicator using the second parameter, the weights of each target indicator, and indicator characteristics, thereby improving the accuracy of the generated risk information.
[0228] Furthermore, in the process of constructing the first and second parameters, based on the sample index values of each target indicator obtained for the sample companies, a multinomial fitting is performed with each target indicator as the independent variable to obtain the weight of each target indicator, and the sample distribution difference value of each bin interval under the target indicator is obtained. Further, the standard risk probability characterization value is determined by combining the weight of the target indicators and the sample distribution difference value of each bin interval. The sample distribution difference value in the expression of such a standard risk probability characterization value can characterize the correlation between the target indicator and the risk of financial report fraud under different bin intervals for the indicator values of different target indicators. Thus, the sample distribution difference value corresponding to each bin interval under each target indicator can characterize the sample distribution characteristics of the sample index values of each target indicator of the negative sample companies and the sample distribution characteristics of the sample index values of each target indicator of the positive sample companies. In this way, by taking the correspondence between the binning intervals under each target indicator and the sample distribution difference value as the sample distribution characteristics of the sample companies under each target indicator, we can obtain the accurate sample distribution characteristics of the sample companies under each target indicator. Then, based on the above sample distribution characteristics, we can generate risk probability representation values, which can more accurately generate risk information representing the risk of financial report fraud.
[0229] Furthermore, the standard risk probability representation value calculated based on the sample distribution difference value that can accurately characterize the sample distribution characteristics can be used to construct a more accurate second parameter and first parameter. Then, based on the obtained standard risk probability representation value, second parameter, first parameter and the weight of each target indicator, the first risk value and second risk value can be accurately generated. Such first risk value and second risk value are essentially a combination of the sample distribution characteristics of positive samples and the distribution characteristics of negative samples, which can more accurately characterize the risk of fraud in the financial reports of the target company.
[0230] In one embodiment of the present invention, the sample indicator values are labeled. The label can identify an event involving the company to which the sample indicator value belongs. Based on the label of the sample indicator value, sample indicator values that can be used to identify negative samples are determined.
[0231] A sample is identified as having a negative metric value based on any of the following events identified by the label:
[0232] I. Financial restatement events of sample companies.
[0233] Specifically, based on at least one of the following information: the restatement time of the financial restatement event of the sample company, the financial report disclosure date of the sample company, the corrected accounting year, the type of correction, and the reason for the restatement, the sample indicator value of the target indicator as the negative sample can be determined from the sample indicator values of the financial restatement events of the sample companies that have financial restatement events before the financial report restatement.
[0234] II. Financial Inquiry Events of Sample Companies
[0235] Specifically, based on at least one of the following information—the disclosure date of the inquiry letter regarding the financial inquiry event of the sample company, and the relevant financial indicators disclosed in the inquiry letter—the sample indicator value that serves as the target indicator for the negative sample can be determined from the sample indicator values of the financial inquiry events that identify the sample company.
[0236] III. Tax violations by the sample companies.
[0237] Specifically, based on at least one of the following information regarding the tax violation of the sample company: the date of disclosure of the tax violation, the accounting year of the non-compliant financial report, the issuing agency of the announcement, the type of violation, and the method of punishment, the sample indicator value of the target indicator for the negative sample can be determined from the sample indicator values of the tax violation of the sample company.
[0238] IV. Financial fraud cases of sample companies.
[0239] Specifically, based on at least one of the following information regarding the financial fraud incident of the sample company, namely the date of disclosure of financial fraud, the accounting year of the fraudulent financial statement, the type of financial fraud, and the method of punishment, the sample indicator value of the target indicator as the negative sample can be determined from the sample indicator values of the label identifying the financial fraud incident of the sample company.
[0240] Alternatively, sample index values other than the negative samples mentioned above can be used as positive samples.
[0241] In this way, using multiple different events to determine the sample index values belonging to negative samples can balance the number of positive and negative samples, reduce the impact of sample imbalance on the accuracy of determining the first and second parameters, and improve the accuracy of the generated risk information.
[0242] In one embodiment of the present invention, the target indicator for risk assessment can be determined from among the candidate indicators according to the following steps:
[0243] Step A: For each candidate indicator, obtain the sample indicator values belonging to positive samples and the sample indicator values belonging to negative samples for the sample companies.
[0244] The candidate indicators include: candidate financial indicators, candidate company attribute indicators, and candidate rule indicators.
[0245] The method for determining positive and negative samples in step A can be found in the above embodiment, and will not be described in detail here.
[0246] The method of obtaining sample indicator values in step A is similar to that of obtaining indicator values in step S101. The difference lies in the different names of candidate indicators and target indicators, and the different names of sample indicator values and indicator values. These details will not be elaborated here.
[0247] Step B: Determine the impact value of each candidate indicator on the risk probability representation value based on the sample indicator values corresponding to each candidate indicator.
[0248] In one implementation, when the number of candidate indicators is m, the influence of the i-th candidate indicator on the risk probability representation value can be determined according to the following expression:
[0249]
[0250] Among them, Good pct_i Let be the ratio of the number of positive samples in the i-th candidate indicator to the number of positive samples in all candidate indicators. Let e′ be the ratio of the number of negative samples in the i-th candidate indicator to the number of negative samples in all candidate indicators. i Let be the sample distribution difference value of the i-th target indicator. Good num_iLet Good be the number of positive samples in the i-th candidate indicator. sum Bad represents the number of positive samples among all candidate metrics. num_i Let Bad be the number of negative samples in the i-th candidate metric. sum This represents the number of negative samples among all candidate metrics.
[0251] In another implementation, when the number of candidate indicators is m, the influence of the i-th candidate indicator on the risk probability representation value can be determined according to one of the following expressions:
[0252]
[0253] Among them, S i Let S' be the influence value of the i-th candidate indicator on the risk probability representation value, and S0 be the standard value of the degree of influence of each candidate indicator on the risk probability representation value. i Let f(x′) represent the i-th candidate index. i f(x′) represents the numerical value that characterizes the degree of influence of the i-th candidate indicator on the risk probability representation value. i Determined according to the following expression:
[0254]
[0255] Where Q is the set of indicators for each candidate indicator, and R∈{Q / x′} i Let} be the set of indicators after removing the i-th candidate indicator from the set of candidate indicators, and q be the set of indicators selected from the set of candidate indicators. v(x′) represents the probability of financial reports being fraudulent, predicted using different combinations of candidate indicators. R∪{i} )-v(x′ R Let x′ be the value of predicting the risk probability representation using an indicator set that includes the i-th candidate indicator, relative to predicting the risk probability representation using an indicator set that does not include the i-th candidate indicator. R∪{i} Let x′ be the set of indicators that includes the i-th candidate indicator. R Let v be the set of indicators excluding the i-th candidate indicator, and v be the value function, which is a function that determines the sample indicator values belonging to positive samples and the sample indicator values belonging to negative samples under each candidate indicator.
[0256] Step C: Based on the determined impact characterization values of each candidate indicator, select the target indicator from the candidate indicators.
[0257] Specifically, candidate indicators whose impact values are greater than a preset impact threshold can be used as target indicators.
[0258] In this way, candidate indicators that have a high impact on the probability of financial reporting fraud can be used as target indicators, which can further improve the accuracy of risk information generation by using the indicator values of target features.
[0259] In one embodiment of the present invention, the information push platform can update the constructed first and second parameters in the following manner: obtain sample indicator values under each candidate indicator; if the preset model update conditions are met, determine the target indicator for risk assessment among the candidate indicators according to the above steps AC. Then, construct the first and second parameters according to steps S401-S407; based on the newly constructed first and second parameters, use the sample indicator values of the target indicator in the test set again to generate risk information according to steps S101-S106; and obtain the accuracy, recall, etc. of the generated risk information based on the positive and negative samples in the test set. If the accuracy and recall of the risk information generated based on the newly constructed first and second parameters are both higher than the first and second parameters before construction, then use the newly constructed first and second parameters as the pre-constructed first and second parameters to generate risk information.
[0260] In the above process, the first and second parameters updated each time can be stored for parameter rollback. This allows for continuous adjustment and optimization of the risk information generation process using new data, improving the stability and real-time performance of risk information generation.
[0261] In one embodiment of the present invention, a user can send the identifier of a target company to an information push platform using a user terminal. Based on the identifier, the information push platform can obtain the indicator values of various target indicators of the target company, then generate risk information and feed it back to the user terminal. Additionally, the information push platform can also generate a visual chart of historical anomalies of the target company in chronological order, in a timeline format. For example, historical anomalies may include financial restatement events, financial inquiry events, tax violations, and financial fraud events, etc.
[0262] In the technical solutions provided by the embodiments of the present invention, the operations of obtaining, storing, using, processing, transmitting, providing and disclosing user personal information and company information are all carried out with authorization.
[0263] Corresponding to the above-mentioned risk information push method, this embodiment of the invention also provides a risk information push device.
[0264] In one embodiment of the present invention, see Figure 5 A schematic diagram of a risk information push device is provided, the device comprising:
[0265] The indicator value acquisition module 501 is used to obtain the indicator values of each financial indicator based on the text data of the target company's financial report; and to obtain the indicator values of each company attribute indicator and each rule indicator of the target company based on the text data of the target company's related information.
[0266] The indicator feature acquisition module 502 is used to determine the binning interval to which the indicator value of each target indicator belongs in the binning interval of each target indicator, and to obtain the indicator features of the target indicator based on the determined binning interval. The target indicators include: each financial indicator, each company attribute indicator and each rule indicator.
[0267] The representation value generation module 503 is used to generate risk probability representation values based on the weight and characteristics of each target indicator. The risk probability representation value represents the ratio between the probability of financial reporting fraud and the probability of no fraud.
[0268] The risk value generation module 504 is used to generate a first risk value for the financial report to have a risk of fraud based on the risk probability characterization value, the pre-built first parameter and the second parameter, and to generate a second risk value for the financial report to have a risk of fraud in each target indicator based on the second parameter, the weight of each target indicator and the indicator characteristics.
[0269] The abnormal indicator determination module 505 is used to determine the abnormal indicators in the target indicators based on the obtained second risk value.
[0270] The risk information generation module 506 is used to generate risk information that characterizes the financial report as having a risk of being falsified, based on a first risk value and anomaly indicators.
[0271] The risk information push module 507 is used to push the risk information to the target user terminal of the user.
[0272] The risk information push device provided in this invention obtains the indicator values of target indicators, including various financial indicators, company attribute indicators, and rule indicators. By determining the indicator characteristics corresponding to the binning interval to which the indicator value of the target indicator belongs, the impact of abnormal indicator values of the target indicators on the generated risk information can be reduced. Furthermore, by using indicators of multiple dimensions, risk probability representation values can be generated more accurately based on the sum and characteristics of each target indicator. Based on the accurate risk probability representation value and the pre-constructed first and second parameters, a first risk value that can represent the probability of fraud in financial reports can be generated. For each target indicator, a second risk value can also be generated based on the second parameter, the weight of the target indicator, and the indicator characteristics. The second risk value can accurately represent the probability of fraud in financial reports under different target indicators. Based on such a second risk value, abnormal indicators among the target indicators can be identified. Thus, the risk information of financial reports generated based on the first risk value and abnormal indicators includes the risk value that the financial report as a whole has a risk of fraud, as well as the abnormal indicators that have a risk of fraud and the risk value that the financial report has a risk of fraud in abnormal indicators. Such risk information can more accurately represent the risk of fraud in the target company's financial reports.
[0273] In one embodiment of the present invention, the characterization value generation module 503 is specifically used to determine the risk probability benchmark value corresponding to the indicator characteristics of each target indicator based on the indicator characteristics of the target indicator, calculate the product of the determined risk probability benchmark value and the weight of the target indicator as the probability score corresponding to the target indicator, and use the sum of the probability scores of each target indicator as the risk probability characterization value.
[0274] In this way, the risk probability benchmark value corresponding to the indicator characteristics of the target indicator can be adjusted by using the weights corresponding to different target indicators. This allows for a more accurate determination of the probability scores corresponding to different target indicators, resulting in a more accurate calculated risk probability representation value.
[0275] In one embodiment of the present invention, the risk value generation module 504 is specifically used to multiply the pre-constructed second parameter by the risk probability representation value as the risk increase value of the financial report, use the pre-constructed first parameter as the benchmark risk value, and determine the first risk value of the financial report as having a risk of fraud based on the difference between the benchmark risk value and the risk increase value; in this way, the risk increase value of the financial report can be accurately calculated, and the first risk value of the financial report as having a risk of fraud can be generated using the difference between the benchmark risk value and the risk increase value.
[0276] In one embodiment of the present invention, the risk value generation module 504 is specifically used to determine the risk probability benchmark value corresponding to each target indicator based on the indicator characteristics of the target indicator, and to use the product of the determined risk probability benchmark value, the weight of the target indicator and the second parameter as the second risk value characterizing the risk of fraud in the financial report for each target indicator.
[0277] In this way, a corresponding second risk value can be generated for different target indicators. When the first risk value indicates a high probability of financial reporting fraud, attribution analysis can be performed based on the value of the second risk value of different target indicators. This analysis can identify the indicators that indicate a high probability of financial reporting fraud and identify abnormal indicators with high risk.
[0278] In one embodiment of the present invention, the apparatus further includes:
[0279] Audit Opinion Obtaining Module. Used to obtain third-party audit opinions on financial reports;
[0280] The risk level determination module is used to determine the risk level described in the third-party audit opinion;
[0281] The risk information generation module is specifically used to generate risk information for financial reports based on risk level, first risk value, and abnormal indicators.
[0282] As can be seen from the above, using the risk level described in the third-party audit opinion to generate risk information for financial reports can further correct the probability of fraud in financial reports based on the first risk value and anomaly indicators, thereby improving the accuracy of the generated risk information.
[0283] In one embodiment of the present invention, the target indicator further includes: public opinion indicators;
[0284] The index values of public opinion indicators are obtained in the following ways:
[0285] Obtain public opinion data about the target company;
[0286] Determine the sentiment direction of the obtained public opinion data, where the sentiment direction is positive or negative;
[0287] Using time units as the unit, determine the first quantity of positive public opinion data and the second quantity of negative public opinion data in each time unit covered by the obtained public opinion data;
[0288] The index values of public opinion indicators are generated based on the first quantity, the second quantity, and the time difference between each time unit and the current time.
[0289] In this way, by combining representative and diverse public opinion data, the risk of fraud in the target company's financial reports can be analyzed from different perspectives, generating more accurate and comprehensive risk information about the financial reports.
[0290] In one embodiment of the present invention, the first parameter and the second parameter are constructed in the following manner:
[0291] Sample indicator values for each target indicator were obtained for the sample companies;
[0292] Based on the sample index values, a multinomial fitting is performed with each target index as the independent variable to obtain the weight of each target index.
[0293] Within the binning intervals of each target indicator, determine the binning interval to which the sample indicator value of each target indicator belongs;
[0294] For each target indicator, determine the sample distribution difference value for each bin interval under that target indicator. The sample distribution difference value represents the difference between the number of positive samples and the number of negative samples in the sample indicator value belonging to the bin interval.
[0295] Based on the sample distribution differences of each bin interval under each target indicator and the weight of each target indicator, the standard risk probability characterization value is determined.
[0296] A second parameter is constructed based on a preset risk increase value and a standard risk probability representation value.
[0297] The first parameter is constructed based on the preset standard risk value, standard risk probability representation value, and second parameter.
[0298] This allows for more accurate construction of the first and second parameters, improving the accuracy of generating the first risk value for financial reporting fraud risks using the first and second parameters, and improving the accuracy of generating the second risk value for financial reporting fraud risks in each target indicator using the second parameter, the weights of each target indicator, and indicator characteristics, thereby improving the accuracy of the generated risk information.
[0299] In one embodiment of the present invention, the sample index values are labeled;
[0300] A sample is identified as having a negative metric value based on any of the following events identified by the label:
[0301] Financial restatement events of the sample companies;
[0302] Financial inquiry events at the sample companies;
[0303] Tax violations by the sample companies;
[0304] Financial fraud cases involving sample companies.
[0305] In this way, using multiple different events to determine the sample index values belonging to negative samples can balance the number of positive and negative samples, reduce the impact of sample imbalance on the accuracy of determining the first and second parameters, and improve the accuracy of the generated risk information.
[0306] In one embodiment of the present invention, the characteristic indicators used for risk assessment are determined from among the candidate indicators in the following manner:
[0307] For each candidate indicator, the sample indicator values belonging to positive samples and the sample indicator values belonging to negative samples are obtained for the sample companies. The candidate indicators include: candidate financial indicators, candidate company attribute indicators and candidate rule indicators.
[0308] Based on the sample index values corresponding to each candidate index, determine the impact value of each candidate index on the risk probability characterization value;
[0309] Based on the determined impact values of each candidate indicator, the target indicator is selected from the candidate indicators.
[0310] In this way, candidate indicators that have a high impact on the probability of financial reporting fraud can be used as target indicators, which can further improve the accuracy of risk information generation by using the indicator values of target features.
[0311] This invention also provides an electronic device, such as... Figure 6 As shown, it includes a processor 601, a communication interface 602, a memory 603, and a communication bus 604, wherein the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604.
[0312] Memory 603 is used to store computer programs;
[0313] The processor 601, when executing the program stored in the memory 603, implements any of the above-mentioned risk information push methods.
[0314] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not indicate that there is only one bus or one type of bus.
[0315] The communication interface is used for communication between the aforementioned electronic devices and other devices.
[0316] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0317] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0318] In another embodiment of the present invention, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of any of the above-described risk information push methods.
[0319] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the risk information push methods described in the above embodiments.
[0320] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0321] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0322] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, electronic devices, storage media, and computer program products are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0323] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A method for pushing risk information, characterized in that, The method includes: Based on the textual data of the target company's financial reports, the values of each financial indicator are obtained. Based on the text data of the target company's associated information, the index values of each company attribute indicator and each rule indicator of the target company are obtained. Within the binning intervals of each target indicator, the binning interval to which the indicator value of each target indicator belongs is determined. Based on the determined binning intervals, the indicator characteristics of the target indicators are obtained. The target indicators include: each financial indicator, each company attribute indicator, and each rule indicator. Based on the weight and characteristics of each target indicator, a risk probability representation value is generated. The risk probability representation value represents the ratio between the probability that the financial report has a risk of fraud and the probability that it does not have a risk of fraud. A first risk value for the financial report to be fraudulent is generated based on the risk probability characterization value, the pre-constructed first parameter and the second parameter; and a second risk value for the financial report to be fraudulent in each target indicator is generated based on the second parameter, the weight of each target indicator and the indicator characteristics. Based on the obtained second risk value, identify the abnormal indicators in the target indicators; Based on the first risk value and the abnormal indicators, risk information is generated that indicates the risk of fraud in the financial report. The risk information is pushed to the target user's device.
2. The method according to claim 1, characterized in that, The step of generating a risk probability representation value based on the weight and characteristics of each target indicator includes: for each target indicator, determining the risk probability benchmark value corresponding to the indicator characteristics of the target indicator based on the indicator characteristics of the target indicator, calculating the product of the determined risk probability benchmark value and the weight of the target indicator as the probability score corresponding to the target indicator, and using the sum of the probability scores of each target indicator as the risk probability representation value. and / or The step of generating a first risk value for the financial report to have a risk of fraud based on the risk probability characterization value, a pre-constructed first parameter, and a second parameter includes: using the product of the pre-constructed second parameter and the risk probability characterization value as the risk increase value of the financial report, using the pre-constructed first parameter as a benchmark risk value, and determining the first risk value for the financial report to have a risk of fraud based on the difference between the benchmark risk value and the risk increase value. and / or The step of generating a second risk value for the financial report in each target indicator based on the second parameter, the weight of each target indicator, and the indicator characteristics includes: for each target indicator, determining the risk probability benchmark value corresponding to the target indicator based on the indicator characteristics of the target indicator, and using the product of the determined risk probability benchmark value, the weight of the target indicator, and the second parameter as the second risk value characterizing the financial report in each target indicator.
3. The method according to claim 1, characterized in that, The method further includes: Obtain a third-party audit opinion on the aforementioned financial reports; Determine the risk level described in the third-party audit opinion; The risk information generated from the first risk value and anomaly indicators to produce the financial report includes: The risk information in the financial report is generated based on the risk level, the first risk value, and the anomaly indicator.
4. The method according to claim 1, characterized in that, The target indicators also include: public opinion indicators; The index values of the aforementioned public opinion indicators are obtained in the following manner: Obtain public opinion data about the target company; Determine the sentiment direction of the obtained public opinion data, wherein the sentiment direction is positive or negative; Using time units as the unit, determine the first quantity of positive public opinion data and the second quantity of negative public opinion data in each time unit covered by the obtained public opinion data; The index values of public opinion indicators are generated based on the first quantity, the second quantity, and the time difference between each time unit and the current time.
5. The method according to any one of claims 1-4, characterized in that, The first and second parameters are constructed as follows: Sample indicator values for each target indicator were obtained for the sample companies; Based on the sample index values, a multinomial fitting is performed with each target index as the independent variable to obtain the weight of each target index. Within the binning intervals of each target indicator, determine the binning interval to which the sample indicator value of each target indicator belongs; For each target indicator, the sample distribution difference value for each bin interval under that target indicator is determined, wherein the sample distribution difference value represents the difference between the number of positive samples and the number of negative samples in the sample indicator value belonging to the bin interval. Based on the sample distribution differences of each bin interval under each target indicator and the weight of each target indicator, the standard risk probability characterization value is determined. The second parameter is constructed based on the preset risk increase value and the standard risk probability characterization value; The first parameter is constructed based on the preset standard risk value, the standard risk probability representation value, and the second parameter.
6. The method according to claim 5, characterized in that, The sample indicator values are labeled; The sample index value is determined to be a negative sample based on any of the following events identified by the label: Financial restatement events of the sample companies; Financial inquiry events at the sample companies; Tax violations by the sample companies; Financial fraud cases involving sample companies.
7. The method according to any one of claims 1-4, characterized in that, The characteristic indicators used for risk assessment are determined from the candidate indicators as follows: For each candidate indicator, the sample indicator values belonging to positive samples and the sample indicator values belonging to negative samples are obtained for the sample companies. The candidate indicators include: candidate financial indicators, candidate company attribute indicators and candidate rule indicators. Based on the sample index values corresponding to each candidate index, determine the impact value of each candidate index on the risk probability characterization value; Based on the determined impact values of each candidate indicator, the target indicator is selected from the candidate indicators.
8. A risk information push device, characterized in that, The device includes: The indicator value acquisition module is used to obtain the indicator values of each financial indicator based on the text data of the target company's financial report; and to obtain the indicator values of each company attribute indicator and each rule indicator of the target company based on the text data of the target company's related information. The indicator feature acquisition module is used to determine the binning interval to which the indicator value of each target indicator belongs in the binning interval of each target indicator, and to obtain the indicator features of the target indicator based on the determined binning interval. The target indicators include: each financial indicator, each company attribute indicator and each rule indicator. The representation value generation module is used to generate a risk probability representation value based on the weight and characteristics of each target indicator. The risk probability representation value represents the ratio between the probability that the financial report has a risk of fraud and the probability that it does not have a risk of fraud. The risk value generation module is used to generate a first risk value for the financial report to have a risk of fraud based on the risk probability characterization value, a pre-constructed first parameter and a second parameter, and to generate a second risk value for the financial report to have a risk of fraud in each target indicator based on the second parameter, the weight of each target indicator and the indicator characteristics. The abnormal indicator determination module is used to determine the abnormal indicators in the target indicators based on the obtained second risk value. The risk information generation module is used to generate risk information that characterizes the financial report as having a risk of being falsified, based on the first risk value and the abnormal indicators. The risk information push module is used to push the risk information to the target user terminal.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the method described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method described in any one of claims 1-7.