Accounting anomaly detection device, accounting anomaly detection method, and accounting anomaly detection program

The accounting anomaly detection device and method address the limitation of existing technologies by using a model learned from teacher data to detect anomalies in datasets that do not follow Benford's law, enhancing the detection of fraudulent transactions.

JP2025093170APending Publication Date: 2025-06-23OBIC CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023208746
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-11
Publication Date
2025-06-23

AI Technical Summary

Technical Problem

Existing accounting anomaly detection methods struggle with datasets that do not follow Benford's law, limiting their applicability and accuracy in detecting fraudulent transactions.

Method used

An accounting anomaly detection device and method that uses a model learned from teacher data to perform anomaly detection, even when the dataset does not follow Benford's law. This involves acquiring evaluation data, calculating expected frequencies, and performing a statistical hypothesis test to determine anomalies.

Benefits of technology

The solution enables effective anomaly detection in datasets that deviate from Benford's law, providing a more comprehensive approach to identifying fraudulent activities by using Bayesian statistics and teacher data to create a suitable prediction model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025093170000001_ABST
    Figure 2025093170000001_ABST
Patent Text Reader

Abstract

To provide an accounting anomaly detection device, an accounting anomaly detection method, and an accounting anomaly detection program that can perform accounting anomaly detection on accounting datasets that do not follow the Benford's law using a model learned from teacher data.SOLUTION: A determination result for evaluation data is obtained by: obtaining the evaluation data in which an expense amount to be evaluated and an observed frequency of each leading digit of the expense amount to be evaluated are set; obtaining an anomaly determination target dataset in which an expected frequency of each leading digit of the expense amount to be evaluated and a deviation degree between the observed frequency of each leading digit of the expense amount to be evaluated and the expected frequency are linked and set based on a teacher dataset and the evaluation data; and obtaining a test statistical amount for each leading digit of the expense amount to be evaluated and a p-value by statistical hypothesis testing of the anomaly determination target dataset.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an accounting anomaly detection device, an accounting anomaly detection method, and an accounting anomaly detection program.

Background Art

[0002] Patent Document 1 discloses a configuration for confirming fraudulent journal entries by comparing the leading digit of the transaction amount of a journal entry or the actual values for every two leading digits with Benford's law.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the invention described in Patent Document 1 above, since not all numbers existing in nature follow the theoretical distribution of Benford's law, there is a problem that analysis cannot be performed on a dataset for which application is difficult.

[0005] The present invention has been made in view of the above problems, and an object thereof is to provide an accounting anomaly detection device, an accounting anomaly detection method, and an accounting anomaly detection program capable of performing accounting anomaly detection using a model learned from teacher data on an accounting dataset that does not follow Benford's law.

Means for Solving the Problems

[0006] In order to solve the above-described problems and achieve the object, an accounting anomaly detection device according to the present invention is an accounting anomaly detection device including a storage unit and a control unit, wherein the storage unit includes an expense storage unit that stores teacher data in which a teacher's expense amount is set, and teacher data sets in which expected probabilities of the leading digit numbers of the teacher's expense amounts are set. The control unit includes an evaluation acquisition unit that acquires evaluation target expense amounts and evaluation data in which observed frequencies of the leading digit numbers of the evaluation target expense amounts are set, and based on the teacher data sets and the evaluation data, an expected frequency of each leading digit number of the evaluation target expense amount and an anomaly determination target data set in which a degree of deviation between the observed frequency of each leading digit number of the evaluation target expense amount and the expected frequency is associated and set are acquired. A determination result acquisition unit that acquires a determination result for the evaluation data by acquiring a test statistic and / or a p-value of each leading digit number of the evaluation target expense amount by performing a statistical hypothesis test on the anomaly determination target data set is provided.

[0007] Further, in the accounting anomaly detection device according to the present invention, the control unit further includes a teacher acquisition unit that acquires observed frequencies of the leading digit numbers of the teacher's expense amounts based on the teacher data and acquires the teacher data sets in which the expected frequencies of the leading digit numbers of the teacher's expense amounts are set.

[0008] Further, in the accounting anomaly detection device according to the present invention, the teacher acquisition unit acquires observed frequencies of the leading digit numbers of the teacher's expense amounts based on the teacher data, and when there is an observed frequency of 0 and / or when the total of the observed frequencies is smaller than a predetermined size, acquires the teacher data sets in which the expected frequencies of the leading digit numbers of the teacher's expense amounts are set based on Bayes' theorem.

[0009] Further, in the accounting anomaly detection device according to the present invention, in the teacher data set, the expected probability of each leading digit number of the teacher's expense amount for each predetermined group is further set, and in the evaluation data, the observed frequency of each leading digit number of the expense amount to be evaluated for each predetermined group is further set. The anomaly determination target data set is further characterized in that the expected frequency of each leading digit number of the expense amount to be evaluated for each predetermined group and the degree of deviation between the observed frequency of each leading digit number of the expense amount to be evaluated and the expected frequency are associated and set.

[0010] Further, in the accounting anomaly detection device according to the present invention, the storage unit further includes a setting master in which a determination item indicating the type of the expense amount to be evaluated, the unit of the predetermined group, a Benford non-use flag indicating that the teacher data set or Benford's law is used for anomaly determination, a significance level, and a prediction model type that is the unit of the anomaly determination target are associated and set. The evaluation data acquisition means acquires the evaluation data based on the setting master, and the determination target acquisition means acquires the anomaly determination target data set based on the setting master, the teacher data set, and the evaluation data when the Benford non-use flag indicates that the teacher data set is used for the anomaly determination. The determination result acquisition means acquires the determination result based on the setting master.

[0011] Further, in the accounting anomaly detection device according to the present invention, the determination target acquisition means further acquires the anomaly determination target data set in which the Benford expected frequency of each leading digit number of the expense amount to be evaluated and the degree of deviation between the observed frequency of each leading digit number of the expense amount to be evaluated and the Benford expected frequency are associated and set based on the setting master and the evaluation data when the Benford non-use flag indicates that Benford's law is used for the anomaly determination.

[0012] Also, in the accounting anomaly detection device according to the present invention, the expense storage means further stores pre - processed data in which the expense amount of the evaluation target is set, and the evaluation acquisition means totals the appearance frequency of the leading - digit number of the expense amount of the evaluation target based on the pre - processed data, and acquires the observed frequency of each leading - digit number of the expense amount of the evaluation target, thereby acquiring the evaluation data in which the expense amount of the evaluation target and the observed frequency of each leading - digit number of the expense amount of the evaluation target are set.

[0013] Also, in the accounting anomaly detection device according to the present invention, the teacher data is expense data in which the expense amount of the previous year of the evaluation target is set.

[0014] Also, the accounting anomaly detection method according to the present invention is an accounting anomaly detection method for causing an accounting anomaly detection device including a storage unit and a control unit to execute. The storage unit includes expense storage means for storing teacher data in which a teacher - used expense amount is set, and a teacher - used data set in which the expected probability of each leading - digit number of the teacher - used expense amount is set. The method includes an evaluation acquisition step of acquiring evaluation data in which the expense amount of the evaluation target and the observed frequency of each leading - digit number of the expense amount of the evaluation target are set, which is executed in the control unit; a determination target acquisition step of acquiring an anomaly determination target data set in which the expected frequency of each leading - digit number of the expense amount of the evaluation target and the degree of deviation between the observed frequency and the expected frequency of each leading - digit number of the expense amount of the evaluation target are associated and set based on the teacher - used data set and the evaluation data; and a determination result acquisition step of acquiring a determination result for the evaluation data by obtaining a test statistic and / or a p - value of each leading - digit number of the expense amount of the evaluation target by performing a statistical hypothesis test on the anomaly determination target data set.

[0015] The accounting anomaly detection program according to the present invention is an accounting anomaly detection program for causing an accounting anomaly detection device including a storage unit and a control unit to execute. The storage unit includes expense storage means for storing teacher data in which a teacher's expense amount is set, and teacher data sets in which expected probabilities of leading digit numbers of the teacher's expense amount are set. In the control unit, an evaluation acquisition step of acquiring evaluation data in which an expense amount to be evaluated and observed frequencies of the leading digit numbers of the expense amount to be evaluated are set, and based on the teacher data set and the evaluation data, an expected frequency of each leading digit number of the expense amount to be evaluated, and an anomaly determination target data set in which a degree of deviation between the observed frequency and the expected frequency of each leading digit number of the expense amount to be evaluated is associated and set are obtained. A determination result acquisition step of obtaining a determination result for the evaluation data by obtaining a test statistic and / or a p-value of each leading digit number of the expense amount to be evaluated by performing a statistical hypothesis test on the anomaly determination target data set is executed.

Effect of the Invention

[0016] According to the present invention, the observation probability of teacher data, which is reliable past data, is treated as the expected probability, and a chi-square test is performed using the distribution of the leading digit numbers of the teacher data, achieving the effect that abnormal detection can be performed. Further, according to the present invention, a model can be created from data and prior information (Benford's law) using Bayesian statistics, and the possibility of fraud in a numerical data set to be evaluated can be detected. Further, according to the present invention, there is an effect that the leading digit number causing the abnormality can be grasped. Further, according to the present invention, even a person in charge without knowledge of statistics can detect data suspected of accounting fraud from numerical business data, which has an effect. Further, according to the present invention, it is applicable to data assumed not to follow Benford's law, expanding the range of analysis target data, and even for data not following Benford's law, by providing teacher data reflecting the characteristics of the enterprise, a model that is considered to be more suitable for that enterprise can be created, which has an effect. Conventionally, it was only possible to analyze whether there is a possibility of fraud in a data set, and it was not possible to draw a lead for performing a detailed analysis of which data is suspicious. However, according to the present invention, by drawing a lead for analysis based on the numbers considered to be the cause of the abnormality, it is possible to reach the details suspected of fraud, which has an effect. Further, according to the present invention, by expanding Benford's law used for the purpose of detecting expense fraud, data with a possibility of expense fraud can be widely detected and utilized for strengthening internal control, which has an effect. Further, according to the present invention, detection is performed using the prior distribution and the posterior distribution calculated from the business data of the enterprise, and details can be confirmed, which has an effect. Further, according to the present invention, even when there is little past data available for the abnormal determination target group (the size of the teacher data is insufficient) and a model capturing the characteristics of the population cannot be created, or in the case of a division-by-zero problem, abnormal detection can be performed, which has an effect.Also, if there is no data in the teacher data that contains even one of the first-digit numbers 1 to 9, the observation probability of that number is 0. Therefore, if the observation probability is treated as the expected probability, a division by zero occurs in the calculation of the divergence (since "divergence of number i = data size × (observation probability of number i - expected probability of number i)^2 / expected probability of number i", a division by zero occurs). However, according to the present invention, by using Bayesian statistics to provide the information of Benford's law in advance, since all the expected probabilities of Benford's law are 0 or more for 1 to 9, the effect of being able to avoid division by zero can be achieved. Further, according to the present invention, by using Bayesian statistics, in addition to the prior information of Benford's law (information that seems to be somewhat reliable), business data can be utilized, and it can play a role as a complement to model creation. Also, when the size of the teacher data set is small, the characteristics of the distribution of the population cannot be accurately captured. However, according to the present invention, in order to avoid this, by providing the information of Benford's law in advance, when the data size is small, a prediction model can be created while using the information of Benford's law. Further, according to the present invention, it is possible to obtain a tuned distribution of expected probabilities that is considered to be suitable for the company from the information of Benford's law and the teacher data (reliable data in the past). Also, according to the present invention, when the size of the teacher data is small (there is little reliable information), by using prior information (Benford), conversely, when there is a large amount of teacher data, it is possible to create a prediction model with little use of prior information.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

DETAILED DESCRIPTION OF THE INVENTION

[0018] Embodiments of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to these embodiments.

[0019] [1. Overview] First, with reference to FIGS. 1 and 2, an overview of the present invention will be described. FIG. 1 is a diagram showing an example of expected probability based on Benford's law. FIG. 2 is a diagram showing an example of a dataset based on Benford's law.

[0020] Conventionally, as a macro statistical analysis, there is a fraud detection method that focuses on the distribution of the frequencies of the leading digit numbers. The frequency of the leading digit numbers generally follows a certain distribution, which is called Benford's law. Here, in the world of accounting processes, Benford's law is utilized by paying attention to the frequency of the leading digit numbers appearing in accounting books to check whether the process is appropriate and whether there is no forgery. Here, as shown in FIG. 1, according to Benford's law, macroscopically, for the leading digit values of many numbers appearing in nature, such as the amounts on accounting vouchers, the smaller the leading digit number, the higher the frequency of appearance. The frequency (expected probability) of the number d from 1 to 9 as the leading digit number appearing is expressed as "the frequency of the number d (leading digit number) = log 10 (1 + 1 / d)".

[0021] Therefore, conventionally, by using the distribution of Benford's law, it is checked whether there is no artificial fraud by whether the accounting numbers of an organization tend to be close to the Benford distribution. Also, conventionally, as an index for measuring the similarity of distributions, a statistical hypothesis test called a goodness-of-fit test (chi-square test) is used (null hypothesis: the observed data set follows Benford's law, alternative hypothesis: the observed data set does not follow Benford's law). The number of appearances of each leading digit number in the journal vouchers of a certain department is tabulated, and it is determined by the chi-square test whether "the observed data follows Benford's law". If the observed data (evaluation data) does not follow Benford's law, it is assumed that some kind of fraud operation may have been carried out. Here, in the chi-square test, the state assumed to be normal is set as the null hypothesis, and the opposite abnormal state is set as the alternative hypothesis. A probability (significance level) serving as a criterion for rejecting the null hypothesis is set, and the test is performed. When the P-value in the test is lower than the significance level, the null hypothesis is rejected and the alternative hypothesis is adopted. Here, the P-value refers to the probability that the test statistic becomes that value under the null hypothesis in a statistical hypothesis test. The smaller the P-value, the less likely it is for the test statistic to become that value.

[0022] For example, as shown in FIG. 2, in conventional observation data, "expected frequency of digit i = expected probability of digit i (FIG. 1) × evaluation data size", "degree of deviation of digit i = (observed frequency of digit i - expected frequency of digit i)^2 / expected frequency of digit i = data size × (observed probability of digit i - expected probability of digit i)^2 / expected probability of digit i", "chi-square estimated value = sum of degrees of deviation of digit i (i = 1, 2,..., 9) = 19.2272", "degrees of freedom: 8 (= 9 - 1)", "rejection region (threshold) of chi-square distribution with significance level: 0.05 = 15.5 or more", and since "chi-square estimated value = 19.2272 > 15.5", the null hypothesis is rejected, and since it does not follow Benford's law, it is determined that there is some abnormality.

[0023] On the other hand, conventionally, since not all numbers existing in nature follow the theoretical distribution of Benford, it has been difficult to analyze data sets that are difficult to apply. Also, conventionally, when aggregating the leading digit numbers by department rather than the entire company, there is a possibility that the characteristics may differ by department and may not all follow Benford's law.

[0024] Therefore, in the present embodiment, a mechanism is provided that can perform a chi-square test on a data set that does not follow Benford's law using a model learned from a teacher data set. Also, in the present embodiment, a mechanism is provided that can detect fraud in changing the amounts of expenses and receipts at the person or department level, etc.

[0025] Here, in the present embodiment, Benford's law is a law stating that the occurrence probabilities of the leading digits of many numerical values appearing in nature follow a certain distribution. The expected (theoretical) probability is the probability of the model learned from the teacher data or the probability of Benford's law. The observed probability is the occurrence probability for each leading digit number aggregated from the data. The expected frequency is the expected probability × the size of the teacher data, and the observed frequency is the observed probability × the size of the evaluation data. A data set is a collection of data (for example, a teacher data set or an evaluation data set, etc.). A (prediction) model is something that is learned from Benford's law or the teacher data and has the feature of "being able to calculate the probability of a certain event". Evaluation means performing a chi-square test from the evaluation data and the model to perform anomaly detection. The leading digit number is the digit of the largest digit (the leftmost) of a number (1 to 9 excluding 0). The evaluation data is the data for which anomaly detection is to be performed. The learning data is the data necessary for learning the prediction model. The teacher data may be a type of learning data (for example, data in a state where the correct answer is given to the learning data (learning data), etc.).

[0026] Also, in the present embodiment, Bayesian statistics is a method by which a prediction model can be estimated in the form of a probability distribution from certain information known in advance (distribution information of Benford's law) and data (teacher data). More precisely, it utilizes the fact that the occurrence frequencies of the numbers 1 to 9 follow a multinomial distribution and that the multinomial distribution has a conjugate prior distribution (Dirichlet distribution: a prior distribution set so that the prior distribution and the posterior distribution have the same type of probability distribution). By determining the theoretical probability of Benford and the magnitude of the prior information, a prior distribution (Dirichlet distribution) is defined, the likelihood of the multinomial distribution is calculated using the teacher data, and the posterior distribution is obtained (posterior distribution ∝ likelihood × prior distribution). Here, in the present embodiment, for each number, the expected value of the posterior distribution is taken as the expected probability.

[0027] [2. Configuration] An example of the configuration of the accounting anomaly detection device 100 according to this embodiment will be described with reference to FIG. 3. FIG. 3 is a block diagram showing an example of the configuration of the accounting anomaly detection device 100 in this embodiment.

[0028] As shown in FIG. 3, the accounting anomaly detection device 100 is a commercially available desktop personal computer. Note that the accounting anomaly detection device 100 is not limited to a stationary information processing device such as a desktop personal computer, and may be a portable information processing device such as a commercially available notebook personal computer, PDA (Personal Digital Assistants), smartphone, or tablet personal computer.

[0029] The accounting anomaly detection device 100 includes a control unit 102, a communication interface unit 104, a storage unit 106, and an input / output interface unit 108. Each unit included in the accounting anomaly detection device 100 is communicably connected via an arbitrary communication path.

[0030] The communication interface unit 104 communicably connects the accounting anomaly detection device 100 to the network 300 via a communication device such as a router and a wired or wireless communication line such as a dedicated line. The communication interface unit 104 has a function of communicating data with other devices via a communication line. Here, the network 300 has a function of communicably connecting the accounting anomaly detection device 100 and the server 200 to each other, and is, for example, the Internet or a LAN (Local Area Network).

[0031] An input device 112 and an output device 114 are connected to the input / output interface unit 108. As the output device 114, in addition to a monitor (including a touch panel), a speaker or a printer can be used. As the input device 112, in addition to a keyboard, a mouse, and a microphone, a monitor that realizes a pointing device function in cooperation with the mouse can be used. In the following, the output device 114 may be described as the monitor 114 or the printer 114, and the input device 112 may be described as the keyboard 112 or the mouse 112.

[0032] The storage unit 106 stores various databases, tables, files, and the like. The storage unit 106 records a computer program for giving instructions to the CPU (Central Processing Unit) to perform various processes in cooperation with the OS (Operating System). As the storage unit 106, for example, a memory device such as a RAM (Random Access Memory) or a ROM (Read Only Memory), a fixed disk device such as a hard disk, a flexible disk, an optical disk, or the like can be used. The storage unit 106 includes an expense database 106a and a setting master 106b.

[0033] The expense database 106a stores accounting data including expense data (business data). Here, the expense database 106a may store teacher data in which the expense amount for the teacher is set, and a teacher data set in which the expected probability of each leading digit number of the expense amount for the teacher is set. Here, in the teacher data set, the expected probability of each leading digit number of the expense amount for the teacher may be set for each predetermined group. Further, the expense database 106a may store pre-processed data in which the expense amount to be evaluated is set. Further, the teacher data may be expense data in which the expense amount of the previous year of the evaluation target is set. Further, the expense database 106a may store evaluation data (post-processed data), an abnormal determination target data set, pre-processed data, a prediction model, update data, and range target data.

[0034] The setting master 106b is a master in which setting values for performing abnormal determination are set. Here, the setting master 106b may be set with a determination item indicating the type of the expense amount to be evaluated, a unit of a predetermined group, a Benford non-use flag indicating that the teacher data set or the Benford's law is used for abnormal determination, a significance level, and a prediction model type that is the unit of the abnormal determination target.

[0035] The control unit 102 is a CPU or the like that comprehensively controls the accounting anomaly detection device 100. The control unit 102 has an internal memory for storing control programs such as an OS, programs defining various processing procedures, and required data, and executes various information processes based on these stored programs. Conceptually in terms of functions, the control unit 102 includes a teacher data acquisition unit 102a, an evaluation data acquisition unit 102b, a determination target acquisition unit 102c, and a determination result acquisition unit 102d.

[0036] The teacher data acquisition unit 102a acquires a teacher data set. Here, the teacher data acquisition unit 102a may acquire a teacher data set in which, based on the teacher data, the observed frequency of each leading digit number of the teacher's expense amount is acquired and the expected frequency of each leading digit number of the teacher's expense amount is set. Also, the teacher data acquisition unit 102a may acquire a teacher data set in which, based on the teacher data, the observed frequency of each leading digit number of the teacher's expense amount is acquired, and when there is an observed frequency that is 0 and / or when the total of the observed frequencies is smaller than a predetermined size, the expected frequency of each leading digit number of the teacher's expense amount is set based on Bayes' theorem. Further, the teacher data acquisition unit 102a may register the teacher data set in the expense database 106a.

[0037] The evaluation data acquisition unit 102b acquires evaluation data. Here, the evaluation data acquisition unit 102b may acquire evaluation data in which the expense amount of the evaluation target and the observed frequency of each leading digit number of the expense amount of the evaluation target are set. Here, the evaluation data may have the observed frequency of each leading digit number of the expense amount of the evaluation target set for each predetermined group. Also, the evaluation data acquisition unit 102b may acquire the evaluation data based on the setting master 106b. Further, the evaluation data acquisition unit 102b may acquire evaluation data in which the expense amount of the evaluation target and the observed frequency of each leading digit number of the expense amount of the evaluation target are set by tabulating the appearance frequency of the leading digit number of the expense amount of the evaluation target based on the pre-processed data and acquiring the observed frequency of each leading digit number of the expense amount of the evaluation target.

[0038] The determination target acquisition unit 102c acquires an abnormal determination target data set. Here, the determination target acquisition unit 102c may acquire an abnormal determination target data set in which the expected frequency of each leading digit number of the expense amount to be evaluated and the degree of deviation between the observed frequency of each leading digit number of the expense amount to be evaluated and the expected frequency are associated based on the teacher data set and the evaluation data. Here, in the abnormal determination target data set, the expected frequency of each leading digit number of the expense amount to be evaluated for each predetermined group and the degree of deviation between the observed frequency of each leading digit number of the expense amount to be evaluated and the expected frequency may be associated and set. Further, when the Benford non-use flag indicates that the teacher data set is used for abnormal determination based on the setting master 106b, the teacher data set, and the evaluation data, the determination target acquisition unit 102c may acquire an abnormal determination target data set. Further, when the Benford non-use flag indicates that the Benford's law is used for abnormal determination based on the setting master 106b and the evaluation data, the determination target acquisition unit 102c may acquire an abnormal determination target data set in which the Benford expected frequency of each leading digit number of the expense amount to be evaluated and the degree of deviation between the observed frequency of each leading digit number of the expense amount to be evaluated and the Benford expected frequency are associated and set.

[0039] The determination result acquisition unit 102d acquires a determination result for the evaluation data. Here, the determination result acquisition unit 102d may acquire a determination result for the evaluation data by obtaining a test statistic and / or a p-value of each leading digit number of the expense amount to be evaluated through a statistical hypothesis test on the abnormal determination target data set. Further, the determination result acquisition unit 102d may acquire a determination result based on the setting master 106b. Further, the determination result acquisition unit 102d may display the determination result.

[0040] [3. Specific Example] A specific example of this embodiment will be described with reference to FIGS. 4 to 28.

[0041] [Accounting Abnormality Detection Process] Here, referring to FIG. 4, an example of the accounting anomaly detection process in the present embodiment will be described. FIG. 4 is a flowchart showing an example of the process of the accounting anomaly detection device 100 in the present embodiment.

[0042] As shown in FIG. 4, the teacher data acquisition unit 102a acquires the observed frequency of the leading digit of each teacher expense amount based on the teacher data stored in the expense database 106a, obtains a teacher data set in which the expected frequency of the leading digit of each teacher expense amount is set, and registers the teacher data set in the expense database 106a (step SA-1).

[0043] Then, the evaluation data acquisition unit 102b acquires evaluation data in which the expense amount to be evaluated and the observed frequency of the leading digit of each expense amount to be evaluated are set based on the setting master 106b (step SA-2).

[0044] Then, the determination target acquisition unit 102c determines whether the Benford non-use flag is set to True (using the teacher data set for anomaly determination) based on the setting master 106b (step SA-3).

[0045] Then, when the determination target acquisition unit 102c determines that the Benford non-use flag is set to True (step SA-3: Yes), the process proceeds to step SA-4.

[0046] Then, the determination target acquisition unit 102c acquires an anomaly determination target data set in which the expected frequency of the leading digit of the expense amount to be evaluated and the degree of deviation between the observed frequency and the expected frequency of the leading digit of the expense amount to be evaluated are associated based on the teacher data set stored in the expense database 106a and the evaluation data (step SA-4).

[0047] Then, based on the setting master 106b, the determination result acquisition unit 102d obtains the test statistic and / or p-value of the leading digit of each expense amount to be evaluated by performing a statistical hypothesis test on the abnormal determination target data set, thereby obtaining the determination result for the evaluation data (step SA-5).

[0048] Then, the determination result acquisition unit 102d causes the output device 114 to display the determination result (step SA-6) and ends the process.

[0049] On the other hand, when the determination target acquisition unit 102c determines that the Benford non-use flag is set to False (not set to True) (step SA-3: No), the process proceeds to step SA-7.

[0050] Then, the determination target acquisition unit 102c obtains an abnormal determination target data set in which the Benford expected frequency of the leading digit of each expense amount to be evaluated and the degree of deviation between the observed frequency of the leading digit of each expense amount to be evaluated and the Benford expected frequency are associated based on the evaluation data (step SA-7), and the process proceeds to step SA-5.

[0051] Here, with reference to FIGS. 5 to 9, an example of the premise of the accounting anomaly detection process in the present embodiment will be described. FIG. 5 is a diagram showing an example of evaluation data in the present embodiment. FIG. 6 is a diagram showing an example of teacher data in the present embodiment. FIGS. 7 to 9 are diagrams showing an example of the setting master 106b in the present embodiment.

[0052] In the present embodiment, for the purpose of detecting anomalies from the number of occurrences of the leading digits of a numerical data set, expense data (business data) is used as target data, and an anomaly determination process for detecting the possibility of fraud for each department is performed about once a year from the expense data (subset: subset of data) for each department. Therefore, as shown in FIGS. 5 and 6, in the present embodiment, expense data for teachers and evaluation is prepared, and the previous year's data is used as teacher data, and the current year's data is used as evaluation data.

[0053] Also, in the present embodiment, a set value for performing abnormality determination is registered in the setting master 106b. Since detection is performed in units of departments, the group unit is set with department codes. As shown in FIG. 7, as a scenario assuming compliance with Benford's law, (A) a scenario of not preparing teacher data (using Benford's law) is set. As a scenario assuming non-compliance with Benford's law, as shown in FIG. 8, (B) a scenario of creating a prediction model for the entire dataset (preparing evaluation data and teacher data) is set. As shown in FIG. 9, (C) a scenario of creating a prediction model for each group (preparing evaluation data and teacher data) is set. Here, the determination item is target data for performing abnormality determination, the group unit is the unit when dividing the evaluation dataset into subsets, the Benford non-use FLG (flag) indicates that when it is TRUE, learning is performed using teacher data instead of Benford's law. The range target is used in the notification message after abnormality detection and can express the range of min to max by specifying the period and range of the target data. The type of the prediction model can be set when the Benford non-use FLG is True and the group unit is set (it is possible to select whether to create a prediction model using all teacher datasets or create one for each group unit (subset)). The size of the prior information can be set when the Benford non-use FLG is True (it is possible to set how much information of Benford's law to give (default 10)).

[0054] Also, with reference to FIGS. 10 to 18, an example of the accounting abnormality detection process in the present embodiment will be described. FIGS. 10 to 18 are diagrams showing an example of the accounting abnormality detection process in the present embodiment.

[0055] First, as shown in FIG. 10, in the present embodiment, for the purpose of arranging data in a form that allows a chi-square test to be performed as a preprocessing, the number of occurrences (observed frequencies) of the first-digit numbers (1 to 9) of the data set set in the determination items of the setting master 106b is obtained for both the teacher data set and the evaluation data set. Here, the Benford's expected frequencies are calculated from the theoretical probabilities. Note that FIG. 10 shows the preprocessing results for Department 1, but this preprocessing is performed for each group (subset of Departments 1 to 9).

[0056] Then, as shown in FIG. 11, in the present embodiment, as a model learning process, when it is predicted that the numerical data to be evaluated does not follow Benford's law, for the purpose of creating a prediction model (expected probability for each digit) from the pre-prepared teacher data, the observed frequencies for each first-digit number are obtained in the same way as in the preprocessing, and the expected probability is calculated using them. Here, as shown in FIG. 11, in the present embodiment, when (B) is selected in the setting master 106b, the preprocessing is performed using all the teacher data (ignoring the group unit), the observed frequencies for each first-digit number are obtained (that is, the observed frequencies are obtained for the data set including all departments), and the expected probability is calculated based on the prior information (model: Benford, data size: 10). Note that in the present embodiment, Bayesian statistics is used for the purpose of preventing division by zero and creating an appropriate model even when the size of the teacher data set is insufficient. And, as shown in FIG. 11, in the present embodiment, when (C) is selected in the setting master 106b, the observed frequencies are obtained using the teacher data set for each group (for each department code), and the expected probability is obtained based on the prior information (Benford, data size 10).

[0057] And in this embodiment, a chi-square test is performed on the dataset to be evaluated, and an anomaly determination process (dataset) for anomaly detection is executed. When groups (for example, department units, etc.) are set, anomaly detection is performed for each subset (for example, for each department, etc.) according to the setting items. Here, in this embodiment, a chi-square test is performed at a significance level of 5%, and an evaluation dataset (or its subset when group units are set) for which the null hypothesis is rejected is determined to be an anomaly. Here, in this embodiment, the null hypothesis may be that the evaluation data follows Benford's law or the learned model.

[0058] That is, as shown in FIG. 12, in this embodiment, when using Benford's law, the degree of deviation from Benford's expected frequency is obtained using the observed frequencies processed in the preprocessing. Here, the observed frequency (i) represents the observed frequency of the leading digit number i. Also, the chi-square estimated value is defined as the sum of the degrees of deviation for each digit. Also, as shown in FIG. 12, in this embodiment, since the rejection threshold of the upper 5% chi-square test is 15.5 for a total degree of deviation of 18.1, this dataset is determined to be an anomaly because 18.1 > 15.5. Also, as shown in FIG. 12, in this embodiment, since the number of categories is 9 (numbers 1 to 9), the degrees of freedom 8 can be obtained by subtracting 1 from it.

[0059] Also, in this embodiment, when using the teacher data, the user selects whether to create a prediction model for each department or create one prediction model for the whole in the setting master 106b (type of prediction model).

[0060] That is, as shown in FIG. 13, in this embodiment, when (B) is selected in the setting master 106b, a prediction model is created for the entire dataset. Here, as shown in FIG. 13, in this embodiment, after dividing the evaluation data by department units, the observed frequencies are obtained by processing in the preprocessing, and anomaly determination is performed for departments 001 to 009 using the learned expected probabilities (learned using all the teacher data) in Table 2-2.

[0061] Also, as shown in FIG. 14, in the present embodiment, when (C) is selected in the setting master, a prediction model is created for each group. As shown in FIG. 14, in the present embodiment, when department B010 does not exist in the teacher data set, a prediction model is created for the entire data set, and anomaly detection is performed using the prediction model. Also, in the present embodiment, even if prediction model creation is set for each group, a prediction model may be created in advance for the entire data set. That is, as shown in FIG. 14, in the present embodiment, after dividing the evaluation data by department, it is processed in preprocessing to obtain the observed frequency, and for each department, an anomaly determination is made using the learned expected probability in Table 3-1 (learned for each department, in this case department B001). In the present embodiment, an anomaly determination is made for department code: B002 using the expected probability in Table 3-2. Also, in the present embodiment, for a department that does not exist in the teacher data set but exists for evaluation (for example, department code: B010), an anomaly determination is made using the expected probability in Table 2-2. That is, the expected probability in Table 2-2 is used for "departments without teacher data and those for which the entire data set has been learned", while the expected probabilities in Tables 3-1 to 3 are used for "departments with teacher data and those learned for the corresponding departments".

[0062] And in the present embodiment, the abnormality determination process (first digit number) is performed only on the data set determined to be abnormal in the abnormality determination process (data set). Here, as shown in FIG. 15, in the present embodiment, after performing a chi-square test to determine the abnormality of the data set, it is difficult to draw a lead line to analyze where the cause is, so a hypothesis test of the binomial distribution is performed for each first digit number. That is, in the present embodiment, a hypothesis test of the binomial distribution is performed for each first digit number for all of the scenarios (A), (B), and (C) registered in the setting master 106b (Tables 1' to 3'). Specifically, as shown in FIG. 15, in the present embodiment, it can be seen from Table 1' that the numbers 1 and 9 contribute greatly to the value of the chi-square estimated value (judged from the deviation degree values of the numbers 1 or 9), so there seems to be fraud in the expense details starting with 1 or 9, but it is difficult to draw a line, such as whether there is a problem with the number 8. Therefore, by performing a hypothesis test of the binomial distribution for each first digit number, a specific threshold can be given. For example, as shown in FIG. 15, in the present embodiment, according to Benford's law, the expected probability of the first digit number 1 is 0.301. Therefore, it can be assumed that the number of details with the first digit number 1 among n expense details follows the binomial distribution B(n, p = 0.301). If n is large, the binomial distribution can be approximated by the normal distribution N(μ = np, σ2 = np(1 - p)) (for example, a guideline for whether approximation is possible: n > 25 to 30). In the present embodiment, the approximation can be performed by the central limit theorem. The distribution of the population following the mean μ and the variance σ2 approaches the normal distribution N(μ, σ2) as the sample size n extracted increases. The mean of the binomial distribution B(n, p) is np, and the variance is np(1 - p).

[0063] Then, as shown in FIG. 16, in this embodiment, the binomial distribution B(465, 0.301) can be approximated by a normal distribution N(μ = 465×0.301, σ2 = 465×0.301(1 - 0.301)), and a z-test is performed on this. Here, in this embodiment, the null hypothesis is set as "the observed frequency of the leading digit 1 = the expected frequency of the leading digit 1", and it is determined whether this hypothesis can be rejected. That is, as shown in FIG. 16, in this embodiment, in order to determine outliers using the statistic (z-value), the approximated normal distribution is standardized, and the obtained value is used as the test statistic (z-value). By performing a two-sided test at a significance level of 5%, from P(|z|≧1.96)=0.05, if the absolute value of this z is greater than 1.96, the null hypothesis is rejected. Here, the p-value is α / 100 of the significance level α%, and is defined by P(|z|)=p-value. In this embodiment, when obtaining the rejection region for the leading digit 1, the width of the rejection region (1.96×(expected probability×(1 - expected probability) / evaluation data size)^(1 / 2)×evaluation data size = 1.96×(0.301×(1 - 0.301) / 465)^(1 / 2)×465 = 19.39) is obtained. When the observed frequency < the expected frequency of digit 1 - 19.39, the expected frequency of digit 1 + 19.39 < the observed frequency, the observed frequency < 120.59, and 159.37 < the observed frequency, digit 1 is determined to be abnormal. Therefore, the digit 1 (observed frequency is 119) in Table 1 is determined to be abnormal.

[0064] And, as shown in FIG. 17, in this embodiment, as a result update process, the information necessary for display of the determination result is updated in a table. Here, as shown in FIG. 17, the chi-square estimated value is obtained by an abnormality determination (data set), and whether the determination result is True or False is determined by whether it exceeds the chi-square threshold value of 15.5 at the upper 5% of the significance level. When it exceeds 15.5, it becomes True. Also, as shown in FIG. 17, the chi-square P-value can be obtained from the chi-square estimated value, and is calculated by integrating from the chi-square estimated value to infinity with respect to the x-axis of the chi-square distribution (the horizontal axis direction of the chi-square distribution curve). The range target Min and range target Max respectively obtain the minimum value and maximum value of the range target data set in the setting master 106b of (A) to (C).

[0065] And, as shown in FIG. 18, in the present embodiment, determination display ((A) in FIG. 17 update data (without data division)) is performed. From the upper graph and the results, it can be read that there are many details with the leading digit number 9 and few details with the leading digit number 1. From the lower two graphs, the ratios of expense items regarding the leading digit numbers 1 and 9 can be compared between employee A and all employees, and the cause of fraud can be investigated.

[0066] Also, referring to FIGS. 19 to 23, it is a diagram showing an example of creating a prediction model using Bayes in the present embodiment. FIGS. 19 to 23 are diagrams showing an example of the prediction model creation process in the present embodiment.

[0067] First, in the present embodiment, Bayesian statistics is a statistics that deals with subjective probabilities. It can derive probabilities even when the data is insufficient. By giving the information known in advance as a prior probability and updating it as a posterior probability every time information is obtained, a model that captures the events of the population can be created. In Bayesian statistics, in order to estimate the posterior probability of unknown parameters or models based on data, Bayes' theorem is applied. Here, in the present embodiment, it is assumed that the observed data follows a multinomial distribution, and a Dirichlet distribution is assumed as the prior distribution. For example, in the present embodiment, for 100 claim amount data, the probability that the leading digit number (1 to 9) appears follows a multinomial distribution with parameters n = 100, P = (p1, p2, ~, p9) (n: number of trials, pi: appearance probability (expected probability) of the leading digit number i).

[0068] Here, in the present embodiment, Bayesian estimation is a method for identifying the population of the source from data obtained under a certain assumption (= prior distribution). A probability distribution outputs the probabilities of all events that can occur in a certain trial. The prior (probability) distribution is the distribution when an approximate distribution is known from past information or experience. The posterior (probability) distribution is a probability distribution obtained as a result of considering the data, which is obtained by combining the prior distribution and the data (likelihood) and calculating the posterior probability according to Bayes' theorem. A conjugate prior distribution is a prior distribution set so that the prior distribution and the posterior distribution have the same type of probability distribution. The prior probability is a probability defined based on past information or experience. The posterior probability is a probability considering certain information or data. The expected probability is the expected value of the posterior probability. The likelihood is an indicator that tells us which preconditions (parameters / models) are reasonable to infer from a certain result (data). It has parameters as variables, and when data is given, the plausibility of the occurrence of the parameters can be expressed by the formula "L(θ|D)=P(D|θ)" (L: likelihood, D: data, θ: parameter, P: probability). The normalization constant is a constant that plays a role in adjusting so that the integral (sum) of the posterior distribution becomes 1 (which can be ignored when deriving the posterior distribution). The subjective probability is the subjective belief or degree of confidence that a person has.

[0069] Here, regarding the prediction model using the multinomial distribution and the Dirichlet distribution, as shown in FIG. 19, in the present embodiment, Bayesian statistics (multinomial distribution and Dirichlet distribution) is applied to the Benford anomaly detection system, and the size of the training dataset is 100 and the size of the prior information (Benford's law) is 50 (note that in this case, since the observed frequency of the number 9 is 0, it would result in a division by zero as it is).

[0070] Then, as shown in FIG. 20, in the present embodiment, in order to obtain the expected probability of each digit, the posterior probability P(θ|D) is calculated and obtained from the prior distribution P(θ) including the likelihood P(D|θ) of the data and the prior information of Benford's law. And as shown in FIG. 20, in the present embodiment, E(θ iFrom the formula of , the expected probability for each digit is obtained and used as the prediction model.

[0071] Here, the graph in Fig. 21 shows the Benford's law probability, the expected probability, and the observed probability (= observed frequency ÷ data size (100)). The expected probability takes a value between the Benford's law probability and the observed probability. As the Benford's information is included as prior information, it can be seen that the expected probability of the digit 9 is not zero and no division by zero occurs.

[0072] Also, regarding the behavior when the prior information size is changed, as shown in Fig. 22, in this embodiment, not only when the prior information size is 50, but also when using the prior information (observed frequency) when the prior information sizes are 10 and 200, when using the teacher dataset in Table T1, as shown in Fig. 23, the posterior distribution can be represented. As can be read from this graph, it can be seen that by increasing the prior information size, it approaches the Benford's law probability. In this embodiment, as tuning when the abnormality degree of the prediction model is too high, the size of the prior information may be changed to approach the Benford's law probability.

[0073] Also, in this embodiment, with reference to Figs. 24 to 28, an example of the scenario setting in this embodiment will be described. Figs. 24 to 28 are diagrams showing an example of the scenario setting in this embodiment.

[0074] As shown in FIG. 24, in the present embodiment, for a data set for which an abnormality is to be determined, when there is no past data (teacher data), the Benford non-use flag is set to FALSE, and an abnormality determination is performed using Benford's law. Here, in the present embodiment, [1] when the determination result shows a significantly abnormal tendency (p-value is less than or equal to α, and α is between 0.001 and 0.01), consider whether teacher data can be collected. If it cannot be collected, consider whether a part of the evaluation data can be assigned as teacher data. If teacher data can be prepared, proceed to [B-2], or consider setting or changing the group unit. If teacher data cannot be prepared, consider setting or changing the group unit. [2] When the determination result shows an abnormality (p-value is between α and 0.05, and α is between 0.001 and 0.01), observe the situation for a while as it is. If this continues, shift to [1]. [3] When the determination result does not show an abnormality, continue the operation as it is.

[0075] Also, as shown in FIG. 25, in the present embodiment, for a data set for which an abnormality is to be determined, when there is past data (teacher data), [B-1] when it can be assumed that Benford's law is followed, the same setting as [A]: the Benford non-use flag is set to FALSE, and an abnormality determination is performed using Benford's law.

[0076] Also, as shown in Fig. 25, in this embodiment, for the data set for which abnormality determination is desired, when there is past data (teacher data), when it is assumed that it is not possible to follow Benford's law, and when [B-2-1] the group unit is not set, a prediction model is created using all of the teacher data sets, and abnormality determination is performed. Here, in this embodiment, [1] when the determination result shows a rather abnormal tendency (p-value is less than or equal to α, and α is 0.001 to 0.01), when considering changing the group unit setting (reviewing the purpose) or when the teacher data is not very reliable, consider increasing the size of the prior information or whether it is possible to increase the size of the teacher data. [2] When the determination result shows abnormality (p-value is α to 0.05), observe the situation like this for a while. If this continues, shift to [1]. [3] When the determination result does not show abnormality, continue the operation as it is.

[0077] Also, as shown in Fig. 25, in this embodiment, for the data set for which abnormality determination is desired, when there is past data (teacher data), when it is assumed that it is not possible to follow Benford's law, and when [B-2-2] there is a group unit setting, and when [B-2-2-1] it can be (temporarily) assumed that there is no significant difference in the distribution of the leading digit numbers for each group, set the Benford non-use flag to TRUE, set the type of the prediction model to "create a prediction model using the entire data set", create a prediction model using the entire data set, and perform abnormality determination for each group. Here, in this embodiment, [1] when the determination result shows a rather abnormal tendency (p-value is less than or equal to α, and α is 0.001 to 0.01), when considering changing the group unit setting (reviewing the purpose) or when the teacher data is not very reliable, consider increasing the size of the prior information, whether it is possible to increase the size of the teacher data, or if there is a possibility that the characteristics are different (have different distributions) for each group, consider whether to perform abnormality determination in [B-2-2-2]. [2] When the determination result shows abnormality (p-value is α to 0.05), observe the situation like this for a while. If this continues, shift to [1]. [3] When the determination result does not show abnormality, continue the operation as it is.

[0078] Also, as shown in FIG. 25, in the present embodiment, for the data set for which abnormality determination is desired, when there is past data (teacher data), when it cannot be assumed to follow Benford's law, and when there is a group unit setting, and when it can be assumed that the distribution of the leading digit numbers is different for each group, or when it is found to be different from [B-2-2-1], the Benford non-use flag is set to TRUE, the type of the prediction model is set to "create a prediction model for each group", a prediction model is created for each group, and abnormality determination is performed for each group. Here, in the present embodiment, when the determination result shows a rather abnormal tendency (p-value is α or less, and α is 0.001 to 0.01), when considering changing the group unit setting (reviewing the purpose) or when the teacher data is not very reliable, consider increasing the prior information size or whether the teacher data size can be increased. When the determination result shows abnormality (p-value is α to 0.05), observe the situation like this for a while, and if this continues, shift to [1]. When the determination result does not show abnormality, continue the operation as it is.

[0079] Also, in the present embodiment, for the data set for which abnormality determination is desired, when there is past data (teacher data), when it cannot be assumed to follow Benford's law, and when it is unknown, it is shifted to [B-2-2-1]. Also, in the present embodiment, for the data set for which abnormality determination is desired, when there is past data (teacher data) and when it is unknown, it is shifted to [A].

[0080] Also, in the present embodiment, when performing anomaly detection for each subset of company departments for the purpose of detection, although there is teacher data (past data), but when it is not known which settings should be made, as shown in Fig. 26, when the situation of [B-3] of [B] is unknown is adopted, the setting master 106b of [A] is set, and under this setting, the results of Table 1' are obtained. Here, in the present embodiment, the chi-square estimated value of Table 1' is 18.1 (chi-square threshold value: 15.5), and the p-value is 0.02 (significance level 0.05 or less), so it is determined to be abnormal.

[0081] Also, in the present embodiment, although the anomaly determination result of department code: B001 is shown in Table 1', when the degree of anomaly is similarly high for other departments, as shown in Fig. 27, from [1] and [2] of [A] (because there is group unit setting), it is shifted to [B-2-2], then shifted to [B-2-2-3], and once, the anomaly determination is performed at [B-2-2-1], and the results of Table 2' are obtained. Here, in the present embodiment, the chi-square estimated value of Table 2' is 21.8 (chi-square threshold value: 15.5), and the p-value is 0.005 (significance level 0.05 or less), so it is determined to be abnormal. Also, in the present embodiment, although the anomaly determination result of department code: B001 is shown in Table 2', when the degree of anomaly is similarly high for other departments, from [1] of [B-2-2-1], regarding the change of group unit, as long as there is a purpose of detecting anomalies for each department, it is not changed, and it is changed if the purpose changes, such as detecting for each "department × expense item". Regarding the size of the prior information, by increasing the size, it will approach the results of Table 1'. While looking at the results of other departments, if it can be adjusted with this, it will be dealt with by the prior information size. Regarding the teacher data, check whether it is appropriate and consider whether the data size can be increased, update the expected probability, and if it is "presumed that the characteristics are different for each group (department)", change the type of the prediction model.

[0082] Also, in this embodiment, if it is "presumed that the characteristics are different for each group (department)", in order to change the type of prediction model, as shown in Fig. 28, it is shifted to [B-2-2-2] for anomaly determination, and the results in Table 3' can be obtained. Here, in this embodiment, the chi-square estimated value in Table 3' is 7.26 (chi-square threshold: 15.5), and the p-value is 0.51 (exceeding the significance level of 0.05), so it is determined that there is no anomaly. Also, in this embodiment, the determination result (no anomaly) of department code: B001 is shown in Table 3'. However, if the degree of anomaly is excessively high in other departments, from [1] of [B-2-2-2], as long as there is a purpose of detecting anomalies for each department regarding the change of group units, without changing, if the purpose changes, such as wanting to detect for each "department × expense item", it can be changed. Regarding the size of the prior information, by increasing the size, it will approach the results in Table 1'. While looking at the results of other departments, if it can be adjusted with this, it can be dealt with by the prior information size. Regarding the teacher data, confirmation of whether it is appropriate and consideration of whether the data size can be increased are executed, and the expected probability is updated. Note that in this embodiment, the expense amount is applied as the business data for detecting anomalies, but it is not limited to the expense amount and can also be applied to numerical data such as financial statements, sales, purchase amounts, or payment amounts, etc.

[0083] [4. Contribution to the United Nations-led Sustainable Development Goals (SDGs)] According to this embodiment, it can contribute to promoting business efficiency and appropriate business judgment of the enterprise, so it is possible to contribute to Goals 8 and 9 of the SDGs.

[0084] Also, according to this embodiment, it can contribute to reducing waste loss and promoting paperless and digitalization, so it is possible to contribute to Goals 12, 13, and 15 of the SDGs.

[0085] Also, according to this embodiment, it can contribute to strengthening control and governance, so it is possible to contribute to Goal 16 of the SDGs.

[0086] [5. Other Embodiments] In addition to the above-described embodiments, the present invention may be implemented in various different embodiments within the scope of the technical idea described in the claims.

[0087] For example, among the processes described in the embodiments, all or part of the processes described as being automatically performed can be manually performed, or all or part of the processes described as being manually performed can be automatically performed by a known method.

[0088] Also, regarding the processing procedures, control procedures, specific names, information including parameters such as registered data and search conditions for each process, screen examples, and database configurations shown in this specification and the drawings, they can be arbitrarily changed unless otherwise specified.

[0089] Regarding the accounting anomaly detection device 100, each illustrated component is a functional concept and does not necessarily have to be physically configured as shown.

[0090] For example, regarding the processing functions provided by the accounting anomaly detection device 100, particularly each processing function performed by the control unit 102, all or any part of them may be realized by a CPU and a program interpreted and executed by the CPU, or may be realized as hardware by wired logic. Note that the program is recorded on a non-transitory computer-readable recording medium including programmed instructions for causing an information processing device to execute the processes described in this embodiment, and is mechanically read by the accounting anomaly detection device 100 as necessary. That is, in a storage unit such as a ROM or an HDD (Hard Disk Drive), a computer program for giving commands to the CPU in cooperation with the OS and performing various processes is recorded. This computer program is executed by being loaded into the RAM and constitutes the control unit in cooperation with the CPU.

[0091] Further, this computer program may be stored in an application program server connected to the accounting anomaly detection device 100 via an arbitrary network, and it is also possible to download all or part of it as necessary.

[0092] Also, a program for executing the processing described in this embodiment may be stored in a non-transitory computer-readable recording medium, or it may be configured as a program product. Here, this "recording medium" includes any "portable physical medium" such as a memory card, a USB (Universal Serial Bus) memory, an SD (Secure Digital) card, a flexible disk, a magneto-optical disk, a ROM, an EPROM (Erasable Programmable Read Only Memory), an EEPROM (registered trademark) (Electrically Erasable and Programmable Read Only Memory), a CD-ROM (Compact Disk Read Only Memory), an MO (Magneto-Optical disk), a DVD (Digital Versatile Disk), and a Blu-ray (registered trademark) Disc.

[0093] Also, the "program" is a data processing method described in any language or description method, regardless of the form such as source code or binary code. Note that the "program" is not necessarily limited to being configured singly, and also includes those that are distributed as a plurality of modules or libraries, or those that achieve their functions in cooperation with another program represented by an OS. Regarding the specific configuration, reading procedure, and installation procedure after reading for reading the recording medium in each device shown in this embodiment, well-known configurations and procedures can be used.

[0094] The various databases and the like stored in the memory unit 106 are storage means such as memory devices like RAM and ROM, fixed disk devices like hard disks, flexible disks, and optical disks, and store various programs, tables, databases, and web page files used for various processes and website provision.

[0095] Further, the accounting anomaly detection device 100 may be configured as an information processing device such as a known personal computer or workstation, or may be configured as the information processing device to which an arbitrary peripheral device is connected. Also, the accounting anomaly detection device 100 may be realized by installing software (including programs or data, etc.) that realizes the processes described in this embodiment in the device.

[0096] Furthermore, the specific form of the distribution and integration of the devices is not limited to that shown in the drawings, and all or part of them can be functionally or physically distributed and integrated in arbitrary units according to various additions or according to the functional load. That is, the above-described embodiments may be arbitrarily combined and implemented, or the embodiments may be selectively implemented.

Industrial Applicability

[0097] The present invention is useful in various industries that require accounting anomaly detection.

Explanation of Signs

[0098] 100 Accounting anomaly detection device 102 Control unit 102a Teacher acquisition unit 102b Evaluation acquisition unit 102c Determination target acquisition unit 102d Determination result acquisition unit 104 Communication interface unit 106 Memory unit 106a Expense database 106b Setting master 108 Input / output interface unit 112 Input device 114 Output device 200 Server 300 Network

Claims

1. An accounting anomaly detection device comprising a memory unit and a control unit, The memory unit Expense storage means for storing teacher data setting the expense amount for teachers and a teacher data set setting the expected probability of each leading digit of the expense amount for teachers, and The control unit Evaluation acquisition means for acquiring evaluation data setting the expense amount to be evaluated and the observed frequency of each leading digit of the expense amount to be evaluated, Based on the teacher data set and the evaluation data, determination target acquisition means for acquiring an anomaly determination target data set in which the expected frequency of each leading digit of the expense amount to be evaluated and the degree of deviation between the observed frequency and the expected frequency of each leading digit of the expense amount to be evaluated are associated, Determination result acquisition means for acquiring a determination result for the evaluation data by obtaining a test statistic and / or a p-value of each leading digit of the expense amount to be evaluated by performing a statistical hypothesis test on the anomaly determination target data set, An accounting anomaly detection device characterized by comprising the above.

2. The control unit Teacher acquisition means for acquiring the observed frequency of each leading digit of the expense amount for teachers based on the teacher data and acquiring the teacher data set setting the expected frequency of each leading digit of the expense amount for teachers, The accounting anomaly detection device according to claim 1, further comprising the above.

3. The teacher acquisition means Based on the teacher data, obtain the observed frequencies of the leading digit numbers of the teacher's expense amounts. If there is an observed frequency of 0 and / or if the sum of the observed frequencies is less than a predetermined size, obtain the teacher data set in which the expected frequencies of the leading digit numbers of the teacher's expense amounts are set based on Bayes' theorem. The accounting anomaly detection device according to claim 2, characterized in that.

4. The teacher data set is Furthermore, the expected probability of each leading digit number of the teacher's expense amount for each predetermined group is set, The evaluation data is Furthermore, the observed frequency of each leading digit number of the expense amount to be evaluated for each predetermined group is set, The anomaly determination target data set is Furthermore, the expected frequency of each leading digit number of the expense amount to be evaluated for each predetermined group, and the degree of deviation between the observed frequency and the expected frequency of each leading digit number of the expense amount to be evaluated are linked and set. The accounting anomaly detection device according to claim 1, characterized in that.

5. The storage unit is A setting master in which a determination item indicating the type of the expense amount to be evaluated, the unit of the predetermined group, a Benford non-use flag indicating that the teacher data set or Benford's law is used for anomaly determination, a significance level, and a prediction model type that is the unit of the anomaly determination target are linked and set, Further includes The evaluation acquisition means is Based on the setting master, obtain the evaluation data, The determination target acquisition means is Based on the setting master, the teacher data set, and the evaluation data, when the Benford non-use flag indicates that the teacher data set is used for the anomaly determination, obtain the anomaly determination target data set, The determination result acquisition means is The accounting anomaly detection device according to claim 4, wherein the determination result is acquired based on the setting master.

6. The determination target acquisition means Furthermore, based on the setting master and the evaluation data, when the Benford non-use flag indicates that the Benford's law is used for the abnormality determination, the Benford expected frequencies of the leading digit numbers of the expense amounts of the evaluation target, and the difference between the observed frequencies of the leading digit numbers of the expense amounts of the evaluation target and the Benford expected frequencies are associated and set. The accounting anomaly detection device according to claim 5, wherein the anomaly determination target data set is acquired.

7. The expense storage means Furthermore, stores pre-processed data in which the expense amount of the evaluation target is set, The evaluation data acquisition means Based on the pre-processed data, the frequency of occurrence of the leading digit number of the expense amount of the evaluation target is tabulated, and the observed frequency of each leading digit number of the expense amount of the evaluation target is acquired, thereby acquiring the evaluation data in which the expense amount of the evaluation target and the observed frequency of each leading digit number of the expense amount of the evaluation target are set. The accounting anomaly detection device according to any one of claims 1 to 6.

8. The teacher data Is expense data in which the expense amount of the previous year of the evaluation target is set. The accounting anomaly detection device according to any one of claims 1 to 6.

9. An accounting anomaly detection method for causing an accounting anomaly detection device including a storage unit and a control unit to execute, The storage unit Expense storage means for storing teacher data in which a teacher's expense amount is set, and a teacher data set in which the expected probability of each leading digit number of the teacher's expense amount is set, Comprises Executed in the control unit, An evaluation acquisition step of acquiring evaluation data in which the expense amount of the evaluation target and the observed frequency of each leading digit number of the expense amount of the evaluation target are set; A determination target acquisition step of acquiring an abnormality determination target data set in which the expected frequency of each leading digit number of the expense amount of the evaluation target and the degree of deviation between the observed frequency and the expected frequency of each leading digit number of the expense amount of the evaluation target are associated based on the teacher data set and the evaluation data; A determination result acquisition step of acquiring a determination result for the evaluation data by obtaining a test statistic and / or a p-value of each leading digit number of the expense amount of the evaluation target by performing a statistical hypothesis test on the abnormality determination target data set; An accounting anomaly detection method, characterized by including the above.

10. An accounting anomaly detection program for causing an accounting anomaly detection device including a storage unit and a control unit to execute, wherein the storage unit, has an expense storage means for storing teacher data in which the expense amount for teachers is set, and a teacher data set in which the expected probability of each leading digit number of the expense amount for teachers is set, and in the control unit, an evaluation acquisition step of acquiring evaluation data in which the expense amount of the evaluation target and the observed frequency of each leading digit number of the expense amount of the evaluation target are set; A determination target acquisition step of acquiring an abnormality determination target data set in which the expected frequency of each leading digit number of the expense amount of the evaluation target and the degree of deviation between the observed frequency and the expected frequency of each leading digit number of the expense amount of the evaluation target are associated based on the teacher data set and the evaluation data; A determination result acquisition step of acquiring a determination result for the evaluation data by obtaining a test statistic and / or a p-value of each leading digit number of the expense amount of the evaluation target by performing a statistical hypothesis test on the abnormality determination target data set; An accounting anomaly detection program for causing the above to be executed.

Citation Information

Patent Citations

  • Centrally managed and accessed system, method for performing data processing on a plurality of independent servers, and dataset

    JP2014182805A

  • Estimation data generating method, estimation data generating program and estimation data generating system

    JP2017016361A

  • Internal audit support device, internal audit support method, and internal audit support program

    JP2019179531A

  • System and method for cloud device collaborative real-time user usage and performance anomaly detection

    JP2020537215A

  • Accounting fraud detection method, associated device, and program

    JP2023044131A