Program and information processing system
The program integrates multiple indicator scores with non-negative weighting coefficients to accurately assess fraudulent transactions, addressing the limitations of existing technologies in detecting circular transactions.
Patent Information
- Application Number
- JP2025084826
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-05-21
AI Technical Summary
Conventional technologies lack sufficient accuracy in detecting fraudulent transactions, particularly for cleverly concealed transactions like circular transactions, as they do not comprehensively assess multiple indicators and the weighting coefficients are not sufficient for the purpose of detecting such transactions, and the existing technologies have not effectively addressed the need to address the challenges of circular transactions, which require a more comprehensive and accurate assessment.
A program that calculates multiple indicator scores based on transaction data, applies non-negative weighting coefficients to these scores, and integrates them to determine the possibility of fraudulent transactions, utilizing a computer system with a processor, memory, and databases to analyze transaction and scenario data.
Enables high-accuracy evaluation of fraudulent transactions by integrating multiple indicators, providing a comprehensive and objective assessment of transaction data.
Smart Images

Figure 0007763988000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to the analysis of data, particularly in audit work. [Background technology]
[0002] While fraudulent transactions in corporate activities, particularly circular transactions, are difficult to detect, they have a significant impact on the company and its stakeholders, so early detection is essential. Traditionally, in accounting audits and advisory services, auditors and other professionals have relied on their specialized knowledge and experience to detect signs of fraud from transaction data and related information.
[0003] On the other hand, technologies have also been proposed that calculate statistical anomalies based on transaction data to assess the possibility of fraudulent transactions. For example, Patent Document 1 discloses a technology that calculates multiple feature quantities, such as anomaly scores and lift values, based on transaction data, and applies weighting factors to these to calculate an integrated score, thereby estimating the possibility of an anomaly. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent No. 7143545 Summary of the Invention [Problem to be solved by the invention]
[0005] However, conventional technologies have not always achieved sufficient accuracy when assessing the possibility of fraudulent transactions. In particular, for cleverly concealed fraud, such as circular transactions, it is important to comprehensively and accurately assess multiple indicators, not just individual indicators, using limited data. However, in the technology of Patent Document 1, the feature values representing the possibility of fraudulent transactions are limited to those expressed as conditional probabilities, which have a narrow range of mathematical expression and are therefore insufficient for quantifying indicators of circular transactions. Furthermore, the weighting coefficients applied to multiple feature values are not always sufficient for the purpose of detecting circular transactions.
[0006] The present invention aims to improve the accuracy of determining the possibility of fraudulent transactions. [Means for solving the problem]
[0007] One aspect of the present invention provides a program for causing a computer to function as a first calculation unit that calculates multiple indicator scores that individually indicate the degree of each of multiple indicators of fraudulent transaction based on transaction data that indicates the content of the transaction, and a second calculation unit that calculates an integrated score that indicates the possibility of the fraudulent transaction by applying a non-negative weighting coefficient to the multiple indicator scores and adding them together. [Effects of the Invention]
[0008] According to the present invention, the possibility of fraudulent transactions can be evaluated with high accuracy. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of a hardware configuration of an information processing apparatus. [Figure 2] FIG. 10 is a diagram illustrating a transaction DB. [Figure 3] FIG. [Figure 4] FIG. 10 is a diagram illustrating a scenario DB. [Figure 5] FIG. 10 is a diagram illustrating a weight file. [Figure 6]FIG. 1 is a diagram illustrating an example of the functional configuration of an information processing device. [Figure 7] 10 is a flowchart illustrating a processing operation for detecting a fraudulent transaction. [Figure 8] FIG. 10 is a diagram showing a specific example of processing for detecting fraudulent transactions. DETAILED DESCRIPTION OF THE INVENTION
[0010] 1. Configuration <Hardware configuration of information processing device> FIG. 1 is a diagram illustrating an example of the hardware configuration of an information processing device 10. The information processing device 10 is a device used by an auditor who audits transactions as a user for auditing work. The information processing device 10 has a processor 11, a memory 12, a communication unit 13, an operation unit 14, and a display unit 15. These components are connected to each other so that they can communicate with each other, for example, by a bus. The processor 11 controls each part of the information processing device 10 by reading and executing programs stored in the memory 12. The processor 11 is, for example, a CPU (Central Processing Unit).
[0011] The operation unit 14 is equipped with operators such as operation buttons, a keyboard, a touch panel, and a mouse for issuing various instructions, and receives operations and sends signals corresponding to the operation content to the processor 11. This operation is, for example, pressing a button or making a gesture on the touch panel. The display unit 15 has a display screen such as a liquid crystal display, and displays images under the control of the processor 11. A transparent touch panel of the operation unit 14 may be placed on top of the display screen of the display unit 15.
[0012] The communication unit 13 is a communication circuit that communicatively connects the information processing device 10 to external devices, etc., via wired or wireless communication. The memory 12 is a storage means that stores an operating system, various programs, data, etc., that are loaded into the processor 11. The memory 12 has a RAM (Random Access Memory) and a ROM (Read Only Memory). The memory 12 may also have a solid state drive, a hard disk drive, etc. The memory 12 stores a transaction DB 121, a scenario DB 122, and a weight file 123.
[0013] <Transaction DB configuration> FIG. 2 is a diagram illustrating a transaction DB 121. The transaction DB 121 is a database that stores transaction tables describing multiple transactions for each transaction table's identification information. The transaction DB 121 shown in FIG. 2 has a data ID list 1211 and a transaction table 1212. The data ID list 1211 is a table that stores a data ID, which is identification information for data describing a transaction, in association with the data name and a type ID indicating the type of data. Each data ID listed in the data ID list 1211 is associated with one transaction table 1212.
[0014] The transaction table 1212 stores transaction data including multiple transaction records that indicate transaction details. FIG. 3 illustrates an example of the transaction table 1212. For example, FIG. 3 shows a transaction table 1212 with a data ID of "D1" and a transaction table 1212 with a data ID of "D2." Each transaction table 1212 includes transaction details, such as the transaction target, the trading partner, and the person in charge. For example, the transaction table 1212 with a data ID of "D1" includes transaction details, such as "time," "location," "seller," "category," "product name," "unit price," "quantity," and "amount." The transaction table 1212 with a data ID of "D2" includes transaction details, such as "time," "client," "department," "person in charge," "approval date," "delivery date," "discount," and "amount." Each row in the transaction table 1212 shown in FIG. 3 represents a specific item. One transaction table 1212 corresponds to one supply chain. A plurality of transaction records contained in one transaction table 1212 represent a plurality of transactions carried out in one commercial stream.
[0015] A transaction record is a record that has values for each of these items. The columns in the transaction table 1212 shown in Figure 3 indicate transaction records. Each transaction record is assigned identification information, such as a serial number, to identify each transaction record. The values of the transaction details items include quantitative variables and qualitative variables. Quantitative variables are variables expressed as numbers indicating quantity, such as "time" and "amount." Quantitative variables are interval scales and ratio scales. Qualitative variables are variables that are not expressed as numbers indicating quantity, such as "seller" and "product name." Qualitative variables are nominal scales and ordinal scales.
[0016] <Scenario DB configuration> FIG. 4 illustrates the scenario DB 122. The scenario DB 122 is a database that defines various situations and patterns, or "signs," that may indicate fraudulent transactions that may be subject to auditing, particularly circular transactions in this embodiment. Here, "fraudulent transactions" refer to intentional misstatements in a company's financial reporting and the misappropriation of assets. Fraudulent transactions include financial statement fraud, such as window dressing, and asset misappropriation, such as embezzlement. While this embodiment focuses on circular transactions as the fraudulent transaction to be detected, the present invention can also be applied to the detection of various other fraudulent transactions, such as the recording of fictitious sales and inflated expense claims. Furthermore, "signs" refer to any clues or circumstantial evidence that suggest the existence of fraudulent transactions. These "signs" are referred to as "scenarios" in this specification. Each scenario is defined based on the auditor's empirical knowledge and professional judgment (domain heuristics). In the example shown in FIG. 4, 11 scenarios are defined, and each scenario is assigned a unique scenario number. The information processing device 10 calculates an indication score indicating the degree of indication of fraudulent transactions from the transaction data stored in the transaction DB 121 based on each scenario defined in the scenario DB 122. Here, an "indication score" is a numerical value that quantitatively indicates the degree of each indication. Specific examples of indication scores include statistical indicators (such as a z-score indicating deviation from the average, or a value indicating the abnormality of occurrence frequency), the degree to which a specific condition is satisfied (for example, the amount of credit limit overrun or the number of days overrun), the magnitude of change in time-series data, or the degree of pattern matching.
[0017] As shown in Figure 4, each scenario is classified into four categories based on its nature. These categories are established to comprehensively capture perspectives in which signs of fraudulent transactions are likely to appear and reflect expert knowledge in audit practice. Specifically, scenarios in the "Business Transactions - Flow Information" category define signs of abnormalities in individual transactions or short-term transaction flows (e.g., transaction timing, amount, profit margin, etc.). Scenarios in the "Business Transactions - Stock Information" category define signs of abnormalities in the status of a company's assets, such as inventory and stock. Scenarios in the "External Environment" category define signs of abnormalities in relationships with external parties such as business partners (e.g., credit limits). Finally, scenarios in the "Internal Environment" category define signs of abnormalities in the company's internal circumstances (e.g., bias toward specific departments or personnel, fluctuations in performance, etc.). Using signs belonging to the above four categories in transaction analysis enables multifaceted analysis. Therefore, it is desirable for the scenarios defined in the scenario DB 122 to include at least one scenario from each of the above four categories.
[0018] In Figure 4, the content of each scenario is described in natural language, such as "large transactions with no continuity in accounting periods." This description is intended to make it easier for users to intuitively understand the content of the scenario and for auditors to define, select, and evaluate scenarios based on their own knowledge. However, each scenario in the scenario DB 122 is not simply a description; it is defined as a predetermined formula, algorithm, or data processing logic that the information processing device 10 uses as input transaction data to calculate specific indicator scores as numerical values. These formulas and algorithms are created by experts, such as accountants or data scientists at auditing firms, based on audit heuristics and statistical methods, and are implemented in the scenario DB 122. Users can also check these definitions and add or update new scenarios as needed.
[0019] In this example, the "Business Transactions - Flow Information" category includes "large transactions with no continuity in accounting periods," "large transactions before and after financial statements that do not achieve targets," and "significantly different margins compared to similar transactions," defining signs of abnormalities in transaction timing, size, and profitability. In this example, the "Business Transactions - Stock Information" category includes "rapid inventory growth," "abnormal inventory growth (relative to total assets)," and "large number of unstarted items," defining signs of abnormal asset management. In this example, the "External Environment" category includes "credit line violation" and "credit line expansion has not kept up with sales growth for some time since the start of new transactions," defining signs of abnormal relationships with external business partners. In this example, the "Internal Environment" category includes "extreme growth in revenue and gross profit margins (compared to each other)," "bias toward specific departments, products, services, or personnel," and "monthly profit and loss fluctuations," defining signs of abnormal internal performance and business execution. By using scenarios from these diverse perspectives, it is possible to detect signs of fraudulent transactions that may be overlooked using a single indicator.
[0020] <Weight file configuration> 5 is a diagram illustrating an example of the weight file 123. The weight file 123 is a file for storing weighting coefficients used by the symptom integration unit 115, which will be described later, when calculating a single integrated score from multiple symptom scores. Here, the "weighting coefficient" is a coefficient used to adjust the degree of influence that each symptom score has on the integrated score when calculating a single integrated score by integrating multiple symptom scores, and in this embodiment, is used as a coefficient in linear combination. In this example, the weighting coefficient is a value determined by learning based on the symptom scores, and is used to more accurately evaluate the possibility of fraudulent transactions.
[0021] As shown in Fig. 5, the weight file 123 includes a data ID list 1231 and a weight coefficient list 1232. The data ID list 1231 stores a data ID that uniquely identifies the transaction table 1212 shown in Fig. 2 as identification information for the transaction data to which a set of weight coefficients is applied. For example, in the example of Fig. 5, identifiers such as "D1" and "D2" are stored. An indication score is calculated for each supply flow corresponding to the transaction table 1212, and a set of weight coefficients for integrating the indication scores is also set for each supply flow.
[0022] The weighting coefficient list 1232 stores a list of specific weighting coefficients associated with each data ID in the data ID list 1231. This list is composed of multiple records, and each record includes fields for "number" and "weighting coefficient." The "number" field stores the number of each scenario defined in the scenario DB 122. The "weighting coefficient" field stores the value of the weighting coefficient by which the symptom score calculated from the scenario with the corresponding number is multiplied. For example, the weighting coefficient list 1232 for the data ID "D1" in FIG. 5 indicates that the weighting coefficient for the scenario with the number "1" is "0.05," the weighting coefficient for the scenario with the number "2" is "0.1," and the weighting coefficient for the scenario with the number "11" is "0.3." All of these weighting coefficients are non-negative values (values greater than or equal to 0), are calculated by the weight learning unit 114 (described later), and are stored in the weight file 123.
[0023] <Functional configuration of information processing device> Next, the functional configuration of the information processing device 10 according to this embodiment will be described with reference to Fig. 6. Fig. 6 is a diagram illustrating an example of the functional configuration of the information processing device 10. The processor 11 of the information processing device 10 functions as various functional units by executing programs stored in the memory 12. That is, the information processing device 10 has, as functional units, a setting unit 111, an acquisition unit 112, a symptom evaluation unit 113, a weight learning unit 114, a symptom integration unit 115, and an estimation unit 116.
[0024] The setting unit 111 receives a scenario setting operation from the auditor via the operation unit 14. The setting unit 111 stores the received scenario in the scenario DB 122.
[0025] The acquisition unit 112 receives an operation to specify transaction data from a user via the operation unit 14. The acquisition unit 112 receives, as the specification of transaction data, the specification of one or more elements for identifying a commercial flow from the transaction subject, the trading partner, and the person in charge. The acquisition unit 112 acquires the transaction data of the commercial flow specified by the specified elements from the transaction DB 121 stored in the memory 12. The acquisition unit 112 is an example of a reception unit in the present invention.
[0026] The symptom evaluation unit 113 reads out the relevant scenario (calculation formula or algorithm) from the scenario DB 122, and calculates multiple symptom scores that individually indicate the degree of each of multiple symptoms (scenarios) related to fraudulent transactions based on the transaction data obtained from the transaction DB 121. The symptom evaluation unit 113 is an example of a first calculation means in the present invention. The weight learning unit 114 determines a weighting coefficient for each symptom score by learning using the symptom scores calculated by the symptom evaluation unit 113 as input, and stores the results in the weight file 123. The weight learning unit 114 is an example of a determination means in the present invention.
[0027] The symptom integration unit 115 calculates an integrated score indicating the possibility of fraudulent transactions by applying a non-negative weighting factor to multiple symptom scores and adding them up. Specifically, the symptom integration unit 115 calculates the integrated score by weighting the symptom scores calculated by the symptom evaluation unit 113 and the weighting factor read from the weight file 123 (or received directly from the weight learning unit 114). Here, the "integrated score" is an index indicating the overall possibility (risk) of fraudulent transactions. The larger the value of this integrated score, the higher the likelihood of fraudulent transactions is evaluated. The symptom integration unit 115 is an example of a second calculation means of the present invention. Based on the integrated scores calculated by the symptom integration unit 115, the estimation unit 116 generates information for the auditor to make a judgment, such as identifying transactions or groups of transactions that are particularly suspected of being fraudulent, or displaying the value of the integrated score as a risk level, and outputs the information to the display unit 15.
[0028] 2.Operation <Fraudulent transaction detection processing> Next, the specific steps of the fraudulent transaction detection process performed by the information processing device 10 according to this embodiment will be described. FIG. 7 is a flowchart illustrating the processing operations in fraudulent transaction detection, and FIG. 8 is a diagram showing a specific example of the process. Below, the process shown in the flowchart in FIG. 7 will be described with reference to FIG. 8 as needed. The process shown in FIG. 7 is started by the processor 11 executing a program in the memory 12, for example, when an auditor issues an instruction to start the process via the operation unit 14 or at a regularly scheduled timing.
[0029] First, in step S001, the acquisition unit 112 acquires the transaction data to be analyzed. This data mainly includes transaction data acquired from the core system of the audited company and stored in the transaction DB 121. The transaction data includes detailed information such as the date and time of each transaction, the trading partner, the product, the quantity, the unit price, and the amount. This corresponds to the "input transaction data" shown on the left side of Figure 8.
[0030] Next, in step S002, the symptom evaluation unit 113 calculates a symptom score corresponding to each scenario based on the transaction data acquired in step S001 and the scenarios predefined in the scenario DB 122 in the memory 12. Each scenario is defined as a formula or algorithm for capturing specific patterns or situations that may suggest fraudulent transactions. The symptom evaluation unit 113 analyzes the transaction data according to these definitions and calculates a numerical value indicating the severity of each symptom. The central part of Figure 8 illustrates this process, showing, for example, how a symptom score of "Score 1 = 6.25" is calculated for scenario numbered "1," which is "large transactions with no continuity in accounting periods," and how a symptom score of "Score 11 = 2.77" is calculated for scenario numbered "11," which is "monthly fluctuations in profit and loss." In this way, symptoms from multiple different perspectives are quantified.
[0031] Next, in step S003, the weight learning unit 114 determines a weighting factor to be used when aggregating each symptom score into an integrated score. This weighting factor indicates the degree of importance of each symptom in indicating the possibility of fraudulent transactions. The learning process for determining the weighting factor will be described in detail below. The determined weighting factor is stored in the weight file 123. Note that step S003 is not necessarily performed each time transaction data is analyzed; it may be learned and stored in the weight file 123 before transaction data is analyzed. In this case, the weighting factor stored in the weight file 123 is read in step S003. The example in FIG. 8 shows a situation in which "0.05" is determined as the weighting factor corresponding to "score 1" and "0.3" is determined as the weighting factor corresponding to "score 11."
[0032] Next, in step S004, the symptom integration unit 115 calculates an integrated score using the multiple symptom scores calculated in step S002 and the weighting coefficients determined in step S003. Specifically, the product of each symptom score and its corresponding weighting coefficient is calculated, and these are added up for all symptoms. The right side of FIG. 8 illustrates an example of a calculation function for the integrated score s defined by specific weighting coefficients, and a calculation formula for the integrated score s in which specific symptom score values are substituted into the calculation function. Generally, the integrated score s is calculated by multiplying the weighting coefficients {w i}(i=1,…,D; D is the number of scenarios) and symptom scores {s i} (i=1,...,D; D is the number of scenarios), and is calculated using the following formula (1).
number
[0033] In step S005, the estimation unit 116 estimates the likelihood of fraudulent transactions based on the integrated score calculated in step S004. This estimation is performed, for example, by comparing the value of the integrated score s with a predetermined threshold. If the integrated score exceeds the threshold, it is determined that there is a high possibility that fraudulent transactions are included in the transaction data group (or the entire analysis period). Alternatively, the integrated score value can be interpreted directly as a risk level, and a ranking such as "high," "medium," or "low" can be performed depending on the score. This step makes it possible to narrow down the targets that require particular attention from among a large amount of transaction data.
[0034] Finally, in step S006, the estimation unit 116 presents the estimation results from step S005 in a format that is easily understandable to the auditor. This presentation is performed on the display unit 15 of the information processing device 10. Possible presentation formats include, for example, displaying a list of transactions in descending order of integrated score, displaying a color-coded graph or table according to risk level, or displaying not only the integrated score but also a breakdown of each symptom score that served as the basis for calculating the score, thereby providing information for the auditor to conduct a detailed analysis. This enables the auditor to make an audit judgment efficiently and effectively using objective information based on statistical evidence.
[0035] <Learning and determining weighting coefficients> Next, the process of step S003 in the flowchart of Fig. 7, that is, the process of determining the weighting coefficients performed by the weight learning unit 114 shown in Fig. 6, will be described in detail. i} (i=1,...,D; D is the number of scenarios) is the symptom score s calculated in step S002. i The weight learning unit 114 quantitatively indicates how much each symptom contributes to the integrated score s calculated in step S004, that is, the importance of each symptom. i The weighting coefficients can be determined by learning based on the above.
[0036] In the present embodiment, the weighting coefficients are determined using an information-theoretic approach. That is, a latent variable z (which cannot be directly observed) indicating the true degree of fraudulent transactions is assumed, and the observed symptom score s i The weighting coefficient w is used to best capture the information of this latent variable z. i Specifically, the weighting coefficient w that maximizes the mutual information I(s;z) between s and z is determined. i Under statistical assumptions, this mutual information maximization can be formulated as an optimization problem to find the individual scores s1, s2,…, s D The variance of the integrated score s under the given condition of the value of σ 2This can be approximated as the problem of minimizing [s]. i and weighting factor w i Therefore, it is necessary that the criteria for judging the integrated score s be the same even if the commercial flow is different. For this reason, the weight coefficient w i In the learning of i There is also a restriction that the sum of the weighting coefficients w i Since corresponds to a scenario based on the knowledge of accountants and auditors, it is difficult to understand why its contribution to the integrated score s is negative. Therefore, the weighting coefficient w i In the learning of each weight coefficient w i is restricted to be non-negative (i.e., 0 or greater). The above restrictions can be expressed as the following equation (2).
number
[0037] The weight learning unit 114 of this embodiment may perform learning in which further constraints are added to the equation (2). For example, the weight coefficient w i A "minimum weight guarantee" for all weight coefficients w i is not zero and a predetermined positive minimum weight value α i (e.g., α=0.01) or more. Adding such a constraint condition corresponds to incorporating the following equation (3) as an additional constraint when solving the optimization problem using the above equation (2).
number
[0038] An additional constraint for the weight coefficient determination process performed by the weight learning unit 114 may be "specifying the magnitude relationship between weight coefficients." This constraint imposes a magnitude relationship constraint on the weight coefficients corresponding to each scenario when, based on the auditor's knowledge and experience, a certain scenario (sign) is considered to be more important in detecting fraudulent transactions than another scenario (sign). The magnitude relationship constraint may not only relate to the magnitude relationship between individual scenarios, but also to the overall magnitude relationship between scenario groups. For example, if a group of signs G1 belonging to a specific category (e.g., a group of weight coefficients for signs belonging to the "internal environment" category) is considered to be more important overall than a group of signs G2 belonging to another category (e.g., a group of weight coefficients for signs belonging to the "commercial activity - flow information" category), a constraint may be imposed that the sum of the weight coefficients for the group of signs G1 is greater than the sum of the weight coefficients for the group of signs G2. Adding such a constraint is equivalent to incorporating the following equation (4) as an additional constraint when solving the optimization problem represented by equation (2) above.
number
[0039] 3. Variations Although the present invention has been described above with reference to an embodiment, it is not limited to the above embodiment and various modifications are possible within the scope of the technical concept of the present invention. Such modifications are described below. The following modifications can be arbitrarily combined as long as they are not inconsistent.
[0040] (Variations regarding input data) In the above embodiment, an example was described in which a sign score was calculated using mainly transaction data as input. However, the data used for analysis is not limited to this. For example, new sign scores may be calculated based on additional information such as balance trend data for each account item, a comparison between budget data and actual data, inventory retention period data, employee data (e.g., years of service, job title, approval authority), and internal control evaluation results. Furthermore, information available from outside the company, such as financial information on competitors, market data (e.g., stock prices, commodity prices), credit investigation information, related news articles, social media reputation, and litigation information, may also be collected and added to the analysis. In particular, unstructured data such as text data and voice data can be analyzed using natural language processing or speech recognition technology to extract and score information suggesting signs of fraud. This enables more multifaceted and in-depth analysis and is expected to improve the accuracy of fraudulent transaction detection. The timing of data acquisition is not limited to batch processing; a system can also be configured to acquire data in real time or near real time in conjunction with a company's systems, enabling continuous monitoring.
[0041] (Modifications regarding symptom scenarios and score calculation) In the above embodiment, an example of calculating a symptom score based on a predefined scenario, such as that shown in FIG. 4, was described. However, the content of the scenario, the method of defining it, and the method of calculating the score can be modified in various ways. For example, new scenarios can be added to address fraud risks other than recursive transactions (e.g., asset misappropriation, inflated expense claims, understatement of reserves, and unrecognized impairment losses). It is also effective to prepare scenario sets specialized for specific industries (e.g., finance, construction, retail, etc.), company sizes, and business models, and use them depending on the audit target. In addition to the formulas and algorithms described in the embodiment, scenarios can also be defined using rule-based descriptions (e.g., "If the transaction amount is 1 million yen or more and the profit margin is 50% or more, add one point to the score") or anomaly detection models built using machine learning (e.g., autoencoders and one-class SVMs). The indicator score can be expressed in a variety of formats depending on the purpose of the analysis, such as not only a single real number but also a probability value indicating the probability of fraud, a confidence interval indicating the reliability of the score, or a categorical value indicating the risk level (e.g., "high," "medium," "low").
[0042] (Modifications regarding determination and learning of weighting coefficients) In the above embodiments, several methods for determining weighting coefficients (e.g., a method based on the amount of information, constrained optimization that reflects the auditor's knowledge, etc.) have been described. However, other methods can also be used. For example, if there are cases (trainee data) in which fraud was actually discovered in past audits, such data can be used to perform supervised learning (e.g., logistic regression, support vector machines, neural networks, etc.) to train weighting coefficients that can best distinguish between fraudulent cases and normal cases. It is also effective to adopt reinforcement learning, which sequentially updates weights using auditor feedback (e.g., judgments of "correct" or "incorrect" for presented results), or ensemble learning, which integrates the results of multiple different weight determination methods. Furthermore, rather than setting weighting coefficients to fixed values, it is also possible to introduce a mechanism that dynamically adjusts weighting coefficients based on the statistical characteristics of the data being analyzed, the period of time, or the auditor's settings. For example, if a specific indicator score shows an abnormally high value during a specific period, adaptive weight adjustment can be performed, such as temporarily increasing the weight of that indicator during that period. The granularity of weight application is not limited to using a weight set for each trade flow, but different weight sets can also be applied at finer units such as the type of transaction, account item, or time period, or conversely, different weight sets can be applied at coarser units such as by business division rather than by person in charge.
[0043] (Modifications regarding abnormality probability estimation) The method for estimating the possibility of an abnormality may not only be a simple threshold judgment, but also a method of clustering targets based on the calculated integrated score and detecting clusters with abnormal patterns, or a change-point detection method of monitoring the transition of the integrated score over time and detecting sudden changes or abnormal trends.In addition to a single integrated score, it is also effective to evaluate the scores by category (business transaction flow, business transaction stock, external environment, internal environment) described in the embodiment and individual symptom scores, evaluate the possibility of an abnormality from multiple perspectives, and present a detailed breakdown of the results, in order to help the auditor understand the situation.
[0044] (Modifications regarding system configuration) In the above embodiment, the information processing device 10 is primarily configured as a single computer. However, the system configuration is not limited to this. For example, the functions of the information processing device 10 (such as the acquisition unit 112, the symptom evaluation unit 113, the weight learning unit 114, the symptom integration unit 115, and the estimation unit 116) may be distributed across multiple computers, forming a client-server system or a distributed processing system that operates over a network. Specifically, a client terminal operated by an auditor may be separated from a server that performs the actual data processing and analysis, with heavy computational processing performed on the server side. Furthermore, these functions may be provided in a cloud computing environment, and the present invention may be implemented as a software-as-a-service (SaaS) system that auditors access via a web browser or the like. Adopting such a configuration can improve the processing capacity of large amounts of data, enhance system scalability and maintainability, and enable simultaneous use by multiple users. Furthermore, a dedicated hardware accelerator (such as an FPGA, GPU, or ASIC) may be incorporated into the information processing device to accelerate specific computational processes (such as symptom score calculation for large amounts of transaction data or weight learning involving complex optimization calculations). In short, the information processing device system of the present invention is composed of one or more information processing devices or computers, and includes: a first step of calculating multiple indicator scores that individually indicate the degree of each of multiple indicators of fraudulent transactions based on transaction data that indicates the content of the transaction; A second step of calculating an integrated score indicating the possibility of fraudulent transaction by applying a non-negative weighting coefficient to the plurality of symptom scores and adding them up may be performed.
[0045] (Other variations) The program of the present invention may be provided in various forms, such as stored on a computer-readable recording medium (e.g., magnetic recording medium such as a magnetic tape or a magnetic disk), an optical recording medium (e.g., an optical disk), a magneto-optical recording medium, or a semiconductor memory, or downloaded via a communication line (e.g., the Internet). The program of the present invention is not limited to being executed by a single processor; multiple processors located in physically separate locations may cooperate to execute the program. Furthermore, the order of operations performed by the processors is not limited to the order described in the above-described embodiment and may be changed as appropriate. The application of the present invention is not limited to detecting circular transactions, but can also be applied to a variety of audit and assurance tasks and risk management tasks, such as detecting fraud in business processes in internal audits, supporting evidence analysis in fraud investigations, detecting money laundering in financial institutions, and checking for regulatory violations in compliance departments. [Explanation of symbols]
[0046] 10...information processing device, 11...processor, 111...setting unit, 112...acquisition unit, 113...sign evaluation unit, 114...weight learning unit, 115...sign integration unit, 116...estimation unit, 12...memory, 121...transaction DB, 1211...data ID list, 1212...transaction table, 122...scenario DB, 123...weight file, 1231...data ID list, 1232...weight coefficient list, 13...communication unit, 14...operation unit, 15...display unit.
Claims
1. On the computer, A first step of calculating multiple indicator scores that individually indicate the degree of each of multiple indicators of fraudulent transactions based on transaction data that indicates the content of the transaction; a second step of calculating an integrated score indicating the possibility of fraudulent transactions by applying a non-negative weighting coefficient to the plurality of symptom scores and adding them up; determining the weighting coefficients based on the plurality of symptom scores to minimize the conditional variance of the integrated score under a constraint; A program to execute.
2. In the determining step, the weighting coefficients are determined by adding a condition that all of the weighting coefficients have a minimum positive weight value instead of or in addition to the condition that the weighting coefficients are non-negative. The program according to claim 1.
3. In the determining step, the weighting coefficients are determined taking into consideration an additional condition regarding the magnitude relationship between a first weighting coefficient group and a second weighting coefficient group included in the plurality of weighting coefficients. The program according to claim 1.
4. A computer, A first step of calculating multiple indicator scores that individually indicate the degree of each of multiple indicators of fraudulent transactions based on transaction data that indicates the content of the transaction; a second step of calculating an integrated score indicating the possibility of fraudulent transactions by applying a non-negative weighting coefficient to the plurality of symptom scores and adding them up; A program for executing In the second step, the integrated score is calculated based on the transaction data to be audited for fraudulent transactions using the weighting coefficient determined based on transaction data other than the transaction data to be audited for fraudulent transactions. program.
5. The fraudulent transactions are circular transactions The program according to claim 1.
6. A computer, A first step of calculating multiple indicator scores that individually indicate the degree of each of multiple indicators of fraudulent transactions based on transaction data that indicates the content of the transaction; a second step of calculating an integrated score indicating the possibility of fraudulent transactions by applying a non-negative weighting coefficient to the plurality of symptom scores and adding them up; A program for executing The plurality of signs include one or more signs belonging to each of four categories: internal environment, external environment, business activity (flow), and business activity (stock). program.
7. A computer, receiving designation of one or more elements for identifying a commercial flow from a transaction target, a transaction partner, and a person in charge; A first step of calculating multiple indicator scores that individually indicate the degree of each of multiple indicators of fraudulent transactions based on transaction data that indicates the content of the transaction; a second step of calculating an integrated score indicating the possibility of fraudulent transactions by applying a non-negative weighting coefficient to the plurality of symptom scores and adding them up; A program for executing In the first step, the plurality of symptom scores are calculated for each of the commercial flows; In the second step, the integrated score is calculated for each of the commercial flows. program.
8. a first calculation means for calculating a plurality of indicator scores that individually indicate the degree of each of a plurality of indicators of fraudulent transactions based on transaction data that indicates the content of the transaction; a second calculation means for calculating an integrated score indicating the possibility of fraudulent transactions by applying a non-negative weighting coefficient to the plurality of indicator scores and adding them up; means for determining the weighting coefficients based on the plurality of symptom scores, such that the conditional variance of the integrated score is minimized under a constraint; An information processing system having the above.
Citation Information
Patent Citations
Illegal transaction detection system
JP2016015000A
Information processing device, information processing method and program
JP2023095063A
Program, and information processing device
JP2023183187A
Analysis program, analysis device, and analysis method
JP2024016300A
Examination work support device, examination work support method and examination work support program
JP2025009653A