Anomaly detection system, anomaly detection method, and anomaly detection program
The anomaly detection system uses an approximation curve and threshold calculation to unify data scales and detect anomalies in project progress, addressing the limitations of existing systems by ensuring accurate and flexible anomaly detection.
Patent Information
- Application Number
- JP2023026405
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-02-22
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2043-02-22
AI Technical Summary
Existing systems fail to accurately detect anomalies in project progress accumulation due to limited sample sizes when comparing projects with matching construction periods and actual results, and are unable to compare when periods and results differ.
Anomaly detection system that uses an approximation curve to unify data scales, calculates a threshold based on dispersion and standard deviation, and determines anomalies by comparing project progress rates with the threshold.
Accurately detects anomalies in project progress with high precision, allowing for flexible anomaly detection and prevention of fraudulent reporting, regardless of construction period and budget differences.
Smart Images

Figure 0007794772000001 
Figure 0007794772000002 
Figure 0007794772000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an anomaly detection system, an anomaly detection method, and an anomaly detection program. [Background technology]
[0002] For example, in industries such as construction and systems development, which use project-based management to accumulate progress over time, there is a demand to "determine whether the accumulation method is appropriate."
[0003] In the past, if the construction period and actual results of past projects differed from the construction period and budget of the project for which an abnormality was to be detected, it was not possible to compare how the results accumulated. Also, when an abnormality was detected based on past projects whose construction period and actual results matched, the number of samples was small, resulting in a judgment result with low validity. For example, Patent Document 1 discloses a conventional system for managing budgets and actual results. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-95122 Summary of the Invention [Problem to be solved by the invention]
[0005] However, Patent Document 1 does not disclose anything about using an approximation curve to detect with high accuracy an abnormality in the accumulation of progress of a target project based on performance data of past projects.
[0006] The present invention has been made in consideration of the above, and aims to provide an anomaly detection system, an anomaly detection method, and an anomaly detection program that can use an approximation curve to detect with high accuracy abnormalities in the way the progress of a target project is accumulating based on performance data of past projects. [Means for solving the problem]
[0007] In order to solve the above-mentioned problems and achieve the object, the present invention provides an anomaly detection system that includes a control unit and detects anomalies in business data, wherein the control unit is configured to be able to access an anomaly detection definition master in which, for each of the anomaly detection definitions, an X-axis item, a Y-axis item, an X-axis scale uniform item, and a Y-axis scale uniform item are set as parameters; monthly performance data of past cases including case identification information, the X-axis item, and the Y-axis item; schedule and performance data of past cases including the case identification information, the X-axis item, and the Y-axis scale uniform item; monthly performance data of cases that are subject to anomaly detection including the case identification information, the X-axis item, and the Y-axis item; and schedule and performance data of cases that are subject to anomaly detection including the case identification information, the schedule period or target end period, and a budget or target; and for the monthly performance data of past cases, the control unit calculates, for each of the case identification information, a progress rate obtained by dividing the X-axis item by the X-axis scale uniform item of the schedule and performance data of the past cases, and a progress rate obtained by dividing the Y-axis item by the X-axis scale uniform item of the schedule and performance data of the past cases The data processing device is characterized by comprising: a data processing means for calculating and processing a progress rate by dividing the schedule and performance data of the past projects by a Y-axis scale unified item; and for calculating and processing, for each project identification information, a progress rate by dividing the X-axis item by the schedule period or target end period of the schedule and performance data of the project to be determined for the abnormality, and a progress rate by dividing the Y-axis item by the budget or target of the schedule and performance data of the project to be determined for the abnormality; an approximation curve calculation means for calculating a performance center approximation curve for all of the processed monthly performance data of the past projects; a threshold calculation means for calculating the dispersion of all of the data from the calculated performance center approximation curve, calculating a standard deviation curve expressing the tendency of the calculated dispersion, and calculating a threshold based on the performance center approximation curve and the standard deviation curve; and an abnormality determination means for comparing the processed monthly performance data of the project to be determined for the abnormality with the calculated threshold to determine an abnormality.
[0008] According to one aspect of the present invention, the X-axis item may include the number of months elapsed, the Y-axis item may include the cumulative actual amount, the X-axis scale unified item may include the planned construction period or implementation period, and the Y-axis scale unified item may include the budget or actual amount.
[0009] According to another aspect of the present invention, the anomaly determination definition master further includes one or more approximate function candidates set as parameters for each anomaly determination definition, and the approximate curve calculation means passes the elapsed time rate and progress rate of the processed monthly performance data of past cases and information on the one or more approximate function candidates set in the anomaly determination definition master to an approximation component to find optimal coefficients, calculates AIC based on the coefficients returned from the approximation component, and determines the approximate function candidate with the smallest AIC as the performance-centered approximate function.
[0010] According to another aspect of the present invention, the one or more approximation function candidates may include at least one of a polynomial function, an exponential function, a logarithmic function, a logistic function, and a Bass model function.
[0011] According to another aspect of the present invention, the threshold calculation means may substitute each progress rate of the processed monthly actual data of the past case into the performance center approximation function to calculate a predicted value of the progress rate, calculate a residual which is the difference between the predicted value and the actual data of the progress rate, calculate the square of the residual, draw an approximation curve of the residual square approximation function for the square of the residual, and take the square root of the residual square approximation function as a standard deviation function, substitute each progress rate of the processed monthly actual data of the case to be determined as abnormal into the performance center approximation function to calculate a predicted value of the progress rate, and substitute into the standard deviation function to calculate the standard deviation, and calculate upper and lower thresholds which have a range above and below the predicted value equal to the value obtained by multiplying the standard deviation by a constant value calculated based on the significance level.
[0012] According to another aspect of the present invention, the abnormality determination means may output an abnormality detection message and a graph thereof to a display unit when it determines that an abnormality has occurred.
[0013] According to another aspect of the present invention, a master maintenance means may be provided for setting data of the abnormality determination definition master in response to an operation by an operator.
[0014] Furthermore, in order to solve the above-mentioned problems and achieve the object, the present invention provides an anomaly detection method executed by an information processing device having a control unit, wherein the control unit is configured to be able to access an anomaly detection definition master in which an X-axis item, a Y-axis item, an X-axis scale uniform item, and a Y-axis scale uniform item are set as parameters for each of the anomaly detection definitions, monthly performance data of past cases including case identification information, X-axis item, and Y-axis item, schedule and performance data of past cases including case identification information, X-axis scale uniform item, and Y-axis scale uniform item, monthly performance data of cases that are subject to anomaly detection including case identification information, X-axis item, and Y-axis item, and schedule and performance data of cases that are subject to anomaly detection including case identification information, a schedule period or target completion period, and a budget or target, and the control unit executes an anomaly detection method to calculate, for each of the case identification information, a process time for calculating a process time for the monthly performance data of past cases by dividing the X-axis item by the X-axis scale uniform item of the schedule and performance data of the past cases. a data processing step of calculating and processing an elapsed time rate and a progress rate obtained by dividing a Y-axis item by a Y-axis scale unified item of the planned and actual data of the past projects, and calculating and processing, for each project identification information, an elapsed time rate obtained by dividing an X-axis item by the planned period or target end period of the planned and actual data of the project to be determined for the abnormality, and a progress rate obtained by dividing the Y-axis item by the budget or target of the planned and actual data of the project to be determined for the abnormality; an approximation curve calculation step of calculating a performance center approximation curve for all of the processed monthly actual data of the past projects; a threshold calculation step of calculating the dispersion of all of the data from the calculated approximation curve, calculating a standard deviation curve that expresses the tendency of the calculated dispersion, and calculating a threshold based on the performance center approximation curve and the standard deviation curve; and an abnormality determination step of comparing the processed monthly actual data of the project to be determined for the abnormality with the calculated threshold to determine an abnormality.
[0015] and a target period or end period, a budget, or a goal. The control unit is configured to access the following: an anomaly detection program to be executed by an information processing device having a control unit; an anomaly detection definition master in which, for each of the anomaly detection definitions, an X-axis item, a Y-axis item, an X-axis scale uniform item, and a Y-axis scale uniform item are set as parameters; monthly performance data of past cases including case identification information, the X-axis item, and the Y-axis item; schedule and performance data of past cases including the case identification information, the X-axis item, and the Y-axis scale uniform item; monthly performance data of cases that are targets for anomaly detection including the case identification information, the X-axis item, and the Y-axis item; and schedule and performance data of cases that are targets for anomaly detection including the case identification information, the schedule period or end period, and a budget or a goal. The control unit is configured to access the following: an anomaly detection definition master in which, for each of the anomaly detection definitions, an X-axis item, a Y-axis item, an X-axis scale uniform item, and a Y-axis scale uniform item are set as parameters; monthly performance data of past cases including case identification information, the X-axis item, and the Y-axis item; a data processing step of calculating and processing the progress rate by dividing the X-axis item by the planned period or target end period of the schedule / performance data of the project to be determined for each project identification information, and the progress rate by dividing the Y-axis item by the budget or target of the schedule / performance data of the project to be determined for each project identification information; an approximation curve calculation step of calculating a performance center approximation curve for all of the processed monthly performance data of the past projects; a threshold calculation step of calculating the dispersion of all of the data from the calculated performance center approximation curve, calculating a standard deviation curve that expresses the tendency of the calculated dispersion, and calculating a threshold based on the performance center approximation curve and the standard deviation curve; and an abnormality determination step of comparing the processed monthly performance data of the project to be determined for each project identification information to determine an abnormality. [Effects of the Invention]
[0016] According to the present invention, it is possible to use an approximation curve to detect with high accuracy an abnormality in the accumulation of progress of a target case based on performance data of past cases. [Brief explanation of the drawings]
[0017] [Figure 1] FIG. 1 is a diagram for explaining the problem to be solved by the present invention. [Figure 2] FIG. 2 is a diagram for explaining the problem to be solved by the present invention. [Figure 3] FIG. 3 is a block diagram showing an example of the configuration of the anomaly detection system according to this embodiment. [Figure 4] FIG. 4 is a diagram showing a processing flow illustrating an example of the overall processing of the control unit of the anomaly detection system according to this embodiment. [Figure 5] FIG. 5 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 6] FIG. 6 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 7] FIG. 7 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 8] FIG. 8 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 9] FIG. 9 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 10] FIG. 10 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 11] FIG. 11 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 12] FIG. 12 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 13A] FIG. 13A is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 13B] FIG. 13B is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 14]FIG. 14 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 15] FIG. 15 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 16] FIG. 16 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 17] FIG. 17 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 18A] FIG. 18A is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 18B] FIG. 18B is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 19] FIG. 19 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 20] FIG. 20 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 21] FIG. 21 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 22] FIG. 22 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 23] FIG. 23 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 24] FIG. 24 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 25] FIG. 25 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 26] FIG. 26 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 27]FIG. 27 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 28] FIG. 28 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 29] FIG. 29 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 30] FIG. 30 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 31] FIG. 31 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 32] FIG. 32 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 33] FIG. 33 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 34] FIG. 34 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 35] FIG. 35 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 36] FIG. 36 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 37] FIG. 37 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 38] FIG. 38 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 39] FIG. 39 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 40] FIG. 40 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 41]FIG. 41 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 42] FIG. 42 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 43] FIG. 43 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 44] FIG. 44 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 45] FIG. 45 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. [Figure 46] FIG. 46 is a diagram for explaining a specific example of the processing of the control unit of the anomaly detection system according to this embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0018] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS An embodiment of the present invention will be described in detail with reference to the accompanying drawings. However, the present invention is not limited to this embodiment.
[0019] [1. Overview] The present invention will be outlined in the following order: background, problems, solution examples, and features of the present invention.
[0020] (1-1. Background) The background of the present invention will be explained. In industries such as the construction industry and the system development industry, which use project-based management that accumulates progress over a period of time, there is a demand to "determine whether the accumulation method is appropriate by comparing with past projects."
[0021] The anomaly detection system of the present invention processes data while taking into account the possibility of exceeding the planned construction period and budget, and by drawing an approximation curve of the progress transition, it is possible to evaluate how the progress accumulates and detect anomalies with high accuracy. In addition, it is possible to set the judgment process according to the characteristics of each task in the master.
[0022] The anomaly detection system of the present invention (1) can flexibly determine anomalies in progress delays and revise plans even when the planned construction period or budget has been exceeded, and (2) can detect fraudulent recording of progress rates and more accurately prevent fraudulent recording of actual results that do not correspond to the actual situation than conventional anomaly detection systems.
[0023] In this specification, "number of months elapsed" refers to the number of months elapsed since the start of a project. "Cumulative actual amount" refers to the actual costs incurred from the start of a project to the number of months elapsed. "Planned construction period" refers to the number of months expected to be required until the project is completed, which is set before the project begins. "Implementation period" refers to the number of months expected to be required until the project is completed, which is determined after the project begins. "Budget" refers to the expected costs expected until the project is completed, which is set before the project begins. "Actual amount" refers to the total amount spent on a project, which is determined after the project is completed. "Percentage of planned construction period elapsed" refers to the number of months elapsed divided by the planned construction period. "Percentage of implementation period elapsed" refers to the number of months elapsed divided by the implementation period. "Percentage of budget elapsed" refers to the cumulative actual amount divided by the budget. "Actual progress rate" refers to the cumulative actual amount divided by the actual amount. "Projects subject to abnormality detection" refers to ongoing projects for which abnormalities in the accumulation of progress are to be detected. "Approximate curve" refers to a curved line when viewed on a graph. An "approximation function" is a formula that represents an approximate curve using letters such as x and y. An "approximation function name" is a name that distinguishes an approximate function by terminology rather than by formula.
[0024] (1-2. Issues) The problems to be solved by the present invention will be described with reference to Figures 1 and 2. Conventionally, there are the following problems.
[0025] 1. When comparing only past projects with matching construction periods and actual results, there is an issue of the small number of samples to compare, resulting in less valid anomaly determination results.
[0026] 2. There is an issue that if the construction periods and actual results of past projects differ, it is not possible to compare the data.
[0027] More specifically, it is as follows.
[0028] (1) When comparing only past projects whose construction periods and actual results match, the number of samples to be compared is small, resulting in low validity of abnormality detection results. Specifically, when collecting data on past projects, if only those whose construction periods and budgets match the project being detected as abnormalities are extracted, the number of samples from past projects will be small. In this way, a small number of samples will result in low validity of abnormality detection results.
[0029] (2) The issue of not being able to compare data when the construction period and actual results of past projects differ will be explained in detail with reference to Figures 1 and 2. This explains the case where you want to determine anomalies in project A0001 based on how the results of project B0001, which has already been completed in the past, have accumulated. As a premise, it is assumed that the actual data and budget data of the project to be determined to be anomaly and the actual data of the past project to be compared are stored as the following data.
[0030] Figure 1(A) shows an example of performance data for a project subject to abnormality detection, including the project code, number of months elapsed, and performance data. Figure 1(B) shows an example of budget data for a project subject to abnormality detection, including the project code, construction period, and budget data. Figure 1(C) shows an example of performance data for a past project, including the project code, number of months elapsed, and performance data.
[0031] Figure 2 shows an image of a graph when comparing the accumulation of a project that has been determined to be abnormal with the accumulation of a comparable project from the past. Figure 2 shows the number of months that have passed and the results for the similar project from the past (B0001) in Figure 1 and the project that has been determined to be abnormal (A0001), with the X axis showing the number of months that have passed (months) and the Y axis showing the results (10,000 yen). Because the construction period and budget for the project that has been determined to be abnormal are different from the construction period and results for the comparable project from the past, it is impossible to determine whether the accumulation of the results for the project that has been determined to be abnormal is abnormal or not.
[0032] (1-3.Solution) Therefore, as a solution, the present invention processes data to unify the scale so that, for example, the scale of the construction period and actual results of a group of past projects matches the scale of the construction period and budget of the project being judged for anomalies.
[0033] 1. When comparing only past projects with matching construction periods and results, the number of samples to compare is small, resulting in anomaly detection results with low validity. However, since it is possible to compare with the project being detected as anomaly regardless of the construction periods and results of past projects, it becomes possible to collect data regardless of construction periods and results.
[0034] 2. In response to the issue of being unable to compare data when the construction periods and actual results of past projects differ, by matching the construction periods and actual results of past projects to the construction period and budget of the project being judged as abnormal, it becomes possible to compare past projects with the project being judged as abnormal.
[0035] (1-4. Features of the present invention) The anomaly detection system of the present invention draws an approximation curve for processed monthly data of past cases, calculates a threshold value, and determines an anomaly based on the threshold value, and has the following features 1 to 3.
[0036] (1) Feature 1 of the present invention It is possible to select what to prioritize when determining anomalies, such as delays in construction schedules or budget overruns. The anomaly detection system of this invention can accommodate the differences in policy regarding how much importance is placed on deviations from the planned construction schedule and budget at the time of project completion for each company (details are explained in "3-4. Options for processing monthly performance data for past projects").
[0037] (2) Feature 2 of the present invention When predicting the progress of a project, multiple approximation functions can be selected. When predicting the progress of a project, polynomial functions, exponential functions, logarithmic functions, logistic functions, and Bass model functions can be specified as approximation function candidates in the anomaly detection definition master. An approximation function can be selected according to the progress of a project that the user of the anomaly detection system knows from experience.
[0038] (3) Feature 3 of the present invention The progress of a project can be predicted with high accuracy. The approximate curve algorithm of the anomaly detection system of the present invention calculates the reference value and threshold value using actual data, so the accuracy is higher than that of conventional methods.
[0039] The anomaly detection system of the present invention can be applied to all business types and industries, and can be suitably applied to projects that involve project-based management, such as the construction industry and the IT media industry, for example.
[0040] [2. Configuration] An example of the configuration of the anomaly detection system 100 according to this embodiment will be described with reference to Fig. 3. Fig. 3 is a block diagram showing an example of the configuration of the anomaly detection system 100.
[0041] The anomaly detection system 100 is a commercially available desktop personal computer. Note that the anomaly detection system 100 is not limited to a stationary information processing device such as a desktop personal computer, and may be a portable information processing device such as a commercially available notebook personal computer, PDA (Personal Digital Assistant), smartphone, or tablet personal computer.
[0042] The anomaly detection system 100 includes a control unit 102, a communication interface unit 104, a storage unit 106, and an input / output interface unit 108. The units included in the anomaly detection system 100 are connected to each other so as to be able to communicate with each other via any communication path.
[0043] The communication interface unit 104 communicatively connects the anomaly detection system 100 to a network 300 via a communication device such as a router and a wired or wireless communication line such as a dedicated line. The communication interface unit 104 has a function of communicating data with other devices via the communication line. Here, the network 300 has a function of communicatively connecting the anomaly detection system 100 with the server 200 and the business system 400, and is, for example, the Internet or a LAN (Local Area Network). The business system 400 creates and updates various types of business data.
[0044] An input device 112 and an output device 114 are connected to the input / output interface unit 108. The output device 114 may be a monitor (including a home television), a speaker, or a printer. The input device 112 may be a keyboard, a mouse, a microphone, or a monitor that cooperates with a mouse to achieve a pointing device function. Note that, hereinafter, the output device 114 may be referred to as the monitor 114, and the input device 112 may be referred to as the keyboard 112 or the mouse 112. Displaying information on the monitor 114 and the user operating the input device 112 may be referred to as a "user operation via a UI."
[0045] The storage unit 106 stores various databases, tables, files, etc. The storage unit 106 stores computer programs that cooperate with the OS (Operating System) to issue commands to the CPU (Central Processing Unit) to perform various processes. The storage unit 106 may be, for example, a memory device such as a RAM (Random Access Memory) or a ROM (Read Only Memory), a fixed disk device such as a hard disk, a flexible disk, or an optical disk. The storage unit 106 includes an abnormality determination definition master 106a, a determination result table 106b, an approximate curve result attribute information table 106c, monthly actual data and schedule / actual data for past cases, monthly actual data and schedule / actual data for cases subject to abnormality determination, etc.
[0046] The abnormality determination definition master 106a can be configured as a table or the like that associates and registers abnormality determination definition identification information (abnormality determination definition ID, abnormality determination definition name), acquisition definition, algorithm used (for example, approximate curve), and parameter settings {X-axis item (for example, number of months elapsed), Y-axis item (for example, cumulative actual amount), X-axis scale unified item (for example, planned construction period), Y-axis scale unified item (for example, actual amount), approximate function candidate name (for example, polynomial function, logistic function, Bass model function), smoothness, significance level)}. The operator can set the data in the abnormality determination definition master 106a.
[0047] The judgment result table 106b is a table for storing judgment result data that is the result of an abnormality judgment. The judgment result data may include the case code, accounting year and month, planned completion period progress rate, actual progress rate, forecast value, standard deviation, upper threshold, lower threshold, and judgment result (FALSE or TRUE). For data that is judged to be abnormal, the judgment result is set to "TRUE", and for other data, it is set to "FALSE".
[0048] The approximate curve result attached information table 106c is a table for storing approximate curve result attached information for displaying an anomaly determination graph. The approximate curve result attached information may include an execution history ID, a row number, X, Y, an upper threshold, and a lower threshold. The "execution history ID" is an ID automatically assigned for each execution of an anomaly determination, the "row number" is a sequential number assigned to each row of display data starting from 0, "X" corresponds to the X-axis value of the display data, "Y" corresponds to the Y-axis value of the display data, the "upper threshold" is the upper threshold of the display data, and the "lower threshold" is the lower threshold of the display data.
[0049] The monthly performance data of past cases may include a case code (case identification information), accounting year and month, an X-axis item (for example, the number of months elapsed), and a Y-axis item (for example, the cumulative performance amount).
[0050] The planned and actual data of past projects may include a project code (project identification information), an X-axis scale standard item (for example, planned construction period), a budget, an implementation period, and a Y-axis scale standard item (for example, actual amount).
[0051] The monthly performance data of the case subject to abnormality determination may include the case code (case identification information), accounting year, X-axis item (for example, the number of months elapsed), and Y-axis item (for example, the cumulative performance amount).
[0052] The planned and actual data for the case to be judged as abnormal may include the case code (case identification information), planned period or target completion period, budget or target, implementation period, and actual amount. The target may include the estimated amount, target retention rate, etc.
[0053] The control unit 102 is a CPU or the like that performs overall control of the anomaly detection system 100. The control unit 102 has an internal memory for storing control programs such as an OS, programs that define various processing procedures, required data, etc., and executes various information processing operations based on these stored programs.
[0054] The control unit 102 receives the abnormality determination definition master 106a, the determination result table 106b, the approximate curve result attached information table 106c, the monthly actual data and the schedule / actual data of past cases, and the monthly actual data and the schedule / actual data of cases to be subjected to abnormality determination, all of which are stored in the storage unit 106. The abnormality determination definition master 106a, the determination result table 106b, the approximate curve result attached information table 106c, the monthly actual data and schedule / actual data of past cases, and the monthly actual data and schedule / actual data of cases subject to abnormality determination may be provided in another location (for example, the server 200) as long as they are accessible by the control unit 102.
[0055] The control unit 102, in terms of functional concept, includes a data acquisition unit 102a, a data processing unit 102b, an approximate curve calculation unit 102c, a threshold calculation unit 102d, an abnormality determination unit 102e, a screen display control unit 102f, and a master maintenance unit 102g.
[0056] The data acquisition unit 102a acquires monthly performance data and schedule / performance data of past cases and monthly performance data and schedule / performance data of cases subject to abnormality determination from the business system 300, and stores the data in the storage unit 106. For example, the data acquisition unit 102a may acquire, from the business system 300, monthly performance data of past cases and cases subject to abnormality determination, including parameter-set X-axis items (e.g., number of months elapsed) and Y-axis items (e.g., cumulative performance amount), and schedule / performance data of past cases and cases subject to abnormality determination, including X-axis scale uniform items (e.g., planned construction period) and Y-axis scale uniform items (e.g., performance amount), in accordance with the acquisition definition of the abnormality determination definition master 106a.
[0057] The data processing unit 102b processes data to unify the scales of monthly performance data of past projects and projects subject to abnormality detection. The data processing unit 102b may process the monthly performance data of past projects by calculating, for each project identification information, a progress rate by dividing the X-axis item by the X-axis scale unifying item of the schedule / performance data of the past projects and a progress rate by dividing the Y-axis item by the Y-axis scale unifying item of the data of the past projects. Furthermore, the monthly performance data of projects subject to abnormality detection may be processed by calculating, for each project identification information, a progress rate by dividing the X-axis item by the planned construction period of the schedule / performance data of the project subject to abnormality detection and a progress rate by dividing the Y-axis item by the budget of the schedule / performance data of the project subject to abnormality detection. The X-axis item may include the number of elapsed months, the Y-axis item may include the cumulative performance amount, the X-axis scale unifying item may include the planned construction period or implementation period, and the Y-axis scale unifying item may include the budget or performance amount.
[0058] The approximate curve calculation unit 102c calculates an actual performance approximate curve for all data of the processed monthly performance data of past cases. The approximate curve calculation unit 102c may pass the elapsed time rate and progress rate of the processed monthly performance data of past cases and information on one or more approximate function candidates set in the anomaly determination definition master 106a to an approximate component to find optimal coefficients, calculate the AIC based on the coefficients returned from the approximate component, and determine the approximate function with the smallest AIC as the actual performance center approximate function. The one or more approximate function candidates may include at least one of a polynomial function, an exponential function, a logarithmic function, a logistic function, and a Bass model function.
[0059] The threshold calculation unit 102d calculates the dispersion of all data of the processed monthly actual data of the past cases from the calculated approximation curve, calculates a standard deviation curve that expresses the tendency of the calculated dispersion, and calculates a threshold based on the actual data center approximation curve and the standard deviation curve. In this case, the threshold calculation unit 102d may substitute each progress rate of the processed monthly actual data of the past cases into the actual data center approximation function to calculate a predicted value of the progress rate, calculate a residual that is the difference between the predicted value and the actual data of the progress rate, calculate the square of the residual, draw an approximation curve of the residual square approximation function for the square of the residual, and take the square root of the residual square approximation function as the standard deviation function. Substitute each progress rate of the processed monthly actual data of the case to be determined as abnormal into the actual data center approximation function to calculate a predicted value of the progress rate, and also substitute into the standard deviation function to calculate the standard deviation, and calculate upper and lower thresholds that have a range above and below the predicted value by the value obtained by multiplying the standard deviation by a constant value calculated based on the significance level.
[0060] The abnormality determination unit 102e compares the progress rate and progress rate of each case identification information of the processed monthly performance data of the case to be determined as abnormal with the calculated threshold value to determine an abnormality. If the abnormality determination unit 102e determines that an abnormality has occurred, it may display an abnormality detection message and an abnormality determination graph on an abnormality determination output screen.
[0061] The screen display control unit 102f controls the display of various screens (for example, an abnormality determination output screen, a master maintenance screen, etc.) displayed on the monitor 114 and the input therefor.
[0062] The master maintenance unit 102g performs settings such as inputting, adding, deleting, and editing data for the abnormality determination definition master 106a in response to an operator's operation on a master maintenance screen displayed on the monitor 114, for example.
[0063] The screen display control unit 102f controls the display of various screens (for example, a store supply alert screen, a master maintenance screen, etc.) displayed on the monitor 114 and the inputs thereto.
[0064] [3. Specific Examples] Specific examples of processing by the control unit 102 of the anomaly detection system 100 according to this embodiment will be described with reference to Fig. 4 to Fig. 46. Fig. 4 is a diagram showing a processing flow for outlining the overall processing by the control unit 102 of the anomaly detection system 100 according to this embodiment. Figs. 5 to 46 are diagrams for explaining specific examples of processing by the control unit 102 of the anomaly detection system 100 according to this embodiment.
[0065] [3-1. Processing flow overview] An overview of the processing flow will be described with reference to Fig. 4. In Fig. 4, the data processing unit 102b executes processing of monthly performance data of past projects (step S1), and processes the monthly performance data of past projects by calculating, for each project identification information, a progress rate obtained by dividing the X-axis item by the X-axis scale uniform item of the planned / performance data of the past project, and a progress rate obtained by dividing the Y-axis item by the Y-axis scale uniform item of the past project data. The X-axis item may include the number of elapsed months, the Y-axis item may include the accumulated performance amount, the X-axis scale uniform item may include the planned construction period or implementation period, and the Y-axis scale uniform item may include the budget or performance amount.
[0066] The data processing unit 102b executes processing of the monthly performance data of the abnormality target project (step S2), and processes the monthly performance data of the abnormality target project by calculating, for each project identification information, the progress rate obtained by dividing the X-axis item by the planned construction period of the planned / actual data of the abnormality target project, and the progress rate obtained by dividing the Y-axis item by the budget of the planned / actual data of the abnormality target project.
[0067] The approximation curve calculation unit 102c executes a performance-centered approximation curve calculation process (step S3) to calculate a performance-centered approximation curve for all data of the processed monthly performance data of past projects. In this case, the approximation curve calculation unit 102c may pass the elapsed time rate and progress rate of the processed monthly performance data of past projects and information on one or more approximation function candidates set in the anomaly determination definition master 106a to the approximation component to find optimal coefficients, calculate the AIC based on the coefficients returned from the approximation component, and determine the approximation function candidate with the smallest AIC as the performance-centered approximation function. The one or more approximation function candidates may include at least one of a polynomial function, an exponential function, a logarithmic function, a logistic function, and a Bass model function.
[0068] The threshold calculation unit 102d executes a process of calculating the dispersion of data from the performance center approximation curve (step S4), and calculates the dispersion of all data of the processed monthly performance data of past cases from the calculated performance center approximation curve. In this case, the threshold calculation unit 102d may substitute each progress rate of the processed monthly performance data of past cases into the performance center approximation function to calculate a predicted value of the progress rate, calculate a residual which is the difference between the predicted value and the actual data of the progress rate, and further calculate the square of the residual.
[0069] The threshold calculation unit 102d executes a process of calculating the tendency of data dispersion (step S5), and calculates a standard deviation curve that expresses the calculated tendency of dispersion. In this case, the threshold calculation unit 102d may draw an approximation curve of the residual square approximation function for the square of the residual, and calculate the standard deviation function by taking the square root of the residual square approximation function.
[0070] The threshold calculation unit 102d executes a process for calculating the predicted value and threshold for the case to be determined as an abnormality (step S6), and calculates the threshold based on the actual performance center approximation curve and the standard deviation curve. In this case, the threshold calculation unit 102d may substitute each progress rate of the processed monthly actual performance data for the case to be determined as an abnormality into the actual performance center approximation function to calculate the predicted value of the progress rate, and may also substitute it into the standard deviation function to calculate the standard deviation, and calculate upper and lower thresholds with a range above and below the predicted value corresponding to the value obtained by multiplying the standard deviation by a constant value calculated based on the significance level.
[0071] The abnormality determination unit 102e executes an abnormality determination process (step S7), and compares the processed monthly performance data of the case to be determined as abnormal with the calculated threshold value to determine an abnormality. In this case, the abnormality determination unit 102e may output and display an abnormality detection message on an abnormality determination output screen when determining that an abnormality has occurred.
[0072] The abnormality determination unit 102e executes calculation processing of the display data (step S8), creates approximate curve result attached information for displaying a graph of the abnormality determination result, stores it in the approximate curve result attached information table 106c, and displays the graph of the abnormality determination result on the abnormality determination output screen based on the approximate curve result attached information.
[0073] [3-2. Processing flow details] The above processing flow will be described in detail with reference to FIGS.
[0074] (Premise 1) The abnormality determination definition master 106a is assumed to be registered in advance in the storage unit 106. The abnormality determination definition master 106a is a master for managing definitions for performing abnormality determination on business data.
[0075] 5 is a diagram showing an example of the configuration of the abnormality determination definition master 106a. The abnormality determination definition master 106a registers abnormality determination definition IDs, abnormality determination definition names, data acquisition definitions, algorithms used, and parameter settings in association with each other.
[0076] In the example shown in the figure, the anomaly determination definition ID is "JD001," the anomaly determination definition name is "Project progress alert," the algorithm used is "Approximation curve," and the parameter settings are {X-axis item: number of months elapsed, Y-axis item: cumulative actual amount, X-axis scale unified item: planned construction period, Y-axis scale unified item: actual amount, approximate function candidate name: [polynomial function, logistic function, Bass model function], smoothness: 12, significance level: 0.1}.
[0077] (Premise 2) It is assumed that monthly performance data and schedule / performance data of past similar cases are stored in the storage unit 106. The data acquisition unit 102a acquires performance data and schedule / performance data of past cases from the business system 400 and stores them in the storage unit 106.
[0078] Figure 6(A) shows an example of monthly performance data for past cases. Monthly performance data for past cases includes the following data fields: case code, accounting year and month, number of months elapsed [months], and cumulative performance amount [yen in millions]. The "accounting year and month" column is not used in the calculations of the processing flow. In this example, the case codes for past cases are "A001," "A002," and "A003." For simplicity, this explanation uses three cases as an example, but in reality, it is assumed that there will be many past cases to ensure the validity of the anomaly detection results.
[0079] Figure 6(B) is a graph of the monthly performance data of the past projects in Figure 6(A), with the horizontal axis showing the number of months elapsed (months) and the vertical axis showing the cumulative performance amount (10,000 yen).
[0080] Figure 6(C) shows an example of the planned and actual data for a past project. The planned and actual data for a past project includes the following fields: project code, planned construction period (months), budget (million yen), implementation period (months), and actual amount (million yen).
[0081] (Premise 3) It is assumed that monthly performance data of a case that is the target of abnormality determination and schedule / performance data before the start of the case are stored in the storage unit 106. The data acquisition unit 102a acquires the monthly performance data of a case that is the target of abnormality determination and schedule / performance data before the start of the case from the business system 400 and stores them in the storage unit 106.
[0082] Figure 7(A) shows an example of monthly performance data for cases subject to abnormality detection. The monthly performance data for cases subject to abnormality detection has the following data fields: case code, accounting year and month, number of months elapsed [months], and cumulative performance amount [yen in millions]. The "accounting year and month" column is not used in the calculations of the processing flow, but is used when displaying a message of the abnormality detection result. In this example, the case codes for the cases subject to abnormality detection are "B001" and "B002". For simplicity, this explanation will use two cases as an example, but it is assumed that in reality there will be many cases subject to abnormality detection.
[0083] Figure 7(B) is a graph of the monthly performance data for the cases in Figure 7(A) that are subject to abnormality detection, with the horizontal axis representing the number of months elapsed (months) and the vertical axis representing the cumulative performance amount (10,000 yen).
[0084] Figure 7(C) shows an example of planned and actual data for a project that is subject to abnormality detection. The planned and actual data for a project that is subject to abnormality detection includes the following fields: project code, planned construction period [months], budget [yen], implementation period [months], and actual amount [yen]. Because the project that is subject to abnormality detection is an ongoing project, there is no data in the implementation period and actual amount columns.
[0085] (Premise 4) The control unit 102 is equipped with a calculation component (hereinafter referred to as an "approximation component") that determines the optimal coefficients of a function. The approximation component returns the optimal coefficients by passing the following information to it. The approximation component will be described below with reference to FIG. 8.
[0086] Figure 8(A) shows an example of information (1) to be passed to the approximating part, which is the data for which you want to draw an approximate curve. The data for which you want to draw an approximate curve is a pair of data such as (X, Y), which corresponds to the value of the horizontal axis X and the vertical axis Y when viewed on a graph.
[0087] Figure 8(B) shows an example of information (2) to be passed to the approximating part, which passes the function to be approximated. In this example, y = ax 2 +bx+c. The letters other than x and y in the function you want to approximate are the coefficients.
[0088] In the processing within the approximation part, for example, the function to be approximated is y=ax 2 In the case of +bx+c, in the above example, the optimal coefficients are calculated as a = 0.5, b = -3.5, c = 7.6 using the approximation components. The function to be approximated is not limited to this example, and can also be y = a or y = ax + b.
[0089] Figure 8(C) shows an example of information returned by the approximating component, which returns the optimal coefficients of the function to be approximated obtained through processing within the approximating component. In this example, the information returned by the approximating component is the coefficients a = 0.5, b = -3.5, and c = 7.6.
[0090] (S1: Processing of monthly performance data for past projects) 9 and 10, the processing of monthly performance data of past cases will be described. In the processing of monthly performance data of past cases, the number of elapsed months and the cumulative performance amount of the monthly performance data of past cases are converted into values in which the X-axis scale unification item and the Y-axis scale unification item are set to 100%. Specifically, the following calculation process is performed based on the data registered in the parameter setting column of the abnormality judgment definition master 106a.
[0091] (1) Divide the value of the X-axis item column of monthly performance data by the value of the X-axis scale unification item. (2) Divide the value of the Y-axis item column of monthly performance data by the value of the Y-axis scale unification item.
[0092] The following explanation will be given using the case code "A003" of the monthly performance data of past cases as an example. The values by which to divide the X-axis item and the Y-axis item of the monthly performance data of past cases can be set in the X-axis scale unification item and the Y-axis scale unification item of the anomaly judgment definition master 106a. Differences in the processed data due to differences in the setting of the scale unification item will be explained later in [3-3. Options for processing monthly performance data of past cases].
[0093] Fig. 9(A) is a diagram showing an example of the settings of the abnormality determination definition master 106a. Fig. 9(B) is a diagram showing an example of monthly performance data for A003. Fig. 9(C) is a diagram showing an example of schedule and performance data for A003.
[0094] FIG. 9(D) shows an example of the processed monthly actual data for A003, and in accordance with the setting example of the abnormality determination definition master 106a, the planned construction period elapsed rate [%] is calculated by dividing the X-axis item "number of months elapsed" by the X-axis scale unified item "planned construction period", and the actual progress rate [%] is calculated by dividing the Y-axis item "cumulative actual amount" by the Y-axis scale unified item "actual amount".
[0095] This data processing is "Pattern B: Flexible scheduled construction period / strict budget" in [3-3. Options for processing monthly performance data of past projects].
[0096] The characteristics of the processed data are as follows. The planned completion rate is calculated by dividing the number of months elapsed by the construction period before the project began, so if the project 1. finishes earlier than planned, the planned completion rate in the final elapsed month will be less than 100%, and 2. if it takes longer than planned, the planned completion rate in the final elapsed month may exceed 100%. Also, since the actual progress rate is calculated by dividing the cumulative actual amount by the actual amount, the actual progress rate in the final elapsed month will always be 100%.
[0097] After processing the data for all past cases "A001" to "A003" as described above for "A003," the data will look like the data shown in Figure 10. Figure 10(A) shows an example of the processed monthly performance data for A001 to A003. Figure 10(B) shows an example of the data in Figure 10(A) graphed (Feature 1 of the present invention described above).
[0098] (S2: Processing of monthly performance data for cases subject to abnormality detection) The details of the processing of monthly performance data for cases subject to abnormality detection will be explained with reference to Figures 11 and 12. In the processing of monthly performance data for cases subject to abnormality detection, the number of months elapsed and the cumulative performance amount of the monthly performance data for cases subject to abnormality detection are divided by the planned construction period and budget, respectively. Specifically, the following processing is performed.
[0099] (1) Divide the number of months elapsed by the planned construction period. (2) Divide the cumulative actual amount by the budget.
[0100] The following explanation will be given using the case code "B001" of the monthly performance data of the case to be judged as an anomaly as an example. Fig. 11(A) is a diagram showing an example of monthly performance data for B001. Fig. 11(B) is a diagram showing an example of the schedule and performance data for B001.
[0101] FIG. 11(C) shows an example of the monthly actual data for B001 after processing, and in accordance with the setting example of the abnormality judgment definition master 106a, the planned construction period elapsed rate [%] is calculated by dividing the X-axis item "number of months elapsed" by the X-axis scale unified item "planned construction period", and the budget progress rate [%] is calculated by dividing the Y-axis item "cumulative actual amount" by "budget" (in past projects, division was done by "actual amount", but here it is divided by "budget").
[0102] After processing the data for all cases "B001" to "B002" that are subject to abnormality detection in the same way as "B001" above, the data will be in the state shown in Figure 12. Figure 12(A) shows an example of the monthly performance data for B001 to B002 after processing. Figure 12(B) shows an example of the data in Figure 12(A) graphed.
[0103] (S3: Calculation process of the actual center approximation curve) The details of the performance centered approximation curve calculation process will be explained with reference to Figures 13A to 15. In the performance centered approximation curve calculation process, data is not distinguished by project code, and an approximation curve is calculated using approximation components based on all project data of past monthly performance data. This process has the following features.
[0104] (1) The options for approximation functions include polynomial functions, exponential functions, logarithmic functions, logistic functions, and Bass model functions, and multiple options can be selected. (2) Among the selected functions, the one with the best AIC (Akaike Information Criterion), an index for measuring the accuracy of approximation, will be the approximate curve used in the subsequent processing flow.
[0105] Hereafter, the approximation function deemed to have the best approximation accuracy will be referred to as "Y = F(X)", the name of this approximation function will be called the "performance-centered approximation function", and the curve that this approximation function shows when plotted on a graph will be called the "performance-centered approximation curve". X represents the planned completion rate, Y represents the actual progress rate, and F represents the performance-centered approximation function. By passing X: the planned completion rate to F, the predicted value of Y: the actual progress rate can be calculated.
[0106] The above "bus model function" is used by the Ministry of Land, Infrastructure, Transport and Tourism to estimate the progress rate of construction work. In addition, polynomial functions and logistic functions are also used for comparison with the bus model function (https: / / www.mlit.go.jp / sogoseisaku / jouhouka / content / 001348995.pdf).
[0107] Specifically, the following processing is performed.
[0108] (1) Among the parameter settings registered in the abnormality determination definition master 106a, information on the X-axis item, Y-axis item, and function for each approximate function candidate name is passed to the approximation component to find the optimal coefficient. (2) Calculate the AIC based on the coefficients returned from the approximate parts. (3) The approximation function with the smallest AIC is used as the performance center approximation function.
[0109] Each process will be described in detail below. (1) Among the parameter settings registered in the abnormality determination definition master 106a, information on the X-axis item, Y-axis item, and function for each approximate function candidate name is passed to the approximation component to find the optimal coefficient.
[0110] 13A and 13B, (A) shows an example of data in the anomaly determination definition master 106a. In this example, the anomaly determination definition ID is "JD001," the anomaly determination definition name is "project progress alert," the algorithm used is "approximate curve," and the parameter settings are {X-axis item: number of months elapsed, Y-axis item: cumulative actual amount, approximate function candidate names: [polynomial function, logistic function, Bass model function], ...}. The options for the approximating function correspond to the approximate function candidates in the parameter settings for the data registered in the anomaly determination definition master 106a, and the X and Y values of the data for which an approximate curve is to be drawn correspond to the processed values in the X value item and the Y value item in the parameter settings for the data registered in the anomaly determination definition master 106a.
[0111] (B) shows an example of monthly performance data for a past project after processing. (C) shows the candidate approximation functions, including polynomial functions, logistic functions, and Bass model functions. If a polynomial function is selected, polynomial function expressions of degrees 0 to 6 will be passed. In this example, polynomial functions, logistic functions, and Bass model functions have been selected as candidate approximation functions, so the nine functions above will be passed to the approximation component. Exponential functions and logarithmic functions can also be selected as candidate approximation functions.
[0112] The following process is performed for each candidate approximation function. As an example, we will explain the case where the function to be approximated is y = a / (1 + ebx + c) (logistic function). (D) shows an example of information (1) (data to be passed as an approximate curve) passed to the approximation component. (E) shows an example of information (2) (function to be approximated) passed to the approximation component. (F) shows an example of the information (coefficients) returned from the approximation component, where a = 100.85, b = -0.069, and c = 3.68.
[0113] (2) Calculate the AIC based on the coefficients returned from the approximate parts. For details on calculating the AIC, see http: / / www.radio3.ee.uec.ac.jp / ronbun / TR_YK_048_AIC.pdf.
[0114] In FIG. 14, (A) shows the formula for calculating AIC, which is AIC = (data size) × log (sum of squares of differences from predicted value) / (data size) + 2 × (number of coefficients). Here, the predicted value of Y is calculated by substituting X, the data for which an approximate curve is to be drawn, into the approximation function. The square of the difference from the predicted value of Y is calculated as the square of (Y value - predicted value of Y) of the data for which an approximate curve is to be drawn. The data size is the number of pairs of data (X, Y).
[0115] AIC is characterized by the fact that the smaller the sum of squares of the difference from the predicted value, the smaller the AIC, and the larger the number of coefficients, the larger the AIC. The AIC value itself has no meaning; comparing AIC values is an indicator for determining the superiority of an approximation function. The AIC formula is explained below. The first term in the AIC formula, "(data size) × log(sum of squares of the difference from the predicted value) / (data size)," represents the accuracy of the data for which the approximate curve is to be drawn, and the second term, "2 × (number of coefficients)," represents the penalty for having a large number of coefficients. By determining the superiority of an approximation function based on AIC, a well-balanced approximation function that does not overly fit the data for which the approximate curve is to be drawn and is robust to data changes is selected. The effect of determining the superiority of an approximation function based on AIC is explained below. Even if the actual data from past projects increases over time, an approximation function that fits past data too well is not selected. This has the following two effects: (1) It prevents the type and order of the approximation function from frequently changing each time an anomaly is detected. (2) Even if the number of projects that deviate significantly from the planned progress is drastically reduced due to thorough project progress management, the approximate curve will shift gradually before and after the thorough management.
[0116] (B) shows the data (X, Y) for which you want to draw an approximate curve. (C) shows the predicted value of Y and the sum of squares of the difference between the predicted value of Y and the data.
[0117] When the data size, the sum of squares of the difference between the predicted value of Y, the coefficient, and the number of coefficients (3) are substituted into the AIC calculation formula, the AIC becomes 94.18, as shown in (D).
[0118] When the above process is completed for all approximate function candidates, data such as that shown in Figure 15(A) can be obtained, and the AICs for all approximate function candidates can be obtained. In this figure, actual values are entered in the "...".
[0119] (3) The approximation function with the smallest AIC is used as the performance center approximation function. Assuming that the AIC of all the candidate approximation functions was calculated and the AIC of the logistic function was the smallest, we will explain the processing flow. If the performance center approximation function is Y = F(X), in this explanation, the performance center approximation function is expressed as y = 100.85 / (1 + e -0.069X+3.68 ) This is the function with the smallest AIC, and the coefficients are set to the function to be approximated. Hereafter, this function will be treated as F(X) (Features 2 and 3 of the present invention).
[0120] (S4: Processing to calculate the dispersion of data from the performance center approximation curve) The process of calculating the dispersion of data from the performance center approximation curve will be described in detail with reference to Figures 16 and 17. The process of calculating the dispersion of data from the performance center approximation curve includes the following steps.
[0121] (1) Calculate the predicted value of the actual progress rate for each scheduled progress rate of the monthly actual data of past projects. (2) Calculate the difference (residual) between the predicted value and the actual data of the progress rate. (3) Calculate the square of the residual (square of the residual).
[0122] A specific calculation will be described with reference to FIG. (1) Calculate the predicted value of the actual progress rate as predicted value = F (progress rate of construction period). (2) The residual, which is the difference between the actual progress rate and the predicted value, is calculated as follows: Residual = Actual progress rate - Predicted value. (3) Square the residual, Square of residual = (residual) 2 Calculate as follows.
[0123] Figure 16(A) shows the planned completion rate [%] and the actual progress rate [%] of the actual data of past projects. Figure 16(B) shows the newly added columns of predicted value [%], residual [%], and squared residual [%]. 2]. The forecast value [%] is calculated by passing the planned completion progress rate to the actual performance centered approximation function F(X). Because the residuals are squared, all residual squares are greater than or equal to 0. By squaring the residuals, the magnitude of the residual square can be expressed as the degree to which the past performance data deviates from the approximation curve.
[0124] Figure 17(A) is a graph showing the predicted value vs. approximate curve of past project data, with the horizontal axis showing the planned progress rate [%] and the vertical axis showing the actual progress rate [%]. Past actual data is plotted, and the predicted value is plotted on the approximate curve. The actual progress rate - predicted value shows the residual.
[0125] FIG. 17(B) is a graph showing the residuals, with the horizontal axis representing the progress rate [%] over the scheduled construction period and the vertical axis representing the residuals [%].
[0126] FIG. 17(C) is a graph showing the square of the residual, where the horizontal axis is the progress rate [%] and the vertical axis is the square of the residual [% 2 ] is shown.
[0127] (S5: Processing to calculate the data dispersion trend) The process of calculating the tendency of data dispersion will be described in detail with reference to Figures 18A to 20. The process of calculating the tendency of data dispersion involves the following steps.
[0128] (1) Draw an approximate curve for the square of the residual, name this approximate function the residual square approximation function, and define this residual square approximation function as σ 2 =V(X). The residual square approximation function is (polynomial function) 2 By passing X (the percentage progress of the planned completion period) to the residual square approximation function, the predicted value of the residual square at that percentage progress of the planned completion period is returned. (2) Take the square root of the residual square approximation function, name this function the standard deviation function, and let this standard deviation function be σ = S(X). By passing X: the planned completion period progress rate to the standard deviation function, the standard deviation (the degree of dispersion of the actual progress rate) for that planned completion period progress rate is returned.
[0129] It will be explained in [3-4. Why calculate the tendency of data dispersion] that a function that returns the standard deviation can be created by taking the square root of the residual square approximation function. The specific process will be explained below.
[0130] (1) Draw an approximate curve using approximate parts for the square of the residual (this curve is called the residual square approximate curve). In FIG. 18A and FIG. 18B, (A) shows the planned construction period progress rate [%], actual progress rate [%], forecast value [%], residual [%], square of residual [%] of the performance data of past projects. 2 (B) shows candidates for the residual square approximation function.
[0131] (Polynomial functions) 2 The reason why we specify the function as the function to be approximated is as follows. 1. Because we are not approximating the progress rate of a project, there is no need to specify special functions such as bus model functions that are said to represent the progress of a project, so we specify functions that are as simple as possible. 2. When approximating the actual progress rate, the actual progress rate increases as the planned progress rate progresses, but the squared residual may increase or decrease as the planned progress rate progresses. Therefore, in order to capture the increase or decrease in the squared residual, a polynomial function is used. 2 We have adopted the following. 3. (Polynomial functions) 2 This allows us to later use the square root of the residual square approximation function (a polynomial function) as the standard deviation function.
[0132] For each candidate residual squared approximation function, the following process is performed. For example, let the function to be approximated be y=(ax 2 +bx+c) 2 18A and 18B, (C) shows information (1) (data for which an approximate curve is to be drawn) passed to the approximating part, and (D) shows information (2) (function to be approximated) passed to the approximating part. (E) shows the information returned from the approximating part, with coefficients a = -0.0028, b = 0.3209, and c = 2.1336.
[0133] Next, the AIC is calculated. Since the same processing as that in (S3: performance center approximation curve calculation processing) in the processing flow is performed, detailed explanation will be omitted.
[0134] When the above process is completed for all candidates of the function that approximates the square of the residual, data such as that shown in FIG. 19(A) is obtained.
[0135] In this example, the function with the smallest AIC and the coefficients set to the function to be approximated is the residual square approximation function y=(-0.0028x 2 +0.3209x+2.1336) 2 From now on, this (quadratic function) 2 The following processing flow will be explained under the assumption that the AIC of the function of the form is the smallest. 2 =V(X). In this example, V(X)=(-0.0028X 2 +0.3209X+2.1336) 2 This becomes:
[0136] Figure 19(C) shows a graph in which the residual square approximation curve V(X) is plotted against the square of the residual. The horizontal axis is the progress rate [%], and the vertical axis is the square of the residual [% 2 In this way, by drawing an approximate curve, it is possible to take into account the variance in the data according to the progress rate of the construction period, such as the deviation from the approximate curve being small from the start of the project to the early and later stages, and the deviation being large in the middle stages.
[0137] (2) Calculate the standard deviation function. When the expected completion time progress rate is passed to X of the residual square approximation function obtained in (1), the output value is in [% 2 ], so we take the square root and convert it into a standard deviation function with units of [%]. The standard deviation function S(X) is calculated as S(X) = √V(X). In this processing flow example, V(X) = (-0.0028X 2 +0.3209X+2.1336) 2 Therefore, S(X)=-0.0028X2 The result is +0.3209X+2.1336.
[0138] Figure 20(A) shows the progress rate, progress rate, forecast value, residual, square of residual, and the absolute value of residual of the actual data of past projects. Figure 20(B) shows a graph that displays the standard deviation function V(X) as a standard deviation curve, with the horizontal axis representing progress rate [%] and the vertical axis representing the absolute value of residual [%]. 2 ]. The absolute value of the residual indicates the distance of deviation between the performance data of past projects and the performance center approximation curve. Note that in actual calculations, the absolute value of the residual is not calculated from the residual between the performance data of past projects and the predicted value using the performance center approximation function, but it is provided here as a reference when displaying the standard deviation curve on a graph.
[0139] The meaning of drawing an approximation curve to the square of the residuals and taking the square root of that approximation function will be explained in detail in [3-4. Why calculate the tendency of data dispersion].
[0140] (S6: Calculation of predicted values and thresholds for cases subject to abnormality detection) The calculation process for the predicted value and threshold for an abnormality detection target case will be explained in detail with reference to Figures 21 to 24. The actual result center approximation function F(X) and standard deviation function S(X) for finding the predicted value of the actual progress rate have been prepared through the processing up to this point. In the calculation process for the predicted value and threshold for an abnormality detection target case, past case data is not used, and the actual data of the abnormality detection target case is used for the calculation. In the calculation process for the predicted value and threshold for an abnormality detection target case, the following processing is performed on the processed actual data of the abnormality detection target case.
[0141] (1) The predicted value is calculated using the estimated progress rate and actual results center approximation function F(X). (2) Calculate the standard deviation (the predicted value of the distance of deviation from the approximate curve) using the scheduled completion period progress rate / standard deviation function S(X). (3) The upper and lower thresholds are set by multiplying the standard deviation by a constant value calculated based on the significance level, leaving a range above and below the predicted value.
[0142] (1) Calculate the predicted value. Calculate the predicted value of the actual progress rate for the scheduled progress rate of the project to be judged as abnormal.
[0143] Figure 21(A) shows the processed actual data of the project to be judged as abnormal, with columns for predicted values added to the project code, fiscal year and month, progress rate, and actual progress rate. The predicted values are calculated by using the actual results center approximation function F(X) where x = progress rate.
[0144] FIG. 21(B) is a graph of the predicted values of FIG. 21(A), with the horizontal axis representing the planned progress rate [%] and the vertical axis representing the actual progress rate [%].
[0145] (2) Calculate the standard deviation. Calculate the standard deviation of the progress rate of the planned completion period for the processed projects to be judged as abnormal.
[0146] Figure 22(A) shows the processed actual data for a project to be judged as abnormal, with the standard deviation column added to the project code, fiscal year and month, progress rate, actual progress rate, and forecast value. The standard deviation is calculated using the standard deviation function S(X) where x = progress rate. The calculated standard deviation means that there is a 68% chance that the actual data will fall within a range that is the standard deviation away from the center performance approximation curve.
[0147] FIG. 22(B) is a graph of the standard deviation of FIG. 22(A), where the horizontal axis is the progress rate [%] and the vertical axis is the standard deviation [% 2 ] is shown.
[0148] (3) Determine the threshold value. Figure 23(A) shows the processed monthly performance data for projects subject to abnormality detection, with columns for upper threshold [%] and lower threshold [%] added to the project code, fiscal year and month, planned completion period progress rate, actual progress rate, forecast value, and standard deviation.
[0149] The threshold is calculated as follows, as shown in FIG. 23(B): Upper threshold = predicted value + Z α / 2 × (standard deviation), lower threshold = predicted value - Z α / 2Calculate by multiplying the standard deviation by Z. α / 2 means the value of the upper α / 2% point of the standard normal distribution. For this α, the value of the parameter setting (significance level) registered in the abnormality determination definition master 106a in premise (1) is used.
[0150] FIG. 24(A) is a diagram showing an example of data in the abnormality determination definition master 106a, in which a parameter of {significance level: 0.1} is set for the abnormality determination definition ID "JD001."
[0151] For example, if the significance level α is 0.1, then α / 2 = 0.05 (5%), and Z α / 2 = 1.645. In this case, in Figure 24(B), for the record of the project B001 to be judged as abnormal where the planned completion period elapsed rate is 40%, the upper threshold = 28.678 + 1.645 x 10.415 and the lower threshold = 28.678 - 1.645 x 10.415. This can be interpreted as meaning that there is about a 90% probability that the actual progress rate will be above the lower threshold and below the upper threshold. Conversely, if this range is exceeded, it will be judged as abnormal.
[0152] Figure 24(C) is a graph of the predicted values, upper threshold, and lower threshold of Figure 23(B), with the horizontal axis representing the planned progress rate [%] and the vertical axis representing the actual progress rate [%]. Note that the graphs of the actual progress center approximation curve and the upper and lower thresholds are shown, but they are only for reference.
[0153] Figure 24(D) is a diagram to explain the significance level and percentile (https: / / ai-trend.jp / basic-study / normal-distribution / normal-distribution / ).
[0154] (S7: Abnormality determination process) The abnormality determination process will be described in detail with reference to Figures 25 to 27. In the abnormality determination process, the following processes are performed.
[0155] (1) Abnormality determination is performed based on the monthly performance data of the case to be determined and the threshold value obtained in the previous process. (2) The data of the determination result is stored in a table (not shown) in the storage unit 206.
[0156] The abnormality determination process will be specifically described below. (1) For cases subject to abnormality detection, an abnormality is detected by comparing the actual progress rate of the case subject to abnormality detection with the threshold value determined in the previous process.
[0157] FIG. 25(A) shows the case code, accounting year and month, progress rate after scheduled completion, progress rate, forecast value, standard deviation, upper threshold, and lower threshold of the performance data of the case to be judged as abnormal.
[0158] Figure 25(B) shows a graph of the anomaly determination in Figure 25(A), with the horizontal axis representing the planned completion period elapsed rate [%] and the vertical axis representing the actual progress rate [%], and plotting the actual performance center approximation curve, upper threshold, lower threshold, B001, and B002. In the figure, B001's planned completion period elapsed rate of "40" and actual progress rate of "60" are determined to be abnormal because they exceed the upper threshold. Also, B002's planned completion period elapsed rate of "57.143" and actual progress rate of "20" are determined to be abnormal because they fall below the lower-upper threshold. For example, if the upper threshold is exceeded, the possibility of fraudulent overstating of actual performance is detected. Also, if the lower threshold is exceeded, the possibility of a delay in progress is detected.
[0159] (2) The judgment result data is stored in the judgment result table 106b. 26 is a diagram showing an example of judgment result data. The judgment result data may include the case code, accounting year and month, progress rate for the planned completion period, actual progress rate, forecast value, standard deviation, upper threshold, lower threshold, and judgment result (FALSE or TRUE). For data judged to be abnormal, the judgment result is "TRUE", and for other data, it is "FALSE".
[0160] If an anomaly is detected, a message indicating this is displayed. Figure 27 shows an example of an anomaly detection message. The anomaly detection message displays the anomaly detection definition name, the case code in which the anomaly was detected, the calculation method, the fiscal year, the actual progress rate, and the deviation from the threshold. In the example shown in the figure, the following is displayed: {Anomaly detection definition: JD001 Case progress alert, Case: B001, Fiscal year / month: 202203, Calculation method: Approximate curve (significance level: 0.1), Actual progress rate: 60%, Automatically detected because the actual progress rate for fiscal year / month 202203 exceeded the upper limit (45.81%). Exceeds the reference value by 14.19.}
[0161] (S8: Calculation processing of display data) The calculation process for the display data will be explained in detail with reference to Figures 28 to 30. In the process up to this point, the performance center approximation function F(X) and the standard deviation function S(X) based on past cases have been calculated. In the calculation process for the display data, data for display (for displaying the approximation curve on the screen) is created and stored in a table for the resultant ancillary information.
[0162] The display data calculation process involves the following steps. (1) Calculate the progress rate of the construction period to display the data for display. (2) The values on the approximation curve of the data to be displayed, the upper and lower threshold values are calculated. (3) The results of the approximation curve algorithm are stored in the attached information table.
[0163] The specific processing contents will be explained below. (1) Calculate the progress rate of the construction period to be displayed in the data for display. The progress rate of the construction period of the displayed data is the minimum and maximum value of the progress rate of the construction period of the past project data divided by the value of the parameter called "smoothness." Figure 28 shows a graph of the actual data of past projects, with the horizontal axis representing the planned progress rate [%] and the vertical axis representing the actual progress rate [%], and A001, A002, and A003 are plotted.
[0164] The range of this planned progress rate from 0 to 120% is divided into 12 equal parts with a smoothness of 12 to prepare the progress rate of the progress of the project for display. For the smoothness, the parameter setting (smoothness) value registered in the abnormality judgment definition master 106a in premise (1) is used.
[0165] 28(B) shows an example of data settings in the abnormality determination definition master 106a. In the example shown in the figure, the parameter "smoothness: 12" is set for the abnormality determination definition ID "JD001".
[0166] As shown in Figure 31, when the display data is displayed on the screen, the predicted value, upper threshold, and lower threshold may be displayed as a line graph, so the larger the "smoothness" value, the smoother the line graph will be, closer to a curve.
[0167] By dividing 0 to 120% into 12 equal parts, the construction progress rate of the display data shown in Figure 28(C) is prepared. In this example, the interval between each construction progress rate is 10%, which is 0 to 120% divided into 12 equal parts, and the number of points in the display data is 13.
[0168] (2) The values on the approximation curve of the data to be displayed, the upper and lower threshold values are calculated. Using F(X) and S(X) obtained when processing the actual data of past projects, the values on the approximate curve for the progress rate of the construction period of the data to be displayed and the values of the upper and lower thresholds are calculated. Since this is exactly the same process as that used to calculate the predicted values and thresholds for projects subject to abnormality detection, a detailed explanation of the calculation process will be omitted.
[0169] After the calculation process, the display data will be as shown in Figure 29(A), with columns for predicted value, upper threshold, and lower threshold added to the planned completion period elapsed rate. Figure 29(B) is a graph (line graph) of the display data in Figure 29(A), with the horizontal axis representing the planned completion period elapsed rate [%] and the vertical axis representing the actual progress rate [%], and the predicted value, upper threshold, and lower threshold are plotted. The values of the white dots on the graph are calculated.
[0170] As can be seen from the graph, the width between the upper and lower thresholds and the performance center approximation curve is calculated using a standard deviation function, so the thresholds reflect the following two characteristics:
[0171] 1. In the early and later stages of a project, there is little chance of deviation from the approximate curve. Therefore, even a slight deviation from the approximate curve is highly abnormal, so the threshold range is narrow. 2. There is a high possibility of deviation from the approximate curve in the middle stage of a project. As a result, even if something is normal, there is a possibility of deviation from the approximate curve to some extent, and even if there is a slight deviation, it is not considered abnormal, so the range of the threshold is wide.
[0172] (3) The approximate curve result is stored in the attached information table 106c. The following display data is stored in an approximate curve algorithm result attached information table (hereinafter referred to as "approximate curve result attached information table"). The approximate curve result attached information table 106c stores data to be displayed as supplementary information when the anomaly determination result is displayed as a graph on the screen. Unlike the data itself of the case to be determined to be anomaly, data to be compared with the data of the case to be determined to be anomaly is stored. The stored data is calculated based only on data from past cases.
[0173] 30(A) shows an example of display data, and FIG. 30(B) shows an example of approximate curve result attached information stored in approximate curve result attached information table 106c based on the display data. The approximate curve result attached information has items such as execution history ID, line number, X, Y, upper threshold, and lower threshold.
[0174] Here, the execution history ID is an ID automatically assigned for each execution of anomaly determination, the row number is a consecutive number assigned to each row of the display data starting from 0, X corresponds to the X-axis value of the display data, Y corresponds to the Y-axis value of the display data, the upper threshold is the upper threshold of the display data, and the lower threshold is the lower threshold of the display data.
[0175] As described above, when an abnormality is detected using the judgment result data (see FIG. 26) and the information attached to the approximate curve result (see FIG. 30(B)), a message indicating the abnormality is detected and a graph of the abnormality judgment result (for example, a line graph) is displayed. FIG. 31 shows an example of the message indicating the abnormality detection (the same as FIG. 27) and an example of a graph of the abnormality judgment result. In the example of the graph of the abnormality judgment result, the horizontal axis is the planned construction period elapsed rate [%] and the vertical axis is the actual progress rate [%], and the predicted value, threshold values (upper and lower), and B001 are plotted. B001 is displayed based on the judgment result data. The predicted value and threshold values (upper and lower) are displayed based on the information attached to the approximate curve result.
[0176] [3-3. Options for processing monthly performance data for past projects] 32 to 41, options for processing monthly performance data of past cases will be described (Feature 1 of the present invention). As described above, scale unification items can be set in the abnormality determination definition master 106a.
[0177] Fig. 32 is a diagram showing a setting example of the abnormality determination definition master 106a. The assumed item set in the X-axis scale unified item is the planned construction period or the implementation period, and the assumed item set in the Y-axis scale unified item is the budget or the actual amount. The setting example of the abnormality determination definition master 106a shown in Fig. 32 is an example for the construction industry, but since it is also assumed that it will be used in other industries, setting examples for other industries are shown in Fig. 33.
[0178] Figure 33(A) shows an example for the system development industry, with the X-axis item being "Elapsed man-hours," the Y-axis item being "Cumulative sales amount," the X-axis unified scale item being "Planned man-hours OR Actual man-hours," and the Y-axis unified scale item being "Estimated amount OR Sales amount."
[0179] Here, [elapsed man-hours [man-days]] is the man-hours required since system development began. [Cumulative sales amount [yen]] is the amount recorded as sales at the point when the elapsed man-hours have elapsed. [Planned man-hours [man-days]] is the planned man-hours required until development is complete, determined before system development begins. [Actual man-hours [man-days]] is the man-hours required until development is complete, determined after system development is complete. [Estimated amount [yen]] is the amount determined before system development begins and planned to be recorded as sales when development is complete. [Sales amount [yen]] is the sales amount recorded at the time of development completion, determined after system development is complete.
[0180] Figure 33(B) shows an example from the corporate training industry, with the X-axis item being "number of days elapsed since training," the Y-axis item being "learning retention rate," the X-axis scale unified item being "planned number of days to complete OR actual number of days required," and the Y-axis scale unified item being "target retention rate OR actual retention rate."
[0181] Here, [number of days elapsed since the start of training [days]] is the number of days that have passed since the start of training. [Learning retention rate [%]] is the retention rate ascertained through check tests, etc., at the point in time when training has elapsed. [Planned completion date [days]] is the number of days planned until the end of training, determined before the start of training. [Actual number of days required [days]] is the number of days required for training, determined after the training has ended. [Target retention rate [%]] is the retention rate of the goal that participants are desired to achieve, determined before the start of training. [Actual retention rate [%]] is the actual retention rate of participants, determined after the training has ended.
[0182] In addition, the X-axis item, Y-axis item, X-axis scale unification item, and Y-axis scale unification item can be freely set according to the industry of the user of this anomaly detection system.
[0183] Fig. 34(A) is a diagram showing an example of monthly performance data of past cases, and Fig. 34(B) shows an example of schedule and performance data of past cases.
[0184] Four methods for processing data of past cases will be described with reference to Fig. 35. Data is classified into four patterns, A, B, C, and D, according to the parameter settings of the abnormality determination definition master 106a.
[0185] AX-axis scale unified item: implementation period, Y-axis scale unified item: actual amount BX-axis scale unified item: planned construction period, Y-axis scale unified item: actual amount In the case of a unified scale item on the C-axis: implementation period, and a unified scale item on the Y-axis: budget DX axis scale unified item: planned construction period, Y axis scale unified item: budget
[0186] The patterns ABCD are named as follows: Pattern A is "strict scheduled construction period and budget type", Pattern B is "flexible scheduled construction period and strict budget type", Pattern C is "strict scheduled construction period and flexible budget type", and Pattern D is "flexible scheduled construction period and budget type".
[0187] "Strict" and "Flexible" mean checking whether the project was completed on schedule and within budget, respectively, in a strict and flexible manner. Below, we will explain examples of each pattern A to D.
[0188] (Pattern A. Strict scheduled construction period and budget type (X-axis scale unified item: implementation period, Y-axis scale unified item: actual amount)) Pattern A, the strict scheduled construction period / budget type, will be explained with reference to Figures 36 and 37. An example of processing A003 will be explained below. Figure 36(A) shows an example of monthly actual data for A003, and Figure 36(B) shows an example of scheduled / actual data for A003. When the monthly actual data for A003 is processed and the implementation period elapsed rate is calculated using the number of months elapsed / implementation period, and the actual progress rate is calculated using the cumulative actual amount / actual amount, the result is as shown in Figure 36(C).
[0189] When the monthly performance data for all past cases is processed, it will look like Figure 37(A). Figure 37(B) is a graph of the monthly performance data for past cases in Figure 37(A), and Figure 38(C) shows a graph after calculation processing of the display data.
[0190] Since the maximum X-axis values for past projects are all 100%, the thresholds calculated in subsequent processes will converge with an X value close to 100%. This makes it possible to strictly check that projects to be judged as abnormal are completed on schedule before they started. Also, since the maximum Y-axis values for past projects are all 100%, the thresholds calculated in subsequent processes will converge with a Y value close to 100%. This makes it possible to strictly check that projects to be judged as abnormal are completed on budget before they started.
[0191] (Pattern B. Flexible construction period and strict budget (X-axis scale unified item: planned construction period, Y-axis scale unified item: actual amount)) Pattern B, flexible scheduled construction period / strict budget type, will be explained with reference to Figures 38 and 39. An example of processing A003 will be explained below. Figure 38(A) shows an example of monthly actual data for A003, and Figure 38(B) shows an example of planned / actual data for A003. When the monthly actual data for A003 is processed and the planned construction period progress rate is calculated using the number of months elapsed / planned construction period, and the actual progress rate is calculated using the cumulative actual amount / actual amount, the result is shown in Figure 38(C).
[0192] When the monthly performance data for all past cases is processed, it will look like Figure 39(A). Figure 39(B) is a graph of the monthly performance data for past cases in Figure 39(A), and Figure 39(C) shows a graph after calculation processing of the display data.
[0193] Since the maximum value of the X-axis for past projects is 120%, the threshold value calculated in subsequent processing will converge to an X value of around 120%. This makes it possible to determine that a project to be judged as abnormal is within the normal range even if it is completed with a slight delay. Also, since the maximum value of the Y-axis for past projects is all 100%, the threshold value calculated in subsequent processing will converge to a Y value of around 100%. This makes it possible to strictly check that projects to be judged as abnormal are completed according to the budget before they started.
[0194] (Pattern C. Strict scheduled construction period, flexible budget (X-axis scale unified item: implementation period, Y-axis scale unified item: budget)) Pattern C, strict scheduled construction period / flexible budget, will be explained with reference to Figures 40 and 41. An example of processing A003 will be explained below. Figure 40(A) shows an example of monthly actual data for A003, and Figure 40(B) shows an example of planned / actual data for A003. When the monthly actual data for A003 is processed and the implementation period progress rate is calculated using the number of months elapsed / implementation period, and the budget progress rate is calculated using the cumulative actual amount / budget, the result is shown in Figure 40(C).
[0195] When the monthly performance data for all past cases is processed, it will look like Figure 41(A). Figure 41(B) is a graph of the monthly performance data for past cases in Figure 41(A), and Figure 41(C) shows a graph after calculation processing of the display data.
[0196] Since the maximum value of the X-axis for past projects is always 100%, it is possible to strictly check that projects subject to abnormality detection are completed on schedule before they start. Also, since the budget progress rate at the time of completion of past projects varies depending on the project, an abnormality will not be detected even if the budget progress rate at the time of project completion is slightly different from 100%.
[0197] (Pattern D. Flexible scheduled construction period and budget (X-axis scale unified item: scheduled construction period, Y-axis scale unified item: budget) Pattern D, flexible planned construction period and budget, will be explained with reference to Figures 42 and 43. An example of processing A003 will be explained below. Figure 42(A) shows an example of monthly actual data for A003, and Figure 42(B) shows an example of planned and actual data for A003. When the monthly actual data for A003 is processed and the planned construction period progress rate is calculated using the number of months elapsed / planned construction period, and the budget progress rate is calculated using the cumulative actual amount / budget, the result is shown in Figure 42(C).
[0198] When the monthly performance data for all past cases is processed, it will look like Figure 43(A). Figure 43(B) is a graph of the monthly performance data for past cases in Figure 43(A), and Figure 43(C) shows a graph after calculation processing of the display data.
[0199] Since the progress rate and budget progress rate at the time of completion of past projects vary depending on the project, the threshold value is also calculated with a range of threshold values in the later stages of the progress rate.
[0200] [3-4. What does it mean to draw an approximate curve to the square of the residual and take the square root of that approximate function?] With reference to FIGS. 44 to 46, the meaning of drawing an approximation curve to the square of the residual and taking the square root of the approximation function will be described.
[0201] Fig. 44 is a diagram for explaining a general method for estimating the dispersion of data. In Fig. 43, it is assumed that the residual follows a standard normal distribution with a mean of 0, and the variance of the residual is σ 2 The estimated value is calculated as σ = (sum of squared residuals) / (number of data). The standard deviation is the square root of the variance, and is calculated as σ = √(sum of squared residuals) / (number of data). As shown in (A), this can be interpreted as meaning that there is a 68% chance that the residual will be in the range of -σ to +σ.
[0202] Figure 45 shows a graph of the residuals, with the horizontal axis representing the progress rate [%] and the vertical axis representing the predicted value difference [%]. This can be interpreted as meaning that there is a 68% probability that the dispersion falls within this range, so by using +σ and -σ, a threshold range can be constructed.
[0203] In Figure 46(A), the following equation appears when estimating the variance: σ 2 =(sum of squared residuals) / (number of data points) is the formula for finding the average of the squared residuals, and when viewed on a graph plotting the squared residuals, it corresponds to the coefficient a when approximating with a straight line graph of the form y=a.
[0204] Since σ, which is used to calculate the range of residuals, is the square root of a, the square root of the approximation function of the square of the residual can be used as the standard deviation function. However, the above calculation does not express the variation in residuals according to the progress rate of the scheduled completion period (it does not express the characteristic that residuals are small in the early and later stages of project progress and large in the middle stages).
[0205] Therefore, by replacing the part that is approximated by a straight line graph of the form y = a with a function like the one shown in (B), it is possible to express the variation in residuals according to the progress rate of the planned construction period. In this way, a unique calculation method has been introduced into the part that was approximated by a straight line graph.
[0206] As described above, according to this embodiment, the data processing unit 102b processes monthly performance data of past cases by calculating, for each case identification information, a progress rate obtained by dividing the X-axis item by the X-axis scale uniform item of the schedule / performance data of the past cases, and a progress rate obtained by dividing the Y-axis item by the Y-axis scale uniform item of the schedule / performance data of the past cases, and also processes monthly performance data of cases subject to abnormality determination by calculating, for each case identification information, a progress rate obtained by dividing the X-axis item by the schedule period or target end period of the schedule / performance data of the case subject to abnormality determination, and a progress rate obtained by dividing the Y-axis item by the budget or target of the schedule / performance data of the case subject to abnormality determination; The system is provided with an approximation curve calculation unit 102c that calculates a performance center approximation curve for all data of the monthly performance data of the case, a threshold calculation unit 102d that calculates the dispersion of all data of the monthly performance data of the past case after processing from the calculated performance center approximation curve, calculates a standard deviation curve that expresses the calculated dispersion trend, and calculates a threshold based on the performance center approximation curve and the standard deviation curve, and an abnormality determination unit 102e that compares the processed monthly performance data of the case that is the target of abnormality determination with the calculated threshold to determine an abnormality, so that it is possible to use the approximation curve to highly accurately detect abnormalities in the way the progress of the target case is accumulating based on the performance data of the past case.
[0207] [4. Contribution to the United Nations-led Sustainable Development Goals (SDGs)] This embodiment can contribute to improving business efficiency and promoting appropriate management decisions by companies, thereby contributing to the achievement of SDGs Goals 8 and 9.
[0208] Furthermore, this embodiment can contribute to reducing waste and promoting paperless and electronic systems, thereby contributing to the achievement of SDGs Goals 12, 13, and 15.
[0209] Furthermore, this embodiment can contribute to strengthening control and governance, which can contribute to the achievement of Goal 16 of the SDGs.
[0210] 5. Other Embodiments The present invention may be implemented in various different embodiments other than those described above within the scope of the technical concept set forth in the claims.
[0211] For example, among the processes described in the embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using known methods.
[0212] Furthermore, the processing procedures, control procedures, specific names, information including parameters such as registered data and search conditions for each process, screen examples, and database configurations shown in this specification and drawings can be changed as desired unless otherwise specified.
[0213] Furthermore, with regard to the anomaly detection system 100, the components shown in the figures are functional concepts, and do not necessarily have to be physically configured as shown in the figures.
[0214] For example, all or any part of the processing functions of the anomaly detection system 100, particularly the processing functions performed by the control unit, may be implemented by a CPU and a program interpreted and executed by the CPU, or may be implemented as hardware using wired logic. The program is recorded on a non-transitory computer-readable recording medium containing programmed instructions for causing an information processing device to execute the processes described in this embodiment, and is mechanically read by the anomaly detection system 100 as needed. That is, a computer program for providing instructions to the CPU in cooperation with an OS and performing various processes is recorded in a storage unit such as a ROM or HDD (Hard Disk Drive). The computer program is executed by being loaded into RAM, and cooperates with the CPU to form the control unit.
[0215] Furthermore, this computer program may be stored in an application program server connected to the anomaly detection system 100 via any network, and all or part of it may be downloaded as needed.
[0216] Furthermore, the program for executing the processes described in this embodiment may be stored in a non-transitory computer-readable recording medium, or may be configured as a program product. Here, the term "recording medium" includes any "portable physical medium" such as a memory card, a Universal Serial Bus (USB) memory, a Secure Digital (SD) card, a flexible disk, a magneto-optical disk, a ROM, an Erasable Programmable Read Only Memory (EPROM), an Electrically Erasable and Programmable Read Only Memory (EEPROM (registered trademark)), a Compact Disk Read Only Memory (CD-ROM), a Magneto-Optical disk (MO), a Digital Versatile Disk (DVD), and a Blu-ray (registered trademark) disc.
[0217] Furthermore, a "program" is a data processing method written in any language or description method, regardless of the format, such as source code or binary code. Note that a "program" is not necessarily limited to a single program, but also includes programs that are distributed as multiple modules or libraries, or programs that achieve their functions by working together with other programs, such as an OS. Note that the specific configurations and reading procedures for reading a recording medium in each device shown in the embodiments, as well as the installation procedures after reading, can use well-known configurations and procedures.
[0218] The various databases stored in the memory unit are storage means such as memory devices such as RAM and ROM, fixed disk devices such as hard disks, flexible disks, and optical disks, and store various programs, tables, databases, and web page files used for various processes and providing websites.
[0219] The anomaly detection system 100 may be configured as an information processing device such as a known personal computer or workstation, or may be configured as the information processing device to which any peripheral device is connected. The anomaly detection system 100 may also be realized by installing software (including programs, data, etc.) that causes the device to perform the processing described in this embodiment.
[0220] Furthermore, the specific form of distribution and integration of the devices is not limited to that shown in the drawings, and all or part of them can be configured by functionally or physically distributing and integrating them in any unit according to various additions or functional loads. In other words, the above-described embodiments can be implemented in any combination, or embodiments can be implemented selectively. [Explanation of symbols]
[0221] 100 Anomaly Detection System 102 Control section 102a Data acquisition section 102b Data Processing Department 102c Approximate curve calculation section 102d Threshold calculation unit 102e Abnormality judgment section 102f Screen display control unit 102g Master Maintenance Department 104 Communication interface unit 106 Storage section 106a Abnormality judgment definition master 106b Judgment result table 106c Approximate curve result additional information table 108 Input / Output Interface Section 112 Input Device 114 Output Device 200 servers 300 Network 400 Business Systems
Claims
1. An anomaly detection system that includes a control unit and detects anomalies in business data, The control unit An abnormality determination definition master in which X-axis item, Y-axis item, X-axis scale unification item, and Y-axis scale unification item are set as parameters for each abnormality determination definition; Monthly performance data of past projects, including project identification information, X-axis items, and Y-axis items; Past project schedule and performance data including project identification information, X-axis scale standard items, and Y-axis scale standard items; Monthly performance data of cases subject to abnormality determination, including case identification information, X-axis items, and Y-axis items; Planned and actual data of the case to be determined as abnormal, including case identification information, planned period or target completion period, budget or target; It is configured to be accessible to data processing means for calculating and processing the monthly performance data of the past cases, for each case identification information, a progress rate obtained by dividing the X-axis item by the X-axis scale uniform item of the schedule / performance data of the past cases, and a progress rate obtained by dividing the Y-axis item by the Y-axis scale uniform item of the schedule / performance data of the past cases; and for calculating and processing the monthly performance data of the abnormality determination case, for each case identification information, a progress rate obtained by dividing the X-axis item by the schedule period or target completion period of the schedule / performance data of the abnormality determination case, and a progress rate obtained by dividing the Y-axis item by the budget or target of the schedule / performance data of the abnormality determination case; an approximation curve calculation means for calculating a performance center approximation curve for all data of the processed monthly performance data of past cases; a threshold calculation means for calculating a dispersion of all the data from the calculated performance center approximation curve, calculating a standard deviation curve that expresses the tendency of the calculated dispersion, and calculating a threshold based on the performance center approximation curve and the standard deviation curve; an abnormality determination means for comparing the processed monthly performance data of the case to be determined as an abnormality with the calculated threshold value to determine an abnormality; An anomaly detection system comprising:
2. 2. The anomaly detection system according to claim 1, wherein the X-axis item includes the number of elapsed months, the Y-axis item includes a cumulative actual amount, the X-axis scale unified item includes a planned construction period or implementation period, and the Y-axis scale unified item includes a budget or actual amount.
3. The abnormality determination definition master further includes one or more approximate function candidates set as parameters for each of the abnormality determination definitions, 2. The anomaly detection system according to claim 1, wherein the approximate curve calculation means obtains optimal coefficients by passing the elapsed time rate and progress rate of the processed past monthly actual data and information on one or more approximate function candidates set in the anomaly judgment definition master to an approximating component, calculates AIC based on the coefficients returned from the approximating component, and determines the approximate function candidate with the smallest AIC as the actual performance center approximate function.
4. The anomaly detection system according to claim 3 , wherein the one or more candidate approximation functions include at least one of a polynomial function, an exponential function, a logarithmic function, a logistic function, and a Bass model function.
5. The threshold calculation means Substituting each progress rate of the processed monthly performance data of the past projects into the performance center approximation function of the performance center approximation curve to calculate a predicted value of the progress rate, calculating the residual which is the difference between the predicted value and the actual data of the progress rate, and further calculating the square of the residual, Draw an approximate curve of the residual square approximation function for the square of the residual, and take the square root of the residual square approximation function to obtain the standard deviation function. The anomaly detection system according to claim 1, characterized in that each progress rate of the processed monthly actual data of the case to be judged as an anomaly is substituted into the actual result center approximation function to calculate a predicted value of the progress rate, and also substituted into the standard deviation function to calculate a standard deviation, and an upper threshold and a lower threshold are calculated by multiplying the standard deviation by a constant value calculated based on the significance level, with a range above and below the predicted value.
6. 2. The anomaly detection system according to claim 1, wherein the anomaly determination means outputs an anomaly detection message and a graph to a display unit when an anomaly is determined to be present.
7. 7. The anomaly detection system according to claim 1, further comprising a master maintenance means for setting data of the anomaly determination definition master in response to an operation by an operator.
8. An anomaly detection method executed by an information processing device including a control unit, The control unit An abnormality determination definition master in which X-axis item, Y-axis item, X-axis scale unification item, and Y-axis scale unification item are set as parameters for each abnormality determination definition; Monthly performance data of past projects, including project identification information, X-axis items, and Y-axis items; Past project schedule and performance data including project identification information, X-axis scale standard items, and Y-axis scale standard items; Monthly performance data of cases subject to abnormality determination, including case identification information, X-axis items, and Y-axis items; Planned and actual data of the case to be determined as abnormal, including case identification information, planned period or target completion period, budget or target; It is configured to be accessible to Executed in the control unit: a data processing step of calculating and processing the monthly performance data of the past projects for each project identification information by calculating a progress rate by dividing the X-axis item by the X-axis scale uniform item of the schedule / performance data of the past projects and a progress rate by dividing the Y-axis item by the Y-axis scale uniform item of the schedule / performance data of the past projects, and also calculating and processing the monthly performance data of the abnormality determination target projects for each project identification information by calculating a progress rate by dividing the X-axis item by the schedule period or target completion period of the schedule / performance data of the abnormality determination target projects and a progress rate by dividing the Y-axis item by the budget or target of the schedule / performance data of the abnormality determination target projects; an approximation curve calculation step of calculating a performance center approximation curve for all data of the processed monthly performance data of past projects; a threshold calculation step of calculating the dispersion of all the data from the calculated approximation curve, calculating a standard deviation curve that represents the tendency of the calculated dispersion, and calculating a threshold based on the performance center approximation curve and the standard deviation curve; an abnormality determination step of comparing the processed monthly performance data of the case to be determined to be abnormal with the calculated threshold value to determine an abnormality; An anomaly detection method comprising:
9. An abnormality detection program to be executed by an information processing device having a control unit, The control unit An abnormality determination definition master in which X-axis item, Y-axis item, X-axis scale unification item, and Y-axis scale unification item are set as parameters for each abnormality determination definition; Monthly performance data of past projects, including project identification information, X-axis items, and Y-axis items; Past project schedule and performance data including project identification information, X-axis scale standard items, and Y-axis scale standard items; Monthly performance data of cases subject to abnormality determination, including case identification information, X-axis items, and Y-axis items; Planned and actual data of the case to be determined as abnormal, including case identification information, planned period or target completion period, budget or target; It is configured to be accessible to In the control unit, a data processing step of calculating and processing the past monthly actual performance data for each case identification information by calculating a progress rate by dividing an X-axis item by an X-axis scale uniform item of the schedule / performance data of the past case and a progress rate by dividing a Y-axis item by a Y-axis scale uniform item of the schedule / performance data of the past case, and also calculating and processing the monthly actual performance data of the abnormality determination target for each case identification information by calculating a progress rate by dividing an X-axis item by a scheduled period or a target completion period of the schedule / performance data of the abnormality determination target and a progress rate by dividing a Y-axis item by a budget or target of the schedule / performance data of the abnormality determination target; an approximation curve calculation step of calculating a performance center approximation curve for all data of the processed monthly performance data of past projects; a threshold calculation step of calculating the dispersion of all data from the calculated performance center approximation curve, calculating a standard deviation curve that represents the tendency of the calculated dispersion, and calculating a threshold based on the performance center approximation curve and the standard deviation curve; an abnormality determination step of comparing the processed monthly performance data of the case to be determined to be abnormal with the calculated threshold value to determine an abnormality; An anomaly detection program to execute the above.
Citation Information
Patent Citations
Management system for controlling progress of construction work
JP2004094511A
System for managing process progress
JP2004326279A
Development process evaluation management system, management server device and method
JP2012027590A
Construction work management support system
JP2015095122A
Shape measurement method and shape measurement device
JP2017078641A