Model building support system and model building support method
The predictive model construction support system addresses the challenge of diverse equipment failure modes by dividing explanatory variables into groups and calculating identification feature scores, resulting in improved prediction accuracy and system efficiency.
Patent Information
- Application Number
- JP2021133298
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-08-18
- Publication Date
- 2025-05-07
- Estimated Expiration
- 2041-08-18
AI Technical Summary
Existing methods for constructing failure prediction models for equipment, such as storage drives in data centers, face challenges due to diverse and complex failure modes and operation modes, leading to a large workload in analyzing various factors and low prediction accuracy.
A predictive model construction support system that divides explanatory variables into groups to improve prediction accuracy, using an information processing system to calculate and score identification features based on the accuracy and coverage rate of each group.
The system effectively supports the construction of predictive models with high accuracy for equipment failures, improving the efficiency and stability of system operation by optimizing feature quantity search and group division methods.
Smart Images

Figure 0007672301000001 
Figure 0007672301000002 
Figure 0007672301000003
Abstract
Description
[Technical field]
[0001] The present invention relates to a model building support system and a model building support method. [Background technology]
[0002] Patent Document 1 describes a processor implementation method in which deep learning is used to extract features from sensor data mapped to a knowledge base, and a machine learning model is generated to analyze the sensor data based on the extracted features, thereby performing predictive analysis.
[0003] Patent Document 2 describes an information processing device (computer) configured for the purpose of ensuring the diversity of experiments through trial and error and making experiments more efficient. The information processing device generates an experiment plan from existing data using a regression model, displays a graph showing the consistency between features in each experiment plan, and allows the user to select a feature.
[0004] Patent Document 3 describes a system configured for the purpose of reducing the trial and error required by an analyst to select a set of data items to be analyzed in multidimensional data analysis using an OLAP (Online Analytical Processing) tool. The above system performs a process of recommending an analysis axis for multidimensional data analysis, a process of calculating the degree of association between data items of multidimensional data, a process of extracting a set of data items suitable for the analysis target based on the degree of association, and a process of presenting the extracted set of data items as an analysis axis to be recommended to the analyst when analyzing multidimensional data.
[0005] Patent Document 4 describes a data mining system that is configured to enable a user to change the granularity of attributes without trial and error. The data mining system selects attributes from data including a plurality of attributes and attribute values, groups the attribute values of the selected attributes based on a classification hierarchy that hierarchically represents classifications corresponding to the attributes stored in advance, calculates a test amount indicating the strength of association between the grouped attribute values and the attribute to be analyzed, determines whether the grouped attribute values are characteristic in relation to the attribute to be analyzed based on the calculated test amount, and if it is determined that the grouped attribute values are not characteristic, executes grouping processing again based on a classification in a higher hierarchy than the hierarchy used in the previous grouping in the classification hierarchy. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] U.S. Pat. No. 1,066,4698 [Patent Document 2] U.S. Pat. No. 10,318,674 [Patent Document 3] JP 2012-103841 A [Patent Document 4] JP 2011-034457 A Summary of the Invention [Problem to be solved by the invention]
[0007] In sites where a large number of devices are operated, there is a strong need to accurately predict the occurrence of device failures. If the occurrence of events such as device failures can be predicted with high accuracy, problems with the devices can be prevented and the devices can be operated efficiently. For example, in sites such as data centers and system centers, a large number of storage drives (Flash Module Drives (FMDs) and the like), numbering in the tens of thousands, are operated. If a highly accurate failure prediction model could be constructed for these storage drives, problems could be prevented and the occurrence of This will enable efficient and stable system operation.
[0008] However, when constructing such a failure prediction model, the following problems must be solved. For example, in the actual operation site of equipment, when multiple failure modes and operation modes exist and these are diverse and complex, it is necessary to perform cumbersome progress management for various factor analyses in order to search for appropriate feature values, which generates a large workload. In addition, when the frequency of equipment failure is low, it is difficult to grasp the signs of failure due to differences in product types and operation modes (such as when the frequency of use increases at the end of the month), and conventional identification methods such as decision trees, random forests, and XGBoost cannot solve this problem. In addition, in general, the formula for calculating feature values is often unclear, and when using search methods such as AutoML, genetic algorithms, and reinforcement learning, it is necessary for a person to prepare feature values in advance.
[0009] In Patent Documents 1 to 3, feature values are searched for by calculating feature value scores, but the diversity of failure modes and operation modes is not taken into consideration. In Patent Document 4, it is determined whether grouped attribute values are characteristic in relation to the attribute to be analyzed, but a hierarchy of groups must be prepared manually in advance.
[0010] The present invention has been made in consideration of the above background, and aims to provide a predictive model building support system and a predictive model building support method that support the building of a predictive model for accurately predicting events that occur in equipment. [Means for solving the problem]
[0011] To achieve the above object, one aspect of the present invention provides a method for determining whether a prediction model is capable of predicting an event of an equipment based on input explanatory variables, and a method for determining whether the prediction model is capable of predicting an event of an equipment based on input explanatory variables. prediction An information processing system for assisting in a search for a method of dividing the explanatory variables into groups to improve accuracy, the information processing system being configured using an information processing device having a processor and a memory element, and dividing the explanatory variables into a plurality of groups, and setting the discriminant feature amount based on the explanatory variables of each of the groups, Prediction of the dependent variable Calculate the accuracy C, and Said prediction Based on the accuracy C and the coverage rate S, which is the proportion of the explanatory variables of each group to the total explanatory variables before division, a score SF of the discrimination feature for each group is calculated, and information based on the calculated score SF is generated.
[0012] Problems, configurations and effects other than those described above will become apparent from the following description of the preferred embodiment of the invention. Effect of the Invention
[0013] According to the present invention, it is possible to assist in building a prediction model for accurately predicting events occurring in equipment. [Brief description of the drawings]
[0014] [Figure 1A] FIG. 13 is a diagram for explaining a case where division of explanatory variables into groups needs to be taken into consideration. [Figure 1B] FIG. 13 is a diagram for explaining a case where division of explanatory variables into groups needs to be taken into consideration. [Diagram 2] FIG. 13 is a diagram illustrating an example in which explanatory variables are divided into groups. [Figure 3A] FIG. 13 is a diagram illustrating an example of score calculation. [Figure 3B] FIG. 13 is a diagram illustrating an example of score calculation. [Figure 3C] FIG. 13 is a diagram illustrating an example of score calculation. [Figure 4A] FIG. 13 is a diagram showing another example of calculation of the score. [Figure 4B] FIG. 13 is a diagram showing another example of calculation of the score. [Figure 5A] FIG. 13 is a diagram showing another example of calculation of the score. [Figure 5B] FIG. 13 is a diagram showing another example of calculation of the score. [Figure 6] 1 is an example of model building support information. [Figure 7A] FIG. 1 is a diagram illustrating an example of a system configuration of a model building support system. [Figure 7B] 1 is an example of an information processing device used in the configuration of a model building support system. [Figure 8] FIG. 2 is a diagram showing main information handled in the model building support system. [Figure 9] FIG. 2 is a diagram showing main functions of the model building support system. [Figure 10] 1 is an example of time series data. [Figure 11A] 13 is an example of group information. [Figure 11B] 13 is an example of group information. [Figure 11C] 13 is an example of group information. [Figure 12] 1 is an example of a feature library. [Figure 13] 1 is an example of a feature amount table. [Figure 14A] FIG. 1 is a UML diagram showing an example of a process tree. [Figure 14B] FIG. 1 illustrates the construction of components of a process tree. [Figure 15] 11 is a flowchart illustrating a main process. [Figure 16] FIG. 13 is a diagram showing a list of main operations on a process tree. [Figure 17] 1A to 1E are diagrams illustrating changes in a process tree in response to operations (processing). [Figure 18] 13 is a flowchart illustrating details of a data registration process. [Figure 19A] 13 is a flowchart illustrating details of a target variable registration process. [Figure 19B] This is an example of a response variable. [Figure 19C] FIG. 13 is a diagram showing the structure of a process tree after a target variable is registered. [Figure 20] 11 is a flowchart illustrating details of a discrimination feature (DFS) registration process. [Figure 21] 13 is a flowchart illustrating details of a group division process. [Figure 22] 11 is a flowchart illustrating details of a group split feature (GFS) registration process. [Figure 23A] 11 is a flowchart illustrating details of registration of a group splitting feature (GFS) and group splitting processing using the group splitting feature (GFS). [Figure 23B] 13 is an example of a process tree structure after registration of a group splitting feature (GFS) and execution of a group splitting process using the group splitting feature (GFS). [Figure 24A] 13 is a flowchart illustrating details of a score calculation process. [Figure 24B] 24B is a flowchart illustrating details of the best child selection process of the group in FIG. 24A. [Diagram 25] 13 is a flowchart illustrating details of a process for acquiring the latest result. [Figure 26A] FIG. 13 is a diagram illustrating an example of a process tree structure before execution of a process for obtaining the latest result. [Figure 26B] FIG. 13 is a diagram illustrating an example of a process tree structure after execution of a process for obtaining the latest result. [Figure 27A] 13 is a flowchart illustrating details of a remuneration calculation process. [Figure 27B] 27B is a flowchart illustrating details of the calculation process of the influence degree in FIG. 27A. [Figure 27C] 27B is a flowchart illustrating details of the difficulty level calculation process in FIG. 27A. [Figure 28A] 13 is a diagram showing an example of a case where time series data is divided into groups based on a result of classification of the time series data by an analyst or the like through visual inspection. [Figure 28B] 28B is a flowchart illustrating a preparation process for classification in FIG. 28A. [Figure 28C] 28B is a flowchart illustrating a group information generating process in FIG. 28A. [Figure 29A] FIG. 13 is a diagram illustrating an example of a method for generating a feature amount table. [Figure 29B] FIG. 13 is a diagram illustrating an example of a method for generating a feature amount table. [Diagram 30] FIG. 13 is a diagram illustrating an example of description of feature amounts. [Diagram 31] 11 is an example of progress confirmation information. [Diagram 32] FIG. 13 is a diagram illustrating a case where a feature amount is searched for by trial and error. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. The present invention can be implemented in various other forms. The following description and drawings are merely examples for explaining the present invention, and are omitted and simplified as appropriate for clarity of explanation. Each component described below may be singular or plural unless otherwise specified.
[0016] In the following description, various types of information may be described using expressions such as "information," "data," "table," and the like, but the various types of information may be expressed using data structures other than these. When describing identification information, expressions such as "identifier" and "ID" are used, but these are interchangeable.
[0017] In the following description, the letter "S" before a reference symbol means a processing step. In the following description, the same reference symbol may be used for identical or similar configurations, and duplicate explanations may be omitted. In the following description, for convenience of explanation, different reference symbols may be used for information with the same content.
[0018] In the following, an information processing system (hereinafter referred to as "model construction support system 1") will be described that supports the construction of a model (hereinafter referred to as "prediction model") that predicts information (objective variables) related to events such as equipment failures, based on explanatory variables acquired from each piece of equipment operated in the field. Note that the type of prediction model is not necessarily limited, and may be a machine learning model such as a DNN (Deep Neural Network), or a rule-based model. It can also be something like that.
[0019] In the following, an example will be described in which the site is a data center or a system center, the equipment operated at the site is a storage drive (e.g., a flash module drive (FMD)), the explanatory variables are time series data obtained for the storage drive, and the prediction model is a model that predicts failures of the storage drive.
[0020] In the operation of storage drives in data centers and system centers, there are multiple failure modes (memory failure, deterioration, etc.) and operation modes. In addition, signs of failure may be hidden due to differences in product types and operation modes (such as an increase in frequency of use at the end of the month). In order to solve this problem, it may be effective to not only search for features that collectively include all explanatory variables (time-series data) of a large number of devices in operation, but also to divide the explanatory variables into multiple groups and search for features on a group-by-group basis. Therefore, the model building support system 1 provides information that is useful for searching for features that collectively include all explanatory variables (time-series data) of a large number of devices in operation, and also provides information on a method of dividing explanatory variables into multiple groups that is effective in improving the prediction accuracy of the objective variable (failure).
[0021] Figures 1A and 1B are diagrams specifically showing cases where it is necessary to consider dividing explanatory variables into groups in order to improve the prediction accuracy of storage drive failures. For example, as shown in Figure 1A, when a storage drive becomes blocked in the field, it can be due to "failure" or "strategic replacement." "Strategic replacement" here refers to cases where, for example, "storage drives with a degradation level of 95% or more are replaced even if they are not faulty," and "storage drives that make up the same RAID are replaced at that time." In this case, the explanatory variables for "failure" and "strategic replacement" are classified into different groups, and features are searched for separately for each group. By doing so, it is expected that the prediction accuracy of the objective variable will be improved.
[0022] As shown in the figure, causes of "failure" of a storage drive include "write failure", "read failure", "communication failure", "no cause description", etc. In this case, by classifying explanatory variables into groups according to causes and searching for features for each group individually, it is expected that the prediction accuracy of the objective variable will be improved. Also, for example, as shown in FIG. 1B, the calculation formula for the measurement value (e.g., the calculation formula for data capacity) may differ depending on the difference between the old and new models of storage drives. In this case, by classifying explanatory variables into groups for the old and new models and searching for features for each group individually, it is expected that the prediction accuracy of the objective variable will be improved.
[0023] 2 is a diagram showing an example of dividing explanatory variables into groups. In this example, when explanatory variables (time-series data) of all devices (storage drives) are not divided into groups, the prediction accuracy C of the objective variable (presence or absence of a failure sign) is 20%.
[0024] Division example 1 in the figure is a case where the explanatory variables of all devices are divided into group 1 and group 2. Here, the division is performed using a feature for performing the division (hereinafter referred to as a "group division feature"), and the accuracy D of the division using the group division feature is set to 100% in this example. Furthermore, due to this division, the coverage rate S (Support Ratio) of group 1 with respect to all explanatory variables is 30%, and the coverage rate S of group 2 is 70%. In this example, for group 1, the prediction accuracy C of the dependent variable using a prediction model based on the features (hereinafter referred to as "discrimination features") discovered for that group is 70%, and for group 2, the prediction accuracy C of the dependent variable based on the discrimination features discovered for that group is 40%.
[0025] Division example 2 in the figure is a case where the explanatory variables of all devices are divided into group 3 and group 4. Here, the division accuracy D by the group division features used for this division is 20%. Furthermore, due to this division, the coverage rate S of group 3 with respect to all explanatory variables is 20%, and the coverage rate S of group 4 is 80%. In this example, for group 3, the prediction accuracy C of the objective variable by the prediction model based on the discrimination features searched for that group is 80%, and for group 4, the prediction accuracy C of the objective variable by the prediction model based on the discrimination features searched for that group is 10%.
[0026] As described above, the model building support system 1 provides information useful for searching for discriminative features, and also provides information on a method for dividing explanatory variables into multiple groups, which is useful for improving the prediction accuracy of the objective variable. Therefore, analysts, domain experts, and service business designers who build predictive models can efficiently build prediction models with high prediction accuracy based on the provided information.
[0027] For example, the analyst considers improving the prediction accuracy of the objective variable by improving the discrimination feature based on the information provided by the model building support system 1. For example, in the example of Fig. 2, the analyst considers which discrimination feature should be improved, that of the whole or groups 1 to 4, whether the whole and groups 1 to 4 should be further divided into other groups, whether the discrimination feature used for division should be improved, etc.
[0028] Furthermore, for example, based on the information provided by the model building support system 1, the domain expert may give the analyst an insight or a hint about a new identification feature (e.g., remembering that the meaning of a variable differs depending on the manufacturer), provide a new division method (obtain fault diagnosis results from a repair center, visually classify charts (linear / quadratic curves, etc.), change the problem settings (exclude equipment that has been replaced due to reaching the end of its life from faulty equipment, etc.), etc.
[0029] Furthermore, for example, a designer of a service business can take action such as considering requesting the design department to take measures to address the root cause of a failure mode discovered based on the information provided by the model building support system 1, or changing the target of the predictive model (concentrating the target of failure identification only on new models for which failure identification can be performed reliably, and giving up on failure identification for older models, etc.).
[0030] The model building support system 1 supports the analyst or the like in constructing a predictive model by repeating trial and error of "division into groups" and "generation of discrimination features for each group" to obtain an appropriate objective variable. To this end, the model building support system 1 calculates a "score" when an explanatory variable is given to a predictive model constructed based on the groups and discrimination features for each group set in the trial and error process, and also calculates a "reward" that is information that serves as a guideline for improving the group and discrimination features in a more appropriate direction. Then, information (hereinafter referred to as "model building support information") is generated that lists the calculated "score" and "reward" in a graph that shows the groups, the type of division, and the features in a tree structure, and is provided to the analyst or the like.
[0031] In this example, the "score" is defined as follows: First, the score SF of the discrimination feature is defined by the following equation. [Number 1] SF=S*C Equation 1 In the above formula, S (Support Ratio) is the coverage ratio of a group to its higher-level group. C (Confidence) is the accuracy of the classification feature F (Feature) (for example, F in machine learning). value (F-measure).
[0032] Also, the score SG of group G is defined as follows: [Number 2] SG=max({SF in G},{SD*D in G}) ···Equation 2
[0033] The above SD is the score SD for the group division method, and is defined as follows: D is the accuracy of the group division feature. [Number 3] SD = sum({SG in D})...Equation 3
[0034] In this embodiment, the "reward" is defined by the following formula. [Number 4] Reward = Influence * Probability of Success Equation 4
[0035] In the above formula, the impact (the degree of improvement in the score of the whole group by improving the score of an individual group) is calculated by dividing the difference in the score of the whole group (difference due to improvement) by the difference in the score of the individual group (= difference in the score of the whole group / individual group) The success probability is the probability that a good feature can be discovered, and is obtained by, for example, the number of features that have been explored (registered) so far (if a large number of trial and error attempts have already been made, the success probability is low because it is assumed that everything has already been considered), the amount of measurement values used in previous attempts (coverage rate) (if the amount is large, the range that has already been verified is large and the success probability is low), etc.
[0036] 3A to 3C show examples of score calculation. FIG. 3A is a bar graph showing the ratio of faulty devices and normal devices, and FIGS. 3B and 3C are tree-structured graphs showing examples of score calculation for the bar graph of FIG. 3A. As shown in FIG. 3A, this example uses explanatory variables (time-series data) FIG. 3B shows a case where a classification feature F1 is found that provides a prediction accuracy of the objective variable of 20%, and the score SF of the classification feature F1 is 20%, the score SG of the group is 20%, and the score SD of the division method is 20%. FIG. 3C shows a case where a classification feature F2 is found that provides a prediction accuracy of the objective variable of 30%, and the score SF of the classification feature F2 is 30%, the score SG of the group is 30%, and the score SD of the division method is 30%.
[0037] 4A and 4B show other examples of score calculation. FIG. 4A is a bar graph showing the ratio of faulty devices and normal devices, and FIG. 4B is a tree structure graph showing an example of score calculation for the bar graph of FIG. 4A. As shown in FIG. 4A, this example is a case where a division D2 is performed on the explanatory variable (time series data) and the data is divided into two groups (group 1, group 2) at a ratio of 30% and 70%. As shown in FIG. 4B, for example, the scores of the discrimination features F21 and F22 of each group are 80% and 20%, and the group scores SG are 24% for group 1 and 14% for group 2, the score SD of the division D2 is 38%, and the score SG of group G11 is 38%.
[0038] Other calculation examples of the score are shown in FIG. 5A and FIG. 5B. FIG. 5A is a bar graph showing the ratio of faulty devices and normal devices, and FIG. 5B is a tree structure graph showing an example of calculating the score for the bar graph of FIG. 5A. As shown in FIG. 5A, this example is a case where the explanatory variable (time series data) is further divided D3 for the division D1 (group 1, group 2) shown in FIG. 4A, and the time series data of group 2 of division D1 is further divided into two groups (group 5, group 6) at a ratio of 50% and 50%. In this example, the score SD of division D3 is 50%, the score SD of division D2 is 61%, and the score SG of group G11 is 61%. In this example, it can be seen that improving the accuracy D (currently 20%) of the group division feature of division D4 increases the score SD of the entire division method, thereby increasing the influence and the reward.
[0039] FIG. 6 is an example of the model building support information 600 described above. As shown in the figure, the model building support information 600 is a tree structure graph that hierarchically represents the relationships between groups. The nodes that make up the graph include nodes that represent groups, nodes that represent divisions, nodes that represent each group belonging to the division, and nodes that represent the features adopted in each group. Among these, the node that represents the entire group displays the classification accuracy of the objective variable and the reward for the entire group. In addition, the node that represents the division displays the classification criteria (blocked items, by manufacturer, memory error-induced failure, etc.), the classification accuracy of the features used in the division, the classification accuracy of the objective variable for the division, etc. In addition, the node that represents the group displays the coverage rate of the explanatory variables (time-series data) of the original (higher) group, the classification accuracy of the objective variable by the group, and the reward. In addition, the node that represents the feature adopted in each group displays the content of the feature, the classification accuracy of the objective variable of the feature, etc.
[0040] In the case of the illustrated model building support information, for example, in the node group indicated by the reference numeral 611, multiple feature values have been tried in the group without division, and it can be seen that the improvement of the classification accuracy is sluggish. Therefore, the analyst can obtain a suggestion that, for example, feature values should not be searched for any further in the group without division.
[0041] Also, for example, when comparing the node indicated by reference numeral 612a with the node indicated by reference numeral 612b, it is found that the same feature amount can be significantly affected by the difference in manufacturer. Therefore, the analyst can know that the definition of the feature amount may differ depending on the manufacturer, and can obtain a suggestion that dividing the group by manufacturer may be effective in improving the recognition accuracy of the objective variable.
[0042] Also, for example, in the node indicated by the reference numeral 613, the classification accuracy of the objective variable is improved. Therefore, the analyst can obtain a suggestion that, for example, the classification accuracy of the objective variable may be improved by dividing the waveform. Also, the analyst can obtain a suggestion that, for example, the classification accuracy may be efficiently improved by applying software (logic) that mechanically classifies the waveform.
[0043] 3B, 3C, 4B, 5B, etc. may be presented (displayed) to an analyst or the like as model building support information 600. By referring to these graphs, an analyst or the like can visually and easily confirm the accuracy C, coverage rate S, score SF, score SG, score SD, etc. of the classification features, and can efficiently search for classification features and group division methods.
[0044] Next, a specific configuration of the model building support system 1 will be described in detail.
[0045] 7A is a system configuration diagram of a model building support system 1. As shown in the figure, the model building support system 1 includes a large number of devices 7, sensor devices 8, a user terminal 20, a data server 30, and a model building support device 100. These devices are communicatively connected via a communication network 5. The communication network 5 is, for example, a communication infrastructure (communication infrastructure) that realizes wireless or wired communication, and may be, for example, a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, a dedicated line, Various public communication networks, etc.
[0046] The devices 7 are a number of devices operated in the field, and in this example are storage drives (FMD (Flash Module Drive)), flash memory drives (SSD (Solid State Drive), HDD (Hard Disk Drive), etc.).
[0047] The sensor device 8 includes a communication module and various sensors, and transmits information about the device 7 (temperature, rotation speed, data read speed, data write speed, IOPS (Input / Output Per Second), response time, throughput, latency, remaining capacity, etc.) using the various sensors. The sensor 30 measures the temperature, generates time-series data based on the measured values, and transmits the generated time-series data to the data server 30. The various sensors include those realized by hardware such as a temperature sensor, as well as those realized by software such as a program that measures the data read speed and write speed.
[0048] The user terminal 2 is an information processing device (computer) that is communicatively connected to the sensor device 8, the data server 30, and the model building support device 100, and sets various information for these devices, acquires information to be provided, and presents it to the user, monitors it, controls it, etc.
[0049] The data server 30 is configured using an information processing device, and accumulates and manages (stores) various data such as time-series data sent from the sensor device 8.
[0050] The model building support device 100 is configured using an information processing device, and performs analysis of time-series data stored in the data server 30, extraction of features, construction of a prediction model, and provision of information to support construction of the prediction model.
[0051] FIG. 7B shows an example of the hardware configuration of an information processing device used to realize the user terminal 20, the data server 30, and the model building support device 100.
[0052] The illustrated information processing device 10 includes a processor 11, a main memory device 12, an auxiliary memory device 13, an input device 14, an output device 15, and a communication device 16. The information processing device 10 may be, for example, a personal computer, a smartphone, a tablet, an office computer, a general-purpose computer, or the like. The user terminal 20, the data server 30, and the model building support device 100 may be realized, for example, by using a plurality of information processing devices 10 connected to each other so as to be able to communicate with each other.
[0053] The information processing device 10 may be realized, in whole or in part, by using virtual information processing resources provided by using virtualization technology, process space separation technology, etc., such as a virtual server provided by a cloud system. All or in part of the functions provided by the information processing device 10 may be realized, for example, by a service provided by the cloud system via an API (Application Programming Interface) or the like.
[0054] All or part of the functions provided by the information processing device 10 may be realized by using, for example, Software as a Service (SaaS), Platform as a Service (PaaS), Infrastructure as a Service (IaaS), or the like.
[0055] The processor 11 shown in the figure is, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit), It is composed of AI (Artificial Intelligence) chips, etc.
[0056] The main memory device 12 is a device for storing programs and data, and is, for example, a ROM (Read Only Memory). It is composed of memory elements such as non-volatile memory (NVRAM (Non Volatile RAM)), random access memory (RAM), and non-volatile memory (NVRAM).
[0057] The auxiliary storage device 13 is, for example, a solid state drive (SSD) or a hard disk drive. The auxiliary storage device 13 may be a hard disk, an optical storage device (CD (Compact Disc), DVD (Digital Versatile Disc), etc.), a storage system, a reading / writing device for a recording medium such as an IC card, an SD card, or an optical recording medium, a storage area of a cloud server, etc. Programs and data can be read into the auxiliary storage device 13 via a recording medium reading device or a communication device 16. The programs and data stored (memorized) in the auxiliary storage device 13 are read into the main storage device 12 as needed.
[0058] In addition, all or part of the programs and data that realize the functions of the information processing device 10 may be stored in advance in the main memory device 12 or the auxiliary memory device 13, or, if necessary, may be read into the main memory device 12 or the auxiliary memory device 13 from a non-transitory recording medium or a non-transitory memory device provided in another device via a recording medium reading device or a communication device.
[0059] The input device 14 is an interface that accepts input from the outside, and is, for example, a keyboard, a mouse, a touch panel, a card reader, a pen-input tablet, a voice input device, or the like.
[0060] The output device 15 is an interface that outputs various information such as the process progress and the process result. The output device 15 is, for example, a display device (liquid crystal monitor, LCD (Liquid Crystal Display), graphic card, etc.) that visualizes the above-mentioned various information, a device that converts the above-mentioned various information into voice (voice output device (speaker, etc.)), and a device that converts the above-mentioned various information into text (printer, etc.). For example, the information processing device 10 may be configured to input and output information to and from other devices via the communication device 16.
[0061] The input device 14 and the output device 15 constitute a user interface that realizes interactive processing (receiving input of information, presenting information, etc.) with a user (a user or an administrator).
[0062] The communication device 16 is a device that realizes communication with other devices. The communication device 16 is a wired or wireless communication device that realizes communication with other devices via the communication network 5. The communication interface is, for example, a NIC (Network Interface Card), a wireless communication module, a USB module, and the like.
[0063] The information processing device 10 may be implemented with, for example, an operating system, a file system, a DBMS (DataBase Management System) (relational database, NoSQL, etc.), a KVS (Key-Value Store), etc.
[0064] The various functions of the user terminal 20, the data server 30, and the model building support device 100 are realized by the processor 11 included in each of them reading and executing a program stored in the main memory device 12, or by the hardware (FPGA, ASIC, AI chip, etc.) that constitutes each of them.
[0065] The user terminal 20, the data server 30, and the model building support device 100 store various types of information (data), for example, as tables in a database or files managed by a file system.
[0066] FIG. 8 is a diagram showing main information (data) handled in the model building support system 1. As shown in FIG.
[0067] As shown in the figure, the data server 30 manages (stores) time-series data 101 received from the sensor device 8.
[0068] The model building support device 100 stores time-series data 101 , a target variable 102 , group information 103 , a feature library 104 , a feature table 105 , a process tree 106 , and a device ID 107 .
[0069] 9 is a block diagram showing main functions of the model building support system 1. As shown in the figure, the model building support system 1 has the functions of an explanatory / objective variable setting unit 120, a verification processing unit 125, a feature quantity formula registration unit 130, a feature quantity registration unit 133, a feature quantity score calculation unit 135, a feature quantity setting unit 137, a group setting unit 138, a reward calculation unit 140, a model building support information generation unit 143, a group score calculation unit 150, a group registration unit 155, a progress confirmation unit 160, and a result acquisition unit 163.
[0070] Of the above functions, the explanation / objective variable setting unit 120 performs processes related to acquisition (reception) and setting (registration, editing, deletion, etc.) of the time-series data 101 and the objective variable 102.
[0071] The testing processing unit 125 performs processing related to testing of the prediction model (calculation of the discrimination accuracy of the objective variable, etc.).
[0072] The feature quantity formula registration unit 130 performs processing related to the setting (registration, editing, deletion, etc.) of the feature quantity library 104 (provision of a user interface for setting, etc.).
[0073] The feature registration unit 133 performs processing related to the registration of information related to the feature in the process tree 106 .
[0074] The feature amount score calculation unit 135 performs processing related to calculation of the scores of the feature amounts (discrimination feature amounts, group division feature amounts).
[0075] The feature amount setting unit 137 performs processing related to setting (registration, editing, deletion, etc.) of feature amounts (such as providing a user interface for setting).
[0076] The group setting unit 138 performs processing (such as providing a user interface for setting) related to setting (such as registering, editing, deleting, etc.) the group information 103, which is information related to a group.
[0077] The remuneration calculation unit 140 performs processing related to the calculation of remuneration.
[0078] The model building support information generation unit 143 performs processing related to the generation and presentation (display) of model building support information 600.
[0079] The group score calculation section 150 performs processing related to the calculation of the group score.
[0080] The group registration unit 155 performs processing related to the registration of information related to groups in the process tree 106 .
[0081] 10 shows an example of time series data 101. As shown in the figure, the illustrated time series data 101 is composed of a plurality of records each having items of a device ID 1011, a time stamp 1012, and a measurement value group 1013. One record of the time series data 101 corresponds to a measurement value measured by a sensor device 8 at a certain point in time (time stamp) for a certain device 7.
[0082] Of the above items, the device ID 1011 stores a device ID that is an identifier of the device 7. The time stamp 1012 stores information indicating the date and time when the measurement value was acquired. The measurement value group 1013 stores one or more types of measurement values.
[0083] 11A to 11C show examples of group information 103. In the group information 103, information related to groups is managed. In the group information 103a shown in FIG. 11A, the correspondence between a device 7 (device ID 1031a) and a group (GroupSet 1032a) to which the device belongs is managed. In the group information 103b shown in FIG. 11B, a target variable 1032b for each device (device ID 1031b) is managed. In the group information 103c shown in FIG. 11C, a failure reason 1032c for each device (device ID 1031c) is managed.
[0084] FIG. 12 shows an example of the feature library 104. The feature library 104 manages information related to features. As shown in the figure, the illustrated feature library 104 is composed of one or more records having the items of logic 1041, dimension 1042, and lambda expression 1043. One record of the feature library 104 corresponds to one feature (discrimination feature, group division feature). Of the above items, logic 1041 stores information indicating the type of logic of the feature. Dimension 1042 stores the dimension of the feature. Lambda expression 1043 stores information expressing the feature as a lambda expression.
[0085] 13 shows an example of the feature amount table 105. The feature amount table 105 stores information about the results of calculating features based on time-series data. As shown in the figure, the feature amount table 105 includes items of a device ID 1051 and a feature amount data group 1052. Of the above items, the device ID described above is stored in the device ID 1051. The feature amount data group 1052 stores values of one or more types of feature amounts.
[0086] FIG. 14A is a UML (Unified Modeling Language) diagram showing the structure of the process tree 106. 14A and 14B are diagrams showing attributes of components of the process tree. As shown in FIG. 14A and FIG. 14B, the process tree 106 includes a GroupSet 1061, a variable Group1062, group split feature (GFS)1062, and discriminant feature (DFS) 1065 is associated with a device ID 1063 and a feature amount ID 1066.
[0087] 15 is a flowchart illustrating the main processing (hereinafter referred to as "main processing S500") performed by the model building support device 100. Below, the main processing S500 will be described with reference to this figure.
[0088] First, the explanation / objective variable setting unit 120 performs processing for setting the time-series data 101 and the objective variable 102 (data registration processing S511, objective variable registration processing S512).
[0089] Next, the feature amount score calculation unit 135 calculates the feature amount score, and performs a process of inputting the calculated value to the feature amount registration unit 133 and the verification processing unit 125 (score calculation process S513).
[0090] Next, the remuneration calculation unit 140 performs a process of calculating the remuneration based on the process tree 106 (remuneration calculation process S514).
[0091] Next, the model building support information generating unit 143 performs a process of displaying the model building support information based on the calculated remuneration and presenting (displaying) it to the analyst (model building support information display process S515).
[0092] Next, the model building support device 100 waits for an operational input by the analyst (S520).
[0093] Here, for example, when the analyst performs an operation to display the progress check information, the progress check unit 160 performs a process (progress check information display process S521) of generating and displaying a screen on which the progress check information 3100 is described. After that, the process returns to S520.
[0094] Also, for example, when the analyst performs an operation to display the latest result, the result acquisition unit 163 performs a process of generating and displaying a display screen for the latest result (latest result acquisition process S522). After that, the process returns to S520.
[0095] Also, for example, when the analyst performs an operation to register a feature library, the feature registration unit 133 performs a process (feature library registration process S523) of displaying a feature library registration screen and accepting registration of the feature library from the analyst. After that, the process returns to S520.
[0096] For example, when an analyst performs a registration operation of a discrimination feature (DFS), the feature registration unit 133 However, the process (discrimination feature (DFS) registration process S524) is performed to display a registration screen for the discrimination feature (DFS) and to accept registration of the discrimination feature (DFS) from the analyst. After that, the process returns to S513. return.
[0097] Also, for example, when the analyst performs an operation to divide the time series data into groups, the group registration unit 155 performs a process of displaying a group division registration screen and accepting the registration of the group division from the analyst (group division process S525). After that, the process returns to S513.
[0098] For example, when an analyst performs an operation to register a group split feature (GFS), The feature registration unit 133 displays a registration screen for group split feature (GFS) and receives group split feature (GFS) from the analyst. A process for accepting registration of the loop division feature (GFS) (group division feature (GFS) registration process S526) is performed. After that, the process returns to S513.
[0099] For example, an analyst may register group split features (GFS) and perform group splitting. When this is done, the feature registration unit 133 displays a registration screen for group split feature (GFS) and The group registration unit 15 accepts registration of group split features (GFS) from analysts. 5 performs group division on the process tree 106 using the group division feature (GFS). Then, the process returns to S512.
[0100] Also, for example, when the analyst performs an operation to change the objective variable 102, the explanation / objective variable setting unit 120 performs a process (objective variable change process S528) of displaying a screen for changing the objective variable 102 and accepting the change to the objective variable 102. Thereafter, the process returns to S511.
[0101] Also, for example, when the analyst performs an operation to change the time-series data 101, the explanatory / objective variable setting unit 120 performs a process (time-series data change process S529) of displaying a screen for changing the time-series data 101 and accepting the change to the time-series data 101. Thereafter, the process returns to S511.
[0102] Fig. 16 is a list of operations (processing) performed on the process tree 106 in the main processing S500 shown in Fig. 15, and Fig. 17(a) to Fig. 17(h) are specific examples of changes to the process tree. In the figure, reference numeral 1611 indicates a reference numeral in the main processing S500. Operation content 1612 indicates the content of the operation on the process tree 106 corresponding to the reference numeral 1611. Example 1613 of process tree change is the number of a diagram (at least one of Figs. 17(a) to (h)) that explains the change to the process tree for the operation.
[0103] Fig. 18 is a flowchart for explaining details of the data registration process S511 shown in Fig. 15. The data registration process S511 will be explained below with reference to this figure.
[0104] First, the explanatory / objective variable setting unit 120 acquires time series data to be used as explanatory variables from the data server 30, and stores the acquired time series data as the time series data 101 (S5111).
[0105] Next, the explanatory / objective variable setting unit 120 extracts a list of device IDs from the time-series data 101 (S5112).
[0106] Next, the explanatory / objective variable setting unit 120 registers the variable (Group) and calculates the group score. Setting a value (100%) to the variable (Group.S) and the variable for the device ID belonging to the group The device ID acquired in S5112 is set in (Group.objects) (S5113).
[0107] Fig. 19A is a flowchart illustrating the details of the objective variable registration process S512 shown in Fig. 15. The objective variable registration process S512 will be described below with reference to this figure.
[0108] First, the explanation / objective variable setting section 120 acquires the objective variable 102 (S5121). An example of the objective variable 102 is shown in FIG.
[0109] Next, the explanation / objective variable setting unit 120 adds a variable (GroupSet) to the variable (Group). In addition, the column name of the objective variable is set to the variable (ID), and the accuracy of the classification feature is calculated. C (Confidence) is set to 100% (S5122).
[0110] Next, the explanation / objective variable setting section 120 sets the value (GroupSet.ID) to the variable (ObjectiveSetID) (S5123).
[0111] Next, the explanation / objective variable setting unit 120 lists the types X of the Objective columns (S512 4).
[0112] In the following steps S5125 to S5126, the explanation / objective variable setting unit 120 performs the following for all types X: The following processes are performed: registering a variable (Group) under a variable (GroupSet), setting type X to a variable (Group.ID), setting the number of device IDs of type X to a variable (Group.S), and setting a set of device IDs to a variable (Group.objects).
[0113] FIG. 19C shows an example of the structure of the process tree after the objective variable registration process S512 is executed.
[0114] FIG. 20 is a flow chart for explaining the details of the discrimination feature (DFS) registration process S524 shown in FIG. The discrimination feature (DFS) registration process S524 will be explained below with reference to the same figure. Reveal.
[0115] First, the feature registration unit 133 accepts the registration of the discrimination feature (DFS) from the analyst. The assigned discriminative feature (DFS) is added under the specified group variable (Group) (S5241).
[0116] Next, the feature registration unit 133 calculates the accuracy (DFS.DC) of the discrimination feature (DFS) (S5 242). Specifically, the feature registration unit 133 first acquires a higher-level group (Higher Group) of the specified group (S52421), and sets a group belonging to the higher-level Group to a variable (Inspection GroupSet) indicating a group to be inspected (S52422). Then, a value obtained by inspecting the Inspection GroupSet using the feature is set to DFS.DC (S52423).
[0117] Next, the feature registration unit 133 calculates the score (DFS.SF) of the discrimination feature (DFS) (S Specifically, the feature registration unit 133 first acquires a higher-ranking group (higher-ranking Group) of the specified group (S52431), and assigns DFS.DC to the score (Group.S) of the higher-ranking Group. The multiplied value is set as DFS.SF (S52432).
[0118] Next, the feature registration unit 133 updates DFS.DC and DFS.SF to the values calculated as above (S5244).
[0119] Fig. 21 is a flowchart for explaining details of the group division process S525 shown in Fig. 15. Below, the group division process S525 will be explained with reference to this figure.
[0120] First, the group registration unit 155 acquires the group information 103 (S5251).
[0121] Next, the group registration unit 155 adds a variable (GroupSet) under the specified group (Group) (S5252).
[0122] Next, the group registration unit 155 generates a group for each GroupSet (S5253).
[0123] The subsequent processing of S5254 to S5257 is a loop process in which all groups are selected in sequence. In this loop process, the group registration unit 155 registers the Group (S5255 ), and setting the list of device IDs classified into groups to the variable (Group.objects) (S52 56), and the grouped device IDs into a group score variable (Group.S). Set the ratio (S5257).
[0124] FIG. 22 illustrates the details of the group division feature (GFS) registration process S526 shown in FIG. This is a flowchart of the group split feature (GFS) registration process. Let me explain 526.
[0125] First, the feature registration unit 133 adds a group division feature (GFS) under the specified (GroupSet) (S5261).
[0126] Next, the feature registration unit 133 obtains GFS.GC (S5262). Specifically, the feature registration unit 133 first obtains the upper Group of the upper GroupSet of the variable (g1) (S526 21), and obtains all Groups subordinate to the GroupSet superior to the variable (g2s) (S52622). Then, the feature registration unit 133 classifies g2.objects from g1.objects using the group division features (GFS.features) (S52623), and sets the classification accuracy in the variable (GFS.GC) (S52624).
[0127] Fig. 23A is a flow chart for explaining the details of the GFS registration and group division processing S527 by the GFS shown in Fig. 15. Below, the GFS registration and group division processing S527 by the GFS will be explained with reference to this figure.
[0128] First, the feature registration unit 133 adds a group division feature (GFS) under the specified (GroupSet), and sets the variable (GFS.GC) to 100% (S5271).
[0129] Next, the feature registration unit 133 acquires the upper Group of the upper GroupSet of the variable (g1). (S5272), and classify g1.objects into groups using group splitting features (GFS.features) (S5273).
[0130] The subsequent processing of S5274 to S5277 is a loop process in which all groups are selected in sequence. In this loop process, the feature registration unit 133 registers Groups (S5275). , setting the list of device IDs classified into groups to the variable (Group.objects) (S527 6), and the ratio of grouped device IDs to the group score variable (Group.S). Set the rate (S5277).
[0131] FIG. 23B shows an example of the structure of the process tree 106 after the GFS registration and the execution of the group division process S527 by the GFS.
[0132] Fig. 24A is a flowchart illustrating details of the score calculation process S513 shown in Fig. 15. Below, the score calculation process S513 will be described with reference to this figure.
[0133] In the figure, S5131 and S5132 are loop processes for sequentially selecting nodes up to the root for all leaves of the process tree 106. In the loop process, first, the feature score calculation unit 135 determines whether the selected node is a group. If the selected node is a group (S5133: YES), a process for selecting the best child of the group is performed (hereinafter referred to as "process for selecting the best child of the group S5134"). On the other hand, if the selected node is not a group (S5133: NO), the process proceeds to S5136.
[0134] FIG. 24B is a flowchart illustrating the details of the selection process S5134 of the best child of the group in FIG. 24A.
[0135] First, the feature amount score calculation unit 135 secures a storage area for set X (S51341).
[0136] S51342 to S51343 are for all the discrimination features (DFS) under the group. This is a loop process in which selection is performed sequentially. In S51343, the feature score calculation unit 135 sets the selected discrimination feature (DFS) to a variable (Y.Item), sets DFS.FS to a variable (Y.Score), and adds Y to the set X.
[0137] S51344 to S51346 select all GroupSets under the Group in question in sequence. This is a loop process in which Y is added to set X. In S51345, the feature score calculation unit 135 sets the maximum GFS.GC among all group division features (GFS) of GroupSet to a variable (maxGC). In S51346, the feature score calculation unit 135 sets GroupSet to a variable (Y.Item), sets GroupSet.SD*maxGC to a variable (Y.Score), and adds Y to set X.
[0138] In the next step S51347, the feature score calculation unit 135 selects the feature with the highest score from the set X. Get the Item and Score as the return value.
[0139] Returning to FIG. 24A, in S5135, feature amount score calculation section 135 sets the Score of the best child, which is the return value of the selection process of the best child of the group S5134, to a variable (Group.SG).
[0140] In S5136, the feature amount score calculation unit 135 determines whether or not the selected node is a GroupSet. If the selected node is a GroupSet (S5136: YES), the total of the scores SG of all Groups under the GroupSet is set to a variable (GroupSet.SD).
[0141] Fig. 25 is a flowchart for explaining details of the latest result acquisition process S522 shown in Fig. 15. The latest result acquisition process S522 will be explained below with reference to this figure.
[0142] First, the result acquisition unit 163 sets Group.SG of the root group (Group) to a variable (best identification result) (S5221).
[0143] Next, the result obtaining unit 163 duplicates the process tree 106 (S5222).
[0144] The processing from S5223 to S5225 is a loop process that selects nodes in order from the root of the process tree 106 to all the leaves. First, in S5224, the result acquisition unit 163 determines whether or not the selected node is a group. If the selected node is a group (S5224: YES), the above-mentioned selection process S5134 of the best child of the group in FIG. 24B is performed, and the discrimination features (DFS) or GroupSets other than the best child are deleted (S5225). ).
[0145] FIG. 26A shows an example of the structure of the process tree 106 before the latest result acquisition process S522 is executed, and FIG. 26B shows an example of the structure of the process tree 106 after the latest result acquisition process S522 is executed.
[0146] Figure 27A is a flowchart illustrating the details of the remuneration calculation process S514 shown in Figure 15. Below, the remuneration calculation process S514 will be described with reference to this figure.
[0147] First, the remuneration calculation unit 140 performs a calculation process of the influence degree S5241.
[0148] FIG. 27B is a flowchart illustrating details of the influence degree calculation process S5241.
[0149] First, the reward calculation unit 140 sets the variable (Score_all_B) to Group.SG of the root group (S52411).
[0150] The process from S51412 to S51416 in the figure is to This is a loop process in which the reward calculation unit 140 sequentially selects Group.SG in the variable (Score_B) (S51413). Next, the reward calculation unit 140 sets the variable (Score_A) to Score_B + constant X (S15414). Next, the reward calculation unit 140 performs the score calculation process S513 shown in FIG. 24A. Next, the reward calculation unit 140 sets the variable (Score_all_A ) is set to Group.SG of the root group (Group) (S51415). Then, the remuneration calculation unit 140 sets the variable (DC influence degree) to (Score_all_A - Score_all_B) / (Score_A - Score_B) (S51416).
[0151] The process from S51417 to S514102 in the figure is a loop process of sequentially selecting GroupSet from all groups (GroupSet). First, the reward calculation unit 140 sets max(GFS.GC for all GFS) to a variable (Score_B) (S51418). The reward calculation unit 140 also sets Score_B + constant X to a variable (Score_A) (S51419). Next, the reward calculation unit 140 performs the score calculation process S513 shown in FIG. 24A. Next, the reward calculation unit 140 sets the variable (Score_all_A) to Group.SG of the root group (S51420). 14101). Next, the reward calculation unit 140 sets the variable (GC influence degree) to (Score_all_A - Score_all_B) / (Score_A - Score_B) (S51416).
[0152] Returning to FIG. 27A, the reward calculation unit 140 then performs difficulty level calculation processing S5242.
[0153] FIG. 27C is a flowchart for explaining the details of the difficulty level calculation process S5242. The processes from S51421 to S51426 in the figure are a loop process for sequentially selecting the discrimination features (DFS) of groups and children of all groups from all groups. The output unit 140 sets the number of discrimination features (DFS) in the variable (A) (S51423). Next, the remuneration calculation unit 140 sorts the top N DFS.DC of the discrimination features (DFS) in ascending order in the variable (B). The reward calculation unit 140 sets the gradient of DFS.DC as the result of the test (S15424). Next, the reward calculation unit 140 sets the coverage rate of the features used in the discrimination features (DFS) of all children (for example, the coverage rate understood from the history of trial and error of the discrimination features (DFS)) to the variable (C) (S51425). Then, the reward calculation unit 140 sets the variable (difficulty level) to the weighted sum of A, B, and C (S5141 6).
[0154] The process from S51427 to S514203 is a loop process that sequentially selects the group split feature (GFS) of the group and the children of all groups from all groups. The calculation unit 140 sets the number of group split features (GFS) to the variable (A) (S51429). Next, the reward calculation unit 140 sets the gradient of the GFS.GC, which is the result of sorting the top N GFS.GC of the group split features (GFS) in ascending order, to the variable (B) (S154201). Next, The reward calculation unit 140 sets the coverage rate of the features used in the group splitting features (GFS) of all children to the variable (C) (S514202). Next, the reward calculation unit 140 sets the weighted sum of A, B, and C to the variable (difficulty) (S514203).
[0155] Returning to FIG. 27A, the remuneration calculation unit 140 then calculates the Group. influence degree for all groups. * Calculate the Group. difficulty and find the group's reward (Group. reward) (S5145).
[0156] Next, the reward calculation unit 140 calculates GroupSet.influence*GroupSet.difficulty for all groups (GroupSet) to obtain the reward (GroupSet.reward) for the group (GroupSet) (S5146).
[0157] 28A to 28C are diagrams for explaining a case where the group setting unit 138 shown in FIG. 9 divides the time series data into groups based on the result of classification of the time series data by an analyst or the like through visual inspection. As shown in FIG. 28A, in this example, the analyst or the like visually judges the waveform. Based on the result, the time series data is classified into either a parabola or a straight line, and group information 103c is generated.
[0158] FIG. 28B illustrates a case where the group setting unit 138 outputs files for all devices (Group.objects) belonging to a group when performing the operation of reference numeral 2810 shown in FIG. 28A. 20 is a flowchart showing how the group setting unit 138 creates files with the device ID set in the file name for all devices belonging to the group (S2811 to S2813).
[0159] Fig. 28C is a flowchart for explaining the process by which group setting unit 138 generates group information 103c shown in Fig. 28A. The processes of S2821 to S2824 are loop processes for sequentially selecting all files in all subdirectories of a specified directory. In the above loop process, group setting unit 138 sets the selected file name to the device ID (S2823), and generates records of the device ID and subdirectory (S2824). In S2825, group setting unit 138 outputs the generated record group as group information 103c.
[0160] 29A and 29B are diagrams showing an example of a method for generating the feature table 105. Fig. 29A shows a case where the feature is the slope (=(max(y)-min(y)) / (max(x)-min(x))) of the time series data in the section from time min(X) to time max(X). Note that as the period of the time series data used to calculate the feature, for example, in order to detect a sign of a failure, the most recent predetermined period (six days from B to C in the figure) is not used, as shown in Fig. 29B.
[0161] FIG. 30 shows an example of a feature quantity described in a predetermined algorithm description language.
[0162] 31 is an example of information presented to an analyst or the like by the progress confirmation unit 160 of the model building support device 100. The progress confirmation unit 160 manages information related to the status of trial and error (progress of the work of building a prediction model) of the analyst or the like's search for features and division of explanatory variables (time-series data) into groups based on the history of processing performed on the process tree 106. In the progress confirmation information display process S521 described above, the progress confirmation unit 160 generates a graph shown in the figure (hereinafter referred to as "progress confirmation information 3100") based on the above information, and presents (displays) the generated progress confirmation information 3100 to the analyst or the like.
[0163] As shown in the figure, the progress check information 3100 lists the device IDs of each device by group along the horizontal axis, and lists various features acquired for each device along the vertical axis. In addition, a predetermined color (black in the figure) is displayed in each cell located at the intersection of the device ID and the feature, with a density (or color) according to the magnitude of the difference between the feature of the device ID and the average value of the entire device. In the progress check information 3100, the more dense the areas with high density are, the higher the contribution of the feature to improving the identification accuracy of the objective variable, and the more areas there are, the more appropriately the search for the feature and grouping is progressing. Therefore, by referring to the progress check information 3100, analysts and the like can visually and easily check the degree of progress in the search for the feature and grouping for building a prediction model.
[0164] FIG. 32 shows an example of a case where a search for an appropriate feature quantity is performed by trial and error. In trial 1, a judgment (identification) of the objective variable (presence or absence of a fault) is performed by a threshold value (th1). In this case, In trial 2, the threshold value used to judge the second time series data from the left was adjusted to th2, which resulted in an error in the judgment of the second and fourth time series data from the left. The second time series data from the left is judged correctly, but the fourth time series data from the left is judged correctly. The judgment remains incorrect. In trial 3, the four time series data were normalized and then the threshold (th1 ) is applied, but the judgment of the fourth time series data from the left remains incorrect. In trial 4h, the four time series data are normalized and judgment is made based on whether the magnitude of the slope exceeds the threshold. In this example, the correct judgment is made for all the time series data. In this example, the feature and grouping scores for trial 4 are the maximum.
[0165] As described above, according to the model building support system 1 of the present embodiment, various information useful for supporting the search for the feature to be adopted in the prediction model that outputs the objective variable related to the event predicted for the equipment, and the method of dividing the explanatory variables into groups that optimizes the feature is provided from various viewpoints. Therefore, analysts and the like can efficiently build a prediction model for accurately predicting the events that occur in the equipment, such as failures.
[0166] It goes without saying that the present invention is not limited to the above-mentioned embodiment, and various modifications are possible without departing from the spirit of the present invention. For example, the above-mentioned embodiment has been described in detail to explain the present invention in an easy-to-understand manner, and the present invention is not necessarily limited to those having all of the described configurations. In addition, it is possible to add, delete, or replace part of the configuration of the above-mentioned embodiment with other configurations.
[0167] For example, in the above embodiments, the risk prediction model is configured using a linear regression model, but the risk prediction model may also be configured using, for example, other types of statistical models or machine learning models (e.g., DNN (Deep Neural Network)).
[0168] The above-mentioned configurations, functional units, processing units, processing means, etc. may be realized in part or in whole by hardware, for example, by designing them as integrated circuits. Also, the above-mentioned configurations, functions, etc. may be realized in software by a processor interpreting and executing a program that realizes each function. Information such as the programs, tables, and files that realize each function may be stored in a memory, a hard disk, a recording device such as an SSD (Solid State Drive), an IC card, etc. It can be placed on a storage medium such as a notebook, SD card, or DVD.
[0169] The above-described layouts of the various functional units, processing units, and databases of each information processing device are merely examples. The layouts of the various functional units, processing units, and databases can be changed to optimal layouts in terms of the performance, processing efficiency, communication efficiency, and the like of the hardware and software that these devices are equipped with.
[0170] The configuration (schema, etc.) of the database that stores the various types of data described above can be flexibly changed from the viewpoints of efficient use of resources, improved processing efficiency, improved access efficiency, improved search efficiency, and the like. [Explanation of symbols]
[0171] 1 Model construction support system, 5 Communication network, 7 Equipment, 8 Sensor device, 10 Information processing device, 11 processor, 12 main memory device, 13 auxiliary memory device, 14 input device, 15 output device, 16 communication device, 20 user terminal, 30 data server, 100 model building support device, 101 time series data, 102 objective variable, 103 group information, 104 feature library, 105 feature table, 106 process tree, 107 device ID, 120 explanation / objective variable setting unit, 125 test processing unit, 130 feature expression registration unit, 133 feature registration unit, 135 feature score calculation unit, 137 feature setting unit, 138 group setting unit, 140 reward calculation unit, 143 model building support information generation unit, 150 group score calculation unit, 155 group registration unit, 160 progress confirmation unit, 163 Result acquisition unit, 600 model construction support information, S500 main processing, S511 data Registration process, S512 Objective variable registration process, S513 Score calculation process, S514 Reward calculation process, S516 Model building support information display process, S521 Progress check information display process, S522 Latest result acquisition process, S523 Feature library registration process, S524 Discrimination feature (DFS) registration process, S525 Group division process, S526 Group division feature (GFS) Registration process, S527 GFS registration and group division process by GFS, S528 Purpose Variable change processing, S529 Time series data change processing
Claims
1. 1. An information processing system that supports a search for a discrimination feature, which is a feature used to construct a prediction model that outputs a response variable related to an event predicted for an equipment based on input explanatory variables, and a method of dividing the explanatory variables into groups that improves the prediction accuracy of the response variable by the prediction model, comprising: The information processing device includes a processor and a memory device. dividing the explanatory variables into a plurality of groups, and calculating a prediction accuracy C of the objective variable of the discriminant feature when the discriminant feature is set based on the explanatory variables of each of the groups; calculating a score SF of the discrimination feature for each of the groups based on the prediction accuracy C of the discrimination feature and a coverage rate S which is a ratio of the explanatory variables of each of the groups to the entire explanatory variables before division, and generating information based on the calculated score SF; Model building support system.
2. 2. A model building support system according to claim 1, Calculating the accuracy D of division of group division features which are features used in dividing the explanatory variables into the groups; Calculating a score SG for each of the groups based on the score SF and the accuracy D for each of the groups, and generating information based on the calculated score SG. Model building support system.
3. 3. A model building support system according to claim 2, Calculating a score SD, which is a score for the division method, by aggregating the scores SG, and generating information about the score SD. Model building support system.
4. 4. A model building support system according to claim 3, generating information based on at least one of the scores SF, SG, and SD, along with information on a graph that represents the groups, the division method, and the features in a tree structure; Model building support system.
5. 4. A model building support system according to claim 3, Calculating a reward, which is a guideline value for improving the discrimination feature and the group division feature, based on an influence level, which is a value indicating the degree of influence of the score SG or the score SF used in the aggregation of the score SD on the score SD; Model building support system.
6. 6. A model building support system according to claim 5, storing information indicating a search history of the discrimination feature; The reward reflects a success probability, which is a probability that the discrimination feature that improves the prediction accuracy C will be searched for in the future, obtained based on the history. Model building support system.
7. 7. A model building support system according to claim 6, The success probability is calculated based on the number of the features previously searched for or the amount of the explanatory variables previously used in searching for the features. Model building support system.
8. 6. A model building support system according to claim 5, Generate and output information based on at least one of the prediction accuracy C, the coverage rate S, the score SF, the score SG, the score SD, and the reward, along with information on a graph that represents the group, the type of division, and the feature amount in a tree structure. Model building support system.
9. 2. A model building support system according to claim 1, A plurality of devices are listed by group along the horizontal axis, and the feature amounts obtained for the devices are listed along the vertical axis, and information is generated in which each cell located at the intersection of the device and the feature amount is displayed with a density or color according to the magnitude of the difference between the feature amount of the device and the average value of all the devices. Model building support system.
10. 1. An information processing method by an information processing system for supporting a search for a discrimination feature, which is a feature used in constructing a prediction model that outputs a response variable related to an event predicted for an equipment based on input explanatory variables, and a method of dividing the explanatory variables into groups that improves a prediction accuracy of the response variable by the prediction model, comprising: An information processing device having a processor and a memory element, a step of dividing the explanatory variables into a plurality of groups, and calculating a prediction accuracy C of the objective variable of the discriminant feature when the discriminant feature is set based on the explanatory variables of each of the groups; calculating a score SF of the discriminant feature for each group based on the prediction accuracy C of the discriminant feature and a coverage rate S which is a ratio of the explanatory variables of each group to the entire explanatory variables before division, and generating information based on the calculated score SF; A model building support method for carrying out the above.
11. The model building support method according to claim 10, The information processing device, A step of calculating a division accuracy D of a group division feature which is a feature used for dividing the explanatory variables into the groups; calculating a score SG of each of the groups based on the score SF and the accuracy D of each of the groups, and generating information based on the calculated score SG; The model building support method further comprises:
12. The model building support method according to claim 11, The information processing device, A step of calculating a score SD, which is a score for the division method, by aggregating the scores SG, and generating information about the score SD. The model building support method further comprises:
13. The model building support method according to claim 12, generating information based on at least one of the scores SF, SG, and SD in a graph that represents the groups, the division method, and the features in a tree structure by the information processing device; The model building support method further comprises:
14. The model building support method according to claim 12, A step in which the information processing device calculates a reward, which is a value serving as a guideline for improving the discrimination feature and the group division feature, based on an influence level, which is a value indicating the degree of influence of the score SG or the score SF used in the aggregation of the score SD on the score SD. The model building support method further comprises:
15. The model building support method according to claim 14, The information processing device, storing information indicating a search history of the discrimination feature; A step of reflecting a success probability, which is a probability that the discrimination feature that improves the prediction accuracy C will be searched for in the future, based on the history, in the reward; The model building support method further comprises:
Citation Information
Patent Citations
Statistical machine translation device
JP2009294747A
Data mining system, data mining method and data mining program
JP2011034457A
Method, system, and program for recommending analysis axis in data analysis
JP2012103841A
Cluster division evaluation device, cluster division evaluation method and cluster division evaluation program
JP2020154825A
Comparison and selection of experiment designs
US10318674B2
Cited By
Global explanations of machine learning model predictions for input containing text attributes
US12670320B2
Global explanations of machine learning model predictions for input containing text attributes
US20240232526A1