Information Processing Apparatus, Information Processing Method, and Program
The information processing apparatus addresses the challenge of identifying variable influences by calculating and categorizing influence degrees and selection frequencies, ensuring comprehensive identification of influential factors in data-rich environments.
Patent Information
- Application Number
- JP2022036391
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-09
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-03-09
AI Technical Summary
Existing methods for identifying the influence of multiple variables on an output in data-rich environments, such as semiconductor factories and chemical plants, face challenges in accurately ranking variables due to temporary influences or steady but low-impact factors, leading to potential oversight of important factors.
The information processing apparatus calculates the influence degree and selection frequency of each explanatory variable, associating these metrics to categorize variables into four categories: high influence and high selection frequency, high influence but low selection frequency, low influence but high selection frequency, and low influence and low selection frequency, allowing for more effective identification of influential variables.
This approach enables users to easily identify the influence of multiple variables on the output by categorizing them based on their influence degree and selection frequency, thereby reducing the likelihood of overlooking important factors.
Smart Images

Figure 0007693587000016 
Figure 0007693587000017 
Figure 0007693587000018
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to an information processing apparatus, an information processing method, and a program.
Background Art
[0002] For example, in semiconductor factories and chemical plants, various types of products are mass-produced, and a large amount of data can be obtained from sensors and the like installed in each manufacturing process. Further, by analyzing the accumulated data, it has become possible to take measures to suppress variations in quality characteristics. For example, various efforts are being made every day to improve productivity and yield based on the analysis results. As one of the countermeasures, regression analysis using statistics and machine learning is often used. By using data such as sensor values, set values, and device information as explanatory variables of the regression model and using quality characteristics as the target variable, it becomes possible to analyze the causes of variations in quality characteristics.
[0003] In factor analysis using a regression model, highly interpretable models such as linear models, decision trees, and additive models are used. For each explanatory variable, an index representing the degree of influence on the output data (target variable) of the model, such as a regression coefficient and importance, is calculated, and by using this index, it becomes possible to identify factors (explanatory variables) that can explain variations in quality characteristics.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] The problem to be solved by the present invention is to provide an information processing apparatus, an information processing method, and a program that can more easily identify the influence of a plurality of variables on an output.
Means for Solving the Problem
[0006] The information processing apparatus according to the embodiment includes a calculation unit and an output control unit. The calculation unit is a model estimated using, for each of K periods (K is an integer of 2 or more), a plurality of input data including a plurality of variables, and is based on K first models that input input data including a plurality of variables and output output data, and calculates a first influence degree of the plurality of variables on the output data and the frequency with which the plurality of variables are selected as variables that affect the output data. The output control unit outputs the first influence degree and the frequency in association with each other.
Brief Description of the Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Best Mode for Carrying Out the Invention
[0008] Hereinafter, with reference to the accompanying drawings, a preferred embodiment of an information processing apparatus according to the present invention will be described in detail.
[0009] As described above, by using an index indicating the degree of influence of each explanatory variable on the output data of the model, it becomes possible to identify factors (explanatory variables) that can explain the variation in quality characteristics. However, when the number of explanatory variables is large, it is practically impossible to analyze all variables one by one, so it is necessary to narrow down the explanatory variables to be confirmed. One method is to use a regression model with a penalty term. This makes it possible to estimate (construct) a regression model consisting of a small number of important explanatory variables. As another method, there is also a method of estimating a model consisting of a specified number of explanatory variables by sequentially including in the model an explanatory variable having a high correlation with the objective variable or an explanatory variable that improves the accuracy of the model. Conversely, it is also possible to sequentially exclude explanatory variables until the specified number is reached.
[0010] Data in semiconductor factories and chemical plants, etc., often have many explanatory variables and their trends change from moment to moment. In order to always grasp the latest trends, periodic model updates using the latest data are required. At this time, if only the latest data is used, since the number of data is relatively small, the influence of noise appears stronger and it may be difficult to identify the factors. Therefore, in order to perform more accurate factor analysis, it is necessary to grasp not only the latest trends but also medium- and long-term trends. That is, it is necessary to perform an integrated analysis including the models estimated in the past.
[0011] In integrated evaluation, ranking using the average value of the influence degree of each explanatory variable is often used. It is judged that the explanatory variables at the top of the ranking are important factors. The average value of the influence degree is calculated by, for example, the following two methods. (M1) A method of calculating the average value of the influence degree over the entire period. (M2) A method of using the influence degree of an explanatory variable only for the period selected as a variable that affects the output of the model to calculate an average value. For example, when a regression model with a penalty term is used, only some of the explanatory variables are selected by the model in each period. For each explanatory variable, the influence degree of the explanatory variable is used to calculate the average value only for the period selected by the model.
[0012] The above two methods each have the following drawbacks. Regarding (M1): When there is an explanatory variable that has a temporarily strong influence due to a sudden failure or the like, the influence degree of this explanatory variable becomes zero in many periods, so the average value becomes small. As a result, even though the urgency is high, it is difficult to appear in the upper ranks of the ranking. Regarding (M2): An explanatory variable that has a temporarily strong influence is overestimated, and an explanatory variable that always has an influence but has a small influence degree value is relatively underestimated. Such an explanatory variable is difficult to appear in the upper ranks of the ranking even though it has a stable influence.
[0013] In this way, no matter which method is used to calculate the average value of the influence degree for ranking, there is a possibility of overlooking factors.
[0014] In addition to the above two methods, for example, there is also a method of comprehensively evaluating from the frequency of the period when the influence degree is equal to or greater than a threshold value. However, in this method, only an explanatory variable with both a large influence degree and a large frequency is identified as a factor, so there is a possibility of overlooking factors.
[0015] Therefore, the information processing apparatus according to each of the following embodiments calculates, for each explanatory variable, the influence degree in the period selected as a variable that affects the output data and the selected frequency (hereinafter, selection frequency), and associates and displays the calculated influence degree and the selection frequency. For example, the explanatory variables can be classified and displayed in the following four categories. The user can extract and identify appropriate factors without omission while referring to the characteristics of each category. That is, it is possible to more easily identify the influence of a plurality of variables on the output. (C1) Variables with both high influence and high selection frequency: Variables that have a consistently high impact. (C2) Variables with high influence but low selection frequency: Variables that have a high degree of suddenness as they have a temporary impact. (C3) Variables with low influence but high selection frequency: Variables that have a low influence but have a steady impact. (C4) Variables with both low influence and low selection frequency: Variables that can be excluded from the list of factor candidates.
[0016] (First Embodiment) FIG. 1 is a block diagram showing an example of the configuration of an information processing system including the information processing apparatus of the present embodiment. As shown in FIG. 1, the information processing system has a configuration in which an information processing apparatus 100 and a management system 200 are connected via a network 300.
[0017] Each of the information processing apparatus 100 and the management system 200 can be configured as, for example, a server apparatus. The information processing apparatus 100 and the management system 200 may be realized as a plurality of physically independent apparatuses (systems), or the respective functions may be configured within one physical apparatus. In the latter case, the network 300 may not be provided. At least one of the information processing apparatus 100 and the management system 200 may be constructed on a cloud environment.
[0018] The network 300 is a network such as, for example, a LAN (Local Area Network) and the Internet. The network 300 may be either a wired network or a wireless network. The information processing apparatus 100 and the management system 200 may transmit and receive data using a direct wired connection or wireless connection between components without going through the network 300.
[0019] The management system 200 is a system that manages the models processed by the information processing apparatus 100 and the data used for learning (estimation) and analysis of the models. The management system 200 includes a storage unit 221 and a communication control unit 201.
[0020] The storage unit 221 stores various information used in various processes executed by the management system 200. For example, the storage unit 221 stores input data used for model estimation. The storage unit 221 can be composed of any commonly used storage media such as flash memory, memory cards, RAM (Random Access Memory), HDD (Hard Disk Drive), and optical disks.
[0021] The model inputs explanatory variables and outputs the inference result of the target variable. The model is, for example, a linear regression model, a polynomial regression model, a logistic regression model, a Poisson regression model, a generalized linear model, a generalized additive model, a decision tree, and a neural network model. The model is not limited to these, and any model can be used as long as it is a model expressed using parameters.
[0022] The model is estimated by learning using input data including the target variable and explanatory variables. The target variable is, for example, a quality characteristic, a defect rate, and information indicating either a good product or a defective product. The explanatory variables are other sensor values, set values such as processing conditions, and control values.
[0023] The communication control unit 201 controls communication with external devices such as the information processing device 100. For example, the communication control unit 201 transmits input data to the information processing device 100.
[0024] Each of the above units (communication control unit 201) is realized by, for example, one or a plurality of processors. For example, each of the above units may be realized by causing a processor such as a CPU (Central Processing Unit) to execute a program, that is, by software. Each of the above units may be realized by a processor such as a dedicated IC (Integrated Circuit), that is, by hardware. Each of the above units may be realized by using software and hardware in combination. When using a plurality of processors, each processor may realize one of the units or may realize two or more of the units.
[0025] The information processing apparatus 100 includes a storage unit 121, an input device 122, a display 123, a communication control unit 101, a reception unit 102, a model estimation unit 103, a calculation unit 104, and an output control unit 105.
[0026] The storage unit 121 stores various information used in various processes executed by the information processing apparatus 100. For example, the storage unit 121 stores information (input data, etc.) acquired from the management system 200 via the communication control unit 101 and the reception unit 102, model parameters estimated by the model estimation unit 103, and information calculated by the calculation unit 104. The storage unit 121 can be configured by any generally used storage medium such as a flash memory, a memory card, a RAM, an HDD, and an optical disk.
[0027] The input device 122 is a device for inputting information by a user or the like. The input device 122 is, for example, a keyboard and a mouse. The display 123 is an example of an output device that outputs information and is, for example, a liquid crystal display. The input device 122 and the display 123 may be integrated, for example, like a touch panel.
[0028] The communication control unit 101 controls communication with an external device such as the management system 200. For example, the communication control unit 101 receives input data and the like from the management system 200.
[0029] The reception unit 102 receives the input of various information. For example, the reception unit 102 receives a plurality of input data received from the management system 200 via the communication control unit 201 and the communication control unit 101. The plurality of input data are, for example, a plurality of data obtained in K mutually different periods (data periods) (K is an integer of 2 or more). Further, each of the plurality of input data includes at least a plurality of variables (explanatory variables) that are inputs to the model.
[0030] The K periods may be predetermined, or values specified by a user or the like may be used. Further, the period may be determined based on the accuracy of the model estimated by the model estimation unit 103.
[0031] The reception unit 102, for example, requests the management system 200 to transmit data for a specified (determined) period, and receives the input data transmitted from the management system 200 in response to the request. The reception unit 102 or the model estimation unit 103 may be configured to extract the input data for the specified period from the plurality of input data received from the management system 200.
[0032] The model estimation unit 103 estimates a plurality of models using the plurality of input data. For example, for each of the K periods, the model estimation unit 103 estimates a model (first model) that inputs the input data and outputs output data using the plurality of input data obtained within the period. As a result, K models corresponding to the K periods are respectively estimated.
[0033] In addition, when a plurality of models estimated in the past are obtained, the model estimation unit 103 may newly estimate a model only for a period for which the model has not been estimated (for example, the latest period). For example, the model estimation unit 103 may estimate the model for the latest period using a plurality of models estimated in the past and stored in the storage unit 121 or the like.
[0034] The calculation unit 104 calculates an index related to an explanatory variable that affects the output data (target variable) of the model using the estimated model. For example, the calculation unit 104 calculates the frequency (selection frequency) and the influence degree (first influence degree) as indices. The selection frequency represents the frequency with which a plurality of explanatory variables are selected as variables that affect the output data among K periods. The influence degree represents the degree to which a plurality of explanatory variables affect the output data.
[0035] The output control unit 105 controls the output of various information processed by the information processing apparatus 100. For example, the output control unit 105 associates the selection frequency and the influence degree calculated by the calculation unit 104 and displays them on the display 123.
[0036] The output control unit 105 may output information to a device external to the information processing apparatus 100. For example, the output control unit 105 may transmit information for displaying the selection frequency and the influence degree in association with each other to an external device including a display device.
[0037] Each of the above units (communication control unit 101, reception unit 102, model estimation unit 103, calculation unit 104, output control unit 105) is realized by, for example, one or a plurality of processors. For example, each of the above units may be realized by causing a processor such as a CPU to execute a program, that is, by software. Each of the above units may be realized by a processor such as a dedicated IC, that is, by hardware. Each of the above units may be realized by using a combination of software and hardware. When using a plurality of processors, each processor may realize one of the units or two or more of the units.
[0038] Note that in this embodiment, the model is estimated within the information processing apparatus 100, but the model may be estimated by a device external to the information processing apparatus 100. In this case, the information processing apparatus 100 may not include a function (such as the model estimation unit 103) used for estimating the model.
[0039] Next, the model estimation process and visualization process by the information processing apparatus 100 according to the first embodiment configured as described above will be described. FIG. 2 is a flowchart showing an example of the estimation process and visualization process in the first embodiment.
[0040] The reception unit 102 receives a plurality of input data corresponding to a plurality of periods from the management system 200 (step S101). The model estimation unit 103 estimates a model using the plurality of input data acquired during this period for each of the plurality of periods (step S102). Here, it is assumed that the model estimation unit 103 estimates a regression model for each period.
[0041] The calculation unit 104 calculates the selection frequency of each of the plurality of explanatory variables and the influence degree of each of the plurality of explanatory variables using the estimated model (step S103). The output control unit 105 associates the calculated selection frequency and influence degree and displays them on, for example, the display 123 (step S104), and ends the estimation process and the visualization process.
[0042] Next, the details of the estimation process and the visualization process will be further described. Hereinafter, examples of the estimation process of a model applied to quality control in a factory (semiconductor factory) and a plant (chemical plant) and the visualization process based on the estimated model will be mainly described.
[0043] In semiconductor factories and chemical plants, it is required to improve the yield by suppressing variations and fluctuations in quality characteristics and reducing defects. And in order to clarify the factors of variations and fluctuations in quality characteristics, models such as regression models and classification models are used.
[0044] Products are completed through many manufacturing processes. When analyzing the factors of variations in the quality characteristics of the finished products, the model is estimated using information such as the type of manufacturing equipment in each manufacturing process and the sensor values detected by the installed sensors as explanatory variables.
[0045] In addition, since the manufacturing equipment deteriorates over time, the trend of the acquired data also gradually changes. Furthermore, operations that have a drastic impact on the data trend, such as regular maintenance and parts replacement, are also carried out. Therefore, the model is updated in accordance with the change in the data trend.
[0046] As described above, in the model for each period of the present embodiment, the target variable is information indicating, for example, quality characteristics, defect rate, and good product / defective product. The explanatory variables are, for example, other sensor values, set values, and control values. The date and time are the manufacturing start date and time, the manufacturing completion date and time, and the processing date and time at a specific device.
[0047] The explanatory variables may be pre-processed. The pre-processing is, for example, standardization, normalization, conversion by a specific function, addition of interaction terms, time lag, time lead, dummy variable conversion, encoding, outlier processing, and missing value processing.
[0048] The input data including data such as the target variable and the explanatory variables is stored in the storage unit 221 of the management system 200. The reception unit 102 receives the input of the input data received from the management system 200 via the communication control unit 101.
[0049] Hereinafter, it is assumed that the number of input data is n (n is an integer of 1 or more), and each input data includes p explanatory variables x, 1 target variable y, and a numerical value t representing 1 date and time. The i-th (1 ≤ i ≤ n) input data (x i , y i , t i ) is represented by the following formula (1).
Equation
[0050] x i represents a p-dimensional vector of explanatory variables, y i represents a scalar of the target variable, and t i represents a scalar of the date and time. t iThe length of time (number of days, hours, minutes, seconds, etc.) counted from any date and time may be used. Here, for simplicity of notation, let 0 = t1 ≤ t2 ≤ ··· ≤ t n = T. The starting date and time can be determined in any way. Also, if the times are not in order, they may be sorted in advance.
[0051] The indices of the K periods and the K models estimated in the K periods are represented by k = 1, ···, K in chronological order. Let the time point at which each model was estimated be t k (1 ≤ k ≤ K). The model estimation unit 103 estimates the model in period k using the input data Dk = {(x i , y i , t i ) | t k―1 < t i ≤ t k}. For example, in the case of a linear regression model, the model estimation unit 103 estimates the regression model by solving the optimization problem represented by the following equation (2) in each period. β0 is a one-dimensional vector, and ^β (k) represents a p-dimensional vector. The symbol “^” represents the hat attached to the right variable (β (k) in this example). The “T” in β T represents transpose. As a result, K regression models ^β (k) are estimated.
Equation
[0052] The method for estimating the model is not limited to the method using the least squares method as in equation (2), and any method may be used. For example, penalty regression such as Ridge, Lasso, SCAD (Smoothly Clipped Absolute Derivation), MCP (Minimax Concave Penalty), Lq (0 ≤ q < 1) norm, Elastic Net, L1 / 2 norm, etc. may be used. These penalty regressions can be interpreted as methods for estimating the model so that the parameters have sparsity.
[0053] Also, the loss function is not limited to the mean squared error, and any function may be used. For example, among the absolute value loss, quantile loss, Huber loss, epsilon-insensitive loss, logistic loss, exponential loss, hinge loss, and smoothed hinge loss, a loss function applicable to the estimation method of the adopted model may be used.
[0054] Also, the model estimator 103 may use a loss function weighted according to the reliability and time of each input data.
[0055] Also, the model to be estimated is not limited to the linear regression model, and may be a polynomial regression model, logistic regression model, Poisson regression model, generalized linear model, generalized additive model, decision tree, and neural network model, etc.
[0056] The calculator 104 can obtain a coefficient matrix ^β including the regression coefficients (p) of the models for the entire period (K periods) from the K models estimated as described above. (All) The coefficient matrix ^β (All) is represented by, for example, the following equation (3). ^β (All) In this case, zero is set for the elements without information on the corresponding regression coefficients due to the narrowing down of the explanatory variables.
Equation
[0057] Hereinafter, β ~ is the standardized regression coefficient, and m = 1, ···, p is the index of the explanatory variable. The calculator 104 calculates the standardized regression coefficient β ~ by, for example, the following equation (4). Also, the calculator 104 calculates the influence degree e m and the selection frequency g m of the m-th explanatory variable by the following equations (5) and (6), respectively.
Equation
[0058] (5) The value of the denominator of the formula and the numerator of the formula (6) corresponds to the number of non-zero regression coefficients among the K regression coefficients corresponding to the m-th explanatory variable. The fact that the regression coefficient is not zero can be interpreted as the explanatory variable being selected as a variable that affects the output of the model. Therefore, it can be interpreted that the value of the denominator of the formula (5) and the numerator of the formula (6) corresponds to the number of times the m-th explanatory variable is selected as a variable that affects the output data.
[0059] Also, the influence degree e calculated by the formula (5) m can be interpreted as corresponding to the average value of the influence degrees (second influence degrees) for each period in which the m-th explanatory variable is selected as a variable used in the estimation of the regression model. Note that the calculation method of the influence degree e m is not limited to the formula (5). For example, the calculation unit 104 may calculate the median or maximum value of the value of the numerator of the formula (5) as the influence degree e m . Also, when identifying factors by focusing on the more recent influence degree, the calculation unit 104 may calculate the value represented by the following formula (7) as the influence degree e m . [Number]
[0060] Also, the calculation unit 104 may use the value represented by the following formula (8) or formula (9) instead of the value of the numerator of the formula (5). [Number] [Number]
[0061] Also, the calculation unit 104 may use the standardized regression coefficient β~ Alternatively, regression coefficients without standardization may be used. The regression coefficient corresponds to information representing the contribution degree of an explanatory variable to the regression model. As information representing the contribution degree of an explanatory variable, information other than the regression coefficient according to the model to be applied may be used. For example, the importance obtained by a decision tree or the weights of a neural network may be used.
[0062] Also, the calculation method of the selection frequency g m is not limited to the formula (6). For example, the calculation unit 104 may calculate, as the selection frequency, the number of times the m-th explanatory variable is selected, that is, the value corresponding to the numerator of the formula (6).
[0063] The output control unit 105 displays the calculated influence degree and selection frequency in association with each other. For example, the output control unit 105 arranges the influence degree on the vertical axis (an example of the first axis) and arranges the selection frequency on the horizontal axis (an example of the second axis), and displays the influence degree and the selection frequency in association with each other by a two-dimensional matrix diagram.
[0064] FIG. 3 is a diagram showing an example of the displayed matrix diagram. In FIG. 3, an example is shown in which black circles corresponding to seven explanatory variables (flow rate, voltage, tank pressure, concentration, resistance, device temperature, and air temperature) are arranged at positions corresponding to the influence degree and selection frequency of each explanatory variable.
[0065] Also in FIG. 3, an example is shown in which the display target area of the matrix diagram is divided into four areas (upper right, upper left, lower right, and lower left) and displayed. The four areas of upper right, upper left, lower right, and lower left correspond to the areas where the explanatory variables classified into the above (C1), (C2), (C3), and (C4) are arranged, respectively.
[0066] In the above calculation method (M1) of the average value of the influence degree, the value corresponding to e m ×g m in the present embodiment is calculated as the average value of the influence degree. In such a method, it is easy to overlook factors that suddenly have a large impact on the quality characteristics (the explanatory variables classified into the upper left in the example of FIG. 3).
[0067] In contrast, the information processing apparatus 100 of the present embodiment calculates an index for identifying a factor into an influence degree e m and a selection frequency g m and outputs a matrix diagram showing the relationship between the influence degree and the selection frequency. As a result, for example, even a factor that has a sudden and large impact on a quality characteristic can be identified without omission from the output matrix diagram. Whether it is sudden or not, the factors that affect the quality characteristic (output data) are explicitly represented by the influence degree e m for this reason.
[0068] On the other hand, regardless of the magnitude of the influence degree, factors with a high selection frequency are explicitly represented by the selection frequency g m Therefore, for example, even an explanatory variable with a low influence degree but a high selection frequency (such as the explanatory variable classified in the lower right in the example of FIG. 3) can be identified without omission.
[0069] Note that the output methods of the influence degree and the selection frequency are not limited to a matrix diagram such as FIG. 3. For example, a matrix diagram with the influence degree arranged on the horizontal axis and the selection frequency arranged on the vertical axis may be used. Also, instead of the matrix diagram such as FIG. 3, other forms of output information such as a graph, a scatter diagram, and a table may be used.
[0070] Also, when the value represented by equation (9) is used in the numerator of equation (5), when the influence degree can take a negative value, the influence degree may be arranged using an axis indicating both the plus direction and the minus direction. FIG. 4 is a diagram showing an example of a matrix diagram in such a case.
[0071] By using the information classified and displayed in a plurality of categories in this way, the user can extract without omission the factors that affect the quality characteristic (output data). Even when configured to update a regular model using the latest data, the factors can be extracted and identified without omission based on the influence degree and the selection frequency calculated and displayed using the updated model.
[0072] As an example, the process of identifying the factors causing variations in quality characteristics for a device installed outdoors will be described.
[0073] When identifying the factors in a state where the operation of the device is stable, for example, it is considered that the lower right region of FIG. 3 should be preferentially checked. In the case of FIG. 3, since the temperature is in the lower right region, the relationship between the temperature and the variations in quality characteristics is preferentially analyzed. In contrast, when identifying the factors causing the sudden stop of the device, it is considered that the upper left region of FIG. 3 should be preferentially checked. In the case of FIG. 3, since the flow rate and voltage are in this region, for example, the flow rate control components and the power supply are preferentially analyzed.
[0074] Also, according to the present embodiment, for example, not only during the period of sudden stop, but also analysis can be performed using models for a plurality of periods (K periods). Therefore, even in the case of a past sudden stop, it is possible to comprehensively analyze the influence in a plurality of periods including the preceding and subsequent periods and identify the factors.
[0075] (Second Embodiment) The factors causing variations in quality characteristics may include both factors with relatively high urgency for countermeasures and factors with low urgency. For example, if there is an explanatory variable that had a low influence in the past but has suddenly become highly influential recently, due to the high urgency, it is required to take countermeasures preferentially.
[0076] Therefore, the information processing apparatus according to the second embodiment further calculates and outputs information indicating the change in the influence degree for each explanatory variable in order to show the tendency of the influence degree such as the urgency.
[0077] FIG. 5 is a block diagram showing an example of the configuration of an information processing system including the information processing apparatus 100-2 according to the second embodiment. Since the management system 200 and the network 300 are the same as those in the first embodiment, the same reference numerals are given and the description thereof is omitted.
[0078] As shown in FIG. 5, the information processing apparatus 100-2 includes a storage unit 121, an input device 122, a display 123, a communication control unit 101, a reception unit 102, a model estimation unit 103, a calculation unit 104-2, and an output control unit 105-2.
[0079] In the second embodiment, the functions of the calculation unit 104-2 and the output control unit 105-2 are different from those in the first embodiment. Since the other configurations and functions are the same as those in FIG. 1, which is a block diagram of the information processing apparatus 100 according to the first embodiment, the same reference numerals are used and the description thereof is omitted here.
[0080] The calculation unit 104-2 further calculates change information indicating a change in the influence degree in at least two of the K periods. The two periods are, for example, the latest period K and the period (K-1) immediately before the period K. For example, for the m-th explanatory variable, the calculation unit 104-2 calculates the difference d m in the influence degree between the period K and the immediately preceding period (K-1) by the following formula (10).
Equation
[0081] The calculation unit 104-2 may further calculate a comparison result between the difference d m and a threshold value (for example, zero). The difference, or the comparison result between the difference and the threshold value, corresponds to the change information.
[0082] The output control unit 105-2 further outputs the calculated change information. For example, the output control unit 105-2 visualizes arrows with different directions as markers according to the comparison result between the difference d m and the threshold value. For example, when d m >0, the output control unit 105-2 visualizes an upward arrow, and when d m ≦0, a downward arrow as a marker for the explanatory variable. The arrows are examples of change information indicating mutually different directions.
[0083] FIG. 6 is a diagram showing an example of a matrix diagram including such markers. In FIG. 3, black circles corresponding to each explanatory variable were displayed, but in the example of FIG. 6, arrows are displayed as markers instead of black circles.
[0084] The method for calculating the change information is not limited to the above formula (10). For example, when representing the change in trend using the magnitude of the change in the influence degree between period K and period (K - 1), the calculation unit 104-2 calculates the difference d of the influence degree m may be calculated by the following formula (11). sign is calculated by the following formula (12). [Number] [Number]
[0085] Also in this case, the output control unit 105-2 distinguishes the expression of the marker according to the comparison result between the difference d m and the threshold value in the same manner as above.
[0086] Further, when the value represented by formula (9) is used as the numerator of the influence degree e m the calculation unit 104-2 may calculate the difference d of the influence degree m by the following formula (13). [Number]
[0087] Also in this case, the output control unit 105-2 distinguishes the expression of the marker according to the comparison result between the difference d m and the threshold value in the same manner as above. FIG. 7 is a diagram showing an example of the matrix diagram displayed in this case.
[0088] The output method of the change information (marker) is not limited to the above example. Any method may be used as long as it outputs the change information indicating an increase in the influence degree and the change information indicating a decrease in the influence degree in different manners.
[0089] For example, a method may be used in which the shape of the marker is fixed to a simple figure other than an arrow, and the tendency of whether the influence degree has increased or decreased is indicated by color-coding the figure. Note that the method in FIG. 6 is a method in which the color of the marker is fixed and the tendency is indicated by the direction of the arrow corresponding to the shape of the marker.
[0090] In the above example, the threshold value for comparison with the difference d m is set to zero, but a real number other than zero may be used as the threshold value. Also, the output control unit 105-2 may determine at least one of the color and the length of the arrow according to the magnitude of the difference d m .
[0091] Also, the output control unit 105-2 may not calculate the difference d m , and use, as change information, which model of the two periods the m-th explanatory variable is selected in. For example, the output control unit 105-2 may output change information indicating that the m-th explanatory variable is selected in the L-th period (where L is an integer satisfying 1 ≦ L < K) but not in the (L + 1)-th period, and change information indicating that the m-th explanatory variable is not selected in the L-th period but is selected in the (L + 1)-th period.
[0092] Thus, the information processing apparatus according to the second embodiment can further output information indicating a change in the influence degree.
[0093] (Third Embodiment) In the second embodiment, it was described that the parameters of the current model and the model of the immediately preceding period are compared, and the tendency of the change in the influence degree is expressed by the shape or color of the marker. In the third embodiment, an example of outputting the long-term tendency of the influence degree as change information will be described.
[0094] FIG. 8 is a block diagram showing an example of the configuration of the information processing apparatus 100-3 according to the third embodiment. Since the management system 200 and the network 300 are the same as those in the first embodiment, the same reference numerals are given and the description is omitted.
[0095] As shown in FIG. 8, the information processing apparatus 100-3 includes a storage unit 121, an input device 122, a display 123, a communication control unit 101, a reception unit 102-3, a model estimation unit 103-3, a calculation unit 104-3, and an output control unit 105-3.
[0096] In the third embodiment, the functions of the reception unit 102-3, the model estimation unit 103-3, the calculation unit 104-3, and the output control unit 105-3 are different from those in the first embodiment. Other configurations and functions are the same as those in FIG. 1, which is a block diagram of the information processing apparatus 100 according to the first embodiment. Therefore, the same reference numerals are used, and the description thereof is omitted here.
[0097] The reception unit 102-3 further receives a setting of a period for calculating a long-term trend of the influence degree and a setting of a function used by the model estimation unit 103-3.
[0098] The calculation unit 104-3 calculates change information indicating a change in the influence degree in at least two periods set by the user and received by the reception unit 102-3 among the K periods. The period corresponding to the end may be fixed to the latest period K, and only the oldest period corresponding to the start may be set, or two periods corresponding to the start and the end may be set. Hereinafter, the case where only the oldest period is set will be described as an example. For example, the oldest period k is set such that k < K - 1. The setting may use either the index of the period (model) or the date.
[0099] When the date t is set, the calculation unit 104-3 calculates the index τ of the oldest period (model) by the following equation (14).
Equation
[0100] The model estimation unit 103-3 estimates a one-variable regression model using the model of each set period. For example, the model estimation unit 103-3 estimates the coefficient matrix ^β (All)Using this, the regression coefficient ^β of the model in the period set for each explanatory variable (k) m From this, a univariate regression model is estimated by the following formula (15). Here, θ represents the parameter that determines the regression function f m which represents the parameter that determines the regression function f [Equation]
[0101] What regression function f m is used can be set by the user or the like and may be received by the reception unit 102-3. The regression function f m is, for example, linear regression, n(>1)th order approximation, generalized linear regression, exponential function, spline function, and Gaussian process regression, etc.
[0102] Equation (15) is an equation for estimating the regression function using the squared error as the loss function, but the loss function is not limited to this. For example, as the loss function, absolute value loss, quantile loss, Huber loss, epsilon sensitivity loss, logistic loss, exponential loss, hinge loss, and smoothed hinge loss, etc. may be used.
[0103] The calculation unit 104-3 calculates the change information using the regression function estimated in this way. For example, when linear regression is used as the regression function f m the calculation unit 104-3 calculates the slope of the regression function f m as an index indicating the increase or decrease of the influence degree. When a twice-differentiable function such as nth order approximation is used as the regression function f m the calculation unit 104-3 calculates the first derivative and the second derivative in the regression coefficient ^β (k) m as an index indicating the increase or decrease of the influence degree.
[0104] The calculation unit 104-3 may further calculate the comparison result between the index and a threshold value (for example, zero). The index, or the comparison result between the index and the threshold value, corresponds to the change information.
[0105] The output control unit 105-3 further outputs the calculated change information. For example, the output control unit 105-2 visualizes arrows with different directions as markers according to the comparison result between an index (such as slope, first derivative, second derivative, etc.) and a threshold value. For example, when the index > 0, the output control unit 105-3 visualizes an upward arrow, and when the index ≤ 0, a downward arrow, as the marker of the explanatory variable.
[0106] In this way, the information processing apparatus according to the third embodiment can output information indicating the change in the influence degree over a longer period.
[0107] (Fourth Embodiment) In the first embodiment, an example of dividing the display target area into four areas and displaying a matrix diagram has been described. The information processing apparatus according to the fourth embodiment enables adjustment of the division position of the area. For example, the division position of the area is adjusted according to at least one of the settings by the user or the like and the history of past countermeasures.
[0108] When factor analysis is performed regularly, various explanatory variables appear as factor candidates each time, and countermeasures are executed for some of the factors considering priorities and the like. In quality control, the history of executed countermeasures is stored and managed in a database or the like. Therefore, in this embodiment, for example, the user refers to the past history, determines the values of the influence degree and selection frequency corresponding to the position where the area is divided, and sets parameters (area division parameters) for dividing the area. Further, the information processing apparatus of this embodiment calculates and sets the area division parameters by referring to the past history.
[0109] FIG. 9 is a block diagram showing an example of the configuration of the information processing apparatus 100-4 according to the fourth embodiment. Since the management system 200 and the network 300 are the same as those in the first embodiment, the same reference numerals are given and the description is omitted.
[0110] As shown in FIG. 9, the information processing apparatus 100-4 includes a storage unit 121-4, an input device 122, a display 123, a communication control unit 101, a reception unit 102-4, a model estimation unit 103, a calculation unit 104-4, and an output control unit 105-4.
[0111] In the fourth embodiment, the functions of the storage unit 121-4, the reception unit 102-4, the calculation unit 104-4, and the output control unit 105-4 are different from those in the first embodiment. Other configurations and functions are the same as those in FIG. 1, which is a block diagram of the information processing apparatus 100 according to the first embodiment. Therefore, the same reference numerals are used, and the description here is omitted.
[0112] The storage unit 121-4 further stores information regarding the history of countermeasures acquired from, for example, the management system 200. The history of countermeasures includes, for example, an explanatory variable that is the target of countermeasures to suppress variations in quality characteristics, and values of the influence degree and selection frequency calculated for the explanatory variable when the countermeasures are taken. The history of countermeasures may further include information indicating the period or date when the countermeasures were taken.
[0113] The reception unit 102-4 further receives the setting of region division parameters specified by a user or the like. The region division parameters are, for example, a reference value (first reference value) indicating the division position of the influence degree and a reference value (second reference value) indicating the division position of the selection frequency. The setting method may be any method. For example, a method of setting each reference value by using slide bars provided in the vertical and horizontal axis directions of a matrix diagram can be used.
[0114] The information processing apparatus 100-4 may be configured to calculate the region division parameters by referring to the history of countermeasures instead of or together with the user's setting. The calculation unit 104-4 further has a function for such a configuration. That is, the calculation unit 104-4 further has a function of calculating the region division parameters by referring to the history of countermeasures.
[0115] For example, the calculation unit 104-4 reads out the influence degree and selection frequency included in the history of the explanatory variables to be analyzed from the storage unit 121-4, and calculates the average value of the read influence degree and the average value of the selection frequency as region division parameters (first reference value, second reference value). Instead of the average value, the calculation unit 104-4 may calculate the median, maximum value, minimum value, quantiles, etc. as the region division parameters. The calculation unit 104-4 may calculate a reference value for only one of the influence degree and the selection frequency.
[0116] When the information indicating the date or period when the countermeasure was executed for the history is included, the calculation unit 104-4 may refer to this information and read out the history of the countermeasures within the period to be analyzed.
[0117] The output control unit 105-4 outputs a matrix diagram in which the region is divided according to the region division parameters set by the user or calculated by the calculation unit 104-4.
[0118] FIG. 10 is a diagram showing an example of the matrix diagram displayed in the present embodiment. In FIG. 10, an example is shown in which 0.7 is set as the region division parameter of the influence degree and 0.2 is set as the region division parameter of the selection frequency.
[0119] As the region division parameters, two values of the maximum value and the minimum value may be used. FIG. 11 is a diagram showing an example of the matrix diagram in such a case. By dividing the region with the two values of the maximum value and the minimum value, it becomes possible to show the region where the countermeasure was taken in the past (the region between the maximum value and the minimum value).
[0120] As described above, in the fourth embodiment, the division position of the display target region can be adjusted according to the history of past countermeasures and the like.
[0121] As described above, according to the first to fourth embodiments, it is possible to more easily identify the influence of a plurality of variables on the output.
[0122] Next, the hardware configuration of the information processing apparatus according to the first to fourth embodiments will be described with reference to FIG. 12. FIG. 12 is an explanatory diagram showing a hardware configuration example of the information processing apparatus according to the first to fourth embodiments.
[0123] The information processing apparatus according to the first to fourth embodiments includes a control device such as a CPU 51, a storage device such as a ROM (Read Only Memory) 52 and a RAM 53, a communication I / F 54 that connects to a network and performs communication, and a bus 61 that connects each part.
[0124] The program executed by the information processing apparatus according to the first to fourth embodiments is provided by being pre-embedded in the ROM 52 or the like.
[0125] The program executed by the information processing apparatus according to the first to fourth embodiments may be configured to be recorded on a computer-readable recording medium such as a CD-ROM (Compact Disk Read Only Memory), a flexible disk (FD), a CD-R (Compact Disk Recordable), or a DVD (Digital Versatile Disk) in an installable or executable file format and provided as a computer program product.
[0126] Furthermore, the program executed by the information processing apparatus according to the first to fourth embodiments may be configured to be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network. Also, the program executed by the information processing apparatus according to the first to fourth embodiments may be configured to be provided or distributed via a network such as the Internet.
[0127] The program executed by the information processing apparatus according to the first to fourth embodiments can cause a computer to function as each part of the above-described information processing apparatus. This computer can read a program from a computer-readable storage medium into the main storage device and execute it by the CPU 51.
[0128] Although some embodiments of the present invention have been described, these embodiments are presented by way of example and are not intended to limit the scope of the invention. These novel embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention, and are also included in the invention described in the claims and its equivalent scope.
Explanation of Reference Numerals
[0129] 100, 100-2, 100-3, 100-4 Information Processing Apparatus 101 Communication Control Unit 102, 102-3, 102-4 Reception Unit 103, 103-3 Model Estimation Unit 104, 104-2, 104-3, 104-4 Calculation Unit 105, 105-2, 105-3, 105-4 Output Control Unit 121 Storage Unit 122 Input Device 123 Display 200 Management System 201 Communication Control Unit 221 Storage Unit 300 Network
Claims
1. A model estimated using a plurality of input data including a plurality of variables obtained in K (K is an integer of 2 or more) periods, respectively, based on K first models that input the input data including the plurality of variables and output output data, a calculation unit that calculates a first influence degree of the plurality of variables on the output data and a frequency at which the plurality of variables are selected as variables that affect the output data; an output control unit that associates and outputs the first influence degree and the frequency; The output control unit arranges the first influence degree on a first axis and the frequency on a second axis, and outputs a matrix diagram including a plurality of regions divided by a first reference value of the first influence degree and a second reference value of the frequency; The calculation unit calculates at least one of the first reference value and the second reference value based on a history of processing executed on variables selected as variables that affect the output data among the plurality of variables. An information processing apparatus.
2. The calculation unit calculates the first influence degree, which is any one of an average value, a median value, and a maximum value of a second influence degree of the plurality of variables on the output data in each of one or more periods among the K periods in which the plurality of variables are selected as variables that affect the output data. The information processing apparatus according to claim 1.
3. The history includes the first influence degree and the frequency calculated for the variables targeted by the processing, The calculation unit calculates any one of an average value, a median value, a maximum value, a minimum value, and a quantile of the first influence degree included in one or more of the histories as the first reference value, and calculates any one of an average value, a median value, a maximum value, a minimum value, and a quantile of the frequency included in one or more of the histories as the second reference value. The information processing apparatus according to claim 1.
4. The output control unit further outputs change information indicating changes in the first influence degree in at least two of the K periods. The information processing apparatus according to claim 1.
5. The output control unit outputs the change information indicating that the first influence degree has increased and the change information indicating that the first influence degree has decreased in different manners. The information processing apparatus according to claim 4.
6. The output control unit outputs the change information indicating that the first influence degree has increased and the change information indicating that the first influence degree has decreased as information indicating mutually different directions. The information processing apparatus according to claim 5.
7. The output control unit outputs the change information indicating that the first influence degree has increased and the change information indicating that the first influence degree has decreased in mutually different colors. The information processing apparatus according to claim 5.
8. The output control unit further outputs information indicating that any one of the plurality of variables is selected in the L-th period (L is an integer satisfying 1 ≤ L < K) but not selected in the (L + 1)-th period, and information indicating that the variable is not selected in the L-th period but is selected in the (L + 1)-th period. The information processing apparatus according to claim 1.
9. An information processing method executed by an information processing apparatus, Based on K first models (K is an integer of 2 or more) each of which is a model estimated using a plurality of input data including a plurality of variables obtained in K periods, and which input the input data including the plurality of variables and output output data, a calculation step of calculating a first influence degree of the plurality of variables on the output data and a frequency at which the plurality of variables are selected as variables that affect the output data, An output control step of associating and outputting the first influence degree and the frequency. The output control step arranges the first influence degree on the first axis, arranges the frequency on the second axis, and outputs a matrix diagram including a plurality of regions divided by a first reference value of the first influence degree and a second reference value of the frequency. The method further includes a step of calculating at least one of the first reference value and the second reference value based on a history of processing executed on a variable selected as a variable that affects the output data among the plurality of variables. Information processing method.
10. causing a computer to a model estimated using a plurality of input data including a plurality of variables obtained in K (K is an integer of 2 or more) periods, based on K first models that input input data including the plurality of variables and output output data, a calculation step of calculating a first influence degree of the plurality of variables on the output data and a frequency at which the plurality of variables are selected as variables that affect the output data; an output control step of associating and outputting the first influence degree and the frequency; The output control step arranges the first influence degree on the first axis, arranges the frequency on the second axis, and outputs a matrix diagram including a plurality of regions divided by a first reference value of the first influence degree and a second reference value of the frequency. causing the computer to further execute a step of calculating at least one of the first reference value and the second reference value based on a history of processing executed on a variable selected as a variable that affects the output data among the plurality of variables. Program.
Citation Information
Patent Citations
Method for creating equipment model, program for constructing equipment model, method for creating building model, program for creating building model, method for evaluating equipment repair, and method for evaluating building repair
JP2005157829A
Equipment maintenance and management support system, and method for the same
JP2014016691A
Accident prevention management system
JP2018049349A
Factor analysis apparatus and factor analysis method
JP2021149727A