Evaluation method

The method addresses the challenge of verifying the validity of machine learning model inference results by extracting singular data based on prediction errors and using explanatory AI for factor analysis, resulting in efficient and effective model evaluation.

JP2025074701APending Publication Date: 2025-05-14SCREEN HOLDINGS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023185702
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-30
Publication Date
2025-05-14

AI Technical Summary

Technical Problem

Existing methods struggle to efficiently verify the validity of inference results produced by machine learning models, particularly when dealing with large volumes of input and output data from industrial devices.

Method used

A method that involves extracting singular data from the inference results using prediction errors and performing factor analysis with explanatory AI, such as SHAP, to evaluate the validity of the machine learning model.

Benefits of technology

This approach allows for efficient confirmation of the validity of machine learning models by reducing the data to be verified and facilitating easy factor analysis, thereby improving the evaluation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025074701000001_ABST
    Figure 2025074701000001_ABST
Patent Text Reader

Abstract

To provide a technique for efficiently checking the validity of a machine learning model that outputs inference results from sensor data.SOLUTION: An evaluation method for evaluating the validity of inference results output by a machine learning model includes: Steps S11-S13 of extracting anomalous data from the inference results output by the machine learning model; and Steps S21-S22 of performing factor analysis of the machine learning model on the anomalous data. Accordingly, the validity of the machine learning model can be efficiently checked without performing factor analysis on all inference results.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a technique for evaluating the validity of inference results output by a machine learning model. [Background technology]

[0002] In recent years, predictive maintenance, anomaly detection, and failure cause analysis have been carried out for various industrial equipment by predicting and estimating the equipment status using various sensor data and machine learning. When applying machine learning to industrial equipment in this way, there is an issue that the basis for judgment and the relationship between input and output are unclear because many machine learning models are black boxes.

[0003] To address these issues, a technology called explainable AI is used to interpret how an inference result was derived. For example, Patent Document 1 discloses a factor analysis device equipped with a factor calculation unit and a factor visualization unit. This factor analysis device makes it possible to visually grasp the factors that contributed to the inference result. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2023-5037 A Summary of the Invention [Problem to be solved by the invention]

[0005] On the other hand, when trying to verify whether the inference results of a machine learning model are valid, if one were to verify all of the input data for the machine learning model and the inference results output, the amount of data to be verified would be enormous, making it difficult to verify whether the machine learning model has been properly trained.

[0006] The present invention has been made in consideration of the above circumstances, and aims to provide a technology for efficiently confirming the validity of a machine learning model that outputs an estimation result from sensor data. [Means for solving the problem]

[0007] In order to solve the above problems, the first invention of the present application is a method for evaluating the validity of an inference result output by a machine learning model, comprising: a) a step of extracting peculiar data from the inference result output by the machine learning model; and b) a step of performing a factor analysis of the machine learning model on the peculiar data.

[0008] A second invention of the present application is the evaluation method of the first invention, wherein the machine learning model outputs the inference result as time series data based on input time series data, and the step a) includes: a1) a step of predicting the inference result from past data of the inference result; a2) a step of comparing the prediction result of the step a1) with the inference result output from the machine learning model and calculating a prediction error; and a3) a step of extracting the peculiar data from the inference result by using the prediction error.

[0009] A third invention of the present application is the evaluation method of the second invention, in which in the step a1), the inference result is predicted using an autoregressive model or a deep learning model.

[0010] A fourth aspect of the present invention is the evaluation method of the second aspect of the present invention, wherein in the step a3), the inference results in which the absolute value of the prediction error is equal to or greater than a threshold value are extracted as the peculiar data.

[0011] A fifth aspect of the present invention is the evaluation method according to the fourth aspect, wherein the threshold value is three times the standard deviation of the prediction error.

[0012] A sixth aspect of the present invention is an evaluation method according to any one of the first to fifth aspects of the present invention, in which an explainable AI is used for the factor analysis in the step b).

[0013] The seventh invention of the present application is the evaluation method of the sixth invention, wherein the explainable AI is SHAP.

[0014] The eighth invention of the present application is the evaluation method of the seventh invention, wherein in the step b), the SHAP value of each factor in the peculiar data is compared with the SHAP value of each factor in the inference result that is not the peculiar data.

[0015] The 9th invention of the present application is the evaluation method of the 8th invention, wherein in the step b), the SHAP value of each factor in the peculiar data and the SHAP value of each factor in the inference result that is not the peculiar data are displayed in heat map format. Effect of the Invention

[0016] According to the first to ninth aspects of the present application, the validity of a machine learning model can be efficiently confirmed.

[0017] In particular, according to the sixth to ninth aspects of the present application, by using an explainable AI, factor analysis can be easily performed when evaluating a machine learning model.

[0018] In particular, according to the ninth aspect of the present invention, the SHAP values ​​of the various factors can be visually compared, making it easier to carry out factor analysis. [Brief description of the drawings]

[0019] [Figure 1] FIG. 13 is a diagram showing an example of an application device. [Diagram 2] FIG. 13 is a diagram illustrating an example of sensor positions during a learning process of a machine learning model. [Diagram 3] FIG. 2 is a block diagram showing the functions of a computer. [Figure 4] FIG. 1 is a diagram showing input and output of a machine learning model. [Diagram 5] 1 is a flowchart showing the flow of a method for evaluating a machine learning model. [Figure 6] FIG. 13 is a diagram showing an example of a prediction error. [Figure 7] FIG. 13 is a diagram showing an example of a heat map of SHAP values. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0020] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0021] <1. Examples of devices that apply machine learning models> First, a substrate processing apparatus 1 will be described as an example of an apparatus to which a machine learning model M to be evaluated for evaluating the validity of an inference result in the estimation method of the present invention is applied. Fig. 1 is a diagram showing an overview of the substrate processing apparatus 1, which is an example of an apparatus to which a machine learning model is applied.

[0022] The substrate processing apparatus 1 is an apparatus for supplying a processing liquid to a surface of a disk-shaped substrate W (silicon wafer) in a semiconductor wafer manufacturing process, to process the surface of the substrate W. As shown in Fig. 1, the substrate processing apparatus 1 includes a chamber 10, and a substrate holder 20, a rotation mechanism 30, a processing liquid supply unit 40, a processing liquid collector 50, and a blocking plate 60 accommodated in the chamber 10.

[0023] The chamber 10 forms a processing space for processing the substrate W. An airflow supply unit 11 that supplies a downflow airflow is provided at the upper portion of the chamber 10.

[0024] The substrate holding unit 20 is a mechanism that holds the substrate W horizontally inside the chamber 10. The substrate holding unit 20 has a disk-shaped spin base 21 and a plurality of chuck pins 22. The plurality of chuck pins 22 hold the peripheral edge of the substrate W, and position the substrate W at a position with a small gap between the upper surface of the spin base 21. The rotation mechanism 30 is a mechanism for rotating the substrate holding unit 20.

[0025] The processing liquid supply unit 40 is a mechanism that supplies a processing liquid to the upper surface of the substrate W held by the substrate holding unit 20. The processing liquid supply unit 40 has an upper surface nozzle 41 that supplies the processing liquid to the upper surface of the substrate W, and a lower surface nozzle 42 that supplies the processing liquid to the lower surface of the substrate W.

[0026] The processing liquid collecting unit 50 is a mechanism for collecting the processing liquid after use. The processing liquid collecting unit 50 has an inner cup 51, a middle cup 52, and an outer cup 53, each of which is connected to a different liquid discharge path. The inner cup 51, the middle cup 52, and the outer cup 53 can be raised and lowered independently of one another by a lifting mechanism (not shown) between a lower position (position shown in FIG. 1) where the upper end is located below the substrate W, and an upper position where the upper end is located above the upper surface of the substrate W.

[0027] When processing the substrate W, with any one of the three cups 51, 52, 53 of the processing liquid collecting part 50 placed in the upper position, the processing liquid is supplied to the surface of the substrate W by the processing liquid supplying part 40 while the substrate holding part 20 and the substrate W are rotated by the rotating mechanism 30. The processing liquid supplied to the surface of the substrate W and used for surface processing of the substrate W is scattered outward by the centrifugal force generated by the rotation of the substrate W, collected by any one of the cups 51, 52, 53 of the processing liquid collecting part 50, and recovered via the liquid discharge path.

[0028] The blocking plate 60 is a member for suppressing diffusion of gas near the surface of the substrate W when performing some processes, such as drying process of the substrate W after supply of a processing liquid. The blocking plate 60 has a disk-shaped outer shape and is arranged horizontally above the substrate holding part 20. The blocking plate 60 is connected to a lifting mechanism 61. When the lifting mechanism 61 is operated, the blocking plate 60 moves up and down between an upper position that is above and away from the upper surface of the substrate W held by the substrate holding part 20, and a lower position that is closer to the upper surface of the substrate W than the upper position.

[0029] When the processing liquid is supplied from the processing liquid supply unit 40 to the substrate W, the shield plate 60 retreats to the upper position. When the substrate W is dried after the processing liquid is supplied, the shield plate 60 is lowered to the lower position by the lifting mechanism 61. Then, dry gas is blown from the blowing port 62 toward the upper surface of the substrate W. At this time, the shield plate 60 prevents the gas from diffusing. As a result, the dry gas is efficiently supplied to the upper surface of the substrate W.

[0030] <2. About machine learning models> Next, a description will be given of a machine learning model M for estimating an airflow near a substrate W in the above-described substrate processing apparatus 1. The estimation method of the present invention can be used, for example, to evaluate the validity of an inference result of the machine learning model M.

[0031] This machine learning model M estimates airflow near the substrate W from time series data detected by a plurality of airflow sensors in the substrate processing apparatus 1. Fig. 2 is a diagram showing an example of sensor positions during a learning process of the machine learning model M. Fig. 3 is a block diagram showing functions of a computer 90 that executes the learning process of the machine learning model M, the process of predicting airflow around the substrate W by the machine learning model M during substrate processing in the substrate processing apparatus 1, and the process of evaluating the machine learning model M. Fig. 4 is a diagram showing input / output of the machine learning model M.

[0032] 2, the substrate processing apparatus 1 has ten permanent sensors Sp0, Sp1, Sp2, Sp3, Sp4, Sp5, Sp6, Sp7, Sp8, and Sp9 installed in a chamber 10. In addition, during the learning process of the machine learning model M, four temporary sensors St1, St2, St3, and St4 are installed at four points P1, P2, P3, and P4 around the substrate W whose airflow is to be estimated. The permanent sensors Sp0 to Sp9 and the temporary sensors St1 to St4 each measure three values, namely, wind speed, azimuth angle, and depression angle, at each point.

[0033] Here, the time series data of wind speed detected by each of the permanent sensors Sp0 to Sp9 is referred to as wind speed data D10 to D19, the time series data of azimuth angle detected by each of the permanent sensors Sp0 to Sp9 is referred to as azimuth angle data D20 to D29, and the time series data of depression angle detected by each of the permanent sensors Sp0 to Sp9 is referred to as depression angle data D30 to D39.

[0034] Similarly, the time series data of wind speed detected by the provisional sensors St1 to St4 respectively are referred to as wind speed data D41 to D44, the time series data of azimuth angle detected by the provisional sensors St1 to St4 respectively are referred to as azimuth angle data D51 to D54, and the time series data of depression angle detected by the provisional sensors St1 to St4 respectively are referred to as depression angle data D61 to D64.

[0035] The above-mentioned detection values ​​detected by the permanent sensors Sp0 to Sp9 and the temporary sensors St1 to St4 are input to the computer 90, respectively.

[0036] The computer 90 is an information processing device for executing a learning process of the machine learning model M, a process of predicting an airflow around the substrate W by the machine learning model M during substrate processing in the substrate processing apparatus 1, and a process of evaluating the machine learning model M. The computer 90 is electrically connected to the substrate processing apparatus 1. The computer 90 receives at least the detection values ​​of the permanent sensors Sp0 to Sp9 and the temporary sensors St1 to St4, and outputs to the substrate processing apparatus 1 an airflow prediction result around the substrate W output by the machine learning model M, or a control value based on the airflow prediction result. The computer 90 may also function as a control unit that controls the substrate processing apparatus 1.

[0037] 2, the computer 90 has a processor 901 such as a CPU, a memory 902 such as a RAM, and a storage unit 903 such as a hard disk drive. The storage unit 903 stores a computer program Pg for executing a learning process for the machine learning model M, a process for predicting an airflow around the substrate W by the machine learning model M during substrate processing in the substrate processing apparatus 1, and a process for evaluating the machine learning model M.

[0038] 3, the computer 90 has a data acquisition unit 91, a learning unit 92, an airflow prediction unit 93, and an evaluation unit 94. The functions of the data acquisition unit 91, the learning unit 92, the airflow prediction unit 93, and the evaluation unit 94 are realized by a processor 901 of the computer 90 operating in accordance with a computer program Pg.

[0039] During the learning process of the machine learning model M, with the temporary sensors St1 to St4 installed in the chamber 10, the airflow in the chamber 10 is measured by the permanent sensors Sp0 to Sp9 and the temporary sensors St1 to St4 without supplying a processing liquid in the substrate processing apparatus 1, and the measured airflow is input to a data acquisition unit 91 of the computer 90. Then, the data acquisition unit 91 passes the detection results of the permanent sensors Sp0 to Sp9 and the temporary sensors St1 to St4 to a learning unit 92.

[0040] The learning unit 92 performs machine learning of the machine learning model M using the wind speed data D10 to D19, azimuth angle data D20 to D29, and depression angle data D30 to D39 detected by the permanent sensors Sp0 to Sp9 as input variables, and the wind speed data D41 to D44, azimuth angle data D51 to D54, and depression angle data D61 to D64 detected by the temporary sensors St1 to St4 as teacher data.

[0041] 4, the machine learning model M becomes an estimation model that uses the wind speed data D10-D19, azimuth angle data D20-D29, and depression angle data D30-D39 detected by the permanent sensors Sp0-Sp9 as input variables, and outputs estimated wind speeds E11-E14, estimated azimuth angles E21-E24, and estimated depression angles E31-E34 at the four points P1-P4. The estimated wind speeds E11-E14, estimated azimuth angles E21-E24, and estimated depression angles E31-E34 are each time-series data.

[0042] When the learning process of the machine learning model M is completed, the learned machine learning model M is handed over to the airflow prediction unit 93. In addition, the temporary sensors St1 to St4 are removed before substrate processing is started in the substrate processing apparatus 1. When performing substrate processing, the temporary sensors St1 to St4 cannot detect airflow at the four points P1 to P4. In other words, the airflow around the substrate W during substrate processing cannot be detected. For this reason, the state of airflow at the four points P1 to P4 around the substrate W during substrate processing is estimated by the machine learning model M.

[0043] While substrate processing is being performed in the substrate processing apparatus 1, detection values ​​from the permanent sensors Sp0-Sp9 are constantly input to a data acquisition unit 91 of the computer 90 and handed over to an airflow prediction unit 93. The airflow prediction unit 93 uses wind speed data D10-D19, azimuth angle data D20-D29, and depression angle data D30-D39 detected by the permanent sensors Sp0-Sp9, respectively, as input variables, and outputs estimated wind speeds E11-E14, estimated azimuth angles E21-E24, and estimated depression angles E31-E34 at four points P1-P4. These inference results are then output to the substrate processing apparatus 1.

[0044] After the machine learning model M starts operating, if we try to confirm the validity of its inference results (E11-E14, E21-E24, E31-E34), it is difficult to compare them with the actual measurement results because the temporary sensors St1-St4 have been removed. Therefore, we use the evaluation method described below to evaluate the validity of the inference results output by the machine learning model.

[0045] <3. How to evaluate machine learning models> Fig. 5 is a flowchart showing the flow of a method for evaluating the machine learning model M. As shown in Fig. 5, this evaluation method includes a step of extracting peculiar data from the inference result output by the machine learning model M (steps S11 to S13), and a step of performing a factor analysis of the inference result of the machine learning model M for the peculiar data (steps S21 to S22).

[0046] In the evaluation method of FIG. 5, first, the evaluation unit 94 acquires the inference results (E11 to E14, E21 to E24, E31 to E34) output by the airflow prediction unit 93. Then, the evaluation unit 94 predicts the inference result itself from the past data of the inference result, and acquires a self-predicted inference result (step S11). That is, the evaluation unit 94 predicts each value itself from the past data of time series for each of the 12 values ​​of the estimated wind speeds E11 to E14, the estimated azimuth angles E21 to E24, and the estimated depression angles E31 to E34, which are the inference results of the machine learning model M. For example, when the input value and the output value of the machine learning model M are time series data every second, the value of the next estimated wind speed E11 is predicted based on the past data from 20 seconds before to 1 second before the estimated wind speed E11.

[0047] In this embodiment, a VAR model (Vector Auto Regressive model) is used for the self-prediction in step S11. However, other auto-regressive models such as an AR model or an ARMA model, or deep learning models such as LSTM or Teansformer may be used for the self-prediction in step S11.

[0048] Following step S11, the evaluation unit 94 compares the inference result output from the machine learning model M with the self-predicted inference result obtained in step S11, and calculates a prediction error (step S12). FIG. 6 is a diagram showing an example of a prediction error of the estimated wind speed E11. The prediction error of the estimated wind speed E11 is, in other words, an error between the estimated wind speed E11 output from the machine learning model M and the self-predicted result predicted from the past data of the estimated wind speed E11. The horizontal axis of FIG. 6 indicates the time of the substrate processing. The vertical axis of FIG. 6 indicates the prediction error. As shown in FIG. 6, the prediction error of the estimated wind speed E11 is distributed with 0 as the center.

[0049] After calculating the prediction error, the evaluation unit 94 uses the prediction error to extract peculiar data from the inference result (step S13). Specifically, an inference result whose absolute value of the prediction error is equal to or greater than a threshold is extracted as peculiar data. In the example of FIG. 5, the threshold Th is three times the standard deviation of the prediction error.

[0050] In this way, the self-prediction value of the inference result is used to extract anomalous data that should be used to confirm the validity of the inference result, which makes it possible to efficiently reduce the amount of data that requires detailed confirmation by a human.

[0051] After extracting the peculiar data in steps S11 to S13, the evaluation unit 94 then performs a factor analysis on each piece of peculiar data (step S21). Specifically, the explainable AI is used to analyze the contribution of each input variable to the inference result of the machine learning model M. In this embodiment, SHAP (SHapley Additive exPlanations) is used for the explainable AI. Therefore, in step S21, a SHAP value for each input variable is output from SHAP. In this way, by using the explainable AI, factor analysis can be easily performed when evaluating the machine learning model M.

[0052] Next, the evaluation unit 94 causes the display unit 900 connected to the computer 90 to display the SHAP value for each input variable obtained in step S21 (step S22). This allows the user to visually analyze the cause of the decrease in prediction accuracy in the peculiar data and evaluate the machine learning model M.

[0053] In step S21, a factor analysis may be performed on the inference result that is not singular data, and the SHAP value of each factor in the singular data may be compared with the SHAP value of each factor in the inference result that is not singular data. In this case, in step S22, the SHAP value of each factor in the singular data may be displayed in a heat map format and compared with the SHAP value of each factor in the inference result that is not singular data. Comparison in the heat map format makes it easier to perform factor analysis and evaluate the machine learning model M.

[0054] Fig. 7 is a diagram showing an example of a heat map of SHAP values ​​in idiosyncratic data and non-idiosyncratic data. In the heat map of Fig. 7, the vertical axis indicates each input variable, and the horizontal axis indicates the lag number. Lag number 0 indicates the past data one second ago (one second ago in this embodiment), lag number 1 indicates the past data two seconds ago, lag number 2 indicates the past data three seconds ago, and lag number 3 indicates the past data four seconds ago. As shown in Fig. 7, by showing the SHAP values ​​in a heat map, it is easy to visually compare the SHAP values ​​of idiosyncratic data and non-idiosyncratic data. This makes it easier to perform factor analysis.

[0055] <4. Modifications> Although one embodiment of the present invention has been described above, the present invention is not limited to the above embodiment.

[0056] In the above embodiment, the machine learning model M to be evaluated is for estimating airflow near the substrate W in the substrate processing apparatus 1. However, the present invention is not limited to this. The machine learning model to be evaluated may be one that takes some sensor detection result as input and outputs some inference result in other devices such as a printing device or an image processing device. Furthermore, the machine learning model to be evaluated does not necessarily have to be one used in some device.

[0057] Furthermore, the elements appearing in the above-described embodiments and modifications may be combined as appropriate to the extent that no contradiction arises. [Explanation of symbols]

[0058] 90: Computer D10~D19: Wind speed data (permanent sensor) D20~D29: Azimuth angle data (permanent sensor) D30~D39: Depression angle data (permanent sensor) D41~D44: Wind speed data (temporary sensor) D51~D54: Azimuth angle data (temporary sensor) D61~D64: Depression angle data (temporary sensor) E11~E14: Estimated wind speed E21~E24: Estimated azimuth E31~E34: Estimated depression angle M: Machine learning model Sp0~Sp9: Permanent sensors St1~St4: Temporary sensors Th: Threshold

Claims

1. A method for evaluating the validity of an inference result output by a machine learning model, comprising: a) extracting peculiar data from the inference result output by the machine learning model; b) performing a factor analysis of the machine learning model on the peculiar data; The evaluation method has the following features:

2. The evaluation method according to claim 1, The machine learning model outputs the inference result as time-series data based on the input time-series data; The step a) comprises: a1) predicting the inference result from past data of the inference result; a2) comparing the prediction result of the step a1) with the inference result output from the machine learning model to calculate a prediction error; a3) extracting the peculiar data from the inference result using the prediction error; (c) evaluation methods;

3. The evaluation method according to claim 2, An evaluation method in which in step a1), the inference result is predicted using an autoregressive model or a deep learning model.

4. The evaluation method according to claim 2, In the step a3), the inference result in which the absolute value of the prediction error is equal to or greater than a threshold is extracted as the peculiar data.

5. The evaluation method according to claim 4, The evaluation method, wherein the threshold value is three times the standard deviation of the prediction error.

6. The evaluation method according to any one of claims 1 to 5, The evaluation method, wherein an explainable AI is used for the factor analysis in step b).

7. The evaluation method according to claim 6, The evaluation method, wherein the explainable AI is SHAP.

8. The evaluation method according to claim 7, In the step b), the SHAP value of each factor in the singular data is compared with the SHAP value of each factor in the inference result that is not the singular data.

9. The evaluation method according to claim 8, In the step b), the SHAP value of each factor in the idiosyncratic data and the SHAP value of each factor in the inference result that is not the idiosyncratic data are displayed in a heat map format.

Citation Information

Patent Citations

  • Factor analysis device, factor analysis method, and program

    JP2023005037A