A multi-element time sequence missing value imputation method and device
Patent Information
- Application Number
- CN202311240993.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-22
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-09-22
AI Technical Summary
现有的预测模型往往是在完整的多元时序数据的基础上进行推理和预测,然而,在生产实践中,由于存在诸如传感器故障、设备停电等诸多意外情况,实际采集得到的多元时序数据中可能存在数据缺失,导致预测模型无法正常进行预测
[0056]本说明书实施例提出的一种多元时序缺失值插补方法及装置,基于图注意力网络学习多元时序缺失值数据中的图特征,然后基于模型学习到的图特征对多元时序缺失值进行插补,以提高插补的效果。
Smart Images

Figure CN117195962B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of machine learning through one or more embodiments, and more particularly to a multivariate temporal missing value imputation method and apparatus. Background Technology
[0002] Multivariate time-series data has wide applications in many fields. For example, in semiconductor manufacturing, inference and prediction can be made based on multivariate time-series data generated by multiple sensors at multiple time points to predict yield rates, causes of failures, and so on. Existing prediction models often rely on complete multivariate time-series data for inference and prediction. However, in production practice, due to unforeseen circumstances such as sensor failures and power outages, the actual acquired multivariate time-series data may contain missing data, causing the prediction model to fail to make accurate predictions. Therefore, a data imputation method for missing multivariate time-series values is needed. Summary of the Invention
[0003] This specification describes one or more embodiments of a multivariate temporal missing value imputation method and apparatus, which learns graph features in multivariate temporal missing value data based on a graph attention network, and then imputes multivariate temporal missing values based on the graph features learned by the model, so as to improve the imputation effect.
[0004] Firstly, a multivariate temporal missing value imputation method is provided, implemented using an imputation model. The imputation model includes a first graph attention network, a second graph attention network, an encoder, and a decoder. The method includes:
[0005] Acquire first data, which is multi-dimensional time-series data generated by multiple sensors at multiple time points during the semiconductor manufacturing process, and the first data has data missing at one or more data points;
[0006] Based on the first data, a corresponding first missing mask matrix is determined, wherein the first missing mask matrix indicates the location of missing values and observations in the first data;
[0007] The first data is concatenated with the first missing mask matrix and encoded to obtain the first encoding result; the data corresponding to each sensor in the first encoding result is used as graph node data and input into the first graph attention network to obtain the first graph feature matrix;
[0008] The first data is concatenated with the feature matrix of the first image and then input into the encoder to obtain the first interpolated data;
[0009] The first interpolated data is concatenated with the first missing mask matrix and encoded to obtain the second encoding result; the data corresponding to each sensor in the second encoding result is used as graph node data and input into the second graph attention network to obtain the second graph feature matrix;
[0010] The first interpolated data is concatenated with the feature matrix of the second image and then input into the decoder to obtain the second interpolated data;
[0011] Based on the first data, the first interpolation data, and the second interpolation data, the interpolation result data is determined.
[0012] In one possible implementation, the data corresponding to each sensor in the first encoding result is used as graph node data and input into a first graph attention network to obtain a first graph feature matrix, including:
[0013] Based on the first missing mask matrix, a first missing weight matrix is determined. The first missing weight matrix has a value of 0 at the position indicated by the first missing mask matrix as a missing value, and a value of 1 at the position indicated by the first missing mask matrix as an observation value.
[0014] Based on the first missing weight matrix and the first weight matrix of the first graph attention network, the first corrected weight matrix is determined;
[0015] Based on the first modified weight matrix, graph attention is calculated on the first encoding result to obtain the first graph feature matrix.
[0016] In one possible implementation, the data corresponding to each sensor in the second encoding result is used as graph node data and input into a second graph attention network to obtain a second graph feature matrix, including:
[0017] Based on the first missing mask matrix, a second missing weight matrix is determined. The second missing weight matrix has a value less than 1 at the position indicated by the first missing mask matrix as a missing value, and a value of 1 at the position indicated by the first missing mask matrix as an observation value.
[0018] Based on the second missing weight matrix and the second weight matrix of the second graph attention network, determine the second corrected weight matrix;
[0019] Based on the second modified weight matrix, graph attention is calculated on the second encoding result to obtain the second graph feature matrix.
[0020] In one possible implementation, the encoder is an encoder of the Transformer model; the decoder is a decoder of the Transformer model.
[0021] In one possible implementation, the interpolation model further includes a linear layer and a softmax layer; determining the interpolation result data based on the first data, the first interpolation data, and the second interpolation data includes:
[0022] The first data, the first interpolation data, and the second interpolation data are input into the linear layer, and the output results are input into the softmax layer to obtain the interpolation result data.
[0023] In one possible implementation, the interpolation model is trained in the following manner:
[0024] Acquire first training data, which is multi-dimensional time-series data generated by multiple sensors at multiple time points during the semiconductor manufacturing process. The first data has missing data at one or more data points.
[0025] Several observations in the first training data are manually masked and marked as missing data to obtain the second training data;
[0026] Based on the second training data, a corresponding second missing mask matrix and an artificial mask matrix are determined, wherein the artificial mask matrix indicates the location of the observation that is artificially masked.
[0027] The second training data and the second missing mask matrix are input into the interpolation model to obtain the corresponding training interpolation result data;
[0028] The first loss is determined based on the first training data, the training interpolation result data, and the second missing mask matrix;
[0029] The second loss is determined based on the first training data, the training interpolation result data, and the artificial mask matrix;
[0030] The values of the parameters of the interpolation model are adjusted based on the total loss determined by the first loss and the second loss.
[0031] In one possible implementation, several observations in the first training data are artificially masked, including:
[0032] The first training data is divided into multiple time periods in the time dimension. Based on the proportion of missing values in any target time period to all data in that target time period, the data in that target time period is classified into a sparse data set or a dense data set.
[0033] The same number of artificial masks are applied to the observations in both the sparse and dense data sets.
[0034] In one possible implementation, several observations in the first training data are artificially masked, including:
[0035] Based on the distribution of missing values in the first training data in the time dimension and sensor dimension, the target proportion of each missing value type to all missing values is determined. The missing value types include random block missing values, power outage type missing values, and pseudo-random missing values.
[0036] The observations in the first training data are manually masked for each missing type according to the same target ratio;
[0037] Among them, random block missing means that the values of multiple sensors at multiple time points are all missing; power outage type missing means that the values of all sensors at multiple time points are all missing; pseudo-random missing means other missing values besides random block missing and power outage type missing.
[0038] Secondly, a multivariate temporal missing value interpolation device is provided, implemented using an interpolation model. The interpolation model includes a first graph attention network, a second graph attention network, an encoder, and a decoder. The device comprises:
[0039] The acquisition unit is configured to acquire first data, which is multi-dimensional time-series data generated by multiple sensors at multiple time points during the semiconductor manufacturing process, and the first data has data missing at one or more data points.
[0040] The missing mask determination unit is configured to determine a corresponding first missing mask matrix based on the first data, wherein the first missing mask matrix indicates the position of the missing value and the observed value in the first data;
[0041] The first feature determination unit is configured to concatenate and encode the first data with the first missing mask matrix to obtain a first encoding result; and input the data corresponding to each sensor in the first encoding result as graph node data into the first graph attention network to obtain the first graph feature matrix.
[0042] The first interpolation unit is configured to concatenate the first data with the first image feature matrix and input the result into the encoder to obtain the first interpolation data.
[0043] The second feature determination unit is configured to concatenate and encode the first interpolated data with the first missing mask matrix to obtain a second encoding result; and input the data corresponding to each sensor in the second encoding result as graph node data into the second graph attention network to obtain the second graph feature matrix.
[0044] The second interpolation unit is configured to concatenate the first interpolation data with the feature matrix of the second image and input the result into the decoder to obtain the second interpolation data.
[0045] The third interpolation unit is configured to determine the interpolation result data based on the first data, the first interpolation data, and the second interpolation data.
[0046] In one possible implementation, the interpolation model is trained in the following manner:
[0047] Acquire first training data, which is multi-dimensional time-series data generated by multiple sensors at multiple time points during the semiconductor manufacturing process. The first data has missing data at one or more data points.
[0048] Several observations in the first training data are manually masked and marked as missing data to obtain the second training data;
[0049] Based on the second training data, a corresponding second missing mask matrix and an artificial mask matrix are determined, wherein the artificial mask matrix indicates the location of the observation that is artificially masked.
[0050] The second training data and the second missing mask matrix are input into the interpolation model to obtain the corresponding training interpolation result data;
[0051] The first loss is determined based on the first training data, the training interpolation result data, and the second missing mask matrix;
[0052] The second loss is determined based on the first training data, the training interpolation result data, and the artificial mask matrix;
[0053] The values of the parameters of the interpolation model are adjusted based on the total loss determined by the first loss and the second loss.
[0054] Thirdly, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method of the first aspect.
[0055] Fourthly, a computing device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements the method of the first aspect.
[0056] The embodiments of this specification propose a multivariate temporal missing value imputation method and apparatus, which learns graph features in multivariate temporal missing value data based on graph attention network, and then imputes multivariate temporal missing values based on the graph features learned by the model, so as to improve the imputation effect. Attached Figure Description
[0057] To more clearly illustrate the technical solutions of the various embodiments disclosed in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only a few embodiments disclosed in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0058] Figure 1 This diagram illustrates an interpolation model architecture according to one embodiment.
[0059] Figure 2 A flowchart illustrating a multivariate temporal missing value interpolation method according to one embodiment is shown;
[0060] Figure 3 A flowchart illustrating a method for training an interpolation model according to one embodiment is shown.
[0061] Figure 4 A schematic diagram illustrating a scenario of training an interpolation model according to one embodiment is shown.
[0062] Figure 5 A schematic block diagram of a multivariate time-series missing value interpolation apparatus according to one embodiment is shown. Detailed Implementation
[0063] The solution provided in this specification will now be described with reference to the accompanying drawings.
[0064] As mentioned earlier, existing prediction models often rely on complete multivariate time-series data for reasoning and prediction. However, in production practice, due to unforeseen circumstances such as sensor failures and power outages, the actual collected multivariate time-series data may contain missing data, causing the prediction model to fail to make predictions correctly. Therefore, a data imputation method for missing multivariate time-series values is needed.
[0065] To solve the above problems, Figure 1 A model architecture diagram of a multivariate temporal missing value imputation method according to one embodiment is shown. Figure 1In the example, the first data in the lower left corner is multi-dimensional time-series data generated by multiple sensors at multiple time points during the semiconductor manufacturing process. This data has missing data at one or more data points and is the data to be imputed. Then, based on the missing data locations in the first data, a first missing value mask matrix is determined to indicate the missing value locations. The first data and the first missing value mask matrix are concatenated and encoded, and then input into a first graph attention network to obtain a first graph feature matrix. The first data and the first graph feature matrix are concatenated and input into an encoder to obtain the first imputed data for the first data to be imputed. The first imputed data and the first missing value mask matrix are concatenated and encoded, and then input into a second graph attention network to obtain a second graph feature matrix. The first imputed data and the second graph feature matrix are concatenated and input into a decoder to obtain the second imputed data for the first data to be imputed. Finally, the first data, the first imputed data, and the second imputed data are input into a linear layer, and the output is input into a softmax layer to obtain the final imputed data.
[0066] The following describes the specific implementation steps of the above-mentioned multivariate temporal missing value interpolation method with reference to specific embodiments. Figure 2 A flowchart illustrating a multivariate temporal missing value imputation method according to one embodiment is provided. The execution entity of the method can be any platform, server, or device cluster with computing and processing capabilities. Figure 2As shown, the method utilizes an interpolation model, which includes a first graph attention network, a second graph attention network, an encoder, and a decoder. It includes at least the following steps: Step 201, acquiring first data, which is multivariate time-series data generated by multiple sensors at multiple time points during semiconductor manufacturing, wherein the first data has missing data at one or more data points; Step 202, determining a corresponding first missing mask matrix based on the first data, the first missing mask matrix indicating the positions of missing values and observed values in the first data; Step 203, concatenating and encoding the first data with the first missing mask matrix to obtain a first encoding result; and then encoding the data corresponding to each sensor in the first encoding result. The data is used as graph node data and input into the first graph attention network to obtain the first graph feature matrix; in step 204, the first data and the first graph feature matrix are concatenated and input into the encoder to obtain the first interpolated data; in step 205, the first interpolated data and the first missing mask matrix are concatenated and encoded to obtain the second encoding result; the data corresponding to each sensor in the second encoding result are used as graph node data and input into the second graph attention network to obtain the second graph feature matrix; in step 206, the first interpolated data and the second graph feature matrix are concatenated and input into the decoder to obtain the second interpolated data; in step 207, the interpolated result data is determined based on the first data, the first interpolated data and the second interpolated data. The specific execution process of each of the above steps is described below.
[0067] First, in step 201, first data H is obtained. The first data H is multi-dimensional time-series data generated by multiple sensors at multiple time points during the semiconductor manufacturing process. The first data H has data missing at one or more data points.
[0068] The first data H can be in the form of a two-dimensional matrix, with one dimension being the time dimension and the other being the sensor dimension, storing multi-dimensional time-series data generated by multiple sensors at multiple time points. The first data may have missing data at one or more data points.
[0069] Then, in step 202, based on the first data H, a corresponding first missing mask matrix M1 is determined, wherein the first missing mask matrix M1 indicates the position of the missing value and the observed value in the first data H.
[0070] The first missing value mask matrix M1 indicates the location of missing values and observed values in the first data H. In one embodiment, when the value H[i][j] of the first data H at the j-th sensor at the i-th time point is an observed value, the value M1[i][j] of the first missing value mask matrix M1 at the i-th row and j-th column is 1; in another embodiment, when the value H[i][j] of the first data H at the j-th sensor at the i-th time point is a missing value, the value M1[i][j] of the first missing value mask matrix M1 at the i-th row and j-th column is 0.
[0071] Next, in step 203, the first data H is concatenated and encoded with the first missing mask matrix M1 to obtain the first encoding result; the data corresponding to each sensor in the first encoding result is used as graph node data and input into the first graph attention network to obtain the first graph feature matrix G1.
[0072] In one embodiment, the encoding may include: concatenating the first data H with the first missing mask matrix M1 in the sensor dimension to obtain a first concatenated matrix, performing position encoding on the first concatenated matrix, and then summing the first concatenated matrix with the position encoding to obtain a first encoding result.
[0073] Then, the data corresponding to each sensor in the first encoding result is used as graph node data, and a connection edge is established between any two graph nodes to form the first graph data. The first graph data is input into the first graph attention network to obtain the first graph feature matrix G1.
[0074] For example, both the first data H and the first missing mask matrix M1 are matrices of size t*d, where t is the time dimension and d is the sensor dimension. Concatenating them along the sensor dimension yields the first concatenated matrix, which has a size of 2t*d. Then, the data corresponding to each sensor in the first encoding result is used as graph node data, and edges are established between any two graph nodes to form the first graph data. In the first graph data, there are d nodes, and the corresponding vector dimension of each node is 2t.
[0075] In a more specific embodiment, the data corresponding to each sensor in the first encoding result is used as graph node data and input into the first graph attention network to obtain the first graph feature matrix G1, including:
[0076] Based on the first missing mask matrix M1, determine the first missing weight matrix W. e The first missing weight W eThe matrix is set to 0 at positions indicated by the first missing value mask M1, and 1 at positions indicated by the first missing value mask M1. That is, when a value at a certain position in the first data H is missing, it should not contribute to the learning of the graph structure; therefore, the first missing value weight matrix W at the corresponding position is set accordingly. e The value is set to 0.
[0077] Then, based on the first missing weight matrix W e The first corrected weight matrix is determined by combining the first weight matrix W1 of the first graph attention network. Specifically, the first missing weight matrix W can be used to determine the corrected weight matrix. e Left-multiply by the first weight matrix W1 to obtain W e W1 is used as the first corrected weight matrix.
[0078] Finally, graph attention is calculated on the first encoding result based on the first corrected weight matrix to obtain the first graph feature matrix G1. The value of the first graph feature matrix G1 in the i-th row and j-th column represents the first influence weight of the i-th sensor on the j-th sensor.
[0079] In one embodiment, prior to step 203, the method further includes: padding missing values in the first data H with 0.
[0080] Then, in step 204, the first data H is concatenated with the first image feature matrix G1 and input into the encoder to obtain the first interpolation data X1.
[0081] In one embodiment, the encoder may be an encoder for a Transformer model.
[0082] Next, in step 205, the first interpolated data X1 is concatenated and encoded with the first missing mask matrix M1 to obtain the second encoding result; the data corresponding to each sensor in the second encoding result is used as graph node data and input into the second graph attention network to obtain the second graph feature matrix G2.
[0083] In one embodiment, the encoding may include: concatenating the first interpolated data X1 with the first missing mask matrix M1 in the sensor dimension to obtain a second concatenated matrix, performing position encoding on the second concatenated matrix, and then summing the second concatenated matrix with the position encoding to obtain a second encoding result.
[0084] Then, the data corresponding to each sensor in the second encoding result is used as graph node data, and a connection edge is established between any two graph nodes to form the second graph data. The second graph data is input into the second graph attention network to obtain the second graph feature matrix G2.
[0085] In a more specific embodiment, the data corresponding to each sensor in the second encoding result is used as graph node data and input into the second graph attention network to obtain the second graph feature matrix G2, which includes:
[0086] Based on the first missing mask matrix M1, determine the second missing weight matrix W. d The second missing weight matrix W d The first missing value is indicated by a constant less than 1 in the first missing mask matrix M1, for example, 0.8. The value is 1 in the second missing value weight matrix W at the corresponding position. That is, when the value at a certain position in the first imputed data X1 is the imputed value, it should make a small contribution to the learning of the graph structure. Therefore, the second missing value weight matrix W at the corresponding position is set accordingly. d The value is set to the first constant, which is less than 1.
[0087] Then, based on the second missing weight matrix W d The second corrected weight matrix is determined using the second weight matrix W2 of the second graph attention network. Specifically, the second missing weight matrix W can be used to determine the second corrected weight matrix. d Left-multiply by the second weight matrix W2 to obtain W d W2 serves as the second corrected weight matrix.
[0088] Finally, graph attention is calculated on the second encoding result based on the second modified weight matrix to obtain the second graph feature matrix G2. The value of the second graph feature matrix G2 in the i-th row and j-th column represents the second influence weight of the i-th sensor on the j-th sensor.
[0089] Then, in step 206, the first interpolated data X1 is concatenated with the second image feature matrix G2 and input into the decoder to obtain the second interpolated data X2.
[0090] In one embodiment, the decoder may be a decoder of a Transformer model.
[0091] Finally, in step 207, the interpolation result data X3 is determined based on the first data H, the first interpolation data X1, and the second interpolation data X2.
[0092] Specifically, the interpolation model also includes linear layers and softmax layers, and step 207 includes:
[0093] The first data H, the first interpolation data X1, and the second interpolation data X2 are input into the linear layer, and the output results are input into the softmax layer to obtain the interpolation result data X3.
[0094] By using the methods described in steps 201 to 207, the spatial relationships of multivariate time-series missing value data can be explicitly extracted, namely the aforementioned first graph feature matrix G1 and second graph feature matrix G2, thereby improving the interpolation effect.
[0095] In some embodiments, the interpolation model used in steps 201 to 207 is as follows: Figure 3 It was obtained through training. Figure 3 A flowchart illustrating a method for training an interpolation model according to one embodiment is provided. The method can be executed by any platform, server, or device cluster with computing and processing capabilities. Figure 3 As shown, the method includes at least:
[0096] In step 301, first training data T1 is obtained. The first training data T1 is multi-dimensional time-series data generated by multiple sensors at multiple time points during the semiconductor manufacturing process. The first data has missing data at one or more data points.
[0097] In step 302, several observations in the first training data T1 are manually masked and marked as missing data to obtain the second training data T2.
[0098] In one embodiment, step 302 includes: dividing the first training data T1 into multiple time periods in the time dimension; classifying the data in a target time period into a sparse data set or a dense data set according to the proportion of missing values in any target time period to all data in that target time period; and then, performing the same number of manual masking operations on the observations in the sparse data set and the dense data set respectively.
[0099] In another embodiment, step 302 includes: determining the target proportion of each missing value type to all missing values based on the distribution of missing values in the first training data in the time dimension and sensor dimension, wherein the missing value types include random block missing, power outage type missing, and pseudo-random missing; then, manually masking each missing value type for the observations in the first training data according to the same target proportion; wherein, random block missing is the complete missing value of multiple sensors at multiple time points; power outage type missing is the complete missing value of all sensors at multiple time points; and pseudo-random missing is any missing value other than random block missing and power outage type missing.
[0100] For example, suppose there are time points t1-t10 and sensors a1-a5. Then, a missing random block could be due to missing data from sensors a1-a3 at time points t1-t3; a missing power outage type could be due to missing data from sensors a1-a5 at time points t1-t3; and a missing pseudo-random block could be due to missing data from sensor a2 at time point t1 and simultaneously missing data from sensor a5 at time point t4.
[0101] In step 303, based on the second training data T2, the corresponding second missing mask matrix M2 and artificial mask matrix M are determined. a The artificial mask matrix M a Indicates the location of observations that have been artificially obscured.
[0102] The second missing value mask matrix M2 indicates the location of missing values and observed values in the second training data T2. In one embodiment, when the value T2[i][j] of the j-th sensor at the i-th time point of the second training data T2 is an observed value, the value M2[i][j] of the second missing value mask matrix M2 in the i-th row and j-th column is 1; in another embodiment, when the value T2[i][j] of the j-th sensor at the i-th time point of the second training data T2 is a missing value or an artificially masked value, the value M2[i][j] of the second missing value mask matrix M2 in the i-th row and j-th column is 0.
[0103] Artificial mask matrix M a This indicates the location of the observations in the second training data T2 that are artificially masked. In one embodiment, when the value T2[i][j] of the j-th sensor at the i-th time point of the second training data T2 is an artificially masked value, the artificial mask matrix M... a The value M in the i-th row and j-th column a [i][j] is 1; in another embodiment, when the value T2[i][j] of the j-th sensor at the i-th time point of the second training data T2 is an observed value or a missing value, the artificial mask matrix M a The value M in the i-th row and j-th column a [i][j] is 0.
[0104] In step 304, the second training data T2 and the second missing mask matrix M2 are input into the interpolation model to obtain the corresponding training interpolation result data T3.
[0105] In step 305, the first loss L1 is determined based on the first training data T1, the training interpolation result data T3, and the second missing mask matrix M2.
[0106] In one embodiment, the first loss L1 can be determined by formula (1):
[0107]
[0108] Where D and T are the matrix dimensions of the first training data L1 in two dimensions, and ⊙ is the Hadamard product.
[0109] In step 306, based on the first training data T1, the training interpolation result data T3, and the artificial mask matrix M... a The second loss, L2, is determined.
[0110] In one embodiment, the second loss L2 can be determined by formula (2):
[0111]
[0112] Where D and T are the matrix dimensions of the first training data T1 in two dimensions, and ⊙ is the Hadamard product.
[0113] In step 307, the values of the parameters of the interpolation model are adjusted based on the total loss L3 determined by the first loss L1 and the second loss L2.
[0114] In one embodiment, the total loss is the sum of the first loss and the second loss, i.e., L3 = L1 + L2.
[0115] The scenario illustrations for training the interpolation model in steps 301 to 307 can be shown as follows: Figure 4 As shown. Compared to random manual masking, Figure 3 The masking method and the training model method shown can achieve better training results.
[0116] According to another embodiment, a multivariate time-series missing value interpolation device is also provided. Figure 5 A schematic block diagram of a multivariate temporal missing value imputation apparatus according to one embodiment is shown. The apparatus utilizes an imputation model including a first graph attention network, a second graph attention network, an encoder, and a decoder. This apparatus can be deployed in any device, platform, or cluster of devices with computing and processing capabilities. Figure 5 As shown, the device 500 includes:
[0117] The acquisition unit 501 is configured to acquire first data, which is multi-dimensional time-series data generated by multiple sensors at multiple time points during the semiconductor manufacturing process, and the first data has data missing at one or more data points.
[0118] The missing mask determination unit 502 is configured to determine a corresponding first missing mask matrix based on the first data, wherein the first missing mask matrix indicates the position of the missing value and the observed value in the first data;
[0119] The first feature determination unit 503 is configured to concatenate and encode the first data with the first missing mask matrix to obtain a first encoding result; and input the data corresponding to each sensor in the first encoding result as graph node data into the first graph attention network to obtain the first graph feature matrix.
[0120] The first interpolation unit 504 is configured to concatenate the first data with the first image feature matrix and input the result into the encoder to obtain the first interpolation data.
[0121] The second feature determination unit 505 is configured to concatenate and encode the first interpolated data with the first missing mask matrix to obtain a second encoding result; and input the data corresponding to each sensor in the second encoding result as graph node data into the second graph attention network to obtain the second graph feature matrix.
[0122] The second interpolation unit 506 is configured to concatenate the first interpolation data with the feature matrix of the second image and input the concatenation into the decoder to obtain the second interpolation data.
[0123] The third interpolation unit 507 is configured to determine interpolation result data based on the first data, the first interpolation data, and the second interpolation data.
[0124] In one possible implementation, the interpolation model is trained in the following manner:
[0125] Acquire first training data, which is multi-dimensional time-series data generated by multiple sensors at multiple time points during the semiconductor manufacturing process. The first data has missing data at one or more data points.
[0126] Several observations in the first training data are manually masked and marked as missing data to obtain the second training data;
[0127] Based on the second training data, a corresponding second missing mask matrix and an artificial mask matrix are determined, wherein the artificial mask matrix indicates the location of the observation that is artificially masked.
[0128] The second training data and the second missing mask matrix are input into the interpolation model to obtain the corresponding training interpolation result data;
[0129] The first loss is determined based on the first training data, the training interpolation result data, and the second missing mask matrix;
[0130] The second loss is determined based on the first training data, the training interpolation result data, and the artificial mask matrix;
[0131] The values of the parameters of the interpolation model are adjusted based on the total loss determined by the first loss and the second loss.
[0132] According to another embodiment, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed in a computer, causes the computer to perform the methods described in any of the above embodiments.
[0133] According to another embodiment, a computing device is also provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements the method described in any of the above embodiments.
[0134] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0135] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0136] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0137] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0138] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A multivariate temporal missing value imputation method, implemented using an imputation model, wherein the imputation model includes a first graph attention network, a second graph attention network, an encoder, and a decoder, and the method includes: Acquire first data, which is multi-dimensional time-series data generated by multiple sensors at multiple time points during the semiconductor manufacturing process. The first data has missing data at one or more data points, which are data to be interpolated. The reasons for the missing data include: sensor failure and equipment power outage; the first data is in the form of a two-dimensional matrix, with one dimension being the time dimension and the other dimension being the sensor dimension, storing multi-dimensional time-series data generated by multiple sensors at multiple time points; Based on the first data, a corresponding first missing mask matrix is determined, wherein the first missing mask matrix indicates the location of missing values and observations in the first data; The first data and the first missing mask matrix are concatenated and encoded in the sensor dimension to obtain the first encoding result; the data corresponding to each sensor in the first encoding result are used as graph node data and input into the first graph attention network to obtain the first graph feature matrix; the value of the first graph feature matrix in the i-th row and j-th column represents the first influence weight of the i-th sensor on the j-th sensor; The first data is concatenated with the feature matrix of the first image and then input into the encoder to obtain the first interpolated data for the first data to be interpolated. The first interpolated data and the first missing mask matrix are concatenated and encoded in the sensor dimension to obtain the second encoding result; the data corresponding to each sensor in the second encoding result are used as graph node data and input into the second graph attention network to obtain the second graph feature matrix; the value of the second graph feature matrix in the i-th row and j-th column represents the second influence weight of the i-th sensor on the j-th sensor; The first interpolated data is concatenated with the feature matrix of the second image and then input into the decoder to obtain the second interpolated data for the first data to be interpolated. Based on the first data, the first interpolation data, and the second interpolation data, interpolation result data for the first data to be interpolated is determined; the interpolation result data corresponds to data missing data present at one or more data points by multiple sensors; the complete data determined based on the interpolation result data is used to input the prediction model for inference, in order to determine the yield and failure causes in the semiconductor manufacturing field; The interpolation model is trained as follows: First training data is acquired, which is multivariate time-series data generated by multiple sensors at multiple time points during semiconductor manufacturing, and the first data has missing data at one or more data points; several observations in the first training data are manually masked and marked as missing data to obtain second training data; based on the second training data, a corresponding second missing data mask matrix and a manual mask matrix are determined, the manual mask matrix indicating the position of the manually masked observations; the second training data and the second missing data mask matrix are input into the interpolation model to obtain corresponding training interpolation result data; a first loss is determined based on the first training data, the training interpolation result data, and the second missing data mask matrix; a second loss is determined based on the first training data, the training interpolation result data, and the manual mask matrix; and the parameters of the interpolation model are adjusted based on the total loss determined by the first loss and the second loss.
2. The method according to claim 1, characterized in that, The data corresponding to each sensor in the first encoding result are used as graph node data and input into the first graph attention network to obtain the first graph feature matrix, including: Based on the first missing mask matrix, a first missing weight matrix is determined. The first missing weight matrix has a value of 0 at the position indicated by the first missing mask matrix as a missing value, and a value of 1 at the position indicated by the first missing mask matrix as an observation value. Based on the first missing weight matrix and the first weight matrix of the first graph attention network, the first corrected weight matrix is determined; Based on the first modified weight matrix, graph attention is calculated on the first encoding result to obtain the first graph feature matrix.
3. The method according to claim 1, characterized in that, The data corresponding to each sensor in the second encoding result are used as graph node data and input into the second graph attention network to obtain the second graph feature matrix, including: Based on the first missing mask matrix, a second missing weight matrix is determined. The second missing weight matrix has a value less than 1 at the position indicated by the first missing mask matrix as a missing value, and a value of 1 at the position indicated by the first missing mask matrix as an observation value. Based on the second missing weight matrix and the second weight matrix of the second graph attention network, determine the second corrected weight matrix; Based on the second modified weight matrix, graph attention is calculated on the second encoding result to obtain the second graph feature matrix.
4. The method according to claim 1, characterized in that, The encoder is an encoder of the Transformer model; the decoder is a decoder of the Transformer model.
5. The method according to claim 1, characterized in that, The interpolation model further includes a linear layer and a softmax layer; based on the first data, the first interpolation data, and the second interpolation data, the interpolation result data is determined, including: The first data, the first interpolation data, and the second interpolation data are input into the linear layer, and the output results are input into the softmax layer to obtain the interpolation result data.
6. The method according to claim 1, characterized in that, Several observations in the first training data are artificially masked, including: The first training data is divided into multiple time periods in the time dimension. Based on the proportion of missing values in any target time period to all data in that target time period, the data in that target time period is classified into a sparse data set or a dense data set. The same number of artificial masks are applied to the observations in both the sparse and dense data sets.
7. The method according to claim 1, characterized in that, Several observations in the first training data are artificially masked, including: Based on the distribution of missing values in the first training data in the time dimension and sensor dimension, the target proportion of each missing value type to all missing values is determined. The missing value types include random block missing values, power outage type missing values, and pseudo-random missing values. The observations in the first training data are manually masked for each missing type according to the same target ratio; Among them, random block missing means that the values of multiple sensors at multiple time points are all missing; power outage type missing means that the values of all sensors at multiple time points are all missing; pseudo-random missing means other missing values besides random block missing and power outage type missing.
8. A multivariate temporal missing value imputation device, implemented using an imputation model, wherein the imputation model includes a first graph attention network, a second graph attention network, an encoder, and a decoder, and the device comprises: The acquisition unit is configured to acquire first data, which is multi-dimensional time-series data generated by multiple sensors at multiple time points during the semiconductor manufacturing process. The first data has data missing at one or more data points and is data to be interpolated. The reasons for the missing data include: sensor failure and equipment power outage; the first data is in the form of a two-dimensional matrix, with one dimension being the time dimension and the other dimension being the sensor dimension, storing multi-dimensional time-series data generated by multiple sensors at multiple time points; The missing mask determination unit is configured to determine a corresponding first missing mask matrix based on the first data, wherein the first missing mask matrix indicates the position of the missing value and the observed value in the first data; The first feature determination unit is configured to concatenate and encode the first data and the first missing mask matrix in the sensor dimension to obtain a first encoding result; and input the data corresponding to each sensor in the first encoding result as graph node data into a first graph attention network to obtain a first graph feature matrix; the value of the first graph feature matrix in the i-th row and j-th column represents the first influence weight of the i-th sensor on the j-th sensor. The first interpolation unit is configured to concatenate the first data with the first image feature matrix and input the result into the encoder to obtain the first interpolation data for the first data to be interpolated. The second feature determination unit is configured to concatenate and encode the first interpolated data and the first missing mask matrix in the sensor dimension to obtain a second encoding result; input the data corresponding to each sensor in the second encoding result as graph node data into the second graph attention network to obtain a second graph feature matrix; the value of the second graph feature matrix in the i-th row and j-th column represents the second influence weight of the i-th sensor on the j-th sensor; The second interpolation unit is configured to concatenate the first interpolation data with the feature matrix of the second image and input the concatenation into the decoder to obtain the second interpolation data for the first data to be interpolated. The third interpolation unit is configured to determine interpolation result data for the first data to be interpolated based on the first data, the first interpolation data, and the second interpolation data; the interpolation result data corresponds to data missing data present at one or more data points by multiple sensors; the complete data determined based on the interpolation result data is used to input the prediction model for inference to determine the yield rate and failure causes in the semiconductor manufacturing field. The interpolation model is trained as follows: First training data is acquired, which is multivariate time-series data generated by multiple sensors at multiple time points during semiconductor manufacturing, and the first data has missing data at one or more data points; several observations in the first training data are manually masked and marked as missing data to obtain second training data; based on the second training data, a corresponding second missing data mask matrix and a manual mask matrix are determined, the manual mask matrix indicating the position of the manually masked observations; the second training data and the second missing data mask matrix are input into the interpolation model to obtain corresponding training interpolation result data; a first loss is determined based on the first training data, the training interpolation result data, and the second missing data mask matrix; a second loss is determined based on the first training data, the training interpolation result data, and the manual mask matrix; and the parameters of the interpolation model are adjusted based on the total loss determined by the first loss and the second loss.
Citation Information
Patent Citations
Time sequence data missing value interpolation method based on attention mechanism
CN113298131A
Medical missing data completion method based on graph attention mechanism and language large model
CN116598014A