Debiasing missing table end-to-end prediction method and device, and electronic equipment

CN117875287BActive Publication Date: 2026-10-09ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410046931.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-12
Publication Date
2026-10-09
Estimated Expiration
2044-01-12

AI Technical Summary

Technical Problem

然而,它们对于支持不完整表格数据分析并不有效,因为它们未考虑特征和标签相关的缺失状态或信息

Benefits of technology

[0028] Compared with the prior art, the improvements of this invention are as follows: This method adopts a Transformer architecture based on inverse propensity scoring to perform end-to-end prediction of missing table data, overcoming the problem of low prediction efficiency for missing table data; it learns the distribution of missing states through pre-training, removes dataset bias through inverse propensity scoring, and adapts to downstream prediction tasks through semi-supervised fine-tuning, overcoming the problem of inaccurate prediction of missing table data. The prediction accuracy of this method is 25% higher than the current algorithm of first completing and then predicting, and it also has higher prediction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117875287B_ABST
    Figure CN117875287B_ABST
Patent Text Reader

Abstract

The application discloses a kind of to be missing table data end-to-end prediction method, comprising: obtaining missing table data feature and label matrix, and generating corresponding feature and label missing mask matrix;Inverse propensity score is calculated column by column using XGBoost model, determine its corresponding inverse propensity score matrix;Inverse propensity score based on the neural network model of the Transformer of construction;Masking is carried out to missing table data feature matrix and feature missing mask matrix, and self-supervised pre-training is carried out using the way of reconstructing masking information combined with inverse propensity score matrix, to obtain pre-trained neural network model;Marked classification data and continuous data are mapped to high dimension by linear layer;Debiasing semi-supervised fine-tuning module is constructed, and the pre-trained neural network model is fine-tuned using reconstruction error and label error, to obtain debiasing table data prediction model;Missing table data feature matrix is input into prediction model to obtain final prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of tabular data prediction, specifically to an end-to-end prediction method, apparatus, and electronic device for removing biased and missing tables. Background Technology

[0002] As the number of datasets represented by rows and columns increases, tabular data analysis has become fundamental to many practical applications, such as predicting patient health status, energy consumption, and financial modeling. However, due to various factors such as improper collection, lost records, resource constraints, or privacy concerns, some features and labels are often not collected, resulting in missing tabular data. Such omissions directly lead to a decline in data quality and significantly reduce the effectiveness of downstream tasks. Therefore, exploring the problem of incomplete tabular data analysis is both attractive and necessary.

[0003] Existing research on tabular data prediction primarily focuses on the analysis of complete tabular data. Ideally, the preferred methods for analyzing tabular data should effectively identify the complex relationships between various features, thereby achieving accurate predictions. However, they are ineffective for supporting the analysis of incomplete tabular data because they do not consider missing states or information related to features and labels. Even with advanced imputation techniques, the imputed data cannot accurately reflect the true data, especially when complex missing mechanisms exist, and it also significantly reduces prediction efficiency. This significantly impacts the predictive performance of models on incomplete tabular data, as these models heavily rely on accurately capturing the distribution of the underlying data. Therefore, developing a novel end-to-end prediction method for debiased missing tables is crucial and urgent. Summary of the Invention

[0004] To overcome the above technical problems, this application implements and provides an end-to-end prediction method, apparatus, and electronic device for debiased missing tables, in order to improve prediction accuracy and efficiency.

[0005] According to a first aspect of the embodiments of this application, an end-to-end prediction method for biased missing tables is provided, comprising:

[0006] S1: Obtain the feature matrix and label matrix of the missing table data, and generate the corresponding feature missing mask matrix and label missing mask matrix; the feature matrix of the missing table data includes multi-dimensional feature data, which can indicate the value or category of the label;

[0007] S2: Randomly sample the feature matrix of the table data, and calculate the reverse tendency score column by column using the XGBoost model based on the sampling results to determine the corresponding reverse tendency score matrix.

[0008] S3: Construct a Transformer neural network model based on the reverse tendency score matrix;

[0009] S4: Mask the missing table data feature matrix and its missing feature mask matrix. Concatenate the masked missing table data feature matrix, the missing feature mask matrix and its corresponding inverse bias score matrix. Input the concatenated matrix into the Transformer neural network model to obtain the table data representation. Reconstruct and predict the data through a linear layer. Use the method of reconstructing the masking information to combine with the inverse bias score matrix for self-supervised pre-training to obtain the pre-trained Transformer neural network model based on inverse bias score.

[0010] S5: Divide the missing table data feature matrix into a missing table data classification feature matrix and a missing table data continuous feature matrix. After marking them with their corresponding feature missing mask matrices, map the marked missing table data classification feature matrix and missing table data continuous feature matrix to a high dimension through a linear layer, and then concatenate the results to obtain a high-dimensional missing table data feature matrix.

[0011] S6: Based on the high-dimensional missing table data feature matrix, construct a debiased semi-supervised fine-tuning module. Concatenate the high-dimensional missing table data feature matrix, the feature missing mask matrix, and their corresponding inverse bias score matrix. Input the concatenated matrix into the pre-trained Transformer neural network model based on inverse bias score to obtain feature matrix representation and prediction results. Based on the missing table data label matrix, calculate the reconstruction error for samples with missing labels and calculate the debiased error based on inverse bias score for samples with intact labels. Use the total error after weighted summation to fine-tune the pre-trained neural network model to obtain the debiased table data prediction model.

[0012] S7: Input the concatenated missing table data feature matrix, feature missing mask matrix and its corresponding inverse bias score matrix into the debiased table data prediction model to predict the missing table data and obtain the final prediction result.

[0013] According to a second aspect of the embodiments of this application, an end-to-end prediction apparatus for biased missing table is provided, comprising:

[0014] The acquisition module is used to acquire the feature matrix and label matrix of the missing table data, and generate the corresponding feature missing mask matrix and label missing mask matrix; the feature matrix of the missing table data includes multi-dimensional feature data, which can indicate the value or category of the label;

[0015] The reverse tendency scoring module is used to randomly sample the feature matrix of the table data, and calculate the reverse tendency score column by column based on the sampling results using the XGBoost model to determine the corresponding reverse tendency score matrix.

[0016] A construction module is used to construct a Transformer neural network model based on the inverse tendency score according to the inverse tendency score matrix;

[0017] The pre-training module is used to mask the feature matrix of the missing table data and its feature missing mask matrix. The masked feature matrix of the missing table data, the feature missing mask matrix and its corresponding inverse bias score matrix are concatenated and then input into the Transformer neural network model to obtain the table data representation. Reconstruction prediction is performed through a linear layer. Self-supervised pre-training is performed by combining the reconstructed masking information with the inverse bias score matrix to obtain the pre-trained Transformer neural network model based on the inverse bias score.

[0018] The data processing module is used to divide the missing table data feature matrix into a missing table data classification feature matrix and a missing table data continuous feature matrix. After marking them with their corresponding feature missing mask matrices, the marked missing table data classification feature matrix and missing table data continuous feature matrix are mapped to a high dimension through a linear layer, and the results are concatenated to obtain a high-dimensional missing table data feature matrix.

[0019] The fine-tuning module is used to construct a semi-supervised fine-tuning module for de-biasing based on the feature matrix of the high-dimensional missing table data. The high-dimensional missing table data feature matrix, the feature missing mask matrix and its corresponding inverse bias score matrix are concatenated and then input into the pre-trained Transformer neural network model based on inverse bias score to obtain feature matrix representation and prediction results. Based on the missing table data label matrix, the reconstruction error is calculated for samples with missing labels and the de-biasing error based on inverse bias score is calculated for samples with intact labels. The total error after weighted summation is used to fine-tune the pre-trained neural network model to obtain the de-biased table data prediction model.

[0020] The prediction module inputs the concatenated missing table data feature matrix, feature missing mask matrix and its corresponding inverse bias score matrix into the debiased table data prediction model to predict the missing table data and obtain the final prediction result.

[0021] According to a third aspect of the embodiments of this application, an electronic device is provided, comprising:

[0022] One or more sensors;

[0023] One or more processors;

[0024] Memory, used to store one or more programs;

[0025] When the one or more programs are executed by the one or more processors, the one or more processors perform the method as described in the first aspect.

[0026] According to a fourth aspect of the embodiments of this application, a computer-readable storage medium is provided that stores computer instructions thereon, which, when executed by a processor, implement the steps of the method as described in the first aspect.

[0027] The technical solutions provided by the embodiments of this application may include the following beneficial effects:

[0028] Compared with the prior art, the improvements of this invention are as follows: This method adopts a Transformer architecture based on inverse propensity scoring to perform end-to-end prediction of missing table data, overcoming the problem of low prediction efficiency for missing table data; it learns the distribution of missing states through pre-training, removes dataset bias through inverse propensity scoring, and adapts to downstream prediction tasks through semi-supervised fine-tuning, overcoming the problem of inaccurate prediction of missing table data. The prediction accuracy of this method is 25% higher than the current algorithm of first completing and then predicting, and it also has higher prediction efficiency.

[0029] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0030] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0031] Figure 1 This is a flowchart of the end-to-end prediction method for removing biased missing tables in this application.

[0032] Figure 2 This is a pre-training flowchart of the end-to-end prediction method for debiased missing tables in this application.

[0033] Figure 3 This is a block diagram of the end-to-end prediction method apparatus for removing biased missing tables according to this application. Detailed Implementation

[0034] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0035] Existing research on tabular data prediction primarily focuses on the analysis of complete tabular data. Ideally, the preferred methods for analyzing tabular data should effectively identify the complex relationships between various features, thereby achieving accurate predictions. However, they are ineffective for supporting the analysis of incomplete tabular data because they do not consider missing states or information related to features and labels. Even with advanced imputation techniques, the imputed data cannot accurately reflect the true data, especially when complex missing mechanisms exist, and it also significantly reduces prediction efficiency. This significantly impacts the predictive performance of models on incomplete tabular data, as these models heavily rely on accurately capturing the distribution of the underlying data.

[0036] To overcome the above technical problems, embodiments of this application provide an end-to-end prediction method, apparatus, and electronic device for removing biased missing tables, thereby improving prediction accuracy and efficiency.

[0037] The end-to-end prediction method for biased missing tables provided in this application is applicable to tabular data prediction in almost all scenarios, such as patient health status prediction, energy consumption, and financial modeling, and provides a good solution for data prediction problems under data missing conditions. For ease of description, this embodiment uses patient health status prediction as an example.

[0038] Figure 1 This is a flowchart of an end-to-end prediction method for removing biased missing tables provided in an embodiment of this application, as follows: Figure 1 As shown, the method may include the following steps:

[0039] S1: Obtain the feature matrix and label matrix of the missing table data, and generate the corresponding feature missing mask matrix and label missing mask matrix; this step may include the following sub-steps:

[0040] S11: Obtain the feature matrix and label matrix of the missing table data collected by the sensor;

[0041] Specifically, data corresponding to a certain disease can be selected, and the data can have multiple fields, such as a patient's gender, age, height, weight, blood pressure and other feature fields.

[0042] S12: Construct corresponding training, validation, and test sets for the missing table data feature matrix and missing table data label matrix using five-fold cross-validation;

[0043] Specifically, the feature matrix of the missing table data input can be represented as X = (X1, ..., X...). N ) T Each data sample X i =(x i1 , ..., x id ) T Where N is the number of samples and d represents the number of features. Specifically, the missing table data feature matrix and the missing table data label matrix... It can be divided into a labeled set D s ={X s ,Y}, including labeled samples X S It includes its label Y and an unlabeled set D. u ={X u}, containing unlabeled sample X u , where N=s+u.

[0044] S13: Based on the missing status of the missing table data feature matrix and the missing table data label matrix, generate the corresponding feature missing mask matrix and label missing mask matrix.

[0045] Specifically, due to issues such as improper collection, lost records, resource limitations, or privacy concerns, samples may have missing data. Therefore, it is necessary to obtain the missing data matrix M = (m1, ..., m) corresponding to sample X based on the missing data situation. d )∈{0,1} N×d , where m i =(m i1 , ..., m id ), m ij =1 means x ij A value of 0 indicates an observable state, while a value of 0 indicates a missing state.

[0046] S2: Random sampling is performed based on the feature matrix of the table data. Based on the sampling results, the reverse propensity score is calculated column by column using the XGBoost model to determine the corresponding reverse propensity score matrix. This step may include the following sub-steps:

[0047] S21: Based on the table data feature matrix and its corresponding feature missing mask matrix, randomly sample from the table data feature matrix column by column, and then concatenate the sampling results;

[0048] Specifically, for a certain feature dimension F in the feature matrix X of the missing table data... jRandomly sample N values ​​for this dimension that are missing. m Use 10 samples to construct the missing sample set for that dimension. Using the missing sample set The remaining observations are used to estimate the missing values, thus obtaining the missing data set. From this feature dimension F j Randomly sample N values ​​for this dimension that are not missing. m Use 1 sample to construct the complete sample set for this dimension. The sample set is obtained by concatenating the two sample sets.

[0049] S22: Input the concatenated matrix into XGBoost, use whether the current column is missing as the label and the other columns as features for training, use the trained model to calculate the inverse bias score of the remaining samples, and concatenate the inverse bias score matrix to obtain the inverse bias score matrix corresponding to all features and labels.

[0050] Build an XGBoost model Using the spliced ​​sample set The samples in F and their F j Using missing states as labels for training To obtain optimization pass Estimate F j The missing probability matrix P for that column is obtained by calculating the missing probabilities of the remaining samples. j The remaining columns are then calculated using the method described above to determine their missing probability, i.e., the inverse bias score. The matrices are then concatenated to obtain the inverse bias score matrix P for all samples.

[0051] S3: Construct a Transformer neural network model based on the reverse tendency score matrix;

[0052] Specifically, a Transformer neural network model based on inverse bias scoring is constructed. The neural network model includes an embedding layer and multiple Transformer blocks based on inverse bias scoring. The embedding layer consists of multiple learnable multilayer perceptrons, while each Transformer block based on inverse bias scoring consists of a probability-driven multi-head attention mechanism, a feedforward neural network, residual connections, and layer normalization.

[0053] The calculation formula for the probability-driven multi-head attention mechanism is as follows:

[0054] Q = W q ⊙X, K = W k ⊙X, V = W v ⊙X

[0055] Among them Wq W k and W v It is a learnable linear layer, where ⊙ represents element-wise multiplication.

[0056]

[0057] in, It is a modified version of the inverse propensity score matrix P, where the zero values ​​in P are replaced with negative infinity. Softmax is the softmax function. d k It is the dimension of Q and K.

[0058] S4: Mask the missing table data feature matrix and its missing feature mask matrix. Concatenate the masked missing table data feature matrix, the missing feature mask matrix, and their corresponding inverse bias score matrix. Input the concatenated matrix into the Transformer neural network model based on inverse bias score to obtain the table data representation. Reconstruct and predict the data through a linear layer. Perform self-supervised pre-training by combining the reconstructed masking information with the inverse bias score matrix to obtain the pre-trained Transformer neural network model based on inverse bias score. This step may include the following sub-steps:

[0059] S41: Apply a masking rate of γ to the feature matrix of the missing table data at the cell level to mask information at the cell level, thereby generating a masking matrix;

[0060] Specifically, a masking matrix is ​​generated for the missing table samples using a masking rate of γ. in A value of 0 indicates that the cell was manually masked, while a value of 1 indicates that the cell is observable. This application enhances the robustness of the pre-trained model by using a cell-level masking method to drive the model to learn the distribution of observed data and the complex missing data mechanisms in incomplete tabular data.

[0061] S42: Fill the masked information in the missing table data feature matrix with 0, and set the state of the masked cells in the feature missing mask matrix to missing, so as to obtain the masked missing table data feature matrix and feature missing mask matrix.

[0062] Specifically, for cells that are manually masked, all their information will be manually filled with 0 to obtain the masked feature matrix, and the masked cells in the feature missing matrix will be set as missing to obtain the masked missing matrix.

[0063] S43: After concatenating the masked missing table data feature matrix, feature missing mask matrix and inverse tendency score matrix, the feature representation is obtained by inputting it into the Transformer neural network model based on inverse tendency score in an independent manner between features.

[0064] Specifically, the masked missing table data feature matrix, feature missing mask matrix, and inverse scoring matrix are concatenated and used as input to obtain the model input. Where || represents a join operation.

[0065] First, the input is passed through the embedding layer of the Transformer neural network model based on inverse tendency scoring in an independent manner among the feature variables. The embedding layer uses multiple multilayer perceptrons to map the input to the δ dimension to obtain the embedding matrix. The specific process is as follows:

[0066]

[0067] The embedding matrix is ​​then input into a neural network consisting of L layers of Transformer blocks based on inverse propensity scoring to represent the tabular data. Specifically, for the embedding matrix... Each layer of the Transformer block will output an intermediate representation H. l ,in In each layer, the model first captures the dependencies between different locations through a probability-driven multi-head attention mechanism and layer normalization, and then outputs the results through fully connected layers and residual connected layers, as follows:

[0068] A l =H l-1 +PMHA(LN(H l-1 ))

[0069] H l =A l +MLP(LN(A l ))

[0070] Here, PMHA, LN, and MLP represent the probability-driven multi-head attention operation, layer normalization operation, and fully connected layer operation, respectively. This leads to the temporal representation H. L This application concatenates tabular data and missing information before inputting it into the model for representation, enabling the model to perceive the missing state information of the tabular data, thereby improving prediction accuracy.

[0071] S44: The feature representation is mapped to the feature value through a linear layer, and the model is trained by reconstructing the manual masking information to obtain a pre-trained Transformer neural network model based on inverse tendency scoring.

[0072] Specifically, in the pre-training phase, multiple linear layers are used to project different feature representations, resulting in X' = LL(LN(H L ), where LL represents a linear operation. The model is trained using batch gradient descent with a reconstruction error loss function, as follows:

[0073]

[0074] Where P is the inverse propensity score matrix and M is the missing data matrix. To manually mask the missing matrix, l ij The loss function is the mean absolute error function or cross-entropy loss function for the j-th feature of the i-th sample (depending on the category of that dimension; cross-entropy loss function is used for categorical data, and mean absolute error function is used for continuous data). A pre-trained model is obtained through continuous optimization and training. This application pre-trains the model by reconstructing the masked visible information, learning the distribution of missing states, and removing potential biases, thus enabling the model to obtain robust initial parameters.

[0075] S5: Divide the missing table data feature matrix into a missing table data classification feature matrix and a missing table data continuous feature matrix. After labeling them with their corresponding feature missing mask matrices, map the labeled missing table data classification feature matrix and the missing table data continuous feature matrix to a high dimension through a linear layer, and concatenate the results to obtain a high-dimensional missing table data feature matrix. This step may include the following sub-steps:

[0076] S51: Divide the missing table data feature matrix into a missing table data classification feature matrix and a missing table data continuous feature matrix according to the data category;

[0077] Specifically, the missing table data feature matrix is ​​divided into categorical data and continuous data according to the number of unique values ​​in the data column. Data columns with unique values ​​less than or equal to 20 are classified as categorical data, and those with unique values ​​greater than 20 are classified as continuous data. The categorical data is encoded starting from 1 for all categorical data.

[0078] S52: Based on the feature missing mask matrix, fill the missing cells in the missing table data classification feature matrix with -1, and fill the missing cells in the missing table data continuous feature matrix with 0, to obtain the labeled missing table data classification feature matrix and missing table data continuous feature matrix.

[0079] Specifically, missing cells in the categorical data of the missing table data feature matrix are filled with -1, and missing cells in the continuous data of the table data are filled with 0, thus obtaining labeled categorical data and continuous data. In this way, the missing values ​​of the categorical data and continuous data are labeled with different special values.

[0080] S53: Map the labeled missing table data classification feature matrix and the missing table data continuous feature matrix to a high dimension through different linear layers, and then concatenate the results to obtain the high-dimensional missing table data feature matrix.

[0081] S6: Based on the feature matrix of the high-dimensional missing table data, construct a semi-supervised fine-tuning module for de-biasing. Concatenate the high-dimensional missing table data feature matrix, the feature missing mask matrix, and their corresponding inverse bias scoring matrix. Input the concatenated matrix into the pre-trained Transformer neural network model based on inverse bias scoring to obtain feature matrix representation and prediction results. Based on the missing table data label matrix, calculate the reconstruction error for samples with missing labels and the de-biasing error based on inverse bias scoring for samples with intact labels. Fine-tune the pre-trained neural network model using the total error after weighted summation to obtain the de-biased table data prediction model. This step may include the following sub-steps:

[0082] S61: Based on the feature matrix of high-dimensional missing table data, construct a semi-supervised fine-tuning module to remove bias. This module consists of a Transformer layer based on inverse propensity score and multiple linear layers, which can map the feature matrix of high-dimensional missing table data into a prediction matrix for sample labels.

[0083] Specifically, a one-dimensional all-zero CLStoken is added to the end of the classification data matrix. This column will integrate the remaining values ​​of the feature matrix in the representation to form a label prediction representation.

[0084] S62: The high-dimensional missing table data feature matrix, missing matrix and inverse tendency scoring matrix are concatenated and input into the pre-trained neural network model in a feature-independent manner to obtain the feature representation matrix and label representation matrix.

[0085] Specifically, a one-dimensional full-1D is added to the tail of the missing matrix and the inverse tendency score to embed CLStoken, and the continuous data matrix and the categorical data matrix are concatenated to form the feature matrix. As input;

[0086] S63: Map the feature representation matrix to feature prediction results through multiple linear layers, and calculate the reconstruction error between the features of the label missing sample and the feature prediction results based on the label missing mask matrix;

[0087] Specifically, multiple linear layers are used to project different feature representations to obtain X′=LL(LN(H L Where LL represents a linear layer and LN represents layer normalization. The reconstruction error calculation formula is as follows:

[0088]

[0089] Where u is the number of samples with missing labels, P is the inverse bias rating matrix, and M is the missing data matrix. To manually mask the missing matrix, l ij It is the mean absolute error function or cross-entropy loss function for the j-th feature of the i-th sample (depending on the category of that dimension; cross-entropy loss function for categorical data, and mean absolute error function for continuous data);

[0090] S64: Map the label representation to the label prediction result through a linear layer, and calculate the supervision error between the label of the sample with no missing label and the label prediction result based on the label missing mask matrix;

[0091] Specifically, a linear layer is used to project the CLStoken representation, resulting in Y′=LL(LN(CLStoken) ...)=LL(LN(CLStoken))=LL(LN(CLStoken))=LL(LN L Where LL represents a linear layer and LN represents layer normalization. The formula for calculating the supervision error is as follows:

[0092]

[0093] Where s is the number of samples with no missing labels, C is the number of label categories, and P l The reverse tendency score for the label is calculated using the above module;

[0094] S65: The reconstruction error and the supervision error are weighted and summed. The error after weighted summation is minimized to fine-tune the pre-trained Transformer neural network model based on inverse bias scoring, thereby obtaining the biased missing table data prediction model.

[0095] Specifically, the reconstruction error and the supervision error are weighted and summed to obtain the semi-supervised fine-tuning error. The pre-trained neural network model is fine-tuned by minimizing the semi-supervised fine-tuning error, with the following specific objectives:

[0096]

[0097] Where β is a hyperparameter used to control the supervision loss. and reconstruction loss The balance between features is achieved by using a semi-supervised prediction module to predict the future, thereby utilizing all observable information and considering the correlation between features to further improve prediction accuracy and enhance the ability to perform prediction tasks.

[0098] S7: Input the concatenated missing table data feature matrix, feature missing mask matrix and its corresponding inverse bias score matrix into the debiased table data prediction model to predict the missing table data and obtain the final prediction result.

[0099] Specifically, the feature matrix and the missing matrix are concatenated to obtain [X||M], which is then input into the biased missing table prediction model to obtain the prediction result. The CLStoken is then input into a linear layer to obtain the final label prediction result.

[0100] Corresponding to the aforementioned embodiments of the end-to-end prediction method for debiased and missing table data, this application also provides embodiments of an end-to-end prediction apparatus for debiased and missing table data.

[0101] Figure 3 This is a block diagram of the end-to-end prediction method apparatus for removing biased missing tables according to this application. (Refer to...) Figure 3 The device includes:

[0102] Module 1 is used to obtain the feature matrix and label matrix of the missing table data, and generate the corresponding feature missing mask matrix and label missing mask matrix.

[0103] The reverse tendency scoring module 2 is used to randomly sample the feature matrix of the table data, and calculate the reverse tendency score column by column based on the sampling results using the XGBoost model to determine the corresponding reverse tendency score matrix.

[0104] Module 3 is used to construct a Transformer neural network model based on the reverse tendency score according to the reverse tendency score matrix;

[0105] Pre-training module 4 is used to manually mask the feature matrix of the missing table data and its feature missing mask matrix. The masked feature matrix of the missing table data, the feature missing mask matrix and its corresponding inverse bias score matrix are concatenated and then input into the Transformer neural network model based on inverse bias score to obtain the table data representation. Reconstruction prediction is performed through a linear layer. Self-supervised pre-training is performed by combining the reconstructed masking information with the inverse bias score matrix to obtain the pre-trained Transformer neural network model based on inverse bias score.

[0106] Data processing module 5 is used to divide the missing table data feature matrix into a missing table data classification feature matrix and a missing table data continuous feature matrix. After marking it with its corresponding feature missing mask matrix, the marked missing table data classification feature matrix and missing table data continuous feature matrix are respectively mapped to high dimension through a linear layer, and the results are concatenated to obtain a high-dimensional missing table data feature matrix.

[0107] The fine-tuning module 6 is used to construct a semi-supervised fine-tuning module for de-biasing based on the feature matrix of the high-dimensional missing table data. It concatenates the feature matrix of the high-dimensional missing table data, the feature missing mask matrix and its corresponding inverse bias score matrix, and inputs the concatenated matrix into the pre-trained Transformer neural network model based on the inverse bias score to obtain the feature matrix representation and prediction results. Based on the label matrix of the missing table data, it calculates the reconstruction error for samples with missing labels and the de-biasing error based on the inverse bias score for samples with intact labels. It then uses the total error after weighted summation to fine-tune the pre-trained neural network model to obtain the de-biased table data prediction model.

[0108] Prediction module 7 inputs the spliced ​​missing table data feature matrix, feature missing mask matrix and its corresponding inverse bias score matrix into the debiased table data prediction model to predict the missing table data and obtain the final prediction result.

[0109] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0110] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0111] Accordingly, this application also provides an electronic device, comprising: one or more sensors; one or more processors; and a memory for storing one or more programs; wherein when the one or more programs are executed by the one or more processors, the one or more processors perform the method as described in the first aspect.

[0112] Accordingly, this application also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the end-to-end prediction method for biased missing table data as described above.

[0113] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.

[0114] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. An end-to-end prediction method for biased missing tables, characterized in that, include: S1: Obtain the feature matrix and label matrix of the missing table data, and generate the corresponding feature missing mask matrix and label missing mask matrix; S2: Randomly sample the feature matrix of the table data, and calculate the reverse tendency score column by column using the XGBoost model based on the sampling results to determine the corresponding reverse tendency score matrix. S3: Construct a Transformer neural network model based on the reverse tendency score matrix; S4: Mask the missing table data feature matrix and its missing feature mask matrix. Concatenate the masked missing table data feature matrix, the missing feature mask matrix and its corresponding inverse bias score matrix. Input the concatenated matrix into the Transformer neural network model to obtain the table data representation. Reconstruct and predict the data through a linear layer. Use the method of reconstructing the masking information to combine with the inverse bias score matrix for self-supervised pre-training to obtain the pre-trained Transformer neural network model based on inverse bias score. S5: Divide the missing table data feature matrix into a missing table data classification feature matrix and a missing table data continuous feature matrix. After marking them with their corresponding feature missing mask matrices, map the marked missing table data classification feature matrix and missing table data continuous feature matrix to a high dimension through a linear layer, and then concatenate the results to obtain a high-dimensional missing table data feature matrix. S6: Based on the high-dimensional missing table data feature matrix, construct a debiased semi-supervised fine-tuning module. Concatenate the high-dimensional missing table data feature matrix, the feature missing mask matrix, and their corresponding inverse bias score matrix. Input the concatenated matrix into the pre-trained Transformer neural network model based on inverse bias score to obtain feature matrix representation and prediction results. Based on the missing table data label matrix, calculate the reconstruction error for samples with missing labels and calculate the debiased error based on inverse bias score for samples with intact labels. Use the total error after weighted summation to fine-tune the pre-trained neural network model to obtain the debiased table data prediction model. S7: Input the concatenated missing table data feature matrix, feature missing mask matrix and its corresponding inverse bias score matrix into the debiased table data prediction model to predict the missing table data and obtain the final prediction result.

2. The method according to claim 1, characterized in that, S1 specifically includes: S11: Obtain the feature matrix and label matrix of the missing table data collected by the sensor; S12: Construct corresponding training, validation, and test sets for the missing table data feature matrix and missing table data label matrix using five-fold cross-validation; S13: Based on the missing status of the missing table data feature matrix and the missing table data label matrix, generate the corresponding feature missing mask matrix and label missing mask matrix.

3. The method according to claim 1, characterized in that, S2 specifically includes: S21: Based on the table data feature matrix and its corresponding feature missing mask matrix, randomly sample from the table data feature matrix column by column, and then concatenate the sampling results; S22: Input the concatenated matrix into XGBoost, use whether the current column is missing as the label and the other columns as features for training, use the trained model to calculate the inverse bias score of the remaining samples, and concatenate the inverse bias score matrix to obtain the inverse bias score matrix corresponding to all features and labels.

4. The method according to claim 1, characterized in that, The neural network model includes an embedding layer and multiple Transformer blocks based on inverse bias scoring. The embedding layer consists of multiple learnable multilayer perceptrons, while each Transformer block based on inverse bias scoring mainly consists of a probability-driven multi-head attention mechanism, a feedforward neural network, residual connections, and layer normalization.

5. The method according to claim 1, characterized in that, S4 specifically includes: S41: Apply a masking rate of γ to the feature matrix of the missing table data at the cell level to mask information at the cell level, thereby generating a masking matrix; S42: Fill the masked information in the missing table data feature matrix with 0, and set the state of the masked cells in the feature missing mask matrix to missing, so as to obtain the masked missing table data feature matrix and the feature missing mask matrix. S43: After concatenating the masked missing table data feature matrix, feature missing mask matrix and inverse tendency score matrix, the feature representation is obtained by inputting it into the Transformer neural network model based on inverse tendency score in an independent manner between features. S44: The feature representation is mapped to the feature value through a linear layer, and the model is trained by reconstructing the manual masking information to obtain a pre-trained Transformer neural network model based on inverse tendency scoring.

6. The method according to claim 1, characterized in that, S5 specifically includes: S51: Divide the missing table data feature matrix into a missing table data classification feature matrix and a missing table data continuous feature matrix according to the data category; S52: Based on the feature missing mask matrix, fill the missing cells in the missing table data classification feature matrix with -1, and fill the missing cells in the missing table data continuous feature matrix with 0, to obtain the marked missing table data classification feature matrix and missing table data continuous feature matrix. S53: Map the labeled missing table data classification feature matrix and the missing table data continuous feature matrix to a high dimension through different linear layers, and then concatenate the results to obtain the high-dimensional missing table data feature matrix.

7. The method according to claim 1, characterized in that, S6 specifically includes: S61: Based on the feature matrix of high-dimensional missing table data, construct a semi-supervised fine-tuning module to remove bias. This module consists of a Transformer layer based on inverse propensity score and multiple linear layers, which can map the feature matrix of high-dimensional missing table data into a prediction matrix for sample labels. S62: The high-dimensional missing table data feature matrix, missing matrix and inverse tendency scoring matrix are concatenated and input into the pre-trained neural network model in a feature-independent manner to obtain the feature representation matrix and label representation matrix. S63: Map the feature representation matrix to feature prediction results through multiple linear layers, and calculate the reconstruction error between the features of the label missing sample and the feature prediction results based on the label missing mask matrix; S64: Map the label representation to the label prediction result through a linear layer, and calculate the supervision error between the label of the sample with no missing label and the label prediction result based on the label missing mask matrix; S65: The reconstruction error and the supervision error are weighted and summed. The error after weighted summation is minimized to fine-tune the pre-trained Transformer neural network model based on inverse bias scoring, thereby obtaining the biased missing table data prediction model.

8. An end-to-end prediction device for multivariate missing time series data, characterized in that, include: The acquisition module is used to acquire the feature matrix and label matrix of missing table data, and generate the corresponding feature missing mask matrix and label missing mask matrix. The reverse tendency scoring module is used to randomly sample the feature matrix of the table data, and calculate the reverse tendency score column by column based on the sampling results using the XGBoost model to determine the corresponding reverse tendency score matrix. A construction module is used to construct a Transformer neural network model based on the inverse tendency score according to the inverse tendency score matrix; The pre-training module is used to mask the feature matrix of the missing table data and its feature missing mask matrix. The masked feature matrix of the missing table data, the feature missing mask matrix and its corresponding inverse bias score matrix are concatenated and then input into the Transformer neural network model to obtain the table data representation. Reconstruction prediction is performed through a linear layer. Self-supervised pre-training is performed by combining the reconstructed masking information with the inverse bias score matrix to obtain the pre-trained Transformer neural network model based on the inverse bias score. The data processing module is used to divide the missing table data feature matrix into a missing table data classification feature matrix and a missing table data continuous feature matrix. After marking them with their corresponding feature missing mask matrices, the marked missing table data classification feature matrix and missing table data continuous feature matrix are mapped to a high dimension through a linear layer, and the results are concatenated to obtain a high-dimensional missing table data feature matrix. The fine-tuning module is used to construct a semi-supervised fine-tuning module for de-biasing based on the feature matrix of the high-dimensional missing table data. The high-dimensional missing table data feature matrix, the feature missing mask matrix and its corresponding inverse bias score matrix are concatenated and then input into the pre-trained Transformer neural network model based on inverse bias score to obtain feature matrix representation and prediction results. Based on the missing table data label matrix, the reconstruction error is calculated for samples with missing labels and the de-biasing error based on inverse bias score is calculated for samples with intact labels. The total error after weighted summation is used to fine-tune the pre-trained neural network model to obtain the de-biased table data prediction model. The prediction module inputs the concatenated missing table data feature matrix, feature missing mask matrix and its corresponding inverse bias score matrix into the debiased table data prediction model to predict the missing table data and obtain the final prediction result.

9. An electronic device, characterized in that, include: One or more sensors; One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-7.

10. A computer-readable storage medium storing computer instructions thereon, characterized in that, When executed by the processor, this instruction implements the steps of the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Table image cell region identification method and system and medium

    CN116778512A

  • Data processing method and related device

    CN116910358A