Missing data interpolation method and system based on generative adversarial network

By clustering and eigenvector modeling of the missing data matrix and combining it with a generative adversarial network framework, we overcome the limitations of traditional methods in complex data processing, achieve highly accurate and robust data interpolation, and adapt to different missing mechanisms and environments.

CN120763646APending Publication Date: 2025-10-10QINGHAI NORMAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510930415.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

When dealing with complex and nonlinear data, traditional statistical methods are unable to cope with the existing technologies. GAN-based interpolation methods are insufficient in capturing local structures and lack guiding information modeling for the interpolation process, resulting in a lack of interpretability and consistency in interpolation results in cases of high missing rates, strong heterogeneity or small samples.

Method used

By clustering the missing data matrix, using K-means clustering and DTW algorithm to identify local structures, combining logistic regression and Gaussian mixture model for feature vector classification and probability distribution modeling, a generative adversarial network framework is constructed, self-attention mechanism and deep denoising autoencoder are used for data interpolation, and the interpolation loss function is used to optimize the generator and discriminator to realize the training of the interpolation model.

Benefits of technology

It significantly improves the accuracy and robustness of interpolation in complex time series data environments, can effectively capture the local structure and global distribution of data, maintain the interpretability and consistency of interpolation results, adapt to different missing mechanisms, and improve the interpolation accuracy in high missing rate scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763646A_ABST
    Figure CN120763646A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and provides a missing data interpolation method and system based on a generative adversarial network. The method comprises the following steps: clustering a missing data matrix to obtain a clustering cluster containing a cluster label; based on the clustering cluster, performing classification prediction on a data feature vector corresponding to the missing data matrix through a logistic regression algorithm to obtain a cluster label prediction model; performing probability distribution modeling on the data feature vector through a Gaussian mixture model to obtain a probability model; training a generative adversarial network framework based on the cluster label prediction model and the probability model to obtain an interpolation model; and interpolating data to be interpolated through the interpolation model to obtain an interpolation data matrix. According to the invention, the interpolation precision and stability of nonlinear data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a missing data interpolation method and system based on a generative adversarial network. Background Art

[0002] Missing data is a widespread problem in many fields, including healthcare, finance, industrial manufacturing, and environmental science, severely impacting the performance of data-driven models and the accuracy of data analysis. Traditional data interpolation methods are primarily categorized into two types: simple statistical interpolation and machine learning-based interpolation. Simple statistical methods, such as mean interpolation and regression interpolation, use statistics to fill missing values, while machine learning methods, such as k-nearest neighbor interpolation and decision tree interpolation, leverage correlations between data to estimate missing values.

[0003] However, existing technologies still have significant shortcomings: First, traditional statistical methods often assume that data are linearly correlated, which makes them incapable of handling complex and nonlinear data, and may change the original distribution of the data; Second, although GAN-based interpolation methods can learn the global data distribution, they still have deficiencies in capturing local structures. They usually focus more on fitting the global distribution and have poor interpolation effects in local areas; Third, many existing methods lack guiding information modeling for the interpolation process, treating missing values ​​as completely unknown variables, ignoring the potential structural information in existing observations, and easily causing the model to generate pseudo samples that deviate from the true distribution in scenarios with high missing rates or complex missing mechanisms; Finally, most existing methods perform poorly in dealing with high missing rates, strong heterogeneity or small samples, and it is difficult to effectively model the complex correlations and conditional dependencies between different features, making the interpolation results lack interpretability and consistency. Summary of the Invention

[0004] The present invention provides a missing data interpolation method and system based on a generative adversarial network to address the shortcomings of the prior art.

[0005] The present invention provides a missing data interpolation method based on a generative adversarial network, comprising:

[0006] S1: Cluster the missing data matrix to obtain clusters containing cluster labels;

[0007] S2: Based on the clusters, classify and predict the data feature vectors corresponding to the missing data matrix using a logistic regression algorithm to obtain a cluster label prediction model;

[0008] S3: Performing probability distribution modeling on the data feature vector using a Gaussian mixture model to obtain a probability model;

[0009] S4: Based on the cluster label prediction model and the probability model, a generative adversarial network framework is trained to obtain an interpolation model;

[0010] S5: Interpolate the data to be interpolated using the interpolation model to obtain an interpolation data matrix.

[0011] According to a missing data interpolation method based on a generative adversarial network provided by the present invention, step S1 further includes:

[0012] S11: Input a missing data matrix, calculate the distances between multiple groups of time series data points in the missing data matrix, and obtain a distance matrix;

[0013] S12: Iteratively optimizing the distance matrix based on K-means clustering to obtain multiple cluster centers;

[0014] S13: performing cluster attribution determination on the data points in the missing data matrix according to the cluster centers to obtain a plurality of clusters, wherein each cluster includes a corresponding cluster label.

[0015] According to a missing data interpolation method based on a generative adversarial network provided by the present invention, step S2 further includes:

[0016] S21: perform standardization preprocessing on the data eigenvector corresponding to the missing data matrix to obtain a standardized eigenvector;

[0017] S22: performing probability mapping on the standardized feature vector through a logistic regression model to obtain a preliminary prediction model;

[0018] S23: Optimizing the regression coefficients of the preliminary probability model through maximum likelihood estimation to obtain a cluster label prediction model.

[0019] According to a missing data interpolation method based on a generative adversarial network provided by the present invention, step S3 further includes:

[0020] S31: Based on the probability density function, perform multi-Gaussian component modeling on the data distribution of the data feature vector to obtain preliminary model parameters;

[0021] S32: Iteratively estimating the preliminary model parameters by an expectation maximization algorithm to obtain updated model parameters;

[0022] S33: Based on the updated model parameters, a probability model for outputting characteristic conditional distribution parameters is established.

[0023] According to a missing data interpolation method based on a generative adversarial network provided by the present invention, step S4 further includes:

[0024] S41: constructing a generator and a discriminator to obtain a generative adversarial network framework;

[0025] S42: initializing the generative adversarial network framework based on the cluster label prediction model and the probability model;

[0026] S43: performing adversarial training optimization on the initialized adversarial network framework according to an imputation loss to obtain an imputation model.

[0027] According to the imputation method for missing data based on the generative adversarial network provided by the application, the expression of the imputation loss in step S43 is:

[0028] L=L D +L G ;

[0029]

[0030] Wherein, L is an imputation loss function, L G is a generator loss, L D is a discriminator loss, b i is a batch index variable, m i is the i th mask vector, is the imputed estimated data corresponding to the mask vector, is the imputed estimated data corresponding to the i th mask vector, alpha is a hyperparameter for balancing the generator loss and the discriminator loss weight, j is a dimension number index of data, d is a dimension number of data, L obs (x i ,x i ′) is an observation loss function, x i is a real observation value, and x i ′ is a prediction output value of the generator.

[0031] According to the imputation method for missing data based on the generative adversarial network provided by the application, step S5 further comprises:

[0032] S51: inputting the data to be imputed;

[0033] S52: performing standardization processing on the data to be imputed to obtain predicted data;

[0034] S53: inputting the predicted data into the imputation model to obtain an imputation data matrix.

[0035] According to the imputation method for missing data based on the generative adversarial network provided by the application, step S53 specifically comprises:

[0036] generating missing values for the predicted data by the generator in the imputation model to obtain imputation data;

[0037] The interpolation data is screened for authenticity through the discriminator in the interpolation model to obtain an interpolation data matrix.

[0038] According to a missing data interpolation method based on a generative adversarial network provided by the present invention, the expression of the interpolation data is:

[0039]

[0040] in, is the generated interpolation data, x is the incomplete data, m is the mask vector, G is the generator, z is the noise added to the input data, and ⊙ is the element-by-element multiplication.

[0041] The present invention also provides a missing data interpolation system based on a generative adversarial network, comprising:

[0042] Clustering module: used to cluster the missing data matrix to obtain cluster clusters containing cluster labels;

[0043] Prediction module: used to classify and predict the data feature vectors corresponding to the missing data matrix based on the cluster clusters through a logistic regression algorithm, and train and generate a cluster label prediction model;

[0044] Modeling module: used to perform probability distribution modeling on the data feature vector through Gaussian mixture model, and train and generate a probability model;

[0045] Training module: used to train the generative adversarial network framework based on the cluster label prediction model and the probability model to obtain an interpolation model;

[0046] The interpolation module is configured as the interpolation model, and is used to interpolate the data to be interpolated to obtain an interpolation data matrix.

[0047] The present invention provides a missing data interpolation method and system based on a generative adversarial network. The missing data matrix is ​​clustered by the K-means clustering algorithm of the DTW indicator, which can effectively identify local similar patterns in time series data, overcome the limitation of the traditional Euclidean distance that cannot handle the time distortion problem, and enable similar time series segments to be accurately classified into the same cluster, providing more accurate local structural information for subsequent interpolation, and significantly improving the interpolation accuracy in a complex time series data environment; secondly, the present invention combines the logistic regression algorithm to classify the data feature vector, and establishes a nonlinear mapping relationship between the feature vector and the cluster label through a probability mapping mechanism, which can not only capture the classification boundary characteristics of the data, but also provide supervised guidance information for the generative adversarial network, effectively solving the interpolation bias problem caused by the lack of category supervision in the traditional GAN ​​method, making the interpolation process more stable and controllable; secondly, the introduction of the Gaussian mixture model of the present invention can accurately characterize the probability distribution characteristics of complex data, provide the generator with rich prior distribution information, so that the generated interpolation value is more consistent with the statistical characteristics of the real data, especially in processing multimodal distribution. The present invention performs well when adapting to data, avoiding the interpolation bias caused by the single distribution assumption; the present invention is also based on the generator of the self-attention mechanism to dynamically capture the dependencies between different positions in the input sequence, and simultaneously focus on local and global features through the multi-head attention mechanism. Combined with the generalization ability of the deep denoising autoencoder, the interpolation process can fully utilize the temporal dependencies and spatial correlations in the observed data, and can still maintain a high interpolation accuracy in high missing rate scenarios; the subsequent mask prediction discriminator accurately distinguishes between real components and false components through the cross-entropy loss function, which not only provides the gradient signal for adversarial training, but also enhances the model's understanding of missing patterns through the mask vector prediction mechanism, enabling the generator to adjust the interpolation strategy according to different missing patterns, significantly improving the adaptability and robustness of the model under different missing mechanisms; the weighted combination design of the basic generation loss and the mask reconstruction loss in the generator loss function can ensure the authenticity of the generated data while ensuring the consistency of the observed data, and balances the requirements of global distribution fitting and local structure preservation through the multi-objective optimization mechanism, effectively avoiding the problem of degraded interpolation quality caused by the imbalance of global and local information in traditional methods.

[0048] In general, the time series similarity calculation capability of the DTW distance metric in the present invention makes the clustering results more consistent with the inherent laws of time series data. The probabilistic output characteristics of logistic regression provide interpretable classification confidence for the subsequent generation process. The multimodal modeling capability of the Gaussian mixture model can handle complex data distribution forms. The long-distance dependency capture capability of the self-attention mechanism solves the gradient vanishing problem of traditional RNN in long sequence processing. The minmax game mechanism of generative adversarial training ensures the high quality and diversity of the generated data. The synergistic effect of multiple algorithm features enables the overall solution to perform excellent performance in processing high-dimensional, high-missing-rate, and nonlinear complex data. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0050] Figure 1 A schematic flow chart of a missing data interpolation method based on a generative adversarial network provided in an embodiment of the present invention;

[0051] Figure 2 A schematic diagram of the structure of a missing data interpolation system based on a generative adversarial network provided by an embodiment of the present invention;

[0052] Figure 3 A schematic diagram comparing the MAE values ​​of various model configurations under different missing rates provided by an embodiment of the present invention;

[0053] Figure 4 This is a schematic diagram comparing the RMSE values ​​of various model configurations under different missing rates provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the drawings in the present invention. Obviously, the embodiments described are part of the embodiments of the present invention, not all of the embodiments, and they should not be understood as limitations on the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. In the description of the present invention, it should be understood that the terms used are only for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0055] In order to better understand the present invention, the research background of the present invention is first explained in detail below.

[0056] Data missing is a common and serious problem that widely exists in many fields such as medical, financial, industrial manufacturing, environmental science, etc., and seriously affects the effectiveness of data-driven models and the accuracy of data analysis. Studies have shown that missing values not only weaken the quality of data mining, but also may introduce bias in the modeling stage, reduce prediction accuracy, and even cause model failure. Therefore, how to scientifically and effectively perform data imputation has become a research hotspot in the field of data science in recent years.

[0057] Traditional data imputation methods can be roughly divided into two categories: simple statistical imputation methods such as mean imputation and regression imputation; and machine learning-based imputation methods such as k-nearest neighbor imputation and decision tree imputation. Although these methods can provide effective solutions in certain cases, they also have obvious limitations. First, simple statistical methods often assume that data is linearly related, which is not sufficient when dealing with complex and nonlinear data. Second, machine learning-based methods can capture complex patterns in data, but they often require a large amount of complete data for training and may change the original distribution of the data.

[0058] In recent years, the introduction of generative adversarial networks (GAN) has provided a new approach to data imputation. GAN-based methods use generators and discriminators to learn the distribution of data and generate missing data. These methods perform well in enhancing feature expression and following data distribution. However, GAN-based imputation methods still face two major challenges: one is the lack of local structure capture ability. GANs usually focus more on fitting the global distribution and perform poorly in local area imputation; the second is the lack of modeling of guiding information in the imputation process. Many GAN-based imputation methods treat missing values as completely unknown variables, ignoring the potential structural information in existing observations, which can easily lead to the generation of pseudo samples that deviate from the true distribution in scenarios with high missing rates or complex missing mechanisms. For example, in tabular structured data, there are often complex correlations and conditional dependencies between different features. If these internal relationships are not effectively modeled, the imputation results will lack interpretability and consistency.

[0059] However, most existing methods still perform poorly in handling high missing rates, strong heterogeneity, or small sample sizes. This limitation is particularly common in real-world data, such as high proportions of missing data in medical data due to ethical or operational restrictions, or large areas of data voids in industrial systems due to sensor failures. Therefore, there is an urgent need for a generative imputation method that can both preserve data distribution information and model local structural features to improve the robustness and accuracy of imputation in complex environments.

[0060] In addition to traditional data interpolation methods, deep learning methods have also emerged in recent years. Each of these methods has its own unique advantages and applicable scenarios. With the advancement of technology, deep learning methods have shown increasing capabilities in processing complex data structures and large-scale datasets.

[0061] GAIN is a generative adversarial network model for data interpolation. GAIN demonstrates its effectiveness in dealing with missing data and the potential of generative adversarial networks in data interpolation tasks. Through adversarial training of the generator and discriminator, GAIN is able to generate realistic interpolated values ​​on incomplete datasets.

[0062] In practical applications, GAIN has been used for data interpolation tasks in various fields, such as medical data, transportation data, and environmental data. Research has shown that GAIN can effectively handle missing data in these fields and improve the accuracy and reliability of data analysis.

[0063] Recently, clustering and classification techniques have been increasingly used in data interpolation, particularly in conjunction with generative adversarial network (GAN) models. Several studies have attempted to improve the accuracy and robustness of interpolation by combining clustering and classification techniques with GAN models. A common approach to GAN clustering is to first transform the complex original data into a more tractable latent space and then perform clustering in this space. For example, methods such as Deep Adversarial Gaussian Mixture (DAGC), DAGAN-GMM, and Information Maximization Deep Generative Clustering (IMDGC) all employ this strategy. In addition to these GAN clustering methods derived from GMMs, several other approaches exist. For example, Classification Generative Adversarial Network (CATGAN) is based on Regularized Information Maximization (RIM), while InfoGAN and ClusterGAN focus on learning interpretable latent spaces. However, in some cases, the generated latent spaces of these methods may not fully capture the complexity of the original data, resulting in inaccurate clustering results. They also sometimes make overly strict assumptions about the data distribution, limiting their applicability to diverse datasets.

[0064] Several studies have improved the GAIN model and applied it to missing data imputation. For example, a novel semi-supervised generative adversarial network model, called SSGAN, introduces a temporal reminder matrix to help the discriminator better distinguish between observed and imputed components. It is used for missing value imputation in multivariate time series data. Another study introduced WGAIN as a Wasserstein correction to GAIN. It is the best imputation model when the missingness is less than or equal to 30%, but its performance may not be as good as expected at higher missingness rates. Another study proposed three data imputation methods based on generative adversarial networks (GANs): SGAIN, WSGAIN-CP, and WSGAIN-GP, for Missing Completely at Random (MCAR) datasets. These methods perform comparable to or better than existing state-of-the-art imputation methods. However, these methods are specifically targeted at Missing Completely at Random (MCAR) and may not perform as well as expected for other missingness mechanisms (such as non-random missingness). Another end-to-end generative model, E2GAN, is also used to impute missing values ​​in multivariate time series. With the help of discriminative loss and squared error loss, E2GAN can estimate incomplete time series through the complete time series generated most recently in a certain stage, but this method is mainly aimed at multivariate time series data and is not applicable to other types of data sets or missing patterns.

[0065] Furthermore, the improved generative adversarial network models mentioned above have not yet considered using implicit category information to improve interpolation performance. When working with complex datasets, these methods often struggle to effectively capture the data's inherent structure and relationships between features, resulting in insufficient interpolation accuracy and classification performance. Furthermore, they struggle to achieve precise pattern classification and feature extraction.

[0066] To this end, this paper proposes a generative adversarial interpolation network based on contrastive predictive coding (CPC-GAIN), which aims to improve the accuracy of data interpolation by accurately simulating the data generation process, enabling the model to adapt to data sets of different types and structures, and providing transparency of model decision-making through causal reasoning and cluster analysis.

[0067] The embodiments of the present invention are described below with reference to the accompanying drawings.

[0068] like Figure 1 As shown, the present invention provides a missing data interpolation method based on a generative adversarial network, comprising:

[0069] S1: Cluster the missing data matrix to obtain clusters containing cluster labels.

[0070] Furthermore, the present invention first applies the K-means clustering algorithm and a clustering algorithm based on Dynamic Time Warping (DTW) to identify potential groups in the data. The K-means algorithm automatically classifies data points by dividing the dataset into K clusters. The DTW-based clustering algorithm can more accurately process the changes and differences in time series data, and by calculating the time series similarity between data points, it helps identify groups that are similar along the time axis.

[0071] Wherein, step S1 further includes:

[0072] S11: Input a missing data matrix, calculate the distances between multiple groups of time series data points in the missing data matrix, and obtain a distance matrix.

[0073] Furthermore, the input missing data matrix is ​​a multidimensional time series data set containing some missing values, wherein each row represents multiple feature observations at a time point, and each column represents a sequence of values ​​of a certain feature at different time points. In step S11, the present invention measures the similarity of the two time series by the DTW distance. Finally, a complete distance matrix can be obtained, which contains the DTW distance values ​​between all time series data points in the missing data matrix.

[0074] S12: Iteratively optimize the distance matrix based on K-means clustering to obtain multiple cluster centers.

[0075] Furthermore, the K-means clustering algorithm in step S12 of the present invention divides the data points into K clusters, so that the distance between the data points in each cluster is as small as possible, and the distance between different clusters is as large as possible. Specifically, when processing the distance matrix obtained in step S11, the DTW distance value is used as the similarity metric, and the cluster center position is gradually adjusted through an iterative optimization process, and finally K optimal cluster center points are determined. The obtained cluster centers can represent the typical characteristics of different data patterns in the missing data matrix.

[0076] S13: performing cluster attribution determination on the data points in the missing data matrix according to the cluster centers to obtain a plurality of clusters, wherein each cluster includes a corresponding cluster label.

[0077] Furthermore, in step S13, the present invention classifies data points based on the distance minimization principle. Specifically, each data point in the missing data matrix is ​​traversed, the DTW distance between the data point and all cluster centers is calculated, and the data point is assigned to the cluster corresponding to the cluster center with the smallest distance. The cluster label refers to the unique identifier assigned to each cluster. Each data point will obtain a corresponding cluster label after the cluster affiliation determination is completed.

[0078] S2: Based on the clustering clusters, the data feature vectors corresponding to the missing data matrix are classified and predicted by a logistic regression algorithm to obtain a cluster label prediction model.

[0079] Wherein, step S2 further includes:

[0080] S21: Perform standardization preprocessing on the data eigenvector corresponding to the missing data matrix to obtain a standardized eigenvector.

[0081] Furthermore, the standardized preprocessing eliminates the negative impact of differences in feature dimensions and numerical ranges on model training through mathematical transformation. The standardized data flow first calculates the statistics of each feature dimension, and then applies the standardized formula sample by sample and feature by feature, and finally generates a standardized feature matrix of the same dimension as the original data matrix, avoiding excessive influence of features with larger values ​​on model training. At the same time, data preprocessing can also ensure that the input data is standardized before using logistic regression, ensuring that the value ranges of all features are within the same order of magnitude. The data preprocessing can not only accelerate the convergence speed of the model, but also improve the accuracy of the prediction.

[0082] S22: Probability mapping is performed on the standardized feature vector through a logistic regression model to obtain a preliminary prediction model.

[0083] Furthermore, the above-mentioned logistic regression model maps the output results of linear regression to a probability space between 0 and 1 through the sigmoid activation function, thereby realizing probability prediction in the classification task. In the present invention, logistic regression is used as a classifier, which undertakes the task of predicting cluster labels from feature vectors, and finally normalizes the probabilities of each category through the softmax function so that their sum is 1.

[0084] S23: Optimizing the regression coefficients of the preliminary probability model through maximum likelihood estimation to obtain a cluster label prediction model.

[0085] Furthermore, the maximum likelihood estimation determines the optimal value of the model parameters by maximizing the likelihood function of the probability of occurrence of the sample data. Specifically, during the training process, the model adjusts the regression coefficient by maximizing the log-likelihood function, learns the mapping relationship between features and cluster labels, and outputs the predicted value of each data point as a probability. The model updates the parameters based on the deviation between this probability and the actual label. After the training is completed, the present invention uses the cross-validation method to evaluate the performance of the model. Cross-validation can effectively avoid the overfitting problem and provide a reliable estimation index for the generalization ability of the model. The final prediction model is based on the logistic regression model, with the characteristic vector of the data as the input variable and the cluster label provided by K-Means as the target variable.

[0086] S3: Performing probability distribution modeling on the data feature vector using a Gaussian mixture model to obtain a probability model.

[0087] Wherein, step S3 further includes:

[0088] S31: Based on the probability density function, perform multi-Gaussian component modeling on the data distribution of the data feature vector to obtain preliminary model parameters.

[0089] Furthermore, multi-Gaussian component modeling decomposes complex data distribution into a weighted combination of multiple Gaussian distributions through a probability density function, where each Gaussian distribution is called a Gaussian component, representing a potential subgroup or cluster in the data. This initialization method makes full use of the prior knowledge provided by the clustering results and provides a reasonable starting point for subsequent parameter optimization.

[0090] S32: Iteratively estimating the preliminary model parameters through an expectation maximization algorithm to obtain updated model parameters; S33: Establishing a probability model for outputting characteristic conditional distribution parameters based on the updated model parameters.

[0091] Since each Gaussian distribution corresponds to a cluster in the data, the model assumes that the data points are generated from multiple different Gaussian distributions. Therefore, in steps S32 to S33, the present invention can model the potential distribution of the data through GMM and use the expectation maximization algorithm to solve the parameters of the model.

[0092] The goal of GMM is to estimate the parameters of the model, namely the mean, covariance matrix and weight of each Gaussian distribution, by maximizing the log-likelihood function based on a given data set. It is usually solved by the expectation-maximization (EM) algorithm. The EM algorithm consists of two main steps: expectation step (E step): based on the current model parameters, calculate the posterior probability that each data point belongs to each Gaussian distribution; maximization step (M step): based on the posterior probability, update the model parameters (mean, covariance matrix and weight).

[0093] The final updated Gaussian probability model is used to estimate the conditional distribution of features, which can enhance the accuracy of generated data. The above model can provide more accurate prior information for the generation model, so that the generator can better follow the law of data distribution when generating data, thereby improving the quality of generated data.

[0094] S4: Based on the cluster label prediction model and the probability model, the generative adversarial network framework is trained to obtain an interpolation model.

[0095] Wherein, step S4 further includes:

[0096] S41: Build the generator and discriminator to obtain the generative adversarial network framework.

[0097] Furthermore, the generator of the present invention adopts an autoencoder neural network architecture based on the self-attention mechanism, which compresses the input incomplete time series data into a low-dimensional potential representation through the encoder, and then reconstructs the potential representation into complete time series data through the decoder. The self-attention mechanism can capture the long-distance dependencies between different time steps in the time series, thereby improving the accuracy of interpolation; the discriminator adopts a multi-layer perceptron structure, with the input being the complete time series data and the corresponding mask vector, and the output being the predicted mask vector. Each element of the predicted mask vector indicates whether the data at the corresponding position is a true observation value or a generated interpolated value.

[0098] S42: Initializing the generative adversarial network framework based on the cluster label prediction model and the probability model.

[0099] Furthermore, in step S42, the present invention first uses the output result of the cluster label prediction model, i.e., the logistic regression classifier, as conditional information to input into the generator. The logistic regression classifier calculates the probability that each data point belongs to a specific cluster. These probability values ​​are input into the generator together with the original data as conditional vectors to guide the data generation process. At the same time, a Gaussian mixture model is used as a probability model to estimate the conditional distribution of features, and these parameters are estimated by the expectation maximization algorithm. In the expectation step, the posterior probability of each data point belonging to each Gaussian component is calculated. In the maximization step, the model parameters are updated based on the posterior probability. The probability distribution information is used to initialize the weight parameters of the generator, so that the generator can better learn the potential distribution characteristics of the data.

[0100] S43: Based on the interpolation loss, the initialized adversarial network framework is optimized through adversarial training to obtain the interpolation model.

[0101] Furthermore, in step S43, the process of optimizing the initialized adversarial network framework through adversarial training based on the interpolation loss to obtain the interpolation model involves the joint optimization of the generator loss and the discriminator loss. During the adversarial training process, the generator attempts to minimize the generator loss function so that the generated data can deceive the discriminator, while the discriminator attempts to minimize the discriminator loss function to improve its ability to distinguish between real data and generated data. The adversarial process is performed alternately using the gradient descent algorithm until a Nash equilibrium is reached.

[0102] The interpolation loss in step S43 is expressed as:

[0103] L=L D +L G ;

[0104]

[0105] Among them, L is the interpolation loss function, L G is the generator loss, L D is the discriminator loss, b i is the batch index variable, which indicates the i-th sample in the currently processed data batch, m i is the i-th mask vector, is the estimated data after interpolation corresponding to the mask vector, is the estimated data after interpolation corresponding to the i-th mask vector, α is the hyperparameter for balancing the generator loss and the discriminator loss weight, j is the dimension index of the data, d is the dimension of the data, L obs (x i ,x i ′) is the observation loss function, x i is the true observation value, that is, the actual measurement value or labeled value of the i-th position in the original data, x i ′ is the predicted output value of the generator, that is, the estimated value or predicted value generated by the generator network G at the i-th position.

[0106] Among them, L obs (x i ,x i The expression of ′) is:

[0107]

[0108] Furthermore, the loss of the generator consists of the basic generation loss and the mask reconstruction loss. The purpose is to ensure that the data distribution learned by the generator is as similar as possible to the true distribution. The basic generation loss is used to evaluate the authenticity of the generated data, thereby prompting the generator to generate data that the discriminator cannot distinguish. Therefore, the loss function of the generator is divided into two parts. The first part is the loss of the estimated value, and the second part is the loss of the observed value. Finally, these two parts are combined in the loss function L G As shown in the above expression.

[0109] While for the discriminator, there is only one output, which is either completely real or completely fake, in the interpolation setting, however, the output of the discriminator contains multiple components, some of which are real and some of which are fake. Therefore, the discriminator tries to distinguish between the observed real components and the missing fake components, which is achieved by predicting a mask vector m. This predicted mask can then be compared with the original mask M.

[0110] Specifically, the discriminator can be defined as D:X×Y→[0,1] d , where [0,1] drepresents the predicted mask vector m, X represents the feature space of the input data, which contains the complete time series data or multidimensional feature vectors, and Y represents the auxiliary information space. In the objective function optimization formula, the loss function of D can be expressed as the cross entropy equation, as shown above.

[0111] S5: Interpolate the data to be interpolated using the interpolation model to obtain an interpolation data matrix.

[0112] Wherein, step S5 further includes:

[0113] S51: Input the data to be interpolated.

[0114] Furthermore, in step S51, the original data set containing missing values ​​is first received, where the missing values ​​are represented by special marks in the data matrix, such as NaN or empty values, and a corresponding mask matrix is ​​generated to identify the data integrity status of each position. The mask matrix generated based on the original data set has the same dimensional structure as the original data matrix, where an element value of 1 indicates that there is real observation data at the corresponding position, and an element value of 0 indicates that there is missing data at the corresponding position that needs to be interpolated, providing clear missing pattern information for subsequent data processing.

[0115] S52: performing standardization processing on the data to be interpolated to obtain data to be predicted.

[0116] Furthermore, normalization ensures that features of different dimensions and numerical ranges are processed on the same scale, preventing large features from dominating the model training process. This accelerates the convergence of the neural network and improves the stability of numerical calculations. For data in missing locations, the normalization process maintains its missing state and only applies the normalization transformation to the locations marked as 1 in the mask matrix. The normalized data is then combined with the original mask matrix to form the dataset to be predicted.

[0117] S53: Inputting the data to be predicted into the interpolation model to obtain an interpolation data matrix.

[0118] Wherein, step S53 specifically includes:

[0119] The generator in the interpolation model generates missing values ​​for the data to be predicted to obtain interpolation data; the discriminator in the interpolation model screens the interpolation data for authenticity and obtains an interpolation data matrix.

[0120] Furthermore, in step S53, the data to be predicted is input into the interpolation model to obtain an interpolated data matrix, and the data interpolation task is completed through the collaborative work of the generator and the discriminator. The generator receives the standardized incomplete data, the mask vector, and the random noise vector as input, where the noise vector is sampled from the standard normal distribution N(0,1) and its dimension matches the input data. The addition of noise enhances the randomness and generalization ability of the generator. The self-attention mechanism within the generator captures the dependencies between different time steps in the time series data by calculating the attention weights between the query matrix, the key matrix, and the value matrix, and finally obtains the attention output. The discriminator's process of distinguishing the authenticity of interpolated data is implemented through a multi-layer perceptron network structure. The discriminator receives interpolated data and auxiliary information as input, where the auxiliary information contains the cluster label probability distribution obtained from the previous clustering and classification steps and the parameter information of the Gaussian mixture model. The forward propagation process of the discriminator transforms the input data nonlinearly through multiple fully connected layers and activation functions. The output layer of the discriminator generates a prediction mask vector, each element of which represents the probability that the discriminator predicts that the data at the i-th position is the true observation value. The discriminator evaluates the interpolation quality by comparing the difference between the predicted mask and the true mask. When the two are close, it means that the discriminator has correctly identified the authenticity of the data at that position.

[0121] The expression of the interpolation data is:

[0122]

[0123] in, is the interpolated data generated, x is the incomplete data, m is the mask vector, G is the generator, z is the noise added to the input data, ⊙ is the element-by-element multiplication, based on the above generator interpolated data expression, m⊙G[x,m,(1-m)⊙z] in the formula represents the generator output retaining the observed position, (1-m)⊙x represents the original data retaining the missing position, and the two parts are combined into complete interpolated data through element-by-element addition.

[0124] like Figure 2 As shown, the present invention also provides a missing data interpolation system based on a generative adversarial network, comprising:

[0125] Clustering module: used to cluster the missing data matrix to obtain cluster clusters containing cluster labels;

[0126] Prediction module: used to classify and predict the data feature vectors corresponding to the missing data matrix based on the cluster clusters through a logistic regression algorithm, and train and generate a cluster label prediction model;

[0127] Modeling module: used to perform probability distribution modeling on the data feature vector through Gaussian mixture model, and train and generate a probability model;

[0128] Training module: used to train the generative adversarial network framework based on the cluster label prediction model and the probability model to obtain an interpolation model;

[0129] The interpolation module is configured as the interpolation model, and is used to interpolate the data to be interpolated to obtain an interpolation data matrix.

[0130] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0131] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0132] The following describes a missing data interpolation method and system based on a generative adversarial network, and a corresponding generative adversarial interpolation model, CPC-GAIN, provided by the present invention, in conjunction with specific embodiments.

[0133] This example was evaluated on five real-world air quality datasets. The data primarily came from environmental monitoring systems, environmental protection departments, and air monitoring projects. The datasets cover different climate zones and have representative spatial and temporal distributions, making them suitable for evaluating the generalization and adaptability of multivariate time series prediction models.

[0134] Dataset A contains environmental data from multiple sites, sampled hourly. Key variables such as PM2.5, temperature, and humidity are recorded in the data. Its main features are dense sensor density and wide geographical distribution, making it suitable for evaluating model performance under spatial proximity. The data covers different seasons and climate conditions, helping the model learn about periodic fluctuations in air quality.

[0135] Dataset B covers 12 monitoring stations and contains hourly data for air pollutants such as PM2.5, PM10, and SO2, as well as meteorological variables such as temperature and humidity. The data is high-quality and subject to significant fluctuations, making it a popular benchmark for air quality forecasting tasks. It supports both single-site modeling and multi-site joint forecasting experimental setups.

[0136] The C dataset contains 21 variables, covering air pollution indicators and meteorological factors. The dataset has obvious seasonal characteristics and is suitable for evaluating the model's adaptability to cyclical and trend-type changes. The distribution of variables is balanced across sites, facilitating spatial generalization analysis.

[0137] The monitoring parameters of the D dataset include PM2.5, PM10, CO, NO2, SO2, and meteorological data. This area is the core of the urban area, where pollutants fluctuate frequently, making it a typical sample for short-term mutation detection and prediction tasks.

[0138] Dataset E has the same data structure as dataset D above, but is located in an industrial area and has significantly different trends in pollutant concentrations. The two datasets can be used together for cross-regional modeling and transfer learning experiments to verify the model's generalization ability across different regions.

[0139] The final Table 1 summarizes the basic information of each data set.

[0140] Table 1 Dataset description

[0141] Dataset name Data quantity Number of features A 1048576 5 B 504 6 C 218641 18 D 14682 15 E 15918 15

[0142] In this example, multiple baseline models are compared to evaluate the interpolation performance. The comparison models include Mean, MICE, SAITS, BRITS, and Transformer. To eliminate the influence of the training framework, all models are run under the same data preprocessing and evaluation process.

[0143] Specifically, the baseline model includes: Mean: the simplest static imputation method, directly filling in the missing values with the historical mean of each variable, as a reference baseline without model learning ability; MICE (Multiple Imputation by Chained Equations): a multiple chained regression imputation method, constructing a regression model for each variable to impute, suitable for handling missing data with strong correlation between multiple variables; BRITS (Bidirectional Recurrent Imputation for Time Series): an imputation model based on bidirectional recurrent neural network (Bi-RNN), using end-to-end training to simultaneously impute and predict, and introducing backpropagation error signals to improve temporal consistency and stability; SAITS (Self-Attention-based Imputation for Time Series): a self-attention time series imputation model based on the Transformer structure, capturing temporal dependencies through a double attention mechanism, suitable for multi-variable time series; GAIN: a generative adversarial network (GAN) model using a fully connected layer, used for data imputation.

[0144] By comparing the above baseline methods with different modeling assumptions and complexities, the imputation ability of the proposed model under different missing mechanisms and data structures can be comprehensively evaluated.

[0145] In order to comprehensively evaluate the performance of the model in the time series missing value imputation task, the present application selects two widely used indicators: root mean squared error (Root Mean Squared Error, RMSE) and mean absolute error (Mean Absolute Error, MAE), both of which are calculated on the time steps where the true value is observable, measuring the difference between the imputation results and the true data, RMSE is a common evaluation indicator for estimating the deviation between predicted values and original values.

[0146] The present application analyzes the imputation performance of CPC-GAIN and other baseline models on multiple air quality data sets under different missing rates. In order to ensure the fairness of the experiment, all models are run under the same data preprocessing process and evaluation framework, and the imputation effect is measured by root mean squared error (RMSE) and mean absolute error (MAE). The programming language used in the model experiment is Python3.6.13, and the deep learning framework is TensorFlow1.15.0. Tables 2-6 below are the experimental results of the six methods under different data sets and missing rates.

[0147] From the experimental results in the table, it can be seen that CPC-GAIN performs well under different missing rates in all datasets, especially showing obvious advantages over other baseline methods under high missing rates. The specific analysis is as follows:

[0148] On the A, B, C, D, and E region datasets, the MAE and RMSE of CPC-GAIN are always better than other methods, especially when the missing rate reaches more than 50%, the interpolation accuracy of CPC-GAIN is significantly better than traditional interpolation methods (such as Mean, MICE) and deep learning methods (such as BRITS, SAITS).

[0149] In the B dataset, when the missing rate is as high as 80%, the MAE of CPC-GAIN is only 0.141, which is much lower than other methods. For example, the MAE of the Mean method under all missing rates is close to 0.734, showing its limitations in complex time series data.

[0150] In the D dataset, CPC-GAIN performs well under all missing rates, especially under high missing rates, its interpolation effect still maintains low MAE and RMSE.

[0151] In the D dataset, when the missing rate is 80%, the MAE of CPC-GAIN drops to 0.095, which is significantly better than SAITS (0.578) and MICE (0.424), further verifying its reliability under extreme missing conditions.

[0152] Baseline method comparison analysis: The Mean method always performs poorly due to its simplicity, its MAE and RMSE hardly change with the missing rate, mainly serving as a benchmark method to verify the effectiveness of other methods.

[0153] The MICE method performs well when the missing rate is low (such as 10%-40%), but its performance drops significantly as the missing rate increases, especially when the missing rate is above 50%, the MAE and RMSE increase sharply.

[0154] BRITS and SAITS perform well on more complex time series datasets, but their performance is usually not as good as CPC-GAIN under high missing rates. Even in some scenarios with strong time series features, CPC-GAIN can still maintain good stability.

[0155] The GAIN model, as a variant of the generative adversarial network (GAN), also performs well, but its effect is slightly inferior to CPC-GAIN, especially on the Vietnam D and E region datasets, the interpolation error is higher than CPC-GAIN.

[0156] Table 2 Experimental results of different methods in the A dataset

[0157]

[0158] Table 3 Experimental results of different methods on dataset B

[0159]

[0160]

[0161] Table 4 Experimental results of various methods on the C dataset

[0162]

[0163] Table 5 Experimental results of different methods on the D dataset

[0164]

[0165] Table 6 Experimental results of different methods on dataset E

[0166]

[0167]

[0168] From the comprehensive experimental results, it can be seen that CPC-GAIN has excellent imputation performance in multiple air quality data sets, especially in the case of high missing rate (more than 50%). Compared with traditional imputation methods (such as Mean and MICE) and some advanced deep learning methods (such as BRITS, SAITS and GAIN), CPC-GAIN shows significant advantages in MAE and RMSE indicators. For example, in the B data set, when the missing rate is as high as 80%, the MAE of CPC-GAIN is significantly lower than that of other methods; in the D data set, its MAE is also significantly better than that of competitors. Even in the face of 80% missing rate on A and B data sets, the CPC-GAIN model can still maintain a low error, showing its excellent robustness and accuracy. Its excellent performance benefits from the combination of the generative adversarial architecture and the clustering and classification mechanism, which can effectively capture the temporal structure and data distribution, providing a precise and stable solution for data imputation under complex missing patterns. Specifically, on the B data set, compared with the traditional Mean method under 80% missing rate, the MAE of CPC-GAIN is reduced by about 82.4%, and the RMSE is reduced by about 79.5%; compared with the MICE method, the MAE is reduced by about 71.3%, and the RMSE is reduced by about 68.2%. In addition, on the D data set, when the missing rate reaches 80%, the MAE of CPC-GAIN is reduced by about 83.6% compared with SAITS, and the RMSE is reduced by about 81.4%; compared with MICE, the MAE is reduced by about 77.6%, and the RMSE is reduced by about 74.9%. This excellent performance reflects the strong generalization ability and robustness of CPC-GAIN, which can maintain a stable low error level under different data characteristics and missing patterns.

[0169] To further verify the effectiveness of the method proposed in the application, an ablation experiment was conducted on the B data set to evaluate the influence of different components on the performance of the model. The experiment includes four different configurations: a, only use the GAIN model for imputation, remove the clustering, classification and probability modeling part; b, remove clustering and classification, use the probability modeling module, and keep other components for imputation; c, use the GAIN model for imputation after clustering and classification; d, use the complete CPC-GAIN model, including GAIN, clustering and classification components, and the final results are shown in Figure 3 and Figure 4 .

[0170] As shown in Figure 3 , Figure 4 , under different missing rates, the four model configurations show significant differences in MAE and RAMSE indicators, indicating that each component has an important influence on the performance of the model.

[0171] As can be seen from the results of configuration a "GAIN (Vanilla)", the MAE and RAMSE values are generally higher when the missing values are only imputed by the original GAIN model, especially when the missing rate exceeds 50%, the error increases significantly, indicating that the GAIN structure alone is difficult to maintain the accuracy of imputation in the high missing rate scenario.

[0172] Comparing configuration b with configuration a, it can be found that the overall performance is improved after introducing the probability modeling module (even if the clustering and classification are removed), especially in the case of low to moderate missing rate (10%-50%), the MAE and RAMSE are significantly reduced, which shows that the probability modeling module has a positive effect on guiding the GAN to learn the data distribution. However, when the missing rate is high (such as 70%, 80%), the performance decreases significantly, which shows that the independent role of probability modeling still has certain limitations.

[0173] Configuration c, i.e. after introducing the clustering and classification module, the performance is more stable compared to configurations a and b, which shows that clustering can effectively distinguish the internal structure of the data, and the classifier enhances the modeling ability of the missing data pattern, especially in the case of high missing rate, it can still maintain a relatively low error, which proves that the combination has strong robustness.

[0174] The complete configuration d "CPC-GAIN" achieves the best results under all missing rates, both MAE and RAMSE values are significantly lower than the other three configurations, which shows the effectiveness of the synergistic effect of each component, especially in the range of 10%-50% missing rate, the MAE value is maintained at 0.037-0.07, and the RAMSE is also controlled at 0.074-0.127, which reflects the stability and high accuracy of the model in the wide missing scenario.

[0175] The ablation experiment shows that each sub-module (clustering, classifier, probability modeling) in CPC-GAIN contributes to improving the imputation effect, and the complete CPC-GAIN structure enhances the model's ability to capture the structure of missing data and the accuracy of imputation through multi-module collaboration.

[0176] In response to the widespread problem of missing data in multiple fields, this paper proposes a new generative adversarial interpolation model, CPC-GAIN, which integrates clustering, classification, and probabilistic modeling mechanisms. The model aims to improve the interpolation accuracy and stability of complex and nonlinear data. By designing a multi-module collaborative structure, the model fully utilizes the correlation between local structure and global dependency in data samples under the adversarial training framework of generators and discriminators, significantly reducing the error between the model's interpolated value and the true value. The clustering module effectively identifies the local similarity of data, the classification module enhances the supervisory role of category information, and the probabilistic modeling module optimizes the learning process of global information, so that the model still maintains high interpolation accuracy and stability under complex missing patterns.

[0177] Test results on multiple real-world datasets demonstrate that CPC-GAIN achieves significant performance improvements even with high missingness rates. In particular, when dealing with nonlinear and high-dimensional data, its interpolation errors (such as MAE and RMSE values) are significantly lower than those of traditional methods and other generative adversarial network (GAN) models. This demonstrates that CPC-GAIN has a clear advantage in capturing complex relationships between features and the spatial distribution of samples, enabling it to better fit the original data. Ablation experiments further validated the effectiveness of each submodule. The results show that removing any module leads to a significant decrease in interpolation performance, demonstrating the synergy and complementarity of clustering, classification, and probabilistic modeling mechanisms in CPC-GAIN.

[0178] In summary, the CPC-GAIN model demonstrates strong performance and adaptability in addressing missing data imputation, demonstrating its high practical application value. Future research will continue to explore its performance on larger and more diverse datasets, and further improve the model's computational efficiency and generalization capabilities.

[0179] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A missing data interpolation method based on generative adversarial networks, characterized in that: include: S1: Cluster the missing data matrix to obtain clusters containing cluster labels; S2: Based on the clusters, classify and predict the data feature vectors corresponding to the missing data matrix using a logistic regression algorithm to obtain a cluster label prediction model; S3: Performing probability distribution modeling on the data feature vector using a Gaussian mixture model to obtain a probability model; S4: Based on the cluster label prediction model and the probability model, a generative adversarial network framework is trained to obtain an interpolation model; S5: Interpolate the data to be interpolated using the interpolation model to obtain an interpolation data matrix.

2. A missing data interpolation method based on a generative adversarial network according to claim 1, characterized in that: Step S1 further comprises: S11: Input a missing data matrix, calculate the distances between multiple groups of time series data points in the missing data matrix, and obtain a distance matrix; S12: Iteratively optimizing the distance matrix based on K-means clustering to obtain multiple cluster centers; S13: performing cluster attribution determination on the data points in the missing data matrix according to the cluster centers to obtain a plurality of clusters, wherein each cluster includes a corresponding cluster label.

3. The missing data interpolation method based on generative adversarial network according to claim 1, characterized in that: Step S2 further comprises: S21: perform standardization preprocessing on the data eigenvector corresponding to the missing data matrix to obtain a standardized eigenvector; S22: performing probability mapping on the standardized feature vector through a logistic regression model to obtain a preliminary prediction model; S23: Optimizing the regression coefficients of the preliminary probability model through maximum likelihood estimation to obtain a cluster label prediction model.

4. The missing data interpolation method based on generative adversarial network according to claim 1, characterized in that: Step S3 further comprises: S31: Based on the probability density function, perform multi-Gaussian component modeling on the data distribution of the data feature vector to obtain preliminary model parameters; S32: Iteratively estimating the preliminary model parameters by an expectation maximization algorithm to obtain updated model parameters; S33: Based on the updated model parameters, a probability model for outputting characteristic conditional distribution parameters is established.

5. The missing data interpolation method based on generative adversarial network according to claim 1, characterized in that: Step S4 further comprises: S41: Construct the generator and discriminator to obtain the generative adversarial network framework; S42: Initializing the generative adversarial network framework based on the cluster label prediction model and the probability model; S43: Based on the interpolation loss, the initialized adversarial network framework is optimized through adversarial training to obtain the interpolation model.

6. A missing data interpolation method based on generative adversarial network according to claim 5, characterized in that: The expression of the interpolation loss in step S43 is: L=L D +L G ; Among them, L is the interpolation loss function, L G is the generator loss, L D is the discriminator loss, b i is the batch index variable, m i is the i-th mask vector, is the estimated data after interpolation corresponding to the mask vector, is the estimated data after interpolation corresponding to the i-th mask vector, α is the hyperparameter for balancing the generator loss and the discriminator loss weight, j is the dimension index of the data, d is the dimension of the data, L obs (x i ,x i ′ ) is the observation loss function, x i is the actual observation value, x i ′ is the predicted output value of the generator.

7. The missing data interpolation method based on generative adversarial network according to claim 5, characterized in that: Step S5 further comprises: S51: input data to be interpolated; S52: performing standardization processing on the data to be interpolated to obtain data to be predicted; S53: Inputting the data to be predicted into the interpolation model to obtain an interpolation data matrix.

8. The missing data interpolation method based on generative adversarial network according to claim 7, characterized in that: Step S53 specifically includes: Generate missing values ​​for the data to be predicted by the generator in the interpolation model to obtain interpolation data; The interpolation data is screened for authenticity through the discriminator in the interpolation model to obtain an interpolation data matrix.

9. The missing data interpolation method based on generative adversarial network according to claim 8, characterized in that: The expression of the interpolation data is: in, is the generated interpolation data, x is the incomplete data, m is the mask vector, G is the generator, z is the noise added to the input data, and ⊙ is the element-by-element multiplication.

10. A missing data interpolation system based on generative adversarial networks, characterized in that: include: Clustering module: used to cluster the missing data matrix to obtain cluster clusters containing cluster labels; Prediction module: used to classify and predict the data feature vectors corresponding to the missing data matrix based on the cluster clusters through a logistic regression algorithm, and train and generate a cluster label prediction model; Modeling module: used to perform probability distribution modeling on the data feature vector through Gaussian mixture model, and train and generate a probability model; Training module: used to train the generative adversarial network framework based on the cluster label prediction model and the probability model to obtain an interpolation model; The interpolation module is configured as the interpolation model, and is used to interpolate the data to be interpolated to obtain an interpolation data matrix.

Citation Information

Cited By

  • Ore grinding granularity prediction method combining missing value completion and multi-model collaboration

    CN121579936A

  • A grinding granularity prediction method combining missing value completion and multi-model cooperation

    CN121579936B