Generative adversarial network missing data filling method based on similarity matching and classification

Through the generation and adversarial network missing data filling method based on similar matching and classification, the problem of inaccurate filling of missing data in numerical simulation in the prior art is solved, and a higher filling accuracy and reliability of numerical simulation are achieved.

CN120045853APending Publication Date: 2025-05-27NANHUA UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510096501.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When the prior art uses data with missing values ​​in numerical simulation, the filling method is simple, which affects the reliability and accuracy of numerical simulation, and it is difficult to deal with the reasons for data loss.

Method used

The method of filling missing data of the generative adversarial network based on similar matching and classification is adopted. Through preliminary data processing, screening matching, importing the generative adversarial network model, filling result analysis and other steps, combining the generative adversarial network and classification model, the filling accuracy of discrete data is improved.

Benefits of technology

It significantly improves the accuracy of missing data filling, improves the reliability and accuracy of numerical simulation, especially when processing high missing rate data, which can significantly improve the quality of filling data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045853A_ABST
    Figure CN120045853A_ABST
Patent Text Reader

Abstract

The invention discloses a missing data filling method for a generative adversarial network based on similarity matching and classification. The method comprises the steps of data acquisition and preliminary analysis, data preliminary processing, data screening and matching, importing data into a generative adversarial network model, filling continuous data and discrete data in the data, and analyzing a filling result; if not, discrete data prediction and discrete data filling are conducted on the filled complete data through the classification model, and then filling result analysis is conducted again till the result meets the expectation. According to the method, through data testing under different missing rates, the stability and superiority of the method under various conditions are shown, the filling performance of continuous data is effectively improved, meanwhile, the filling precision of discrete data is remarkably improved in combination with a classification model, and the problem of data missing in the long-term real-time monitoring process can be effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of filling missing data, and particularly to a method for filling missing data based on similar data matching and classification using a generative adversarial network. Background Art

[0002] Before conducting numerical simulation analysis and verification, it is necessary to obtain long-term real-time monitoring data as input parameters. The reliability and integrity of the data directly determine the accuracy of the numerical simulation results. During the long-term real-time monitoring process, the obtained monitoring data will all contain more or less missing values. If directly using the data with missing values for numerical simulation, most numerical simulation software only simply fills the data, such as taking the average value, mode, or directly using the default value for filling. Such filling methods will affect the reliability and accuracy of the numerical simulation. Usually, the data obtained by researchers are all data that have been collected by other professionals, so it is difficult to know the specific reasons for the missing values in these data. Therefore, a method for filling missing data based on similar matching and classification using a generative adversarial network is needed to solve the above problems. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for filling missing data based on similar matching and classification using a generative adversarial network to solve the problems existing in the above-mentioned prior art.

[0004] To solve the above technical problems, the present invention adopts the following technical solution: A method for filling missing data based on similar matching and classification using a generative adversarial network defines the obtained data as follows:

[0005] Obtain a set of data X with a size of n×m (n, m>0), where n represents that the data has n rows and m represents that the data has m columns, denoted as X nm . Use x ij to represent a value at any position in the data, where i∈(0, n], j∈(0, m]. Denote any column of the data as a feature vector, represented by F j , then X nm =(F 1 ,..., F m ). Based on the above data definition, its main filling method includes the following steps:

[0006] (1) Data acquisition and preliminary analysis;

[0007] (2) Preliminary data processing;

[0008] (3) Data screening and matching;

[0009] (4) Import the data into the generative adversarial network model;

[0010] (5) Fill in the continuous data and discrete data in the data;

[0011] (6) Analysis of the filling results;

[0012] (7) If the filling result of the discrete data does not meet the expectation (the filling accuracy rate is low), then use the classification model to predict and fill in the discrete data of the complete data after filling, and then return to step (6) for analysis of the filling results until the results meet the expectation.

[0013] Preferably, the data acquisition and preliminary analysis include obtaining the data X containing missing values nm , and conducting preliminary analysis on it, including checking whether the data are all numerical data. If there are text data in the data, it is necessary to consider numericalizing the text data, recording the discrete data and continuous data in the data, and summarizing the rules existing in the data.

[0014] Preferably, the preliminary data processing includes obtaining the data containing missing values, and conducting preliminary processing on the data, including but not limited to data normalization processing, deleting some useless or redundant data, adjusting the data format, dividing the data seasonally, etc., and recording any part of the data that needs to fill in missing values after processing as Y ab , where 0 < a ≤ n, 0 < b ≤ m.

[0015] More preferably, the data screening and matching include, based on the data Y that needs to fill in missing values after the regular division ab , dividing the data into Y Miss and Y Complete in two parts, where each row in Y Miss contains at least one missing value, and there is no missing value in each row of Y Complete . At the same time, create an empty data table Y Similar to store the data after screening and matching. By conducting sensitivity analysis on the obtained original data X nm , according to the results, rank the importance of any column in the data, that is, each feature vector, so as to obtain the weight of the feature vector. Here, the situation where the entire row of data is missing is not considered, that is, there is at least one item in the data that is not missing. According to the weight of the feature vector, approximate data without missing values are matched in Y Complete . The matching calculation method is as shown in formula (1):

[0016]

[0017] where A represents in Y MissThe non-missing data value of a certain eigenvector in, B represents in Y Complete the data value of the same eigenvector as A in the absolute value representing the difference magnitude between two numerical values.

[0018] The specific data matching and screening method is as follows:

[0019] (3.1) Find the eigenvector value in the first row of Y according to the weight of the eigenvector. If this value is missing, continue to find the next important eigenvector value, and so on until the first non-missing eigenvector value is found. Miss

[0020] (3.2) Calculate the similarity of the found eigenvector value and the corresponding eigenvector values in each row of Y according to formula (1), and screen out Complete the data row with the smallest value.

[0021] (3.3) If this eigenvector value is not the last one in the importance ranking, go back to (3.1) to continue finding the eigenvector value. After finding the eigenvector value, continue to calculate the similarity with the data row screened out in (3.2), and further screen the data. Repeat this process until the matching of the last eigenvector is completed.

[0022] (3.4) Take the first row of the data row screened out after the matching of the last eigenvector as the finally matched data and put it into Y Similar , and delete the first row of data in Y Miss .

[0023] (3.5) Repeat steps (3.1), (3.2), (3.3), (3.4) until the data in Y Miss is empty.

[0024] Through this method, a complete data row without missing values can be matched for each row in the original Y Miss that contains missing values.

[0025] More preferably, the importing the data into the generative adversarial network model and data filling includes importing the obtained Y above Similar into the generative adversarial network model. After iterative training of the model, input random noise with the same dimension size as Y Similar to obtain the final fake data denoted as Y Fake . Finally, use Y Fake as the filling value for the missing data in Y ab , and obtain the completely filled data denoted as Y Imputed .

[0026] ​​More preferably, the filling result analysis includes using methods such as model effect evaluation indicators and drawing statistical charts to analyze and evaluate the filling results. The specific analysis content includes whether the filled values are close to the values of the real data, and in which aspects the main errors exist, etc. Through specific experiments, it is found that the generative adversarial network model cannot generate discrete data close to the original data well, mainly because it is unable to eliminate the influence of the numerical magnitude in the discrete data representation.

[0027] More preferably, the specific methods for discrete data prediction and discrete data filling of the filled complete data through a classification model include: obtaining the filled complete data, importing the data into the classification model, discrete data prediction, discrete data filling, and filling result analysis. It mainly utilizes that the generative adversarial network model can effectively learn other data except discrete data, that is, continuous data, to obtain fake data close to the real data, and then imports the fake data into the classification model to classify and predict discrete data, so as to obtain better results.

[0028] Among them, the obtaining the filled complete data and importing the data into the classification model include obtaining the finally generated filled complete data Y Imputed , and performing further processing to obtain a training data set suitable for importing into the classification model, a test data set for verification, and a prediction data set for prediction.

[0029] Preferably, the further processing includes recording the row numbers containing missing values in Y ab , and dividing Y Imputed into two parts according to the row numbers. One part is the data with the rows containing missing values removed, denoted as and the other part is the data composed of all rows containing missing values, denoted as

[0030] More preferably, the training data set suitable for importing into the classification model and the test data set for verification are obtained by performing a correlation analysis on the discrete feature vectors to be filled and the remaining feature vectors in , screening out the feature vectors with strong correlation with the feature vectors to be filled, and then dividing into a training set and a test set according to a ratio of 7:3.

[0031] More preferably, the prediction data set for prediction is the above data.

[0032] More preferably, the discrete data filling includes predicting the feature vectors to be filled in by taking the time with the best classification effect on the test set after parameter tuning, and finally The unfilled data in the feature vector to be filled is replaced with the predicted result to obtain the final filled data.

[0033] More preferably, the filling result analysis includes analyzing and evaluating the filling result by methods such as model effect evaluation indicators and statistical charts, and analyzing its filling effect and main error reasons, etc.

[0034] Compared with the prior art, the present invention predicts and fills discrete data through a classification model, so that the accuracy of the discrete data filling result can be significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is a schematic flowchart of the method in the embodiment;

[0036] Figure 2 is a scatter plot of the filled wind speed and time at a missing rate of 10% in the embodiment;

[0037] Figure 3 is a scatter plot of the filled rainfall and time at a missing rate of 10% in the embodiment;

[0038] Figure 4 is a scatter plot of the filled stability and time at a missing rate of 10% in the embodiment;

[0039] Figure 5 is a scatter plot of the unfilled wind direction and time without predicting and filling discrete data at a missing rate of 10% in the embodiment;

[0040] Figure 6 is a scatter plot of the filled wind direction and time after predicting discrete data at a missing rate of 10% in the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0041] For the convenience of understanding by those skilled in the art, the present invention will be further described below in conjunction with the embodiments and the drawings. The content mentioned in the embodiments is not a limitation to the present invention. The present invention takes meteorological data (long-term real-time monitoring data) as an example. The meteorological data parameters that can be processed in the present invention include but are not limited to the meteorological parameters mentioned in the following embodiments.

[0042] As Figure 1 shown, the method for filling missing data based on a generative adversarial network of similarity matching and classification includes the following steps:

[0043] Step (1), data acquisition and preliminary analysis. In this embodiment, in order to better verify the final filling effect, the data will be artificially set to be randomly missing subsequently. The original data does not contain missing values, with a total of 8760 rows of data. Its feature vectors include days, hours, wind direction, wind speed, stability, and rainfall. Among them, days, hours, wind direction, and stability are discrete data, and wind speed and rainfall are continuous data. After analysis, it can be seen that the data within the same season has certain similarities.

[0044] Step (2), preliminary data processing. In this embodiment, the above-obtained original data is processed. After analysis, it can be seen that the meteorological data has obvious seasonal patterns, and through experiments, it can be obtained that the results obtained by dividing the original data into a seasonal division model are better than those without division. Therefore, in this embodiment, the data is divided into four seasons: spring, summer, autumn, and winter according to the regularity. This example takes the spring data as an example for illustration, and the processing methods for other seasons are similar and will not be elaborated too much. In this example, by observing the data, it is found that the two feature vectors of days and hours are not likely to generate missing values, so they are not subjected to random missing processing, and the remaining data is subjected to random missing processing. After processing, there are 2184 rows of data, and a column of indexes is added to this data on the basis of the original data to facilitate subsequent data processing.

[0045] Step (3), data screening and matching. Through the sensitivity analysis of the data, according to the results, an importance ranking is given to each feature vector, the data rows containing missing values are found, and approximate data without missing values are matched in the original data according to the weights of the feature vectors. In this embodiment, the weights of the feature vectors from large to small are wind speed, wind direction, stability, and rainfall. If a certain row of data has a missing value and the missing value is wind speed, then the data is matched according to the weights of the feature vectors. Because the wind speed feature vector is missing, the wind direction, the feature vector of the next-level weight, is used to match the rows in the original data without missing values. There may be multiple rows matching this result. At this time, the data is continued to be matched according to the weights of the feature vectors until all the data is matched. If there are still multiple rows after the matching is completed, the first row is taken as the matching result. However, the situation where the entire row of data is missing is not considered here, that is, at least one item of the data is not missing. Through this method, a complete row of data without missing values can be matched for each row in the original data table containing missing values, and the data dimension size is the same as that of the data containing missing values in the original data table.

[0046] Step (4), import the data into the generative adversarial network model. In this embodiment, the matched complete data is imported into the generative adversarial network model. This generative adversarial network model is composed of a generator and a discriminator. The internal network both uses transposed convolutional networks. The fully connected layer uses the ReLu activation function, and the output layer uses the Sigmoid activation function. The batch training size is 50, and the iterative training is 1500 times.

[0047] Step (5), after training is completed, input random noise with the same dimension size as the training data to obtain the final fake data, and finally use the fake data as the filling value for the missing original data.

[0048] Step (6), filling result analysis. In this embodiment, the complete filled data is obtained through model training. In order to test the filling effect of the model, methods including but not limited to using model effect evaluation indicators and drawing corresponding statistical charts are used.

[0049] In this embodiment, the method is respectively tested under the conditions of 5%, 10%, and 15% data missing rates. The root mean square error (RMSE) is used to measure the similarity between the continuous data of wind speed, stability, and rainfall and the original data. Although the stability feature vector is discrete data, there is a certain magnitude relationship for stability, so it is treated as continuous data here.

[0050] When the data missing rate is 5%, the RMSE of spring data is between 0.48 and 0.59, the RMSE of summer data is between 0.37 and 0.46, the RMSE of autumn data is between 0.58 and 0.75, and the RMSE of winter data is between 0.55 and 0.58.

[0051] When the data missing rate is 10%, the RMSE of spring data is between 0.63 and 0.75, the RMSE of summer data is between 0.43 and 0.52, the RMSE of autumn data is between 0.82 and 0.89, and the RMSE of winter data is between 0.76 and 0.78.

[0052] When the data missing rate is 15%, the RMSE of spring data is about 0.63, the RMSE of summer data is between 0.65 and 0.70, the RMSE of autumn data is between 0.98 and 1.10, and the RMSE of winter data is between 0.95 and 0.96.

[0053] This example gives a scatter plot comparison between the filled data of each feature vector of spring data and the real data under a 10% missing rate, as Figures 2 to 5 shown. It can be known through analysis that the filling method can effectively learn the distribution law of the feature vectors of wind speed, stability, and rainfall in the original data, so as to generate fake data close to the original data.

[0054] Step (7), the classification model predicts discrete data. To solve the problems found in step (5) above, this embodiment proposes a method for filling missing data in a generative adversarial network based on similarity matching and classification. In this embodiment, the data filled in step (5) above is used, the row numbers containing missing values in the original data are recorded, and the data is divided into two parts according to these row numbers. One is the complete original data excluding the rows containing missing values, and the other is the missing original data composed of all rows containing missing values. The feature vector of the complete original data includes days, time, wind direction, wind speed, stability, and rainfall. The feature vector of the missing original data includes days, time, wind speed, stability, and rainfall. The correlation analysis is performed between the wind direction feature vector in the complete original data and other feature vectors, and the feature vectors with strong correlation with this feature vector are selected. The feature vectors selected here are days, time, wind speed, stability, and rainfall. The wind direction in the complete original data is used as the prediction variable for modeling and prediction. In this embodiment, the random forest is selected as the classification model, and the ratio of dividing the training set and the test set is 7:3. After parameter tuning, the prediction is made with the best classification effect of the test set. The prediction data set is the missing original data filled with false data, that is, the missing values included in the feature vectors of wind speed, stability, and rainfall of this data are filled with false data completely. Then, the prediction data set is imported into the model to obtain the predicted wind direction data. Finally, the wind direction data that is not missing in the missing original data is replaced with the predicted wind direction data to obtain the final data.

[0055] Return to step (6), filling result analysis. In this embodiment, the data filled completely is obtained through model training. To test the filling effect of the model, methods including but not limited to using model effect evaluation indicators and drawing corresponding statistical charts are used.

[0056] In this embodiment, the method is respectively tested under the conditions of 5%, 10%, and 15% data missing rates, and the correct rate is used to measure the similarity degree between the discrete data of wind direction and the original data.

[0057] When the data missing rate is 5%, the prediction correct rate of spring data is not less than 0.49, the prediction correct rate of summer data is not less than 0.95, the prediction correct rate of autumn data is not less than 0.75, and the prediction correct rate of winter data is not less than 0.53.

[0058] When the data missing rate is 10%, the prediction correct rate of spring data is not less than 0.48, the prediction correct rate of summer data is not less than 0.93, the prediction correct rate of autumn data is not less than 0.74, and the prediction correct rate of winter data is not less than 0.38.

[0059] When the data missing rate is 15%, the prediction accuracy of spring data is not less than 0.47, the prediction accuracy of summer data is not less than 0.91, the prediction accuracy of autumn data is not less than 0.73, and the prediction accuracy of winter data is not less than 0.37.

[0060] By comparing Figure 5 and Figure 6 the wind direction data in, it can be found that the result of filling discrete data has been greatly improved. Through data tests at different missing rates, the present invention demonstrates its stability and effectiveness in various situations. Especially when dealing with high missing rate data, it can significantly improve the quality of filled data, proving the effectiveness and superiority of this method in filling missing data in numerical simulations.

[0061] In order to enable those of ordinary skill in the art to more conveniently understand the improvements of the present invention over the prior art, some of the drawings and descriptions of the present invention have been simplified, and the above embodiments are the preferred implementation schemes of the present invention. In addition, the present invention can also be implemented in other ways. Any obvious replacement without departing from the concept of the technical solution is within the protection scope of the present invention.

Claims

1. A method for filling missing data in a generative adversarial network based on similarity matching and classification, characterized in that: The following steps are involved: (1) Data acquisition and preliminary analysis; (2) Preliminary data processing; (3) Data screening and matching; (4) Importing data into the generative adversarial network model; (5) Fill in the continuous and discrete data in the data; (6) Filling result analysis; (7) If the discrete data filling result does not meet expectations, the classification model is used to perform discrete data prediction and discrete data filling on the complete data after filling, and then return to step (6) to analyze the filling result until the result meets expectations.

2. The method for filling missing data in a generative adversarial network based on similarity matching and classification according to claim 1, characterized in that: The method of step (7) includes obtaining the complete data after filling, importing the data into the classification model, discrete data prediction, and discrete data filling. The obtaining the complete data after filling and importing the data into the classification model include obtaining the Y generated in step (5). Imputed , and further processing is performed to obtain a training data set suitable for importing into the classification model, a test data set for verification, and a prediction data set for prediction, wherein the further processing includes converting the complete data Y generated in step (5) Imputed , record Y ab contains the row number of the missing value, and Y is converted according to the row number Imputed Divide into two parts, one is the data without the rows containing missing values, recorded as The other set of data consisting of all rows containing missing values ​​is recorded as Will The discrete feature vector to be filled in is analyzed for correlation with the remaining feature vectors, and the feature vector with strong correlation with the feature vector to be filled is selected, and then The feature vector to be filled in is used as the predictor variable for classification model modeling and prediction. The ratio of 7:3 is used to divide the training set and the test set. as a prediction dataset.

3. The method for filling missing data in a generative adversarial network based on similarity matching and classification according to claim 1, characterized in that: The data screening and matching in step (3) includes the data Y that needs to be filled with missing values ​​after regularity division ab Is there any missing value in Y Miss and Y Complete Two parts, one of which is in Y Miss Each row in contains at least one missing value. Complete There are no missing values ​​in each row in the table, and an empty data table Y is created. Similar Used to store the filtered and matched data, by obtaining the original data X nm Perform sensitivity analysis on the data and rank the importance of any column in the data, i.e., each eigenvector, according to the result, so as to obtain the weight of the eigenvector. In this case, the case where the entire row of data is missing is not considered, i.e., at least one item in the data is not missing. According to the weight of the eigenvector, the Y Complete The matching method is shown in formula (1): Among them, A represents the Miss The non-missing data value of a feature vector in Y Complete The data value of the same eigenvector as A, Indicates the absolute value of the difference between two values; The data screening and matching method is as follows: (3.1) Find the first row Y according to the weight of the feature vector Miss If the eigenvector value in is missing, continue to look for the next important eigenvector value, and so on until the first eigenvector value that is not missing is found; (3.2) Compare the found eigenvector value with Y Complete The corresponding eigenvector values ​​of each row in are similarly calculated according to formula (1) to filter out The data row with the smallest value; (3.3) If the eigenvector value is not the last one in the importance sorting, go back to (3.1) to continue looking for the eigenvector value. After finding the eigenvector value, continue to calculate the similarity with the data rows filtered out in (3.2), and further filter the data. Repeat this process until the last eigenvector is matched. (3.4) Put the first row of the data rows filtered out after the last feature vector is matched into Y as the final matched data Similar , and Y Miss Delete the first row of data in; (3.5) Repeat steps (3.1), (3.2), (3.3), (3.4) until Y Miss The data is empty; This method can be used to Miss Each row containing missing values ​​is matched to a complete row without missing values.

4. The method for filling missing data in a generative adversarial network based on similarity matching and classification according to claim 1 or 2, characterized in that: Step (4) is to convert the Y obtained in step (3) into Similar Import the generative adversarial network model and input the Y Similar Random noise of the same dimension size is used to obtain the final false data, which is recorded as Y Fake , and Y Fake As step (5) Y ab The missing data is filled in, and the completed data is recorded as Y Imputed .

5. The method for filling missing data in a generative adversarial network based on similarity matching and classification according to claim 1, characterized in that: In step (6), the filling result analysis includes analyzing and evaluating the filling results using model effect evaluation indicators and statistical chart methods, analyzing the filling effect and main error causes.

6. The method for filling missing data in a generative adversarial network based on similarity matching and classification according to claim 1, characterized in that: In step (2), the preliminary data processing includes normalizing the data, deleting some useless or redundant data, adjusting the data format, dividing the data into regularities, and recording any part of the data that needs to be filled with missing values ​​after processing as Y ab , where 0 <a≤n,0<b≤m。 7. The method for filling missing data in a generative adversarial network based on similarity matching and classification according to claim 1, characterized in that: In step (1), obtain the data X containing missing values nm , and perform a preliminary analysis on it. The preliminary analysis includes checking whether the data are all numerical data. If there is text data in the data, it is necessary to consider digitizing the text data, record the discrete data and continuous data in the data, and summarize the rules in the data.

8. The method for filling missing data in a generative adversarial network based on similarity matching and classification according to claim 2, characterized in that: The discrete data filling includes taking the best classification effect of the test set after parameter tuning. To predict the feature vector to be filled in, The final data can be obtained by replacing the non-missing data in the feature vector to be filled with the predicted results.

9. The method for filling missing data in a generative adversarial network based on similarity matching and classification according to claim 1, characterized in that: The acquired data is defined as follows: Get a set of data X of size n×m (n,m>0), where n means the data has n rows and m means the data has m columns, denoted as X nm , use x ij Represents a value at any position in the data, where i∈(0,n], j∈(0,m], and any column of the data is recorded as a feature vector, and F j Indicates that, then X nm =(F1,...,F m ).

10. The method for filling missing data in a generative adversarial network based on similarity matching and classification according to claim 1, characterized in that: All data processing processes are based on the definition of data in claim 9.

Citation Information

Cited By

  • A method for reconstructing strongly physically constrained time-series data based on a gated multi-scale conditional diffusion model

    CN122734275A