Data preprocessing method and system suitable for big data analysis

Through the dynamic correlation supplement method, combining the data correlation and time series autocorrelation, the missing part of the data is accurately filled, which solves the problem that the characteristic relationship cannot be reflected when filling data is missing in the prior art, and improves the accuracy of supplementary values ​​and the accuracy of subsequent analysis results.

CN120234545AActive Publication Date: 2025-07-01CHENGDU BIG DATA GRP CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510694210.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-07-01
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

When existing data preprocessing methods fill in missing parts of the data, they cannot effectively reflect the relationship between various characteristics of the data, resulting in poor accuracy of filling data, which may introduce deviations, affecting the accuracy of subsequent analysis results.

Method used

The dynamic correlation supplement method is used to calculate the correlation coefficient between other eigenvalues ​​of the data and the missing eigenvalues, determine the correlation weights, and obtain the fitted linear relationship based on historical data, calculate the relevant supplement value, and then supplement the missing eigenvalues. At the same time, consider the autocorrelation part of the time series to further improve the accuracy of the supplementary value.

Benefits of technology

Through the dynamic correlation supplement method, the missing parts of the data can be filled more accurately, the accuracy of supplementary values ​​can be improved, the overall trend and characteristic relationship of the data can be ensured, the deviation can be reduced, and the accuracy of subsequent analysis results can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234545A_ABST
    Figure CN120234545A_ABST
Patent Text Reader

Abstract

The invention discloses a data preprocessing method and system suitable for big data analysis, and relates to the technical field of data preprocessing. Comprising the steps of collecting original data through big data; a dynamic correlation supplementing method is used for calculating a supplementing value, missing feature values of the data are supplemented, the supplementing of the missing feature values better conforms to the overall trend of the data, and the accuracy of the supplementing value is improved; obtaining reliability constants of different sources through a reliability constant calculation method, and extracting data of different sources according to different proportions; and merging, unifying and normalizing the data. According to the method, the supplementary value is calculated through the dynamic correlation supplementary method, the missing feature value of the data is supplemented, according to the correlation between the other feature value of the data and the missing feature value, it is determined that the correlation weight and correlation of the other feature value to the missing feature value are dynamically changed, and supplementary to the missing feature value better conforms to the overall trend of the data; and supplementary value accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data preprocessing, and specifically provides a data preprocessing method and system applicable to big data analysis. Background Art

[0002] Data preprocessing is carried out to improve data quality, enhance model performance, reduce algorithm computational complexity and resource consumption, making the data more suitable for subsequent analysis and modeling tasks to obtain more accurate, reliable and efficient results. By data cleaning, noise, incorrect data and missing values are removed to improve data accuracy and integrity. Standardization, normalization and other operations are used to eliminate data inconsistency to ensure data consistency. Through feature selection and extraction, irrelevant features are removed and key features are extracted to reduce the data dimension. Data normalization and other operations are performed to make the data more regular, thereby improving data quality. Ultimately, it aims to improve model accuracy, accelerate model convergence, enhance model generalization ability, reduce training time and cost, laying a solid foundation for the successful implementation of tasks such as data analysis and machine learning.

[0003] When performing data preprocessing, it is necessary to delete duplicate parts of the data and supplement missing parts. For example, the method for filling missing values of station meteorological data based on the svd algorithm with the patent publication number CN106295175A has the following implementation steps: (1) receiving data; (2) preprocessing the original data; (3) selecting a training set and a test set; (4) training parameters; (5) filling missing values; (6) outputting the filled meteorological station data. The present invention separately extracts each meteorological attribute, forms a data file and performs svd algorithm training respectively to obtain the filled station meteorological data, improving the robustness and the accuracy of missing value filling.

[0004] The role of supplementing data is to fill the missing parts in the data, making the data more complete, enhancing the usability of the data, enabling data analysis and model training to utilize more comprehensive information, thereby improving the accuracy of analysis results and the performance and generalization ability of the model. When filling the original data with the above method, different feature data are predicted and filled separately, which cannot reflect the relationship between various features of the data, and the accuracy of the filled data is poor. Commonly used filling methods also include filling with mean, median, etc., which may introduce biases because this method does not consider the true distribution of the data and the relationship between features, and may cause the data to lose its original features and patterns, resulting in inaccurate subsequent analysis results. Summary of the Invention

[0005] The purpose of the present invention is to provide a data preprocessing method and system applicable to big data analysis to solve the problems raised in the above background art.

[0006] To achieve the above object, the present invention provides the following technical solutions: A data preprocessing method applicable to big data analysis, the method comprising: Collecting raw data through big data and recording the data sources; After removing duplicates from the raw data, using a dynamic correlation supplementation method to calculate supplementation values and supplement missing feature values of the data; Obtaining reliability constants from different sources through a reliability constant calculation method, and performing different proportions of extraction on data from different sources according to the data reliability constants from different sources; Merging and uniformly processing data from different sources and in different formats, and performing operations such as format conversion and encoding unification on the data to make the data consistent in structure and semantics; Performing normalization processing on the integrated data to make the data have the same scale and distribution.

[0007] Preferably, the dynamic correlation supplementation method comprises: S1: Calculating the correlation coefficient of each feature value with respect to the missing feature value according to the remaining feature values of the data, according to the formula: is the number of samples in the obtained data, where is the correlation coefficient between the j-th feature and the k-th feature, is the feature value of the j-th feature of the i-th sample, is the mean of the j-th feature among n samples, is the feature value of the k-th feature of the i-th sample, is the mean of the k-th feature among n samples; S2: Determining the correlation weight between other features and the missing feature according to the correlation coefficient, according to the formula: where represents the correlation weight of the k-th feature for filling the missing value of the j-th feature, is the number of features of each sample; S3: Obtaining the fitted linear relationship according to historical data and calculating the relevant supplementation value, according to the formula: where is the dependent feature value, is the intercept term, represents the feature value of the k-th feature, is the regression coefficient of feature k, is the error term, is the relevant supplementary value of the j-th feature in sample i, is the feature value of the k-th feature of the i-th sample; S4: Use to supplement the j-th feature in sample i.

[0008] Preferably, the dynamic correlation supplementation method further includes, for the autocorrelation part of the time series: G1: Select several consecutive complete samples, calculate the autocorrelation coefficient, according to the formula: where is the autocorrelation coefficient at lag periods, is the number of consecutive data periods selected, is the average value of the numerical values for consecutive periods, is the data value at the t-th period in the selected data periods, is the data value at the t - δ-th period in the selected data periods; G2: Calculate the autocorrelation weight according to the autocorrelation coefficient: where is the autocorrelation weight at lag periods, is the autocorrelation coefficient at lag c periods, is the set maximum lag period number; G3: Calculate the final supplementary value according to the autocorrelation weight and the relevant supplementary value: where is the autocorrelation supplementary value of the j-th feature in sample i, is the final supplementary value of the j-th feature in sample i, is the th feature value of the j-th feature of the sample, is the self-set weight, is the relevant supplementary value of the j-th feature in sample i; G4: Use to supplement the j-th feature in sample i.

[0009] Preferably, the method for obtaining the fitted linear relationship according to historical data includes: L1: Collect historical data, and set the missing feature value as the dependent feature according to the historical data; L2: Set as the intercept term, represents the feature value of the k-th feature, is the regression coefficient for feature k, is the error term, and the formula ; L3: Estimate the regression coefficients by the least squares method to determine , and values, thereby determining the fitted linear relationship.

[0010] Preferably, the calculation method of the reliability constant includes: Obtain the total number of samples, the number of features of each sample, and the number of missing features in each sample, and exclude samples with a missing ratio exceeding 0.5 from a number of samples, and then calculate according to the formula: where is the reliability constant of the data source, is the data defect constant, reflecting the degree of data missing, is the initial credibility of the data source, with a value between 0 and 1, set by the technical personnel, and its numerical value represents the degree of recognition of the authenticity of the data source, is the number of missing features in sample i, is the total number of samples in the obtained data, is the number of features of each sample; When the reliability constant of the data source ≤0.3 or ≥0.7, update the value of the initial credibility of the data source to by the credibility update method to improve the accuracy of the next calculation.

[0011] Preferably, the calculation method of the reliability constant includes: Obtain the total number of samples, the number of features of each sample, and the number of missing features in each sample, and calculate according to the formula: where is the reliability constant of the data source, is the data defect constant, reflecting the degree of data missing, is the credibility of the data source, with a value between 0 and 1, and its numerical value represents the degree of recognition of the authenticity of the data source, is the number of samples lacking feature j, is the total number of samples in the obtained data, The number of features for each sample; When the reliability constant of the data source ≤ 0.3 or ≥ 0.7, the credibility of the data source is updated through the credibility update method to improve the accuracy of the next calculation.

[0012] Preferably, the credibility update method includes: When 0.2 < ≤ 0.3, ; when 0.1 < ≤ 0.2, ; when ≤ 0.1, , and prompt the technician to re - determine the data source; When 0.7 ≤ ≤ 0.8, ; when 0.8 < ≤ 0.9, ; when > 0.9, ; where is the credibility of the updated data source, is the credibility of the data source before update, and where is the initial credibility, set by the technician, is the update speed constant, set by the technician, and the larger the value, the greater the update amplitude.

[0013] Compared with the prior art, the beneficial effects of the present invention are: Calculate the supplementary value through the dynamic correlation supplement method to supplement the missing feature values of the data. Determine the correlation weight of other feature values on the missing feature value according to the correlation between other feature values and the missing feature value of the data. The correlation changes dynamically according to the acquired data, and the supplement of the missing feature value is more in line with the overall trend of the data, improving the accuracy of the supplementary value.

[0014] At the same time, according to the self - change trend of the data in different time series, determine the autocorrelation coefficient of the missing feature, calculate the autocorrelation weight of the feature values of the previous periods of the missing feature on the missing feature value, fully consider the change trend of the missing feature, and determine the final supplementary value by comprehensively considering the influence of other feature values on the missing feature, further improving the accuracy of the supplementary value.

[0015] In addition, according to different data sources, the reliability constant of each data source can be calculated respectively to determine the credibility of the data source. Using the reliability constant of the data source as the screening ratio, the data can be screened to reduce the sample size and improve the calculation speed. Moreover, for the credibility of the data source, it can be iterated through the credibility update method to continuously optimize the recognition degree of the data source during long-term use, further improving the accuracy of the screened samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a schematic flowchart of the data preprocessing method of the present invention; Figure 2 It is a schematic flowchart of the dynamic correlation supplement method of the present invention; Figure 3 It is a schematic flowchart of the reliability constant calculation method (sub-sample calculation) of the present invention; Figure 4 It is a schematic flowchart of the reliability constant calculation method (sub-feature calculation) of the present invention; Figure 5 It is a schematic flowchart of the credibility update method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0018] In this application, for the convenience of understanding, the method steps used do not necessarily need to be executed in the order of the steps in this embodiment during actual operation. In other embodiments, these steps can be carried out synchronously or in a changed order.

[0019] Embodiment 1

[0020] The role of supplementary data is to fill in the missing parts in the data, make the data more complete, improve the usability of the data, enable data analysis and model training to utilize more comprehensive information, thereby improving the accuracy of the analysis results and the performance and generalization ability of the model. When filling in the data, when the feature to be filled has a strong correlation with the other features, it is necessary to calculate the missing feature based on the other features to improve the accuracy of the supplement.

[0021] As Figure 1 shown, the present invention provides a technical solution: a data preprocessing method applicable to big data analysis, including: Collecting raw data through big data and recording the data source; It should be noted that the original data can be collected through various service platforms, social platforms, network platforms, and third-party providers. The data sources need to be manually confirmed by technical personnel to have a certain degree of credibility.

[0022] After deduplicating the original data, the dynamic correlation supplementation method is used to calculate the supplementary value to supplement the missing feature values of the data. Merge and uniformly process the data from different sources and different formats, and perform operations such as format conversion and encoding unification on the data to make the data consistent in structure and semantics. Normalize the integrated data to make the data have the same scale and distribution.

[0023] It should be noted that the technical means such as merging and unifying the format of the data are existing technologies. For example, in Python, the strptime() and strftime() functions of the datetime module can be used to convert date and time data into a unified format; string processing functions in programming languages, such as the upper() or lower() methods in Python, can be used to convert all text into uppercase or lowercase forms. This will not be elaborated here. The normalization processing in the normalization of the integrated data is also an existing technology, and other processing methods such as standardization processing can be performed according to actual needs, and the selection in this embodiment is not used as a limitation.

[0024] Moreover, when extracting different proportions of data from different sources, the reliability constant of the data source can be used as the proportion, and random extraction is performed from the sample according to the proportion, and then the samples extracted from each data source are merged together to form new data. It should be noted that a fixed total number of sample extractions can also be set, calculate the proportion of the reliability constants of different data sources, and extract a certain number of samples from the corresponding data sources according to the proportion to make the total number of samples constant. The extraction method should not be limited to the extraction methods listed in this specific embodiment.

[0025] As Figure 2 shown, when the dynamic correlation supplementation method does not include the autocorrelation part of the time series, the dynamic correlation supplementation method includes: S1: Calculate the correlation coefficient of each feature value with respect to the missing feature value according to the remaining feature values of the data, based on the formula: is the number of samples in the obtained data, where is the correlation coefficient between the j-th feature and the k-th feature, is the feature value of the j-th feature of the i-th sample, is the mean of the j-th feature among n samples, is the eigenvalue of the k-th feature of the i-th sample, is the mean of the k-th feature among n samples; S2: Determine the correlation weight between other features and the missing feature according to the correlation coefficient, according to the formula:

[0026] in represents the correlation weight of the k-th feature for filling the missing value of the j-th feature, is the number of features of each sample; S3: Obtain the fitted linear relationship based on historical data and calculate the relevant supplementary value, according to the formula: where is the dependent eigenvalue, is the intercept term, represents the eigenvalue of the k-th feature, is the regression coefficient of feature k, is the error term, is the relevant supplementary value of the j-th feature in sample i, is the eigenvalue of the k-th feature of the i-th sample; S4: Use the value to supplement the j-th feature in sample i.

[0027] It should be noted that for the convenience of understanding, the following simplified simulated data is set as Table 1: Table 1 Trading Day Data Type Opening Price (yuan) Closing Price (yuan) Trading Volume (lots) Price-to-Book Ratio 1 35.2 35.5 42000 2.9 2 35.8 35.6 40000 2.8 3 36.1 Missing 43000 2.9 4 35.9 36.2 45000 3.0 5 36.3 36.0 40000 2.9 According to the above data (the original data collected), first calculate (the correlation coefficient between the opening price and the closing price), and substitute the above data into the calculation formula: to obtain The value of is approximately 0.66 (when calculating, since the closing price on the third day is missing, when calculating the mean of the closing price and the correlation coefficient between the opening price and the closing price, the data on the third day is removed).

[0028] Similarly, it can be calculated that (the correlation coefficient between the trading volume and the closing price) is approximately 0.50, (the correlation coefficient between the price-to-book ratio and the closing price) is approximately 0.75. According to the formula: It can be calculated that is approximately 0.34, is approximately 0.27, is approximately 0.39.

[0029] For the convenience of calculation, assume that the linear relationship obtained from historical data is: closing price = opening price × 2.5 + trading volume × 0.00004 + price-to-book ratio × 0.7, that is is 0, is 2.5, is 0.00004, is 0.7, is 0, substituting into the formula:

[0030] It can be obtained that (the relevant supplementary value of the closing price on the third day) is approximately 36.1. The supplementary value is calculated by the dynamic correlation supplementary method to supplement the missing characteristic values of the data. According to the correlation between other characteristic values and the missing characteristic values of the data, the relevant weights of other characteristic values for the missing characteristic values are determined. According to the different data obtained, the correlation changes dynamically, and the supplement of the missing characteristic values is more in line with the overall trend of the data, improving the accuracy of the supplementary value.

[0031] In this embodiment, the method for obtaining the fitted linear relationship according to historical data includes: L1: Collect historical data and set the missing characteristic values as the dependent characteristic according to the historical data; L2: Set as the intercept term, represents the characteristic value of the k-th characteristic, is the regression coefficient of characteristic k, is the error term, and determine the formula ; L3: Estimate the regression coefficient by the least squares method to determine , and values, so as to determine the fitted linear relationship.

[0032] It should be noted that the method for determining the regression coefficient by the least squares method is a prior art and will not be elaborated here. Moreover, other methods can also be used to fit the linear relationship, not limited to the least squares method used in this embodiment.

[0033] Embodiment 2: When the missing features of the data have an autocorrelated part in terms of time, the accuracy of supplementing the missing feature values only with the remaining feature values is poor. In the first embodiment, there is a certain influence among the closing prices of each trading day. Calculating the closing price only through the opening price, trading volume, and price-to-book ratio of the third trading day does not fully consider the mutual influence among the closing prices of multiple trading days before and after. In this embodiment, the time correlation of the missing features themselves is calculated, and the supplementary value is more accurate.

[0034] As Figure 2 shown, the difference from the first embodiment is that: the dynamic correlation supplement method further includes: G1: Select a number of consecutive complete samples, calculate the autocorrelation coefficient, according to the formula:

[0035] where is the autocorrelation coefficient at lag period, is the number of consecutive data periods selected, is the average value of the values in consecutive periods, is the data value of the t-th period in the selected data periods, is the data value of the t - δ-th period in the selected data periods; G2: Calculate the autocorrelation weight according to the autocorrelation coefficient:

[0036] where is the autocorrelation weight at lag period, is the autocorrelation coefficient at lag c period, is the set maximum lag period number; G3: Calculate the final supplementary value according to the autocorrelation weight and the relevant supplementary value:

[0037] where is the autocorrelation supplementary value of the j-th feature in sample i, is the final supplementary value of the j-th feature in sample i, is the feature value of the j-th feature of the th sample, is the self-set weight, is the relevant supplementary value of the j-th feature in sample i; G4: Use the value to supplement the j-th feature in sample i.

[0038] It should be noted that for the convenience of calculation and considering only one lag period, the closing prices of the fourth and fifth trading days are selected as continuous data periods for calculation.

[0039] According to the formula:

[0040] It can be calculated is 0.4, and then according to the formula:

[0041] Available (The autocorrelation supplementary value of the closing price on the third day) is 35.6 (for the convenience of calculation, only one lag is considered, so The value is 1. In actual use, the more lag periods are considered, the more accurate the calculation is). Set The value of is 0.5, and we can get (The final supplementary value of the closing price on the third day) is 35.85. According to the changing trend of the data in different time series, the autocorrelation coefficient of the missing feature is determined, and the autocorrelation weight of the eigenvalue of the previous period of the missing feature to the eigenvalue of the missing feature is calculated. The changing trend of the missing feature is fully considered, and the final supplementary value is determined by combining the influence of other eigenvalues ​​on the missing feature, which further improves the accuracy of the supplementary value.

[0042] Embodiment three: Merging data from different sources will generate a lot of computational effort, and the authenticity of data from different sources is different. Therefore, it is necessary to obtain the reliability constants of different sources through the reliability constant calculation method. According to the reliability constants of data from different sources, data from different sources are extracted in different proportions to reduce the computational burden, increase the preprocessing speed, and ensure that the number of samples from data sources with higher accuracy is larger, thereby improving the accuracy of the data.

[0043] like Figure 3 As shown, the reliability constant calculation method includes: Get the total number of samples, the number of features in each sample, and the number of missing features in each sample, and remove samples with a missing ratio of more than 0.5 from several samples, and then calculate according to the formula: , , in is the reliability constant of the data source, is the data defect constant, reflecting the degree of data missing. The credibility of the data source is between 0 and 1, which is set by the technical staff. The value indicates the degree of recognition of the authenticity of the data source. is the number of missing features in sample i, is the total number of samples in the acquired data, is the number of features for each sample.

[0044] It should be noted that for the convenience of calculation, the following simulated data is used in Table 2 below. The data in the table is the attention time (days) of investors on different stocks obtained from a certain stock exchange forum: Table 2 Stock Investors No. 1 No. 2 No. 3 No. 4 No. 5 100 15 25 45 18 22 200 5 Missing 5 Missing Missing 300 0 Missing 15 8 9 According to the above data (the original data collected), in the sample of the second investor, the data missing ratio reaches more than 0.5. Therefore, the second sample is excluded, and then calculate (data defect constant), and substitute the above data into the formula:

[0045] The calculation shows that is 0.83, is between 0 and 1. The larger the value, the more complete the data. Set to be 0.9, indicating a relatively high recognition of the authenticity of the data source. Then, according to the formula:

[0046] Calculate and obtain (reliability constant) is 0.747.

[0047] Example 4: When calculating the reliability constant by dividing samples in Example 3, the influence of a single sample with serious defects can be excluded. However, when the missing features of the data are relatively concentrated, the reliability constant calculated by this calculation method is on the high side and cannot accurately reflect the reliability of the data source. The difference from Example 3 is that this example provides another method for calculating the reliability constant, which is calculated by characteristics and is more sensitive to the missing features and more accurate in calculation.

[0048] As Figure 4 shown, the method for calculating the reliability constant includes: Obtain the total number of samples, the number of features for each sample, and the number of missing features in each sample, and calculate according to the formula: ,

[0049] where is the reliability constant of the data source, is the data defect constant, reflecting the degree of data missing, For the credibility of the data source, the value ranges from 0 to 1 and is set by technicians. The magnitude of the value represents the degree of recognition of the authenticity of the data source. is the number of samples lacking feature j. is the total number of samples in the acquired data. is the number of features of each sample.

[0050] It should be noted that for the convenience of understanding, the data in Embodiment 3 is used for calculation. According to the formula:

[0051] It can be obtained that is 0.73. Set the value of to 0.9, indicating a relatively high degree of recognition of the authenticity of the data source. Then, according to the formula:

[0052] Calculate (reliability constant) is 0.657.

[0053] It should be noted that Embodiment 3 and Embodiment 4 adopt different calculation methods, and the calculation results of the reliability constant of the data source are different. Calculating by sub-samples in Embodiment 3 can exclude the influence of a single sample with serious defects. Calculating by sub-features in Embodiment 4 is more sensitive to the lack of features and the calculation is more accurate. It can be selected according to actual needs, and different calculation methods can be selected for different data sources.

[0054] Embodiment 5: The reliability of the data source is not constant. Affected by various factors, the quasi-reliability of the data source may gradually increase or gradually decrease. In Embodiment 3 and Embodiment 4, the reliability of the data source cannot be updated based on the collected data, and it is difficult to ensure the long-term accuracy of the data. Based on Embodiment 3 and Embodiment 4, this embodiment provides a credibility update method to continuously update the credibility and ensure the reliability of the data source, making the integrated data more accurate.

[0055] As Figure 5 shown, in Embodiment 3 and Embodiment 4, when the reliability constant of the data source ≤0.3 or ≥0.7, the credibility of the data source is updated through the credibility update method to improve the accuracy of the next calculation. The credibility update method includes: When 0.2 < ≤0.3, ; when 0.1 < ≤0.2, ; When ≤0.1, , and prompt the technician to re - determine the data source; When 0.7 ≤ ≤0.8, ; When 0.8 < ≤0.9, ; When >0.9, ; Where is the credibility of the updated data source, is the credibility of the data source before update, and where is the initial credibility, set by the technician, is the update speed constant, set by the technician, and the larger the value, the greater the update amplitude.

[0056] It should be noted that , , are all different names in different calculation cycles. Set to 0.03. In Example 3, calculate (reliability constant) to be 0.747, 0.7 ≤ ≤0.8, and the credibility before update is 0.9, the calculation result of is 0.93. However, assume is set to 0.94, the maximum cannot exceed 0.94. Therefore, the updated is 0.93. When calculating the reliability constant of this data source next time, calculate according to 0.93. Additionally, assume is set to 0.92, the maximum cannot exceed 0.92. Although the calculation result of the updated is 0.93, when calculating the reliability constant of this data source next time,

[0057] calculate according to 0.92. Generally speaking, when the calculated reliability constant is large, the credibility of this data source can be improved; when the calculated reliability constant is small, the credibility of this data source can be reduced; when the credibility of the data source is less than the threshold, the technician can be prompted to abandon this data source and find a new data source to ensure the availability of the acquired data.

[0057] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended embodiments and their equivalents.

Claims

1. A data preprocessing method applicable to big data analysis, characterized in that: Including: Collecting raw data through big data and recording the data sources; After deduplicating the raw data, using the dynamic correlation supplementation method to calculate the supplementary value and supplement the missing feature values of the data; Obtaining the reliability constants of different sources through the reliability constant calculation method, and extracting different proportions of data from different sources according to the data reliability constants of different sources; Combining and uniformly processing the data from different sources and different formats, performing format conversion and encoding unification operations on the data to make the data consistent in structure and semantics; Normalizing the integrated data to make the data have the same scale and distribution.

2. The data preprocessing method applicable to big data analysis according to claim 1, wherein: The dynamic correlation supplementation method includes: S1: Calculating the correlation coefficient of each feature value with respect to the missing feature value according to the remaining feature values of the data, based on the formula: , is the number of samples in the acquired data, where is the correlation coefficient between the j-th feature and the k-th feature, is the eigenvalue of the j-th feature of the i-th sample, is the mean value of the j-th feature among n samples, is the eigenvalue of the k-th feature of the i-th sample, is the mean value of the k-th feature among n samples; S2: Determining the correlation weight between other features and the missing feature according to the correlation coefficient, according to the formula: , wherein represents the relevant weight of the k-th feature for filling the missing value of the j-th feature, is the number of features for each sample; S3: Obtaining the fitted linear relationship from historical data and calculating the relevant supplementary value, based on the formula: , , wherein is the dependent eigenvalue, is the intercept term, represents the eigenvalue of the k-th feature, is the regression coefficient of feature k, is the error term, is the relevant supplementary value of the j-th feature in sample i, is the eigenvalue of the k-th feature of the i-th sample; S4: Use the value to supplement the j-th feature in sample i.

3. A data preprocessing method applicable to big data analysis according to claim 2, characterized in that: The dynamic correlation supplementation method also includes, for the autocorrelation part of the time series: G1: Selecting several consecutive complete samples, calculating the autocorrelation coefficient, according to the formula: , where is the autocorrelation coefficient of the lag period, is the number of consecutive data periods selected, is the average value of the consecutive periods, is the data value of the t-th period among the selected data periods, is the data value of the (t - δ)-th period among the selected data periods; G2: Calculating the autocorrelation weight according to the autocorrelation coefficient: , wherein is the autocorrelation weight of the lag period, is the autocorrelation coefficient of the lag c period, is the set maximum number of lag periods; G3: Calculating the final supplementary value according to the autocorrelation weight and the relevant supplementary value: , where is the autocorrelation supplementary value of the j-th feature in sample i, is the final supplementary value of the j-th feature in sample i, is the eigenvalue of the j-th feature of the th sample, is the correlation supplementary value of the j-th feature in sample i; G4: Use to supplement the j-th feature in sample i.

4. A data preprocessing method applicable to big data analysis according to claim 2, characterized in that: The method for obtaining the fitted linear relationship from historical data includes: L1: Collecting historical data and setting the missing feature value as the dependent feature according to the historical data; L2: Setting is the intercept term, represents the eigenvalue of the k-th feature, is the regression coefficient of feature k, is the error term, and determine the formula ; L3: Estimate the regression coefficients by the least squares method to determine , and values, thereby determining the fitted linear relationship.

5. A data preprocessing method applicable to big data analysis according to claim 1, characterized in that: The reliability constant calculation method includes: Obtaining the total number of samples, the number of features of each sample, and the number of missing features in each sample, and excluding the samples with a missing ratio exceeding 0.5 from several samples, and then calculating according to the formula: , Among them is the reliability constant of the data source, is the data defect constant, reflecting the degree of data missing, is the initial credibility of the data source, with a value between 0 and 1, set by technicians, and its numerical value represents the degree of recognition of the authenticity of the data source, is the number of missing features in sample i, is the total number of samples in the obtained data, is the number of features of each sample; When the reliability constant of the data source ≤ 0.3 or ≥ 0.7, the initial credibility of the data source is updated to by the credibility update method to improve the accuracy of the next calculation.

6. A data preprocessing method applicable to big data analysis according to claim 1, characterized in that: The reliability constant calculation method includes: Obtaining the total number of samples, the number of features of each sample, and the number of missing features in each sample, and calculating according to the formula: , , wherein is the reliability constant of the data source, is the data defect constant, reflecting the degree of data missing, is the credibility of the data source, with a value between 0 and 1, and the numerical value represents the degree of recognition of the authenticity of the data source, is the number of samples lacking feature j, is the total number of samples in the obtained data, is the number of features of each sample; When the reliability constant of the data source ≤ 0.3 or ≥ 0.7, the credibility value of the data source is updated through the credibility update method to improve the accuracy of the next calculation.

7. A data preprocessing method applicable to big data analysis according to claim 5 or 6, characterized in that: The credibility update method includes: When 0.2 < ≤ 0.3, ; when 0.1 < ≤ 0.2, ; when ≤ 0.1, , and prompt the technician to re-determine the data source; When 0.7 ≤ ≤ 0.8, ; when 0.8 < ≤ 0.9, ; when > 0.9, ; Among them is the credibility of the updated data source, is the credibility of the data source before update, and , where is the initial credibility, set by the technical staff, is the update speed constant, set by the technical staff, and the larger the value, the greater the update amplitude.

8. A data preprocessing system applicable to big data analysis, characterized in that: Including: Data collection module: Used to collect raw data through big data and record the data sources; Data processing module: Used for deduplicating raw data, using the dynamic correlation supplementation method to calculate the supplementary value, supplementing the missing feature values of the data, determining the correlation weight of other feature values with respect to the missing feature value according to the correlation between other feature values and the missing feature value of the data, and the correlation changes dynamically according to the obtained data, making the supplementation of the missing feature value more in line with the overall trend of the data and improving the accuracy of the supplementary value; Data screening module: Used to obtain the reliability constants of different sources through the reliability constant calculation method, and extracting different proportions of data from different sources according to the data reliability constants of different sources to reduce the calculation burden and improve the preprocessing speed; Data integration module: Combining and uniformly processing the data from different sources and different formats, performing format conversion and encoding unification operations on the data to make the data consistent in structure and semantics; Data conversion module: Normalize the integrated data to make the data have similar scales and distributions, which can improve the convergence speed of machine learning algorithms and the accuracy of models when the integrated data is used for model training.

Citation Information

Patent Citations

  • Method for filling station meteorological data missing value based on svd (singular value decomposition) algorithm

    CN106295175A

  • Flight data missing supplementing method based on neural network

    CN110782007A

  • Voltage missing value filling method based on historical data auxiliary scene analysis

    CN111507412A

  • Missing value filling method based on attribute dynamic selection and gray correlation analysis

    CN113159194A

  • Intelligent detection and filling method for data missing value of agricultural Internet of Things

    CN119988850A