Comprehensive data preprocessing method, system and equipment and storage medium
Automated data preprocessing through machine learning algorithms solves the problems of low efficiency and unstable quality of manual operations in existing technologies, and realizes an efficient and adaptable data preprocessing process.
Patent Information
- Application Number
- CN202410381402.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-29
- Publication Date
- 2025-09-30
AI Technical Summary
Existing data preprocessing methods rely on manual operations, resulting in heavy workload, low efficiency, unstable quality, and lack of adaptability to variable and complex data sets.
Machine learning algorithms are used to automatically perform data classification, missing value filling, outlier processing, and other data preprocessing functions such as data cleaning, standardization, transformation, and feature extraction, providing a unified framework that adapts to the characteristics of different data sets.
An automated and optimized data pre-processing process is implemented, which improves efficiency and quality, reduces human errors, and adapts to the variability and complexity of data sets.
Smart Images

Figure CN120724271A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of data preprocessing technology, and more particularly to a comprehensive data preprocessing method, system, device, and storage medium for data science and machine learning. Background Art
[0002] With the rise of big data, vast amounts of data have accumulated across all industries. This data is often diverse, high-dimensional, and large in volume. In this context, data preprocessing becomes a critical step in ensuring the accuracy of data analysis and machine learning models. However, most existing preprocessing methods rely on manual data classification and processing, which not only significantly increases the workload for data workers but also limits the efficiency and quality of data preprocessing, especially when the data volume is large. Furthermore, manual preprocessing struggles to maintain consistency and repeatability, potentially introducing human error and affecting the results of subsequent analysis. Existing technologies also often lack overall optimization of the preprocessing process, making them difficult to adapt to the variability and complexity of datasets. Summary of the Invention
[0003] The embodiments of the present disclosure provide a comprehensive data preprocessing method, system, device, and storage medium to solve or alleviate one or more of the above technical problems in the prior art.
[0004] According to one aspect of the present disclosure, a comprehensive data preprocessing method is provided, comprising:
[0005] Collect comprehensive data to be pre-processed;
[0006] performing classification processing on the comprehensive data and outputting the classified processing data;
[0007] Detecting and completing missing values in the comprehensive data, and outputting missing value-completed data;
[0008] Performing outlier processing on the comprehensive data, including detecting and repairing outliers in the comprehensive data, and outputting original values and location information of the outliers;
[0009] Performing a cleaning process on the comprehensive data, wherein the cleaning process includes deduplication, format matching, and format conversion, and outputting the cleaned data;
[0010] Performing standardization on the comprehensive data and outputting the standardized data and parameter information during standardization;
[0011] Performing conversion processing on the comprehensive data to change the distribution characteristics of the comprehensive data, and outputting the converted data and parameter information during the conversion;
[0012] Feature extraction is performed on the comprehensive data to generate new data after feature extraction and selection.
[0013] In a possible implementation, classifying the comprehensive data includes:
[0014] The comprehensive data is classified and processed by using K-means clustering algorithm and support vector machine algorithm.
[0015] In one possible implementation, detecting and filling missing values in the classified comprehensive data include:
[0016] The K nearest neighbor algorithm is used to fill missing values for non-numeric data;
[0017] Multiple imputation algorithms were used to fill missing values in numerical data.
[0018] In a possible implementation, an algorithm for performing outlier processing on the comprehensive data includes a support vector machine algorithm and a local outlier factor algorithm.
[0019] In a possible implementation, cleaning the comprehensive data includes:
[0020] Use regular expressions to match and identify irregular or incorrect formats in comprehensive data, or use cleaning rules to delete incorrect and duplicate data in comprehensive data.
[0021] In a possible implementation, converting the comprehensive data includes:
[0022] The Yeo-Johnson transformation algorithm is used to transform the comprehensive data so that the distribution characteristics of the comprehensive data tend to be normal distribution.
[0023] In a possible implementation, extracting features from the comprehensive data includes:
[0024] The minimum redundancy maximum relevance algorithm is used to extract features from the comprehensive data.
[0025] In one possible implementation, the following steps are included:
[0026] The classified processing data, missing value completion data, original values of outliers and positioning information, cleaned data, standardized data and parameter information during standardization, converted data and parameter information during conversion, and new data after feature extraction and selection are stored.
[0027] According to one aspect of the present disclosure, there is provided a comprehensive data preprocessing system, comprising:
[0028] A data acquisition module, used to collect comprehensive data to be pre-processed;
[0029] A data classification module, configured to classify the comprehensive data and output the classified data;
[0030] A missing value completion module is used to detect and complete missing values in the comprehensive data and output missing value completion data;
[0031] An outlier processing module is used to perform outlier processing on the comprehensive data, including detecting and repairing outliers in the comprehensive data, and outputting the original value and location information of the outliers;
[0032] A data cleaning module, configured to perform cleaning processing on the comprehensive data, wherein the cleaning processing includes deduplication, format matching and format conversion, and output cleaned data;
[0033] A data standardization module is used to standardize the comprehensive data and output the standardized data and parameter information during standardization;
[0034] A data conversion module is used to convert the comprehensive data to change the distribution characteristics of the comprehensive data, and output the converted data and parameter information during the conversion;
[0035] The feature extraction and selection module is used to extract features from the comprehensive data and generate new data after feature extraction and selection.
[0036] According to one aspect of the present disclosure, there is provided a device comprising:
[0037] processor and memory;
[0038] The memory is used to store a computer program, and the processor calls the computer program stored in the memory to execute any one of the above-mentioned comprehensive data preprocessing methods.
[0039] According to one aspect of the present disclosure, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, the processor is enabled to perform any of the above-mentioned comprehensive data preprocessing methods.
[0040] The exemplary embodiments of the present disclosure have the following beneficial effects: The exemplary embodiments of the present disclosure utilize advanced machine learning algorithms to automatically perform data classification, missing value imputation, outlier handling, and multiple other data preprocessing functions, such as data cleaning, standardization, transformation, feature extraction, feature selection, and dimensionality reduction. This method provides a unified framework that can adapt to the characteristics of different datasets and automatically optimize preprocessing steps to achieve optimal performance.
[0041] The details of one or more embodiments of the present application are set forth in the following drawings and description. Other features and advantages of the present application will become apparent from the accompanying drawings. It should be understood that the above general description and the detailed description that follows are merely exemplary and explanatory and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0043] Figure 1 is one of the flow charts of a comprehensive data preprocessing method of this exemplary embodiment;
[0044] Figure 2 This is the second flow chart of a comprehensive data preprocessing method of this exemplary embodiment;
[0045] Figure 3 is a block diagram of a comprehensive data preprocessing system of the present exemplary embodiment;
[0046] Figure 4 FIG. 4 is a schematic structural diagram of a device according to the exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0047] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the present disclosure will be more comprehensive and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or that other methods, components, devices, steps, etc. may be employed. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.
[0048] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware units or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0049] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all steps. For example, some steps may be decomposed, while some steps may be combined or partially combined, so the actual execution order may change according to actual circumstances.
[0050] The terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, such that the embodiments of the application described herein can, for example, be implemented in an order other than that illustrated or described herein.
[0051] In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus that includes a series of steps or submodules is not necessarily limited to those steps or submodules explicitly listed, but may include other steps or submodules not explicitly listed or inherent to such process, method, product or apparatus.
[0052] Figure 1 This is one of the flow charts of a comprehensive data preprocessing method of this exemplary embodiment. Figure 1 As shown, an exemplary embodiment of the present disclosure provides a comprehensive data preprocessing method, comprising:
[0053] S1 collects comprehensive data to be preprocessed;
[0054] S2 classifies the comprehensive data and outputs the classified data;
[0055] S3 detects and completes missing values in the comprehensive data, and outputs missing value-completed data;
[0056] S4 performs outlier processing on the comprehensive data, including detecting and repairing outliers in the comprehensive data, and outputting original values and location information of the outliers;
[0057] S5 performs cleaning processing on the comprehensive data, wherein the cleaning processing includes deduplication, format matching, and format conversion, and outputs the cleaned data;
[0058] S6 performs standardization processing on the comprehensive data and outputs the standardized data and parameter information during standardization;
[0059] S7 performs conversion processing on the comprehensive data to change the distribution characteristics of the comprehensive data, and outputs the converted data and parameter information during the conversion;
[0060] S8 performs feature extraction on the comprehensive data to generate new data after feature extraction and selection.
[0061] For example, the workflow of data preprocessing is as follows Figure 2 As shown in the figure, the data acquisition module first collects data and enters the data preprocessing phase. The second step is carried out in the data classification module, where the classified data is stored and then proceeds to the next process. The third step is missing value imputation, which detects and imputes missing values for each classified data point. The fourth step is outlier processing on the imputed data, detecting and correcting outliers in the data, and storing the original values and location information of the outliers before proceeding to the next process. The fifth step is data cleaning, which performs duplicate removal, format matching, and format conversion based on the predefined regular expression matching and cleaning rules. The cleaned data is then stored. After data cleaning, data standardization can be performed, and the output is the standardized data along with the parameters used for standardization. Data transformation is performed after data standardization, where the standardized data is further transformed to make its distribution closer to a normal distribution. The output is the transformed data along with the parameters used for the transformation. The final step is feature extraction and selection, which can be performed after data standardization or data transformation. New data is generated after feature extraction and selection and stored.
[0062] Specifically, classifying the comprehensive data includes:
[0063] The comprehensive data is classified and processed by using K-means clustering algorithm and support vector machine algorithm.
[0064] Specifically, the missing value detection and completion of the comprehensive data after classification processing include:
[0065] The K nearest neighbor algorithm is used to fill missing values for non-numeric data;
[0066] Multiple imputation algorithms were used to fill missing values in numerical data.
[0067] Specifically, the algorithm for performing outlier processing on the comprehensive data includes a support vector machine algorithm and a local outlier factor algorithm.
[0068] Specifically, cleaning the comprehensive data includes:
[0069] Use regular expressions to match and identify irregular or incorrect formats in comprehensive data, or use cleaning rules to delete incorrect and duplicate data in comprehensive data.
[0070] Specifically, converting the comprehensive data includes:
[0071] The Yeo-Johnson transformation algorithm is used to transform the comprehensive data so that the distribution characteristics of the comprehensive data tend to be normal distribution.
[0072] Specifically, extracting features from the comprehensive data includes:
[0073] The minimum redundancy maximum relevance (mRMR) algorithm was used to extract features from the comprehensive data.
[0074] Specifically, the comprehensive data preprocessing method includes:
[0075] The classified processing data, missing value completion data, original values of outliers and positioning information, cleaned data, standardized data and parameter information during standardization, converted data and parameter information during conversion, and new data after feature extraction and selection are stored.
[0076] The algorithm models in this embodiment are described as follows:
[0077] K-Means Clustering:
[0078] K-means clustering can identify patterns and structures in data in an unsupervised manner, divide the dataset into K cluster centers, and provide preliminary data division for supervised learning tasks.
[0079] K-means clustering assigns data points to K clusters such that the sum of the squared Euclidean distances between each point and the centroid (center point) of its cluster is minimized.
[0080] The objective function of K-means clustering is:
[0081]
[0082] where r ik Is a binary indicator scalar, i is the i-th data in the data point set, if the data point x i is assigned to cluster k, it is 1, otherwise it is 0, and μ k is the center of cluster k.
[0083] The details of the algorithm application are:
[0084] Randomly select k data points as initial cluster centers.
[0085] Assign each point to the nearest cluster center.
[0086] Update the center of each cluster to the mean of all points assigned to that cluster.
[0087] Repeat steps 2 and 3 until the cluster center no longer changes or the preset number of iterations is exceeded.
[0088] Support Vector Machine (SVM):
[0089] Support vector machines perform supervised learning for classification or regression tasks, and perform supervised training based on cluster assignment labels obtained from unsupervised learning algorithms.
[0090] Different categories of data points are distinguished by finding the optimal segmentation hyperplane in the feature space. This hyperplane can maximize the boundary between different categories of data points. The formula principle is:
[0091] Optimization problem (min||w|| 2 ) satisfies (y i (w·x i +b)≥1) for all (i)(2)
[0092] Among them, y i is the data point x i The category label, for the positive class, y i =1; for negative class, y i = -1. w is a weight vector, each component of which corresponds to the weight of a feature. b is a bias value that shifts the hyperplane along the normal direction specified by w, thereby changing the decision boundary.
[0093] This constraint ensures that all positive points are on one side of the hyperplane, and all negative points are on the other side, and neither falls on the boundary of the interval. In practical applications, there are rarely completely linearly separable cases, so soft margins and penalty parameters are introduced to deal with incompletely separable cases.
[0094] The details of the algorithm application are:
[0095] Using kernel techniques, nonlinear decision interfaces can be found in high-dimensional spaces.
[0096] Apply soft margins or penalty parameters to handle cases where the data is not completely separable.
[0097] Use cross-validation to select the best model parameters, such as penalty and kernel parameters.
[0098] K Nearest Neighbors (KNN):
[0099] The K-nearest neighbor algorithm uses the similarity between data points to estimate missing values. For each data point with missing values, the algorithm finds the K neighbors with the shortest Euclidean distance in the feature space (usually based on complete data), and then uses the non-missing values of these neighbors to estimate the missing value.
[0100] The algorithm application process is:
[0101] A suitable K value is determined through cross-validation.
[0102] For each missing value, find the K nearest neighbors in other dimensions. n ) and y=(y1,y2,...,y n The distance calculation formula between ) is:
[0103]
[0104] Where x and y are the coordinates of two points, x1, x2, ..., x n is the x-coordinate, y1,y2,...,y1 is the y-coordinate.
[0105] Use the average or weighted average of the corresponding feature values of these neighbors to fill the missing values. If the average value is used: For a given data point x' with missing values, the missing value of its jth feature can be filled by the following formula:
[0106]
[0107] where x ij is the jth eigenvalue of the ith nearest neighbor, K is the number of nearest neighbors, i represents the ith nearest neighbor, i∈[1,k]; j represents the jth eigenvalue.
[0108] If you use weighted averaging to fill in missing values, you can:
[0109]
[0110] where w i is the weight, which is usually inversely proportional to the distance to the neighbor, that is, the closer the neighbor, the greater the weight.
[0111] Multiple imputation:
[0112] Multiple imputation is used to create multiple different "complete" data sets and estimate missing data multiple times. The imputation of missing values in each data set is based on a random sampling of the existing data to obtain more accurate estimates.
[0113] The algorithm application process is:
[0114] First, estimate the distribution parameter θ of the missing data and obtain the estimated value of the distribution parameter
[0115] Then, according to the estimated distribution, the missing values are sampled:
[0116]
[0117] in is a given parameter The probability distribution of the parameters is , and x′ is a possible imputation for the missing value.
[0118] Repeat the above steps to generate multiple complete data sets.
[0119] The analysis was performed on each complete data set, and the results were then pooled to produce the final estimates and confidence intervals.
[0120] One-Class SVM:
[0121] A first-class support vector machine (SVM) is a method derived from the support vector machine (SVM). Its goal is to find a decision function that can include all normal data points as much as possible while excluding outliers. In feature space, this function is usually defined as a hyperplane with a maximum distance from the origin.
[0122] The algorithm application process is:
[0123] The objective function of a first-class support vector machine can be written as:
[0124]
[0125] Where w represents the normal vector of the hyperplane, ξ i is the slack variable, ρ is the offset term of the decision function, v is a parameter used to control the proportion of outliers and the boundary of the decision function, i represents the index of the summation, which is used to traverse all data points, and m represents the total number of data points.
[0126] Then, the objective function in step 1 also needs to satisfy the following constraints:
[0127] (w·φ(x i ))≥ρ-ξ i , i=1,...,n (8)
[0128] ξ i ≥0
[0129] where ξ i represents the slack variable, φ(x i ) is the mapping function that maps the original data to a high-dimensional space, and n is the total number of data points.
[0130] Finally, the decision function is defined as:
[0131] f(x)=sign(ρ-(w·p(x))) (9)
[0132] When f(x) = 1, x is considered a normal point; when f(x) = -1, it is considered an abnormal point. φ(x) is a mapping function that maps the original data to a high-dimensional space.
[0133] Local Outlier Factor (LOF):
[0134] The Local Outlier Factor is a density-based algorithm for detecting outliers based on local density deviations. The Local Outlier Factor calculates the degree of outlier at each point by comparing the local density of a point with that of its neighbors.
[0135] The algorithm workflow is:
[0136] For each point p, calculate the kth reachability distance of all points in the kth neighborhood of p
[0137] reachability_distance k (p, o) = max(distance k (o), d(o, p)) (10)
[0138] Among them, d k (o) is the kth distance of the neighborhood point o, and d(o, p) is the distance from the neighborhood point o to the point p.
[0139] Then calculate the local k-th local reachability density of each point p:
[0140]
[0141] Among them, N k (p) is the kth distance neighborhood of p.
[0142] Finally, the local outlier factor of each point p is calculated:
[0143]
[0144] Among them, N k (p) is the k-th distance neighborhood of p. If LOF(p) is significantly greater than 1, p may be an outlier.
[0145] For the data points with the largest n local outlier factors, output the outlier value set O:
[0146] O=o1,o2,…,o n
[0147] Among them, o1, o2, …, o n They represent the 1st, 2nd, ..., nth outliers respectively.
[0148] Z-score standardization:
[0149] Z-score normalization (also known as standard deviation normalization) involves rescaling the feature values of the data so that the feature has a mean of 0 and a standard deviation of 1. This transformation is very important for subsequent applications of the dataset, such as data mining, linear regression, machine learning, and neural networks, especially when using distance-based algorithms.
[0150] The algorithm workflow is:
[0151] First, we calculate the mean μ and standard deviation σ of each feature.
[0152] Then, each data point is transformed according to the following formula:
[0153]
[0154] Where x is the original data, μ is the mean, and σ is the standard deviation.
[0155] Yeo-Johnson transform:
[0156] The Yeo-Johnson transformation is a method used to improve data normality and can handle datasets containing zero and negative values. By performing a monotonic transformation on the data, it is used to convert non-normally distributed data into a form that is closer to a normal distribution. This makes some algorithms less sensitive to non-normal characteristics, which generally helps to improve and stabilize the variance and mean.
[0157] The algorithm works as follows:
[0158]
[0159] Where λ is the transformation coefficient, which is determined by maximum likelihood estimation, X represents the original data, and Y represents the transformed data.
[0160] Minimum Redundancy Maximum Relevance (mRMR) algorithm:
[0161] The minimum redundancy maximum relevance algorithm is a feature selection method that aims to simultaneously minimize the redundancy between features and maximize the relevance of features with the target variable.
[0162] Specific application details:
[0163] First, the mutual information between each feature and the target variable is calculated, and the features are ranked.
[0164] Then, features that are highly correlated with the target variable but have low redundancy with the selected feature set are iteratively selected.
[0165] Figure 3 is a block diagram of a comprehensive data preprocessing system of this exemplary embodiment. Figure 3 As shown, an exemplary embodiment of the present disclosure provides a comprehensive data pre-processing system, comprising:
[0166] The data acquisition module 10 is used to collect comprehensive data to be pre-processed;
[0167] The data classification module 20 is used to classify the comprehensive data and output the classified data;
[0168] A missing value filling module 30 is used to detect and fill missing values in the comprehensive data and output missing value filling data;
[0169] An outlier processing module 40 is used to perform outlier processing on the comprehensive data, including detecting and repairing outliers in the comprehensive data, and outputting original values and location information of the outliers;
[0170] A data cleaning module 50 is configured to perform cleaning processing on the comprehensive data, including deduplication, format matching, and format conversion, and output cleaned data;
[0171] A data standardization module 60 is used to standardize the comprehensive data and output the standardized data and parameter information during standardization;
[0172] The data conversion module 70 is used to convert the comprehensive data to change the distribution characteristics of the comprehensive data and output the converted data and parameter information during the conversion;
[0173] The feature extraction and selection module 80 is used to extract features from the comprehensive data and generate new data after feature extraction and selection.
[0174] The framework of this embodiment includes:
[0175] Data acquisition module: This module does not directly use machine learning algorithms. It mainly collects and organizes data, and sets the data format and structure for subsequent modules to ensure that the data quality meets the requirements of subsequent processing processes.
[0176] Data Classification Module: This module consists of two algorithms: K-Means Clustering and Support Vector Machines (SVMs). It primarily classifies raw datasets. This module can switch between unsupervised and supervised learning, using K-Means for exploratory data analysis. When raw data lacks labels, K-Means clustering is used for preliminary data classification. When the target variable is known, SVMs are used for precise classification.
[0177] Missing Value Filler: Data is already classified when it enters this module, so this module primarily helps fill in any missing data to facilitate subsequent analysis or modeling. This module uses both K-nearest neighbor (KNN) and multiple imputation to fill in missing values. K-nearest neighbor is used for filling in missing values in non-numeric data, such as text labels; multiple imputation is used for filling in missing values in numeric data.
[0178] Outlier Processing Module: After data classification and missing value processing, the Outlier Processing Module helps identify and address outliers in the data center, preventing these data from misleading subsequent analysis. This module uses a one-class support vector machine (One-Class SVM) and a local outlier factor (LOF) for outlier detection and processing. One-class support vector machines are supervised methods that require a dataset of both normal and abnormal data for training to detect outliers. The local outlier factor is a statistical method based on the principle that the density around a normal value is similar to the density around its neighborhood, while the density around an outlier is significantly different from the density around its neighborhood. You can choose to use one of these algorithms based on the data quality, data type, and requirements for outlier processing.
[0179] Data Cleaning: This module is performed after outlier processing to ensure data format consistency and accuracy. It does not involve machine learning algorithms, but instead uses regular expressions and cleaning rules for data cleaning. Regular expressions are used to match and identify irregularities or incorrect formats in the data, while cleaning rules are used to remove incorrect and duplicate data.
[0180] Data Normalization: After data cleaning, the data features are converted to a consistent scale to facilitate subsequent comparison and analysis. This module does not involve complex machine learning algorithms but uses Z-score normalization to achieve data normalization. The parameters of the normalization transformation are saved for subsequent denormalization.
[0181] Data transformation module: This module is implemented using the Yeo-Johnson transformation and is performed after data normalization. If subsequent analysis requires statistical properties of the data, this module can be used to improve the distribution characteristics of the data, making the data set closer to a normal distribution, thereby improving certain statistical analyses or the assumptions of model establishment.
[0182] Feature Extraction and Selection: This module uses the Minimum Redundancy Maximum Relevance (mRMR) algorithm and is performed at the end to determine the final set of data features for modeling. After processing, the dataset minimizes redundant information between selected features and maximizes the relevance between the selected features and the target variable.
[0183] The system optimization content proposed in this embodiment is as follows:
[0184] Parallel processing: Supports parallel computing for large data sets. By splitting large data sets into smaller pieces and processing them in parallel on multiple processing units, the overall processing time can be significantly reduced.
[0185] Caching mechanism: During data preprocessing, some calculations may be performed repeatedly. A built-in caching mechanism can store these calculation results to avoid repeated calculations and thus improve processing speed.
[0186] Adaptive algorithm selection: Automatically selects the most suitable preprocessing algorithm and parameter settings by evaluating the characteristics of the dataset, such as data distribution, noise level, and dimensionality, to ensure the optimal balance between preprocessing quality and efficiency.
[0187] Resource management: Monitor hardware resource usage in real time, such as CPU and memory usage, and dynamically adjust algorithm parameters (such as batch size and number of iterations) based on available resources to avoid excessive resource consumption.
[0188] Modular design: This pretreatment method adopts a modular design. Each module can be optimized and upgraded individually, which is easy to maintain and allows users to customize the pretreatment process according to their needs.
[0189] User feedback mechanism: Provides a user feedback interface where users can provide feedback based on the preprocessing results. This information will be used to train and optimize the preprocessing decision model.
[0190] The following examples demonstrate how to apply this disclosure to actual data preprocessing scenarios:
[0191] Scenario: A shopping mall management company wants to analyze the energy consumption data of its multiple shopping malls, perform energy consumption forecasts, optimize energy use, and reduce operating costs.
[0192] Data collection: Various energy consumption data are collected from the commercial plaza's energy management system, including historical records of electricity, water, and natural gas consumption, as well as relevant environmental parameters such as indoor and outdoor temperature, humidity, and mall opening hours.
[0193] Data classification: First, the K-means clustering algorithm is used to perform preliminary clustering of shopping malls to identify different energy consumption patterns. Then, support vector machines are used to subdivide high-, medium-, and low-energy-consuming commercial plazas to provide a basis for further analysis.
[0194] Missing value completion: For the categorized shopping mall energy consumption dataset, multiple interpolation and the K-nearest neighbor algorithm are used to complete the missing energy consumption data to ensure data integrity and provide support for accurate energy consumption analysis.
[0195] Outlier processing: A first-class support vector machine and local anomaly factor algorithm are applied to identify data that does not conform to daily energy consumption patterns, such as abnormal consumption caused by equipment failure or data recording errors, so as to promptly eliminate abnormal factors that affect analysis.
[0196] Data cleaning: De-duplication, format verification, and correction of inconsistent data, as well as cleaning of inappropriate data input and erroneous information.
[0197] Data standardization and transformation: Use Z-score standardization and Yeo-Johnson transformation to standardize data and improve data distribution, preparing data in a suitable format for data mining and machine learning models.
[0198] Feature extraction and selection: The minimum redundancy maximum relevance (mRMR) algorithm is used to select the most relevant features from a number of factors that may affect energy consumption, such as weather conditions and mall opening hours, to improve the accuracy of the energy consumption prediction model.
[0199] By applying this embodiment, the commercial plaza management company can effectively complete the preprocessing of energy consumption data and build a high-precision energy consumption prediction model, thereby achieving reasonable energy allocation and effective cost control, and providing data support for future energy management decisions.
[0200] Figure 4 FIG. 1 is a schematic diagram of the structure of a device of this exemplary embodiment. Figure 4As shown, corresponding to the comprehensive data preprocessing method provided above, the present disclosure also provides a device. Since the embodiment of the device is similar to the above method embodiment, the description is relatively simple. Please refer to the description of the above method embodiment for relevant details. The device described below is only schematic. The device may include: a processor (processor) 1, a memory (memory) 2 and a communication bus (i.e., the above-mentioned device bus) and a search engine, wherein the processor 1 and the memory 2 communicate with each other through the communication bus and communicate with the outside through the communication interface. The processor 1 can call the logic instructions in the memory 2 to execute the comprehensive data preprocessing method.
[0201] In addition, the logic instructions in the above-mentioned memory 2 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes: a memory chip, a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0202] On the other hand, an embodiment of the present disclosure further provides a processor-readable storage medium, on which a computer program 3 is stored. When the computer program 3 is executed by the processor 1, the comprehensive data preprocessing method provided in the above embodiments is implemented.
[0203] The processor-readable storage medium can be any available medium or data storage device that can be accessed by the processor 1, including but not limited to magnetic storage (such as floppy disks, hard disks, tapes, magneto-optical disks (MO)), optical storage (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor storage (such as ROMs, EPROMs, EEPROMs, non-volatile memories (NANDFLASH), solid-state drives (SSDs)), etc.
[0204] The above are merely preferred embodiments of the present disclosure. The scope of protection of the present disclosure is not limited to the above embodiments. All technical solutions based on the principles of the present disclosure are within the scope of protection of the present disclosure. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present disclosure should be considered within the scope of protection of the present disclosure.
Claims
1. A comprehensive data preprocessing method, characterized in that: include: Collect comprehensive data to be pre-processed; performing classification processing on the comprehensive data and outputting the classified processing data; Detecting and completing missing values in the comprehensive data, and outputting missing value-completed data; Performing outlier processing on the comprehensive data, including detecting and repairing outliers in the comprehensive data, and outputting original values and location information of the outliers; Performing a cleaning process on the comprehensive data, wherein the cleaning process includes deduplication, format matching, and format conversion, and outputting the cleaned data; Performing standardization on the comprehensive data and outputting the standardized data and parameter information during standardization; Performing conversion processing on the comprehensive data to change the distribution characteristics of the comprehensive data, and outputting the converted data and parameter information during the conversion; Feature extraction is performed on the comprehensive data to generate new data after feature extraction and selection.
2. The comprehensive data preprocessing method according to claim 1, characterized in that: Classifying the comprehensive data includes: The comprehensive data is classified and processed by using K-means clustering algorithm and support vector machine algorithm.
3. The comprehensive data preprocessing method according to claim 1, characterized in that: The detection and completion of missing values for the comprehensive data after classification processing include: The K nearest neighbor algorithm is used to fill missing values for non-numeric data; Multiple imputation algorithms were used to fill missing values in numerical data.
4. The comprehensive data preprocessing method according to claim 1, characterized in that: The algorithms for processing outliers on the comprehensive data include a support vector machine algorithm and a local outlier factor algorithm.
5. The comprehensive data preprocessing method according to claim 1, characterized in that: Cleaning the comprehensive data includes: Use regular expressions to match and identify irregular or incorrect formats in comprehensive data, or use cleaning rules to delete incorrect and duplicate data in comprehensive data.
6. The comprehensive data preprocessing method according to claim 1, characterized in that: The conversion process of the comprehensive data includes: The Yeo-Johnson transformation algorithm is used to transform the comprehensive data so that the distribution characteristics of the comprehensive data tend to be normal distribution.
7. The comprehensive data preprocessing method according to claim 1, characterized in that: Extracting features from the comprehensive data includes: The minimum redundancy maximum relevance algorithm is used to extract features from the comprehensive data.
8. The comprehensive data preprocessing method according to claim 1, characterized in that: include: The classified processing data, missing value completion data, original values of outliers and positioning information, cleaned data, standardized data and parameter information during standardization, converted data and parameter information during conversion, and new data after feature extraction and selection are stored.
9. A comprehensive data preprocessing system, characterized in that: include: A data acquisition module, used to collect comprehensive data to be pre-processed; A data classification module, configured to classify the comprehensive data and output the classified data; A missing value completion module is used to detect and complete missing values in the comprehensive data and output missing value completion data; An outlier processing module is used to perform outlier processing on the comprehensive data, including detecting and repairing outliers in the comprehensive data, and outputting the original value and location information of the outliers; A data cleaning module, configured to perform cleaning processing on the comprehensive data, wherein the cleaning processing includes deduplication, format matching and format conversion, and output cleaned data; A data standardization module is used to standardize the comprehensive data and output the standardized data and parameter information during standardization; A data conversion module is used to convert the comprehensive data to change the distribution characteristics of the comprehensive data, and output the converted data and parameter information during the conversion; The feature extraction and selection module is used to extract features from the comprehensive data and generate new data after feature extraction and selection.
10. A device, characterized in that include: processor and memory; The memory is used to store a computer program, and the processor calls the computer program stored in the memory to execute the comprehensive data preprocessing method according to any one of claims 1 to 9.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is enabled to execute the comprehensive data preprocessing method according to any one of claims 1 to 9.
Citation Information
Cited By
Laying hen lossless data processing method based on improved SVR
CN121935488A