Abnormal data cleaning method and system based on multi-model fusion

Through the multi-model fusion method, combining isolated forest algorithms and clustering algorithms to identify abnormal data in complex data, and using RF-LSTM prediction model to fill in missing values, it solves the problem of difficult to identify and clean complex data in the prior art, and improves the completeness of data and the accuracy of decision-making.

CN120216878APending Publication Date: 2025-06-27XIAN THERMAL POWER RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510343201.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify and clean discrete and stacked anomaly data in complex data, especially when data volatility and randomness are high, traditional methods cannot accurately fill nonlinear data.

Method used

Multi-model fusion method is adopted, combined with isolated forest algorithm and clustering algorithm, discrete and stacked anomaly data are identified, and missing values ​​are predicted and filled through RF-LSTM prediction model.

Benefits of technology

Improve the accuracy and comprehensiveness of abnormal data identification, ensure the integrity and consistency of data sets, thereby improving the accuracy of business decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216878A_ABST
    Figure CN120216878A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of abnormal data cleaning, and discloses an abnormal data cleaning method and system based on multi-model fusion, and the method comprises the following steps: preprocessing to-be-cleaned data; identifying discrete abnormal data; identifying accumulation abnormal data; establishing a random forest model and a long and short term neural network model; combining the random forest model and the long-short term neural network model based on an artificial fish swarm algorithm to obtain an RF-LSTM prediction model; according to the method, data containing missing values are input into an RF-LSTM prediction model, a prediction result is output and filled into a data set, abnormal data cleaning is completed, abnormal data with discrete distribution and aggregation distribution can be effectively recognized in combination with an isolated forest algorithm and a clustering algorithm, the range of abnormal data recognition is widened, the accuracy and comprehensiveness of recognition are improved, and the method is suitable for large-scale popularization and application. And the missing value is predicted and filled by using the RF-LSTM prediction model, so that the integrity and consistency of the data set can be ensured, and the accuracy of the service decision is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of abnormal data cleaning, and relates to an abnormal data cleaning method and system based on multi-model fusion. Background Art

[0002] In data-driven models, the quality of the original data set has a crucial impact on the operation effect of the model. High-quality data can significantly improve the accuracy and reliability of the model, while low-quality data may lead to a decline in model performance and even incorrect conclusions. However, in the actual data collection process, the quality of data is often affected by various factors, mainly including the following two aspects: First, harsh working environments; sensors may face various harsh working environments in actual applications, such as high temperature, low temperature, high humidity, strong electromagnetic interference, etc. These environmental factors may cause the performance of the sensors to decline or even malfunction, thus affecting the accuracy of data collection. Second, the volatility and randomness of data; the volatility and randomness of the data itself will also affect the collection results. Actual data often has complex characteristics, such as time series data, multi-variable data, and non-linear characteristics.

[0003] Data cleaning includes two aspects of work: identifying errors and correcting anomalies. However, most traditional methods are based on probability distribution models, that is, first statistically analyze the distribution of the data set, determine the probability distribution of the data, and then use this probability distribution model to determine a normal value distribution range. For data that does not conform, linear interpolation, average value, etc. are used to fill or replace it. This method neither considers the information of related quantities at the same time nor has a good prediction result for data with strong non-linearity, and it is difficult to accurately fill the data. Summary of the Invention

[0004] Aiming at the deficiencies of the existing technology, the present invention aims to provide an abnormal data cleaning method and system based on multi-model fusion, starting from two perspectives of time series and related quantities, and using the method of multi-model fusion to jointly complete the prediction, and finally achieve accurate data cleaning work.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions: The present invention provides an abnormal data cleaning method based on multi-model fusion, including the following steps: preprocessing the data to be cleaned; identifying discrete abnormal data; identifying stacked abnormal data; establishing a random forest model and a long short-term neural network model; combining the random forest model and the long short-term neural network model based on the artificial fish swarm algorithm to obtain an RF-LSTM prediction model; inputting the data containing missing values into the RF-LSTM prediction model, and outputting the prediction result to fill the data set to complete the cleaning of abnormal data.

[0006] Further, the preprocessing of the data to be cleaned includes: using the Z-score method to unify the data to be cleaned to the same magnitude; The Z-score method is as follows:

[0007] Where, is the value after standardization, is the i-th original data, is the mean of the normalized data, is the standard deviation of the normalized data.

[0008] Further, the identification of discrete abnormal data includes using the Isolation Forest algorithm to identify discrete abnormal data; the identification of piled-up abnormal data includes using the clustering algorithm to identify piled-up abnormal data.

[0009] Further, using the Isolation Forest algorithm to identify discrete abnormal data includes: using the Isolation Forest algorithm to randomly cut the data space, and then randomly cut each of the two subspaces generated each time until each space is a data point or the tree has reached the maximum depth, calculating the observation score, and the observation score greater than 0.5 is an outlier.

[0010] Further, the observation score is as follows:

[0011] Where, is the mean of the lengths of each row of data in the dataset to be cleaned on each binary tree, is the average length of all binary trees.

[0012] Further, using the clustering algorithm to identify piled-up abnormal data includes the following steps: setting the neighborhood radius and the core point number threshold; randomly selecting k points from the dataset to be cleaned; judging whether each selected point is a core point, if so, marking the point and creating a point cluster; adding the points within the neighborhood of the core point to the point cluster, and judging whether the points within the neighborhood are core points; if the points within the neighborhood are core points, repeat the above steps until all points within the neighborhood radius are visited, and the points not added to the cluster are marked as noise points.

[0013] Further, the modeling of the Isolation Forest algorithm model includes the following steps: screening 80% of the data to be cleaned as the training set, and the rest as the test set; creating a random forest model; inputting the training set into the random forest model for training, and using the automatic hyperparameter optimization framework Optuna to search for the optimal hyperparameters of the random forest model.

[0014] Further, for the clustering algorithm model, the modeling includes the following steps: constructing a multi-layer long short-term memory network (LSTM); determining the loss function and optimizer; taking x consecutive sequences of data in the training set as a group and sequentially inputting them into the neural network model to search for the optimal hyperparameters of the multi-layer long short-term memory network (LSTM) using the automatic hyperparameter optimization framework Optuna.

[0015] Further, to establish the RF-LSTM prediction model, the following steps are included: initialization settings; calculating the initial individual fitness value of the fish school and recording the best state; randomly performing one of the three behaviors of foraging, schooling, and chasing by the individuals of the fish school to obtain a new fish school; evaluating all individuals of the new fish school and marking the optimal individual.

[0016] The present invention also provides an abnormal data cleaning system based on multi-model fusion, including: a preprocessing module: used for preprocessing the data to be cleaned; a first identification module: used for identifying discrete abnormal data; a second identification module: used for identifying piled-up abnormal data; an establishment module: used for establishing a random forest model and a long short-term neural network model; a combination module: used for combining the random forest model and the long short-term neural network model based on the artificial fish school algorithm to obtain an RF-LSTM prediction model; a cleaning module: used for inputting the data containing missing values into the RF-LSTM prediction model and outputting the prediction result to fill the data set to complete the cleaning of abnormal data.

[0017] Compared with the prior art, the present invention has the following beneficial technical effects: An abnormal data cleaning method based on multi-model fusion according to the present invention combines the isolation forest algorithm and the clustering algorithm, can effectively identify abnormal data with discrete distribution and aggregated distribution, broadens the scope of abnormal data identification, improves the accuracy and comprehensiveness of identification, and uses the RF-LSTM prediction model to predict and fill missing values, which can ensure the integrity and consistency of the data set, thereby further improving the accuracy of business decisions.

[0018] An abnormal data cleaning method based on multi-model fusion according to the present invention uses the isolation forest algorithm to identify discrete abnormal data, detects abnormalities by using the sparsity of data points, and can quickly identify abnormal data points. In addition, no assumptions need to be made about the distribution of the data, and various types of data distributions can be processed, including normal distribution, skewed distribution, multimodal distribution, etc.

[0019] An abnormal data cleaning method based on multi-model fusion according to the present invention uses the clustering algorithm to identify piled-up abnormal data. The data points are divided into multiple clusters, and the data points that do not belong to any cluster or are far from other clusters are identified as abnormal data points, which is suitable for identifying abnormal data with an aggregated distribution.

[0020] The present invention relates to an abnormal data cleaning method based on multi-model fusion. The isolation forest algorithm isolates abnormal points by randomly dividing the data space, while the clustering algorithm identifies abnormal points through the distance differences of data points. Combining the isolation forest algorithm with the clustering algorithm can mutually verify and complement each other.

[0021] The present invention relates to an abnormal data cleaning method based on multi-model fusion. The artificial fish swarm algorithm (AFSA) dynamically adjusts model parameters by simulating the behaviors of fish swarms (foraging, clustering, following), avoiding the model falling into local optima, making the RF-LSTM prediction model more flexible when dealing with complex data and capable of adapting to the proportion changes of different data features. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a flow chart of an abnormal data cleaning method based on multi-model fusion according to the present invention; Figure 2 It is a working schematic diagram of the long short-term neural network model in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0024] Embodiment 1 The present invention provides an abnormal data cleaning method based on multi-model fusion, as Figure 1 shown, including the following steps: preprocessing the data to be cleaned; identifying discrete abnormal data; identifying piled-up abnormal data; establishing a random forest model and a long short-term neural network model; combining the random forest model and the long short-term neural network model based on the artificial fish swarm algorithm to obtain an RF-LSTM prediction model; inputting the data with missing values into the RF-LSTM prediction model, and outputting the prediction result to fill the data set to complete the cleaning of abnormal data.

[0025] Data cleaning refers to the process of identifying and correcting errors, inconsistencies or missing data in a data set. The cleaned data can improve data quality and ensure the effectiveness and reliability of analysis results. The goal of data cleaning is to remove duplicate records, eliminate abnormal data, correct incorrect data, and ensure data consistency.

[0026] Preprocess the data to be cleaned, where the data to be cleaned includes abnormal data and missing data. Since the large difference in the value range of the data to be cleaned will lead to an imbalance in the input data, preprocessing is required. The Z-score method is used for preprocessing to unify the data to be cleaned to the same magnitude.

[0027]

[0028] Among them, is the value after standardization, is the i-th original data, is the data mean, is the standard deviation.

[0029] Identify discrete abnormal data and identify piled-up abnormal data, that is, identify outliers and modify the outliers to null values for subsequent operations. Specifically, use the Isolation Forest algorithm to identify discrete abnormal data and use the clustering algorithm to identify piled-up abnormal data.

[0030] Use the Isolation Forest algorithm to identify discrete abnormal data, including constructing multiple isolated binary trees with the Isolation Forest algorithm to achieve random cutting of the data space, and then randomly cutting the two subspaces generated each time until there is only one data point in each space or the tree has reached the maximum depth. Values with an observation score greater than 0.5 can be regarded as outliers. Among them, the observation score is:

[0031] is the mean of the lengths of each row of data in the dataset to be cleaned on each binary tree, is the average length of all binary trees.

[0032] Use the Isolation Forest algorithm to identify discrete abnormal data, which can detect anomalies by using the sparsity of data points and can quickly identify abnormal data points. In addition, no assumptions need to be made about the data distribution, and various types of data distributions can be processed, including normal distribution, skewed distribution, multimodal distribution, etc.

[0033] The Isolation Forest algorithm divides the data space by randomly selecting a feature and a feature value until each data point is isolated. Since abnormal data points usually have a different distribution from normal data points, they are more likely to be isolated. The Isolation Forest evaluates the abnormality of each data point by calculating its "isolation degree" (i.e., the number of divisions required to reach the isolated state).

[0034] Use a clustering algorithm to identify stacked abnormal data, including setting the neighborhood radius and the core point number threshold, then randomly select k points from the preprocessed data to be cleaned, and judge whether they are core points. The judgment criterion for core points is that the number of neighborhood objects is greater than the core point number threshold. If the condition is met, mark it as a core point and create a point cluster for it. Then, traversingly add the points in its neighborhood to the point cluster and judge whether the point is a core point. If it is judged as a core point, repeat the above operations for core points until all points are visited. After that, the points not added to the cluster are noise points.

[0035] Use a clustering algorithm to identify stacked abnormal data. The data points are divided into multiple clusters, and the data points that do not belong to any cluster or are far from other clusters are identified as abnormal data points. This is suitable for identifying abnormal data with a stacked distribution.

[0036] The Isolation Forest algorithm is suitable for identifying discrete abnormal data, and the clustering algorithm is suitable for identifying stacked abnormal data. Using the Isolation Forest algorithm first and then the clustering algorithm can more effectively complete abnormal identification.

[0037] Build a random forest model and a long short-term neural network model. As Figure 2 shown, it is a working schematic diagram of the long short-term neural network. Building a random forest model includes: screening 80% of the data to be cleaned as the training set and the remaining 20% as the test set; creating a random forest model; using the training set as the input and running hyperparameter optimization with Optuna. Try different combinations of the four hyperparameters of the number of decision trees, the maximum depth of the decision tree, the minimum number of samples in the leaf node, and the maximum number of randomly selected features for each tree according to the strategy of the TPE algorithm, and take SAUC as the adjustment target to make it increase as much as possible to form the final random forest model.

[0038] Among them, Optuna is an open-source framework for automated hyperparameter optimization designed for machine learning and deep learning tasks; the TPE algorithm is one of the core algorithms for hyperparameter optimization in the Optuna framework and is an efficient and adaptable optimization algorithm; SAUC is an evaluation index reflecting the accuracy of the model and has the advantage of being less affected by data distribution bias.

[0039] Building a long short-term neural network model includes: constructing a multi-layer long short-term memory network (LSTM), and adding a Dropout layer to each layer of the long short-term memory network (LSTM). In this embodiment, the Dropout layer is set to 0.2 to prevent overfitting. Using the training set as input and running hyperparameter optimization with Optuna, different combinations of six hyperparameters, namely the number of network layers, the number of hidden units in each layer, the learning rate, the optimizer, the batch size, and the number of training epochs, are tried according to the strategy of the TPE algorithm, and the mean squared error (MSE) is used as the adjustment target to minimize it as much as possible, thus forming a final multi-layer long short-term memory network (LSTM).

[0040] Since missing values may occur in any column of the dataset, and the weights of different data affected by time series and relevant influencing factors often vary, dynamic parameters need to be combined with two models during prediction. Therefore, based on the artificial fish swarm algorithm, a random forest model and a long short-term neural network model are combined to obtain an RF-LSTM prediction model.

[0041] Through the artificial fish swarm algorithm, the RF-LSTM prediction model can dynamically adjust the parameters of the random forest model and the long short-term neural network model to adapt to the weights of different data characteristics, and is more flexible and efficient in dealing with complex data. It not only improves the training efficiency of the RF-LSTM prediction model but also reduces the risk of overfitting.

[0042] It should be noted that there are two sources of missing values. One is the values that are missing during data collection, such as data collection device failures, network problems, or incomplete data records. The other is the outliers that occur during data collection, such as sensor failures, data transmission errors, or environmental interference.

[0043] The RF-LSTM prediction model fuses the random forest model and the long short-term neural network model through the following steps: initializing the settings with the output weights of the two models and the sequence length of the multi-layer long short-term memory network (LSTM) as the coordinates of the artificial fish; calculating the initial individual fitness values of the fish swarm and recording the best state; evaluating the individuals in the fish swarm to obtain a new fish swarm; evaluating all individuals in the new fish swarm and marking the optimal individual.

[0044] The initialization settings include the population size of the fish swarm, the initial positions of the fish swarm individuals, the visual field of the fish swarm, the step size of the fish swarm, the crowding factor of the fish swarm, and the number of cycles of the fish swarm.

[0045] In this embodiment, 20 artificial fish are randomly released, and their initial positions are recorded as Xi, the individual fitness values are Yi, the visual field of the fish swarm is 0.005, the step size of the fish swarm, that is, the distance of each movement, is 0.001, and the crowding factor δ of the fish swarm is taken as 0.2 to avoid the fish swarm being too concentrated. The initial positions need to be restricted by boundary conditions.

[0046] Evaluate each individual and select one of the foraging behavior, schooling behavior, and following behavior to execute, and update the position to obtain a new fish swarm. The foraging behavior enables the fish to explore a new solution space and find a better solution; the schooling behavior allows the fish to gather towards other excellent individuals and utilize the group wisdom to improve the search efficiency; the following behavior enables the fish to follow the current optimal individual and accelerate the convergence towards the optimal solution.

[0047] Specifically, foraging behavior: Randomly select a position within its field of vision, and determine whether the fitness value of this position is better than the current position. If it is better, move towards this position and end the behavior; otherwise, continue to find the next random position. If no better position is found after 5 consecutive times, move randomly one step; Schooling behavior: Calculate the central position of all artificial fish and its corresponding fitness value Yc. If the result of dividing this fitness value by the number n of artificial fish in the field of vision of the fish individual currently executing the behavior is less than the product of the crowding factor and the fitness value of the current position, move one step towards the central position; otherwise, perform the foraging behavior; Following behavior: Find the artificial fish with the best fitness value in the field of vision of the fish individual currently executing the behavior. If the result of dividing its fitness value Ym by the number n of artificial fish in the field of vision of the fish individual currently executing the behavior is less than the product of the crowding factor and the fitness value of the current position, move one step in this direction; otherwise, perform the foraging behavior.

[0048] The artificial fish swarm algorithm AFSA dynamically adjusts the model parameters by simulating the behaviors of the fish swarm (foraging, schooling, following), avoiding the model falling into local optima, making the RF-LSTM prediction model more flexible when dealing with complex data and capable of adapting to the proportion changes of different data features.

[0049] Evaluate all individuals of the new fish swarm and mark the optimal individual; when the fitness value of the optimal individual in the new fish swarm is less than 0.05 or the number of loops reaches half of the number of training samples, it is regarded as the optimization completed; otherwise, continue the loop. The optimal individual is the one with the highest fitness value in the current fish swarm, representing the current known optimal solution.

[0050] When the number of loops reaches half of the number of training samples, it is regarded as the optimization completed, in order to prevent meaningless iteration when already close to the optimal solution, thus wasting computing resources.

[0051] Input the data with missing values into the RF-LSTM prediction model, and output the prediction results to fill the dataset, completing the cleaning of abnormal data. Specifically: The prediction results will be used as reasonable estimates of the missing values, and then these predicted values are filled back into the original dataset, thus completing the cleaning work of abnormal data. It can not only effectively handle the missing data, but also to a certain extent retain the original distribution and characteristics of the data, providing a more complete and reliable data basis for subsequent data analysis and modeling.

[0052] In summary, data cleaning includes two aspects of work: identifying errors and correcting anomalies. The present invention automatically performs consistency processing on the data packets obtained from the control system, removes duplicate values and outliers, and then extracts data according to the user-specified test points and writes it into the training dataset of the neural network. Taking the control of thermal power units as an example, a prediction-based method is selected as the input from two perspectives: time series and related quantities, and anomaly detection and data repair are evaluated.

[0053] Embodiment 2 An abnormal data cleaning system based on multi-model fusion according to the present invention includes a preprocessing module, a first identification module, a second identification module, an establishment module, a combination module, and a cleaning module.

[0054] The preprocessing module is used to preprocess the data to be cleaned; the first identification module is used to identify discrete abnormal data; the second identification module is used to identify stacked abnormal data; the establishment module is used to establish a random forest model and a long short-term neural network model; the combination module is used to combine the random forest model and the long short-term neural network model based on the artificial fish swarm algorithm to obtain an RF-LSTM prediction model; the cleaning module is used to input the data with missing values into the RF-LSTM prediction model, output the prediction results and fill them into the dataset to complete the cleaning of abnormal data.

[0055] An abnormal data cleaning system based on multi-model fusion provided by the present invention can implement method steps consistent with the above method, so details are not repeated here.

[0056] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.

Claims

1. An abnormal data cleaning method based on multi-model fusion, characterized in that: The following steps are involved: Preprocess the data to be cleaned; Identify discrete outlier data; Identify abnormal data accumulation; Build random forest models and long- and short-term neural network models; The random forest model and the long-short term neural network model are combined based on the artificial fish swarm algorithm to obtain the RF-LSTM prediction model; The missing value data is input into the RF-LSTM prediction model, and the output prediction results are filled into the data set to complete the abnormal data cleaning.

2. The abnormal data cleaning method based on multi-model fusion according to claim 1 is characterized in that: The preprocessing of the data to be cleaned includes: Use the Z-score method to unify the data to be cleaned to the same level; The Z-score method is: in, is the standardized value, is the i-th original data, is the mean of the normalized data, is the standard deviation of the normalized data.

3. The abnormal data cleaning method based on multi-model fusion according to claim 1 is characterized in that: The identifying of discrete abnormal data includes using an isolation forest algorithm to identify discrete abnormal data; The identifying of the accumulated abnormal data includes using a clustering algorithm to identify the accumulated abnormal data.

4. The abnormal data cleaning method based on multi-model fusion according to claim 3 is characterized in that: Use the Isolation Forest algorithm to identify discrete anomaly data, including: The isolation forest algorithm is used to randomly cut the data space, and the two subspaces generated each time are randomly cut again until each space is a data point or the tree has reached the maximum depth. The observation score is calculated, and an observation score greater than 0.5 is considered an outlier.

5. The abnormal data cleaning method based on multi-model fusion according to claim 4 is characterized in that: The observed score for: in, is the mean length of each row of the data set to be cleaned on each binary tree, is the average length of all binary trees.

6. The abnormal data cleaning method based on multi-model fusion according to claim 3 is characterized in that: The method of using a clustering algorithm to identify accumulated abnormal data includes the following steps: Set the neighborhood radius and core point threshold; Randomly select k points from the data set to be cleaned; Determine whether each selected point is a core point, if so, mark the point and create a point cluster; Add each point in the neighborhood of the core point to the point cluster, and determine whether the point in the neighborhood is a core point; If the point in the neighborhood is a core point, the above steps are repeated until all points within the neighborhood radius are visited, and the points not added to the cluster are marked as noise points.

7. The abnormal data cleaning method based on multi-model fusion according to claim 1 is characterized in that: The random forest model, modeling includes the following steps: 80% of the data to be cleaned is selected as the training set, and the rest is the test set; Create a random forest model; The training set is input into the random forest model for training, and the automatic hyperparameter optimization framework Optuna is used to search for the optimal hyperparameters of the random forest model.

8. The abnormal data cleaning method based on multi-model fusion according to claim 7 is characterized in that: The long-term and short-term neural network model includes the following steps: Build a multi-layer long short-term memory network LSTM; Determine the loss function and optimizer; The data in the training set are grouped as x continuous sequences and input into the neural network model in sequence, and the automatic hyperparameter optimization framework Optuna is used to search for the optimal hyperparameters of the multi-layer long short-term memory network LSTM.

9. The abnormal data cleaning method based on multi-model fusion according to claim 1 is characterized in that: Establishing the RF-LSTM prediction model includes the following steps: Initialization settings; Calculate the initial individual fitness value of the fish school and record the best state; Individual fish swarms randomly perform one of the three behaviors of foraging, grouping, and chasing to obtain new swarms; Evaluate all individuals in the new fish swarm and mark the best individuals.

10. An abnormal data cleaning system based on multi-model fusion, comprising: Preprocessing module: used to preprocess the data to be cleaned; The first recognition module: used to identify discrete abnormal data; The second recognition module is used to identify the accumulated abnormal data; Building module: used to build random forest models and long- and short-term neural network models; Combination module: used to combine the random forest model and the long-short term neural network model based on the artificial fish swarm algorithm to obtain the RF-LSTM prediction model; Cleaning module: used to input missing value data into the RF-LSTM prediction model, output prediction results to fill the data set, and complete abnormal data cleaning.