Data review method based on distributed computing
Through the distributed computing-based data review method, traditional data review methods are solved, which consumes time, has many human errors, and is unable to cope with big data and high costs, and a fast, accurate and automated data review process is achieved.
Patent Information
- Application Number
- CN202411378063.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional data review methods are time-consuming, prone to human errors, unable to cope with large data scale and cost, and lack flexibility and automation.
Using a distributed computing-based data review method, data preprocessing, feature extraction, classification model construction and real-time monitoring are carried out through distributed storage and machine learning frameworks, and automated review and report generation are realized.
Significantly reduces data review time and labor costs, improves the accuracy and consistency of the review process, enables rapid processing of large-scale data, and supports real-time or near-real-time data review.
Smart Images

Figure CN120104598A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of distributed computing technology, and more specifically, to a data review method based on distributed computing. Background Art
[0002] Data review plays a vital role in today's information society. With the rapid growth of data volume and the increasing maturity of data analysis technology, the quality, accuracy and reliability of data have become increasingly critical for various decision-making, business operations and scientific research. As an important checkpoint to ensure data quality, data review has profound background and reasons behind it.
[0003] First of all, data review stems from the uncertainty and potential risks of the data itself. Data may be affected by various factors during collection, transmission, storage and processing, such as human error, equipment failure, software vulnerabilities, etc. These factors may cause data distortion, missing or abnormal. Therefore, strict review of data is a necessary means to ensure data quality and reliability.
[0004] Secondly, data review also depends on data-driven decision-making and business needs. With the popularization of big data and artificial intelligence technologies, more and more companies and organizations have begun to rely on data for decision-making and business operations. In this context, data review has become an important guarantee to ensure the correctness of decision-making and business continuity. Through a comprehensive review of the data, problems and anomalies in the data can be discovered, and corresponding measures can be taken in a timely manner to correct and improve them, thereby improving the accuracy of decision-making and business efficiency.
[0005] In addition, data review also involves the need for data security and privacy protection. In the information society, data security issues are becoming increasingly prominent, and data leakage, abuse and tampering may damage the interests of individuals and organizations. As a means of ensuring data security and privacy protection, data review can timely discover security vulnerabilities and privacy leakage risks in data, take corresponding measures to prevent and deal with them, and protect the data security and privacy rights of individuals and organizations.
[0006] The above disclosed technical solutions have at least the following technical problems: Time-consuming: Traditional data review usually requires a lot of manpower and time to manually clean, verify and analyze data. This leads to a long review cycle, which is not conducive to real-time decision-making and rapid response of data applications.
[0007] Prone to human errors: Since most of the review process relies on manual operations and judgment, it is easy to be affected by personal subjective factors, resulting in inconsistency or errors in the review results. This risk is more significant, especially in the case of large-scale data sets.
[0008] Unable to cope with the scale of big data: As the amount of data increases and the types of data diversify, it is difficult for traditional methods to effectively handle big data review needs. Manual review may not be able to meet the requirements of real-time processing and analysis of large-scale data.
[0009] High cost: Human resource costs are a significant source of cost for traditional data review, especially for large enterprises or complex data environments, which require a lot of cost and time to maintain data quality and reliability.
[0010] Lack of flexibility and automation: Traditional methods often lack automated tools and technical support, and cannot achieve automated processes and real-time monitoring of data review, which limits the flexibility and efficiency of data review. In view of the above problems, the present invention proposes a solution. Summary of the invention
[0011] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a data review method based on distributed computing, which solves the problems raised in the above-mentioned background technology through distributed computing.
[0012] To achieve the above object, the present invention provides the following technical solutions: A data review method based on distributed computing, comprising: obtaining data to be reviewed, storing the data slices in a distributed storage system and preprocessing and extracting features of the data stored in the distributed storage system; sorting the feature data in chronological order to form a procurement time series data set and normalizing it; constructing an optimal classification model for the procurement time series data set after normalization in the data review process based on a distributed machine learning framework and a convolutional neural network model; introducing a real-time data stream monitoring system to continuously monitor data changes and the performance of the optimal classification model, and timely adjust the review strategy; automatically generating a review report, and presenting the review results through a visualization tool.
[0013] A data review method based on distributed computing, characterized in that it includes the following steps: obtaining data to be reviewed, storing the data slices in a distributed storage system and preprocessing and extracting features of the data stored in the distributed storage system; sorting the feature data in chronological order to form a procurement time series data set and normalizing it; constructing an optimal classification model for the procurement time series data set after normalization in the data review process based on a distributed machine learning framework and a convolutional neural network model; introducing a real-time data flow monitoring system to continuously monitor the changes in the data to be reviewed and the performance of the optimal classification model, and timely adjust the review strategy; automatically generating a review report, and presenting the review results through a visualization tool.
[0014] In a preferred embodiment, the data to be reviewed includes a database, a data warehouse, and a log file; and the preprocessing operations include data cleaning, deduplication, and format unification.
[0015] In a preferred embodiment, the feature extraction generates features based on statistical information of the data, and performs feature analysis on the features to obtain feature data; the feature data includes material purchase quantity, material storage quantity, and material consumption quota.
[0016] In a preferred embodiment, the specific process of the normalization processing is as follows: the feature data are arranged in ascending order according to the time field, the mean and standard deviation of each feature data are calculated, the standardized calculation formula is combined, and the sliding window superposition tomography method is used to make all feature data have the same scale and range of variation.
[0017] In a preferred embodiment, the specific process of training the distributed machine learning framework is as follows: randomly divide the preprocessed procurement time series data into several batches; use data parallelism to assign the data of each batch to different computing nodes for processing; initialize the parameters of the optimal classification model, and apply the initialized optimal classification model to each computing node; the computing node uses a synchronous stochastic gradient descent distributed optimization algorithm to iteratively optimize the local gradient parameters; after each iterative step, the computing node aggregates the local gradient parameters to the parameter server and updates the classification model parameters. After the iteration, the optimal classification model is obtained.
[0018] In a preferred embodiment, the specific steps of introducing the real-time data stream monitoring system are as follows: using the real-time data stream monitoring tool Apache Kafka to process large amounts of data and analyze and process it in real time; establishing a data stream pipeline to import the data generated by the optimal classification model into the monitoring system in real time; implementing real-time data analysis in the monitoring system, and using a machine learning model to analyze the changes in the generated data and the model performance; establishing a continuous optimization process and regularly reviewing the performance and effectiveness of the monitoring system.
[0019] In a preferred embodiment, the visualization tools include Python, Tableau and a bulletin board.
[0020] In a preferred embodiment, the sliding window stacking tomography method is specifically as follows: a window function JH of a fixed length L is selected, and the procurement time series data set is divided with a fixed step length l to obtain data segments of fixed length.
[0021] The technical effects and advantages of the data review method based on distributed computing of the present invention are as follows: 1. The present invention significantly reduces the time and labor costs required for data review through the optimized configuration of distributed storage and computing resources. Compared with traditional methods, it can quickly process large amounts of data and support real-time or near real-time data review needs, thereby promoting the rapid response capabilities of real-time decision-making and data applications.
[0022] 2. By introducing automated tools and algorithms, the distributed computing system can reduce manual intervention and subjective judgment, thereby reducing the risk of inconsistency or errors in review results caused by human error. The system can automatically perform data cleaning, verification and analysis according to pre-set rules and algorithms, improving the accuracy and consistency of the review process. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 This is a structural schematic diagram of a data review method based on distributed computing in the present invention. DETAILED DESCRIPTION
[0024] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0025] Embodiment 1, Figure 1 The present invention provides a data review method based on distributed computing, comprising the following steps: Step S1, obtaining data to be reviewed, storing the data in a distributed storage system, and performing preprocessing and feature extraction on the data stored in the distributed storage system; Specifically, data includes databases, data warehouses, and log files; preprocessing includes data cleaning, deduplication, and formatting conversion; preprocessing not only improves the overall quality and availability of data, but also provides a more reliable and effective data foundation for subsequent data analysis and machine learning models, thereby helping to improve analysis efficiency and model performance.
[0026] The feature extraction is to extract relevant information that can represent the data characteristics from the original data for subsequent analysis and modeling. Features are generated based on the statistical information of the data, and feature analysis is performed on these data to obtain feature data; the feature data includes material procurement volume, material storage volume, and material consumption quota.
[0027] When the system is collecting data, failures in metering equipment or transmission equipment can lead to missing data. Data missing is unavoidable and often occurs, and is an important factor causing incomplete data. For missing data, there are usually two processing methods: filtering and filling. If the missing data is not serious, the third Lagrange interpolation method is used to fill it. If the missing data is serious, the data is filtered out.
[0028] Missing value calculation helps improve the completeness of data. Missing values may affect the accuracy of modeling, so filling or processing these missing values can make the data more complete and improve the stability of the model. For machine learning or modeling tasks, a complete data set is a prerequisite for training accurate models. By filling or processing missing values, model deviations caused by missing data can be reduced and the accuracy and stability of the model can be improved.
[0029] The specific method of obtaining missing values is as follows: First, we identify missing values in the data through data visualization, and select appropriate missing value filling methods according to the distribution and characteristics of the data, including mean filling, median filling, mode filling, and Lagrange interpolation. For each feature column, we use the selected method to fill the corresponding missing values. The available Lagrange interpolation method is used to calculate missing values. The expression of the Lagrange basic polynomial is: It should be noted that i is not equal to j and P is a missing value.
[0030] The benefits and functions of storing the data shards in a distributed storage system include: Scalability and fault tolerance: Distributed storage systems can handle large-scale data and have good horizontal scalability. No matter how fast the amount of data grows, the system can expand storage and processing capabilities by adding more nodes. In addition, distributed storage systems usually have fault tolerance mechanisms that can ensure data reliability and persistence in the event of node failures or network problems.
[0031] High-performance access: Data is stored in shards in a distributed system, and the parallel processing capabilities of multiple nodes can be used to accelerate data access and read and write operations. This parallelism can significantly improve the efficiency of data processing and analysis, especially in the case of large data volumes and high concurrent access.
[0032] Flexibility and load balancing: Distributed storage systems allow data to be sharded and distributed as needed, so they can be flexibly configured and optimized according to the needs of the application. Through dynamic load balancing, the system can automatically adjust data distribution to ensure load balancing on each node, thereby improving the efficiency and stability of the overall system.
[0033] Geographic location awareness and data localization: Some distributed storage systems support geographic location awareness and localized storage of data. This means that data can be stored on nodes closer to users, reducing data access latency and improving response speed and user experience.
[0034] Security and data backup: Distributed storage systems usually provide multi-level data backup and security mechanisms to ensure that data is not lost due to the failure of a single node or device. Data can be stored in multiple locations through replication or distributed redundancy to improve data reliability and security.
[0035] Step S2, sorting the characteristic data in chronological order to form a procurement time series data set and performing normalization processing; Sort the feature data in ascending order according to the time field. Time order sorting is the basis for ensuring that each sample in the time series data set is arranged in ascending order of time; normalize the sorted feature data. Normalization can unify the value range of different feature data and avoid some features having too much impact on model training due to their large value range. Combine the sorted feature data into a time series data set and normalize them. Each sample should contain a set of feature data, which are arranged in time order. Each time series data sample corresponds to a snapshot of feature data at a time point.
[0036] The normalization processing formula is:
[0037] in, is the characteristic data in the procurement time series data set, is the mean of the feature data in the procurement time series data set, FC is the variance of the feature data in the procurement time series data set, and i is a positive integer.
[0038] In order to better learn the characteristics of time series data, long sequences with time tags must be constructed into segments of a certain length. Therefore, the sliding window stacking tomography method is used to ensure the correlation of time series data. This method selects a window function JH with a fixed length L and divides the original time series data set with a fixed step length l to obtain data segments of a certain length. The calculation formula of the window function is:
[0039]
[0040] Among them, JH is the total set of the sliced procurement time series data set, is a dataset of time series fragments, is the first data in the data set. For each sample, the signal difference is the jth time series segment after slicing, and N is the total amount of data in the purchased data set.
[0041] Step S3, constructing an optimal classification model in a data review process based on the procurement time series data set based on a distributed machine learning framework and a convolutional neural network model; Divide the dataset into appropriate parts so that it can be distributed to multiple computing nodes in a distributed environment. These parts can be divided into time periods, with each time period corresponding to a training batch. Select and configure the distributed machine learning framework and environment Apache Spark to ensure that it can support the training of large-scale data and models. Set up and configure a distributed computing cluster, including master nodes and worker nodes. Ensure that the computing resources of the cluster are sufficient to support distributed training tasks. Ensure that the nodes in the cluster can communicate and synchronize data efficiently to maintain synchronization and consistency during the training process.
[0042] The specific steps of distributed training are as follows: The preprocessed procurement time series data is divided into multiple batches; the data of each batch is assigned to different computing nodes for processing using a data parallel method; the same optimal classification model parameters are initialized on each computing node; each computing node independently calculates the local gradient parameters using a synchronous stochastic gradient descent distributed optimization algorithm; after each iteration step, the computing node aggregates the local gradient parameters to the parameter server and updates the optimal classification model parameters; the iteration process is repeated to calculate and update the local gradient parameters based on the current optimal classification model parameters until the preset training round or convergence condition is reached.
[0043] In convolutional neural networks, the hidden layer is the key to feature extraction and feature mapping, including two key operations: convolution and pooling. Convolution is a mathematical operator that uses two functions to produce a third function. In this layer, convolution represents the calculation method between the input signal m and the Gaussian convolution kernel g. The calculation formula is:
[0044] Among them, e is the size of the signal, and a is the width parameter of the function, which takes a value of 1.
[0045] After convolution processing, the feature vector dimension is large. Combining these features, the maximum set method is used to reduce the feature map to the minimum under the condition of minimum window length. Assume that there is a convolution layer The size is The convolution kernel can be expressed as (i=1,2,…, ), then the data obtained from the input layer is convolved through the convolution layer to produce The size is The eigenvectors of are:
[0046] Among them, JH is the time sequence segment, is the bias, conv() is the convolution function, and ReLU() is the activation function, specifically f(x)=max(0,x).
[0047] The step size of the pooling port of the pooling layer is designed to be 2s, and the size is , then the feature vector obtained by convolution is After pooling, we get The window length is The eigenvectors of are:
[0048] in, is the shared weight of this layer, is the bias of this layer, and down() is the downsampling function.
[0049] The pooled feature vector is repeatedly pooled to obtain a new feature vector. (q=1,2,…,n). For the eigenvector (q=1,2,…,n) for raster processing, Get a one-dimensional vector =(m=1,2,…,n), which is the input of the fully connected layer.
[0050] The fifth layer of the convolutional neural network is a fully connected layer, which maps one set of vectors to another set of vectors. As the input of the 5th layer, its features are mapped, and the number of neurons in its hidden layer is r. (r=1,2,…,n), the output of this layer can be expressed as:
[0051] in, is the preset hidden layer weight, is the hidden layer bias, tanh() is the activation function After completing the convolution and pooling of the data, the result should be transmitted as input to the classifier to implement logistic regression. The output category is 1, which means that the procurement data is compliant; the output category is 2, which means that the probability of non-compliant procurement data is P. The category with a large probability is taken as the result of the convolutional neural network classification, and the formula is:
[0052] in is the total number of data samples.
[0053] Use the distributed optimization algorithm synchronous stochastic gradient descent (SGD) to aggregate and update the parameters of the convolutional neural network model, obtain the optimal weights, and establish the optimal classification model for time series data. The specific formula is:
[0054] in, is the input, W is the convolution kernel weight, b is the bias, and m is the activation function.
[0055] Using a distributed machine learning framework to train models has the following benefits and functions: Accelerate training speed: Distributed machine learning frameworks can distribute data and computing tasks to multiple computing nodes or multiple machines for processing. This parallel processing method can greatly reduce training time, especially for large-scale data sets and complex models.
[0056] Processing large-scale data: When processing large-scale datasets, a single machine may face limitations in memory and computing power. Distributed machine learning frameworks allow data to be sharded and processed in parallel on multiple nodes, making it possible to process large datasets that are beyond the capabilities of a single machine.
[0057] Improve the generalization ability of the model: Using a distributed framework for training can increase the model's ability to learn and generalize data. By using more data and more complex models, patterns and relationships in the data can be better captured, thereby improving the model's prediction accuracy.
[0058] Flexibility and scalability: Distributed machine learning frameworks are usually designed as scalable and flexible systems that can dynamically add or remove computing nodes based on demand. This flexibility allows the framework to adapt to problems of different sizes and complexities while making efficient use of existing computing resources.
[0059] Fault tolerance and reliability: Distributed training frameworks are usually fault-tolerant, which means that even if some nodes or machines fail, the training process and results can still be reliable. This feature is especially important in long-running training tasks.
[0060] Support for multiple algorithms and models: Distributed machine learning frameworks usually support the training of multiple machine learning algorithms and models, including deep learning models, large-scale linear models, integrated learning, etc. This makes it possible to handle different types of problems and data under the same framework.
[0061] Step S4: Introduce a real-time data flow monitoring system to continuously monitor changes in new data and model performance. Through the real-time feedback mechanism, timely adjust the model or review strategy to respond to data changes and new review requirements.
[0062] Choose the right monitoring tools and technologies: Use Apache Kafka, a real-time data stream monitoring tool. It can handle large amounts of data and support real-time processing and analysis. Define monitoring indicators and thresholds: Determine the accuracy, recall, response time, etc. of the model. Set reasonable thresholds to determine when feedback and adjustment mechanisms need to be triggered. Establish a data flow pipeline: Establish a data flow pipeline to import data generated in the production environment into the monitoring system in real time. This may involve data collection, transmission, transformation, and loading (ETL process) to ensure that data can enter the monitoring system on time and in full. Implement real-time analysis and feedback mechanisms: Implement real-time data analysis in the monitoring system, and use machine learning models or rule engines to analyze changes in new data and model performance. If performance degradation or changes in the distribution of new data are detected, the system can automatically trigger an alarm or notify relevant responsible personnel. Automated adjustment and optimization: Design automated adjustment mechanisms to respond to feedback from the monitoring system. This can include automatically adjusting model parameters, updating data preprocessing processes, modifying review strategies, or reminding manual intervention to ensure that the system can maintain efficiency and accuracy in a changing environment. Continuous optimization and monitoring: Establish a continuous optimization process and regularly review the performance and effectiveness of the monitoring system. Based on feedback and data trends, adjust monitoring indicators, thresholds, and system workflows to maintain the effectiveness and reliability of the system in a changing environment.
[0063] Step S5, automatically generate an audit report and present the audit results through a visualization tool.
[0064] The benefits and functions of automatically generating audit reports and presenting audit results through visualization tools are as follows: Improve efficiency and accuracy: Automated generation of audit reports can greatly reduce the time and labor costs of manual operations. Through automation, the system can extract and organize information from the data and generate reports with consistent and accurate formats, avoiding errors and omissions that may occur in manual processing.
[0065] Timeliness and real-time updates: Visualization tools can display review results in real time, allowing managers to keep up to date with the latest data and trends. This immediacy is especially important for making decisions quickly, especially when strategies need to be adjusted or problems need to be solved quickly.
[0066] Provide deeper analysis and insights: Visualization tools can display the review results in the form of charts, graphs and dynamic data dashboards, making complex data easier to understand and analyze. Managers can explore the data in depth through interactive interfaces to discover potential correlations and trends, so as to formulate strategies and make decisions more accurately.
[0067] Support decision making and strategic planning: Through the analysis results provided by automatically generated review reports and visualization tools, managers can make decisions based on objective data rather than relying solely on subjective judgment or experience. This data-driven approach can help improve the quality and accuracy of decision making.
[0068] Promote team collaboration and communication: Visualization tools can present complex data and analysis results in a concise and clear way, making it easier for team members to understand and share information. This transparent and easy-to-understand communication method helps teams reach consensus on common goals and work together to solve problems or achieve goals.
[0069] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.
[0070] The above embodiments may be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented by software, the above embodiments may be implemented in whole or in part in the form of a computer program product.
[0071] Those of ordinary skill in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0072] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0073] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0074] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A data review method based on distributed computing, characterized in that: The steps include: Acquire the data to be reviewed, store the data in fragments in a distributed storage system, and perform preprocessing and feature extraction on the data stored in the distributed storage system; After the characteristic data are sorted in chronological order, a procurement time series data set is formed and normalized; The normalized procurement time series data set is used to build the optimal classification model in the data review process based on the distributed machine learning framework and convolutional neural network model; Introduce a real-time data flow monitoring system to continuously monitor changes in the data to be reviewed and the performance of the optimal classification model, and adjust the review strategy in a timely manner; Automatically generate review reports and present review results through visualization tools.
2. The data review method based on distributed computing according to claim 1, characterized in that: The data to be reviewed includes databases, data warehouses, and log files; preprocessing operations include data cleaning, deduplication, and format unification.
3. The data review method based on distributed computing according to claim 2 is characterized in that: The feature extraction generates features based on statistical information of the data, and performs feature analysis on the features to obtain feature data; the feature data includes material purchase quantity, material storage quantity, and material consumption quota.
4. The data review method based on distributed computing according to claim 3 is characterized in that: The specific process of the normalization process is as follows: The characteristic data are arranged in ascending order according to the time field, the mean and standard deviation of each characteristic data are calculated, and the standardized calculation formula is combined, and the sliding window stacking tomography method is used to make all the characteristic data have the same scale and variation range.
5. The data review method based on distributed computing according to claim 4 is characterized in that: The specific process of the distributed machine learning framework training is as follows: The pre-processed procurement time series data is randomly divided into several batches; Use data parallelism to assign each batch of data to different computing nodes for processing; Initialize the parameters of the optimal classification model, and apply the initialized optimal classification model to each computing node; The computing nodes use the synchronous stochastic gradient descent distributed optimization algorithm to iteratively optimize the local gradient parameters; After each iteration step, the computing node aggregates the local gradient parameters to the parameter server and updates the classification model parameters. After the iteration, the optimal classification model is obtained.
6. The data review method based on distributed computing according to claim 5, characterized in that: The specific steps of introducing the real-time data flow monitoring system are as follows: Use Apache Kafka, a real-time data stream monitoring tool, to process large amounts of data and perform real-time analysis and processing; Establish a data flow pipeline to import the data generated by the optimal classification model into the monitoring system in real time; Implementing real-time data analysis in the monitoring system, using machine learning models to analyze changes in the generated data and model performance; Establish a continuous optimization process and regularly review the performance and effectiveness of the monitoring system.
7. The data review method based on distributed computing according to claim 6 is characterized in that: The visualization tools include Python, Tableau, and dashboards.
8. A data review method based on distributed computing according to claim 7, characterized in that: The standardized calculation formula is: in, is the characteristic data in the procurement time series data set, is the mean of the feature data in the procurement time series data set, FC is the variance of the feature data in the procurement time series data set, and i is a positive integer.
9. The data review method based on distributed computing according to claim 8, characterized in that: The sliding window stacking chromatography method is specifically: Select a window function JH with a fixed length L, divide the procurement time series data set with a fixed step length l, and obtain data segments of fixed length. The specific calculation formula of the window function is: Among them, JH is the total set of the sliced procurement time series data set, is a dataset of time series fragments, is the first data in the data set. For each sample, the signal difference is the jth time series segment after slicing, and N is the total amount of data in the purchased data set.
10. The data review method based on distributed computing according to claim 9, characterized in that: The specific calculation formula of the optimal classification model is as follows: in, is the input, W is the preset convolution kernel weight, b is the preset bias, and m is the activation function.