ETL system performance bottleneck prediction and analysis device based on multivariable time series
By constructing a bottleneck prediction and abnormal detection model for multivariate time series, the prediction and detection lag of performance bottlenecks in ETL systems is solved, and the active prediction and real-time detection of ETL systems are realized, task scheduling is optimized, and system performance is improved.
Patent Information
- Application Number
- CN202510298085.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-18
Smart Images

Figure CN120336409A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing, and particularly to an ETL (Extract, Transform, Load) system performance bottleneck prediction and analysis device based on multivariate time series. Background Art
[0002] With the advent of the big data era, the ETL technology plays a crucial role in enterprise data management and analysis. However, ETL tasks often face performance bottleneck problems during execution, especially when resources are limited, and the performance bottleneck problems are particularly prominent when multiple tasks are executed in parallel. The traditional operation and maintenance methods are inefficient and lagging, and cannot cope with performance bottleneck problems in a timely manner. Most of the existing prediction models are for specific fields and are difficult to be directly applied to the ETL scenario. Therefore, there is an urgent need for a technical solution that can effectively predict and analyze the performance bottleneck of the ETL system. Summary of the Invention
[0003] The object of the present invention is to provide an ETL system performance bottleneck prediction and analysis device based on multivariate time series, which realizes the active prediction and real-time detection of the ETL system performance bottleneck by constructing a prediction model and an anomaly detection model, and provides a bottleneck location technology to help operation and maintenance personnel quickly locate and solve performance bottleneck problems.
[0004] (I) Technical Solution 1. Construction of bottleneck prediction model: Dataset generation: Design an ETL task test set containing various load types, and collect key performance indicator data during task execution.
[0005] Feature engineering: Extract task features and time series features, and perform data preprocessing, including standardization, dimensionality reduction, and oversampling processing.
[0006] Model training: Use traditional machine learning algorithms (such as XGBoost) to construct a bottleneck prediction model, and perform parameter tuning through Bayesian optimization.
[0007] Model application: Apply the prediction model to the ETL task scheduling, optimize the task scheduling strategy by predicting the performance bottleneck during task execution, and reduce the occurrence of bottlenecks.
[0008] 2. Construction of bottleneck detection model: EDD model design: Propose a multivariate time series anomaly detection model called EDD (Encoder-Decoder-Discriminator), which combines a graph attention network (GAT) and a long short-term memory network (LSTM) to extract spatial features and time features respectively.
[0009] Model training: By using a small amount of abnormal data and jointly training the encoder, decoder, and discriminator, the model's ability to recognize anomalies is enhanced.
[0010] Model inference: By means of reconstruction error and anomaly score, performance bottlenecks in the ETL system are detected in real time.
[0011] 3. Bottleneck localization technology: Bottleneck resource localization: Based on the reconstruction error output by the EDD model, the types of bottleneck resources (such as CPU, memory, disk I / O, network I / O) are located.
[0012] Bottleneck task localization: According to the correlation between tasks and resource consumption, the bottleneck contribution degree of tasks is calculated. For CPU and memory bottlenecks, the bottleneck contribution degree formula for each task is:
[0013] In the formula represents the number of components of the task; represents the consumption of the bottleneck type resources corresponding to the component. C represents the bottleneck contribution degree. For disk I / O and network I / O bottlenecks, the consumption of these two types of resources is closely related to the data volume, and the bottleneck contribution degree formula for each task is:
[0014] In the above formula represents the number of disk I / O or network I / O components in the task; represents the size of each record processed by the component; represents the number of data rows processed by the component; represents the running time of the component; C represents the bottleneck contribution degree.
[0015] Then, the specific bottleneck tasks are located according to the bottleneck contribution degree.
[0016] (2) Technical effects The present invention is an ETL system performance bottleneck prediction and analysis device based on multivariate time series, which can achieve the following technical effects: Proactively predict performance bottlenecks: By constructing a bottleneck prediction model, it is possible to predict in advance the possible performance bottlenecks that may occur during the execution of ETL tasks, optimize the task scheduling strategy, and reduce the occurrence of bottlenecks.
[0017] Detect performance bottlenecks in real time: Through the EDD model, it is possible to detect performance bottlenecks in the ETL system in real time, and promptly discover and handle anomalies.
[0018] Accurately locate bottleneck resources and tasks: Through the bottleneck localization technology, it is possible to quickly locate bottleneck resources and tasks, helping operation and maintenance personnel to solve problems targeted and improve system performance. Description of the drawings
[0019] Figure 1 It is the overall flowchart for ETL performance bottleneck prediction and analysis.
[0020] Figure 2 It is the structure diagram of the EDD model.
[0021] Figure 3 It is the flowchart for bottleneck location.
[0022] Figure 4 It is the pseudo-code diagram for embodiment model training.
[0023] Detailed description Figure 1 : The input features enter the early warning model, and the classification results are finally output.
[0024] Detailed description Figure 2 : The diagram shows the composition of the encoder, decoder, and discriminator and their data flow. The EDD model structure includes three major parts: the encoder, discriminator, and decoder, as well as the details in each part.
[0025] Detailed description Figure 3 : It shows the complete process from anomaly detection to bottleneck resource and task location. The data is input into the anomaly detection model for detection. After reconstructing the error, the bottleneck resources are located. The bottleneck resources include CPU resource bottleneck, disk I / O resource bottleneck, memory resource bottleneck, and network I / O resource bottleneck. After finding the bottleneck resources, the bottleneck tasks are found according to the resources.
[0026] Detailed description Figure 4 : It shows the process logic of embodiment model training in detail, described using pseudo-code. Specific implementation manner
[0027] As Figure 1 shown, in the present invention, the features are input and enter the prediction model, and finally the classification results are input. The specific implementation steps are as follows: Dataset generation: Design an ETL task test set containing various load types, and collect key performance index data during task execution, such as CPU utilization, memory occupancy, disk I / O, network I / O, etc.
[0028] Feature engineering: Extract task features (such as the number of data processing rows, the number of components) and time series features (such as CPU utilization, memory occupancy, etc.), and perform data preprocessing, including normalization, dimensionality reduction, and oversampling processing.
[0029] Model training: Use the XGBoost algorithm to build a bottleneck prediction model, perform parameter tuning through Bayesian optimization, and train the model to predict the performance bottleneck during task execution.
[0030] Model application: Apply the prediction model to ETL task scheduling, optimize the task scheduling strategy and reduce the occurrence of bottlenecks by predicting the performance bottlenecks during task execution.
[0031] Model training: Using a small amount of abnormal data, the model's ability to identify anomalies is enhanced through joint training of the encoder, decoder, and discriminator.
[0032] Model Reasoning: Detect performance bottlenecks in ETL systems in real time by reconstructing error and anomaly scores.
[0033] Bottleneck location: Based on the reconstruction error output by the EDD model, the bottleneck resource type is located, and according to the correlation between the task and resource consumption, the bottleneck contribution of the task is calculated to locate the specific bottleneck task.
[0034] As an embodiment of the present invention, taking EDD model training as an example, Figure 4 To explain in detail, first, The normal data at the moment and a piece of abnormal data randomly selected from the abnormal data set are input into the encoder to obtain their distribution in the latent space and (Line 5), then based on these distributions, the corresponding sampled data is generated using resampling technology (Line 6). Subsequently, these sampled data are fed into the discriminator to obtain the discrimination results and (Line 7). At the same time, the sampled data of the normal data in the latent space is passed to the decoder to reconstruct the original data and generate the reconstructed output (Line 8). To evaluate the model performance, we use the formula Compute the normal distribution and predefined distribution of encoder output The KL divergence between , to quantify the similarity between the two (line 9). In addition, using the formula Calculate the reconstruction error of normal data , evaluate the accuracy of data reconstruction (line 10). In terms of classification performance, according to the formula and formula Calculate the binary cross entropy loss for normal data and abnormal data separately and , these two loss values reflect the performance of the model in anomaly detection (line 11). Finally, these loss terms are integrated and the model parameters are iteratively updated through the gradient descent algorithm (lines 12 and 13).
[0035] As an embodiment of the present invention, taking the search for the ETL data import service resource bottleneck on the server as an example, after deploying this device, when the import service imports data, this device collects CPU utilization rate, memory occupancy, and disk I / O metric data, applies the trained EDD model, and views the output results. In this embodiment, it is found that the bottleneck resource lies in the disk I / O resource, so the server's disk is optimized.
[0036] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
[0037] The above description of the present invention and its embodiments is not restrictive. What is shown in the drawings is only one of the embodiments of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and design similar structural forms and embodiments without creative efforts without departing from the purpose of the present invention, they should all fall within the protection scope of the present invention.
Claims
1. An ETL system performance bottleneck prediction and analysis device based on multivariate time series, characterized in that It includes the following modules: A feature data input device, a bottleneck prediction model, a performance bottleneck detection device, and a performance bottleneck localization device. The execution flow steps between the devices are as follows: The feature data input device designs an ETL task test set containing multiple load types, collects key performance indicator data during task execution; extracts task features and time series features, and performs data preprocessing, including normalization, dimensionality reduction, and oversampling processing; The bottleneck prediction model constructs a bottleneck prediction model using traditional machine learning algorithms and tunes parameters through Bayesian optimization; The performance bottleneck detection device applies the prediction model to ETL task scheduling, optimizes the task scheduling strategy by predicting performance bottlenecks during task execution, and reduces the occurrence of bottlenecks. The performance bottleneck localization device locates the bottleneck task by analyzing performance bottleneck resources.
2. The device according to claim 1, characterized in that, The bottleneck prediction model uses the XGBoost algorithm and tunes parameters through Bayesian optimization.
3. An ETL system performance bottleneck detection device based on multivariate time series, characterized in that, It includes the following steps: Propose a multivariate time series anomaly detection model called EDD (Encoder-Decoder-Discriminator), combine the graph attention network (GAT) and the long short-term memory network (LSTM) to extract spatial features and time features respectively; Use a small amount of abnormal data to enhance the model's ability to identify anomalies through the joint training of the encoder, decoder, and discriminator; Real-time detect performance bottlenecks in the ETL system through reconstruction error and anomaly score.
4. The device according to claim 3, characterized in that, The EDD model maps normal data to a Gaussian distribution through the encoder, reconstructs the original data through the decoder, and the discriminator is used to distinguish normal data from abnormal data.
5. An ETL system performance bottleneck location device based on resource consumption relevance, characterized in that, It includes the following steps: Based on the reconstruction error output by the EDD model, locate the type of bottleneck resource (such as CPU, memory, disk I / O, network I / O); According to the relevance between tasks and resource consumption, calculate the bottleneck contribution degree of tasks. For CPU and memory bottlenecks, the bottleneck contribution degree formula for each task is: ; where represents the number of components of the task; represents the consumption of bottleneck type resources corresponding to the component. C represents the bottleneck contribution degree. For disk I / O and network I / O bottlenecks, the consumption of these two types of resources is closely related to the data volume. The bottleneck contribution degree formula for each task is: ; In the above formula represents the number of disk I / O or network I / O components in the task; represents the size of each record processed by the component; represents the number of data rows processed by the component; represents the running time of the component; C represents the bottleneck contribution degree. Then, locate the specific bottleneck task according to the bottleneck contribution degree.
6. The device according to claim 5, characterized in that The bottleneck resource localization determines the type of bottleneck resource by calculating the reconstruction error of various resources.
7. The device according to claim 5, characterized in that The bottleneck task localization determines the specific bottleneck task by calculating the contribution degree of the task to the bottleneck resource.