ETL pipeline dynamic resource optimization method, system, equipment and medium
By collecting multi-source timing data and combining ARIMA and LSTM models, ETL scheduling strategies are dynamically adjusted, which solves the problems of resource waste and performance bottlenecks in ETL tools, and realizes the automation optimization of ETL pipeline resources and improves system efficiency.
Patent Information
- Application Number
- CN202511086421.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-08-05
AI Technical Summary
Existing ETL tools rely on fixed resource configuration and static scheduling strategies, and cannot adapt to load fluctuations in real time, resulting in resource waste and performance bottlenecks. They lack the ability to couple multi-dimensional data to predict, making it difficult to achieve intelligent adjustments.
Collect multi-source timing data, combine ARIMA and LSTM models to generate a comprehensive prediction sequence, dynamically adjust the ETL scheduling strategy, including collecting industrial equipment indicators, ETL task scheduling indicators and calculating cluster system load indicators, constructing a linear load prediction model and LSTM through ARIMA to build a multi-dimensional timing load prediction model, generate a comprehensive prediction sequence of system load indicators, identify peaks, troughs and resource tight periods, and dynamically adjust task scheduling.
It realizes the automation optimization of ETL pipeline resources, improves system efficiency and stability, solves the problem of rigid scheduling strategies in the existing technology, and realizes accurate prediction and dynamic adaptation of future resource requirements.
Smart Images

Figure CN120578482A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of ETL resource optimization, and in particular to a method, system, device and medium for dynamic resource optimization of an ETL pipeline. Background Art
[0002] In the current field of big data processing, the ETL (Extract, Transform, Load) process, as a core component of data integration, is widely used in scenarios such as manufacturing and finance. It efficiently integrates data from heterogeneous data sources and loads it into target systems. The performance of the ETL pipeline directly impacts the timeliness of enterprise data processing and business continuity. Especially with the surge in data volumes and increasing real-time requirements, resource optimization of the ETL pipeline has become a focus of industry attention.
[0003] In existing technologies, some ETL tools optimize resource utilization through predefined resource configurations and static scheduling rules. For example, they employ fixed parallelism and memory allocation strategies, combined with task prioritization to mitigate resource conflicts. Some methods also incorporate simple threshold alert mechanisms, triggering manual intervention when system load exceeds preset values to maintain processing efficiency. These methods mitigate the risk of idle or overloaded resources to a certain extent and leverage historical data analysis to aid decision-making.
[0004] However, existing technologies have significant shortcomings: fixed resource configuration cannot adapt to dynamic fluctuations in load, resulting in resource shortages during peak hours and resource waste during off-peak hours; static scheduling strategies lack the ability to respond to real-time load changes and are unable to predict and avoid potential bottlenecks; at the same time, data collection is limited to a single dimension (such as monitoring only CPU or memory), ignoring the coupling effects of industrial equipment indicators, ETL task scheduling indicators and system load indicators, resulting in insufficient prediction accuracy and difficulty in achieving intelligent adjustment. Summary of the Invention
[0005] In response to the technical problems that existing ETL tools rely on fixed resource configuration and static scheduling strategies for resource optimization, cannot adapt to load fluctuations in real time, resulting in resource waste and performance bottlenecks, and lack multi-dimensional data coupling prediction capabilities, this application provides an ETL pipeline dynamic resource optimization method, system, equipment and medium. By collecting multi-source time series data, combining ARIMA and LSTM models to generate a comprehensive prediction sequence, and dynamically adjusting the ETL scheduling strategy, it realizes the automated optimization of ETL pipeline resources and improves system efficiency.
[0006] In a first aspect, the present application provides a method for dynamic resource optimization of an ETL pipeline, comprising the following steps: S1. Collect time series data on the ETL pipeline operating environment, including industrial equipment indicators, ETL task scheduling indicators, and computing cluster system load indicators within the ETL pipeline operating environment. S2. Divide the ETL pipeline runtime environment time series data into time windows and perform normalization to obtain standardized observations of the ETL pipeline runtime environment time series data. This data is then organized into standardized feature vectors for each time window. S3. Input the standardized observations of the computing cluster system load indicator time series data into the linear load forecasting model to generate linear forecast values of the system load indicator with a specified continuous forecast step length, forming a linear forecast sequence of the system load indicator. The linear load forecasting model is constructed based on the ARIMA algorithm. S4. Input the standardized feature vectors of the continuous time window into the pre-trained multi-dimensional time series load forecasting model to generate coupled prediction values of system load indicators with a specified continuous prediction step length, forming a coupled prediction sequence of system load indicators. The multi-dimensional time series load forecasting model is built based on the LSTM algorithm. S5. Integrate the system load index linear prediction sequence and the system load index coupled prediction sequence into a system load index comprehensive prediction sequence, and denormalize the system load index comprehensive prediction sequence to obtain the original system load index comprehensive prediction sequence; S6. Based on the comprehensive prediction sequence of raw system load indicators, identify peak and low load periods, and periods of resource shortage within the ETL pipeline operating environment, and dynamically adjust the ETL task scheduling strategy based on the identification results. S7. Execute the adjusted scheduling strategy.
[0007] It should be further explained that in step S1, the industrial equipment indicators include: Number of gateways operating normally; Number of production equipment operating normally; The number of quality inspection equipment in operation; Number of other equipment in operation; Number of alarms; Quantity of products produced; Quantity of products tested; Number of faulty devices; The number of unqualified products detected by the equipment.
[0008] It should be further explained that in step S1, the ETL task scheduling indicators include: Number of offline tasks; Number of real-time tasks; Number of failed missions; Number of task failures due to external reasons; Number of task failures due to resource reasons; Number of task failures due to internal reasons; Number of parallel tasks; The duration of ETL task execution.
[0009] It should be further explained that in step S1, calculating the cluster system load index includes: CPU utilization; Network traffic; Memory usage.
[0010] It should be further explained that in step S2, the formula for normalization processing is:
[0011] Where, The first time series data in the ETL pipeline running environment i Indicators at a time point Observed values of The first time series data in the ETL pipeline running environment i Indicators at a time point The sample mean within the time window; The first time series data in the ETL pipeline running environment i Indicators at a time point The sample standard deviation within the time window; The first time series data in the ETL pipeline running environment i Indicators at a time point The standardized observation value of .
[0012] It should be further explained that, in step S3, the step of using the linear load prediction model to generate a linear prediction sequence of system load indicators with a specified prediction step size includes: S301. Input the standardized observation values of the computing cluster system load index time series data into the linear load forecasting model and automatically select the optimal parameter combination of the ARIMA algorithm formula in the linear load forecasting model using the AIC criterion. The ARIMA algorithm formula is:
[0013] in, For the i The load index of the computing cluster system at time point The standardized observation value of For the i The load index of the computing cluster system at time point The standardized observation value of for Order difference operator; is a constant term; is the autoregressive coefficient; is the moving average coefficient; is the white noise error term; are the core parameters, where: is the number of autoregressive terms, which indicates the number of lagged observations used in the linear load forecasting model; is the difference order, which indicates the number of differences required to make the time series stationary; is the number of moving average terms, which indicates the number of lagged errors used in the linear load forecasting model; The optimal parameter combination automatically selected using the AIC criterion is expressed as ; S302. Use the optimal parameter combination as the core parameter of the ARIMA algorithm recursive formula. Use the ARIMA algorithm recursive formula to generate a linear prediction value of the system load indicator for a specified continuous prediction step. The ARIMA algorithm recursive formula is:
[0014] in, is the time point of the last known observation; is the prediction step length; For the i The load index of the computing cluster system at time point The linear prediction value of The computing cluster system load index at the time point The standardized observation value of hour, = .
[0015] It should be further noted that, in step S2, the length of the time window is 1 hour; In step S302, the prediction step length =1-24, the unit prediction step is 1 hour.
[0016] It should be further explained that, in step S4, the step of using the pre-trained multi-dimensional time series load forecasting model to generate a system load indicator coupled forecast value with a specified continuous forecast step size includes: S401. Splitting each standardized feature vector into an input feature vector and a corresponding multidimensional label. The standardized feature vector consists of standardized observations of industrial equipment indicators and ETL task scheduling indicators, and the multidimensional label consists of standardized observations of computing cluster system load indicators. S402. Use the principal component analysis algorithm to reduce the dimension of the standardized feature vector to obtain the reduced dimension feature vector , For the After dimensionality reduction, is the dimension after dimensionality reduction; S403. Reduce the dimension of the continuous time window feature vector Input the pre-trained multi-dimensional time series load forecasting model to generate the system load indicator coupled prediction value with the specified continuous prediction step.
[0017] It should be further explained that the training steps of the multi-dimensional time series load prediction model include: S411. Collect historical time series data from the ETL pipeline runtime environment to construct a sample set. Retain 20% of the recent continuous historical data from the sample set as a validation set, and the remaining 80% of the samples as a training set. S412 validation set and training set and step S2, step S401, step S402 the same process, obtain the training set dimensionality reduction feature vector, training set multidimensional label, validation set dimensionality reduction feature vector, validation set multidimensional label; S413. Input the reduced-dimensional feature vector of the training set into the multidimensional time series load forecasting model to generate a preliminary system load index coupling prediction value for a specified continuous forecast step, each preliminary system load index coupling prediction value corresponding to a forecast time point; S414. Use the multi-dimensional label of the training set corresponding to each predicted time point as the true label of the predicted time point, calculate the loss using the mean square error as the loss function, and backpropagate the error to update the function; Repeat steps S413-S414 to perform training. During the training process, the validation set is used to monitor the multi-dimensional time series load prediction model, and the optimal model is retained as the pre-trained multi-dimensional time series load prediction model.
[0018] It should be further explained that, in step S5, the system load index linear prediction sequence and the system load index coupled prediction sequence are integrated into a system load index comprehensive prediction sequence using the weighted average method, and the formula is:
[0019] Where, For the i The load index of the computing cluster system at time point The comprehensive forecast value of For the i The load index of the computing cluster system at time point The linear prediction value of For the i The load index of the computing cluster system at time point The coupling prediction value of is the linear load prediction weight coefficient; Predict weight coefficients for multi-dimensional time series loads.
[0020] It should be further explained that obtaining and The steps are: S501. Calculate the mean square error of the linear load prediction model on the validation set and the mean square error of the pre-trained multi-dimensional time series load forecasting model on the validation set ; S502. Calculate the weight coefficient:
[0021] .
[0022] It should be further explained that the formula used for denormalization is:
[0023] For the i The load index of the computing cluster system at time point The original comprehensive forecast value of .
[0024] It should be further explained that in step S6, the method for identifying the peak hours, low hours, and resource shortage hours of the system load in the ETL pipeline operation environment is as follows: The period when the average value of any cluster system load indicator within the specified time window is greater than or equal to 80% is determined as the peak period; The period when the average value of all cluster system load indicators in the specified time window is lower than 50% is determined as the valley period; The peak period is defined as the period when the highest value of any cluster system load indicator within the specified time window is greater than or equal to 90%.
[0025] It should be further explained that, in step S6, the method for dynamically adjusting the scheduling strategy of the ETL task according to the recognition result includes: During peak hours, implement the following measures: Disable retrying after ETL scheduling failures and adjust the recalculation scheduling to off-peak hours; Non-critical ETL tasks are postponed to off-peak hours; Limit the total number of concurrent users during peak hours to 80% of the original number of concurrent users, and increase the total number of concurrent users during off-peak hours to 120% of the original number of concurrent users. Provide users with new ETL scheduling time configuration suggestions. If the user's scheduling time is set during peak hours, an alarm will pop up and prompt the off-peak time period; During off-peak hours, tasks increase memory and concurrency within tasks to speed up execution. Send alarm notifications to relevant operation and maintenance personnel; During periods of exceptionally tight resource availability, perform the following actions: Immediately cease non-emergency dispatches; Send emergency alert notifications to the operations team; Record abnormal event logs for subsequent analysis; Abnormally high resource scheduling analysis.
[0026] In a second aspect, the present application provides an ETL pipeline dynamic resource optimization system for implementing the above-mentioned ETL pipeline dynamic resource optimization method, including: The time series data collection module is used to collect time series data of the ETL pipeline operation environment, including time series data of industrial equipment indicators, ETL task scheduling indicators and computing cluster system load indicators in the ETL pipeline operation environment; The data processing module is used to divide the time series data of the ETL pipeline operation environment into time windows, and then perform standardization processing to obtain the standardized observation values of the time series data of the ETL pipeline operation environment, and organize them into standardized feature vectors for each time window; The linear load prediction module is used to input the standardized observation values of the computing cluster system load index time series data into the linear load prediction model, generate the system load index linear prediction values with a specified continuous prediction step length, and form a system load index linear prediction sequence; A coupled load prediction module is used to input the standardized feature vectors of the continuous time window into the pre-trained multi-dimensional time series load prediction model to generate a system load indicator coupled prediction value with a specified continuous prediction step length, forming a system load indicator coupled prediction sequence; The prediction sequence integration and denormalization module is used to integrate the system load index linear prediction sequence and the system load index coupled prediction sequence into a system load index comprehensive prediction sequence, and denormalize the system load index comprehensive prediction sequence to obtain the original system load index comprehensive prediction sequence; The load period identification and scheduling strategy adjustment module is used to identify the peak period, low period and resource shortage period of the system load in the ETL pipeline operation environment based on the comprehensive prediction sequence of the original system load indicators, and dynamically adjust the scheduling strategy of the ETL task according to the identification results; The scheduling strategy execution module is used to execute the adjusted scheduling strategy.
[0027] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to implement the steps of the above-mentioned ETL pipeline dynamic resource optimization method when executing the computer program.
[0028] In a fourth aspect, the present application provides a storage medium having a computer program stored thereon, which implements the steps of the above-mentioned ETL pipeline dynamic resource optimization method when executed by a processor.
[0029] It can be seen from the above technical solutions that this application has the following advantages: 1. This application collects time-series data of the ETL pipeline operating environment (including industrial equipment indicators, ETL task scheduling indicators, and computing cluster system load indicators), solving the problem of single data collection dimensions and inability to fully reflect system status in existing technologies. It also establishes a foundation for the integrated analysis of multi-source data and provides support for accurate prediction.
[0030] 2. This application uses the ARIMA algorithm to construct a linear load forecasting model to generate a linear forecast sequence of system load indicators, and combines it with the LSTM algorithm to construct a multidimensional time series load forecasting model to generate a coupled forecast sequence of system load indicators. This solves the problem of a single forecasting model in the existing technology that cannot capture linear and nonlinear load trends, and realizes a comprehensive prediction of future resource demand.
[0031] 3. This application solves the problem of rigid scheduling strategies and inability to dynamically adapt to load changes in existing technologies by integrating the linear prediction sequence and coupled prediction sequence of system load indicators into a comprehensive prediction sequence, and based on this, identifying peak periods, off-peak periods, and periods of abnormal resource shortages, thereby realizing real-time optimization and adjustment of ETL task scheduling strategies.
[0032] 4. By implementing adjusted scheduling strategies (such as limiting the number of concurrent tasks during peak hours and postponing non-critical tasks), the lack of automated execution mechanisms in existing technologies was resolved, enabling automated optimization of resource allocation and improving the operational efficiency and stability of the ETL pipeline. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for the description. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0034] Figure 1 This is a flowchart of the ETL pipeline dynamic resource optimization method in one embodiment of the present application.
[0035] Figure 2 This is a schematic block diagram of an ETL pipeline dynamic resource optimization system in one embodiment of the present application.
[0036] Figure 3 It is a schematic diagram of the hardware structure of an electronic device in one embodiment of the present application. DETAILED DESCRIPTION
[0037] In order to make the application objectives, features, and advantages of this application more obvious and easy to understand, the technical solutions protected by this application will be clearly and completely described below using specific embodiments and drawings. Obviously, the embodiments described below are only part of the embodiments of this application, not all of them. Based on the embodiments in this patent, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this patent.
[0038] The following describes in detail the ETL pipeline dynamic resource optimization method involved in this application. Specific details, such as specific system structures and technologies, are provided for illustrative purposes rather than for limitation, to facilitate a thorough understanding of the embodiments of this application. However, those skilled in the art will appreciate that this application may also be implemented in other embodiments without these specific details.
[0039] In the ETL pipeline dynamic resource optimization method involved in this application, the term "comprising" is used to indicate the presence of the described features, entities, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, entities, steps, operations, elements, components, and / or collections thereof. The terms "including," "comprising," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized.
[0040] To facilitate the clear description of the technical solutions of this application, the words "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that the words "first" and "second" do not limit the quantity or order of execution, and the words "first" and "second" do not necessarily mean different.
[0041] The phrases "one embodiment" or "some embodiments" described in this application mean that the specific features, structures, or characteristics described in the embodiment are included in one or more embodiments of the application. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in other embodiments," etc. that appear in different places in this application do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized.
[0042] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.
[0043] The ETL pipeline dynamic resource optimization method provided in the embodiment of the present application is executed by a computer device, and accordingly, the ETL pipeline dynamic resource optimization system runs in the computer device.
[0044] Figure 1 This is a flowchart of a method for dynamic resource optimization of an ETL pipeline according to an embodiment of the present application. Figure 1 The execution subject can be an ETL pipeline dynamic resource optimization system. According to different requirements, the order of the steps in the flowchart can be changed, and some steps can be omitted.
[0045] like Figure 1 As shown, the ETL pipeline dynamic resource optimization method includes: Step S1: Collect the time series data of the ETL pipeline operation environment, including the time series data of the industrial equipment indicators, ETL task scheduling indicators and computing cluster system load indicators in the ETL pipeline operation environment.
[0046] By limiting the collection of three types of time-series data: industrial equipment indicators, ETL task scheduling indicators, and computing cluster system load indicators, we can achieve comprehensive, multi-dimensional data collection of the ETL pipeline operating environment status (underlying support, task requirements, system resources), providing the necessary data foundation for subsequent analysis and prediction.
[0047] In some specific embodiments, the industrial equipment indicators include: Number of gateways operating normally; Number of production equipment operating normally; The number of quality inspection equipment in operation; Number of other equipment in operation; Number of alarms; Quantity of products produced; Quantity of products tested; Number of faulty devices; The number of unqualified products detected by the equipment.
[0048] By limiting the selection range of industrial equipment indicators to cover key aspects such as equipment operating status, production efficiency, alarms and quality conditions, we achieve a comprehensive quantitative perception of the operating status of the industrial environment supporting the operation of the ETL pipeline, and provide key input data reflecting the intensity of underlying industrial activities for subsequent load forecasting models.
[0049] In some specific embodiments, the ETL task scheduling indicators include: Number of offline tasks; Number of real-time tasks; Number of failed missions; Number of task failures due to external reasons; Number of task failures due to resource reasons; Number of task failures due to internal reasons; Number of parallel tasks; The duration of ETL task execution.
[0050] By limiting the selection scope of ETL task scheduling indicators to cover core dimensions such as task type, execution status (especially the breakdown of failure reasons), parallelism, and idle conditions, we achieve refined monitoring of the execution behavior characteristics of ETL tasks, providing key input data that reflects the ETL task's own needs and status changes for subsequent load prediction models.
[0051] In some specific embodiments, calculating the cluster system load index includes: CPU utilization; Network traffic; Memory usage.
[0052] By limiting the selection of computing cluster system load indicators to focus on the three core performance dimensions of CPU utilization, network traffic, and memory usage, we achieve standardized measurement of key aspects of resource consumption in the computing environment where the ETL pipeline is located, providing a clear target prediction object and system status evaluation benchmark for the load prediction model.
[0053] In step S2, the ETL pipeline operation environment time series data is divided into time windows, and then standardized to obtain standardized observation values of the ETL pipeline operation environment time series data, and organized into standardized feature vectors for each time window.
[0054] By dividing the collected time series data into time windows and performing standardization processing, it is organized into standardized feature vectors for each time window, which achieves the normalization and structured processing of the original heterogeneous data, eliminates dimensional differences, and prepares for the input of subsequent prediction models.
[0055] In some specific embodiments, the formula for normalization is:
[0056] Where, The first time series data in the ETL pipeline running environment i Indicators at a time point Observed values of The first time series data in the ETL pipeline running environment i Indicators at a time point The sample mean within the time window; The first time series data in the ETL pipeline running environment i Indicators at a time point The sample standard deviation within the time window; The first time series data in the ETL pipeline running environment i Indicators at a time point The standardized observation value of .
[0057] By limiting the standardization processing and using the Z-score formula based on the sample mean and sample standard deviation within the time window, the normalization processing of original time series data of different dimensions and magnitudes is achieved, the scale differences between indicators are eliminated, and a data basis is provided for the subsequent construction of feature vectors and model training and prediction.
[0058] Step S3: input the standardized observation values of the cluster system load index time series data into the linear load prediction model to generate the system load index linear prediction value with a specified continuous prediction step, forming a system load index linear prediction sequence. The linear load prediction model is constructed based on the ARIMA algorithm.
[0059] By inputting the standardized data of system load indicators into the linear load forecasting model built based on ARIMA to generate a short-term linear forecast sequence, an efficient and statistically supported extrapolation forecast of the system load short-term trend is achieved, which only relies on the historical data of the load itself.
[0060] In some specific embodiments, the step of using the linear load prediction model to generate a linear prediction sequence of system load indicators with a specified prediction step size includes: S301. Input the standardized observation values of the computing cluster system load index time series data into the linear load forecasting model and automatically select the optimal parameter combination of the ARIMA algorithm formula in the linear load forecasting model using the AIC criterion. The ARIMA algorithm formula is:
[0061] in, For the i The load index of the computing cluster system at time point The standardized observation value of For the i The load index of the computing cluster system at time point The standardized observation value of for Order difference operator; is a constant term; is the autoregressive coefficient; is the moving average coefficient; is the white noise error term; are the core parameters, where: is the number of autoregressive terms, which indicates the number of lagged observations used in the linear load forecasting model; is the difference order, which indicates the number of differences required to make the time series stationary; is the number of moving average terms, which indicates the number of lagged errors used in the linear load forecasting model; The optimal parameter combination automatically selected using the AIC criterion is expressed as ; S302. Use the optimal parameter combination as the core parameter of the ARIMA algorithm recursive formula. Use the ARIMA algorithm recursive formula to generate a linear prediction value of the system load indicator for a specified continuous prediction step. The ARIMA algorithm recursive formula is:
[0062] in, is the time point of the last known observation; is the prediction step length; For the i The load index of the computing cluster system at time point The linear prediction value of The computing cluster system load index at the time point The standardized observation value of hour, = .
[0063] By limiting the specific application steps of the linear load forecasting model, including automatically selecting the optimal parameter combination using the AIC criterion and using a recursive formula for multi-step forecasting, an efficient and interpretable short-term linear extrapolation forecast of system load based on the temporal dependency of historical load data is achieved.
[0064] In some specific embodiments, in step S2, the length of the time window is 1 hour; In step S302, the prediction step length =1-24, the unit prediction step is 1 hour.
[0065] By limiting the time window length to 1 hour and the prediction step range to 1-24 hours, the consistency of the prediction model input data and prediction output in time granularity is achieved, ensuring the feasibility and practicality of short-term system load prediction in hours.
[0066] Step S4, input the standardized feature vector of the continuous time window into the pre-trained multi-dimensional time series load forecasting model, generate the system load indicator coupling prediction value of the specified continuous prediction step, and constitute the system load indicator coupling prediction sequence. The multi-dimensional time series load forecasting model is constructed based on the LSTM algorithm.
[0067] By inputting the standardized feature vector of the continuous time window into the multi-dimensional time series load forecasting model built based on LSTM to generate a coupled forecast sequence, a nonlinear system load forecast is achieved that takes into account the complex interaction relationship between multiple environmental factors.
[0068] In some specific embodiments, the step of using a pre-trained multi-dimensional time series load forecasting model to generate a system load indicator coupled prediction value with a specified continuous prediction step size includes: S401. Splitting each standardized feature vector into an input feature vector and a corresponding multidimensional label. The standardized feature vector consists of standardized observations of industrial equipment indicators and ETL task scheduling indicators, and the multidimensional label consists of standardized observations of computing cluster system load indicators. S402. Use the principal component analysis algorithm to reduce the dimension of the standardized feature vector to obtain the reduced dimension feature vector , For the After dimensionality reduction, is the dimension after dimensionality reduction; S403. Reduce the dimension of the continuous time window feature vector Input the pre-trained multi-dimensional time series load forecasting model to generate the system load indicator coupled prediction value with the specified continuous prediction step.
[0069] By limiting the input of the multidimensional time series load forecasting model to a continuous feature vector that has undergone dimensionality reduction and integrates industrial equipment indicators and ETL task scheduling indicators, a nonlinear system load forecast is achieved that takes into account the complex coupling relationship between multi-source heterogeneous factors in the ETL pipeline operating environment.
[0070] In some specific embodiments, the training step of the multi-dimensional time series load prediction model includes: S411. Collect historical time series data from the ETL pipeline runtime environment to construct a sample set. Retain 20% of the recent continuous historical data from the sample set as a validation set, and the remaining 80% of the samples as a training set. S412 validation set and training set and step S2, step S401, step S402 the same process, obtain the training set dimensionality reduction feature vector, training set multidimensional label, validation set dimensionality reduction feature vector, validation set multidimensional label; S413. Input the reduced-dimensional feature vector of the training set into the multidimensional time series load forecasting model to generate a preliminary system load index coupling prediction value for a specified continuous forecast step, each preliminary system load index coupling prediction value corresponding to a forecast time point; S414. Use the multi-dimensional label of the training set corresponding to each predicted time point as the true label of the predicted time point, calculate the loss using the mean square error as the loss function, and backpropagate the error to update the function; Repeat steps S413-S414 to perform training. During the training process, the validation set is used to monitor the multi-dimensional time series load prediction model, and the optimal model is retained as the pre-trained multi-dimensional time series load prediction model.
[0071] By clarifying the training process of the multi-dimensional time series load prediction model (LSTM), including data set partitioning, standardization and dimensionality reduction, model training and validation set monitoring, a prediction model with stronger generalization capability is constructed based on the complex nonlinear mapping relationship between multi-source features of the historical data learning environment and system load.
[0072] Step S5: integrating the system load indicator linear prediction sequence and the system load indicator coupled prediction sequence into a system load indicator comprehensive prediction sequence, and denormalizing the system load indicator comprehensive prediction sequence to obtain an original system load indicator comprehensive prediction sequence.
[0073] By integrating the linear prediction sequence and the coupled prediction sequence into a comprehensive prediction sequence and performing denormalization, a more robust prediction result that integrates the advantages of different prediction models is achieved, and the prediction value is restored to the original dimension for practical application.
[0074] In some specific embodiments, a weighted average method is used to integrate the system load indicator linear prediction sequence and the system load indicator coupled prediction sequence into a system load indicator comprehensive prediction sequence, and the formula is:
[0075] Where, For the i The load index of the computing cluster system at time point The comprehensive forecast value of For the i The load index of the computing cluster system at time point The linear prediction value of For the i The load index of the computing cluster system at time point The coupling prediction value of is the linear load prediction weight coefficient; Predict weight coefficients for multi-dimensional time series loads.
[0076] By limiting the use of the weighted average method to integrate the linear prediction sequence and the coupled prediction sequence, and calculating the dynamic weight coefficient based on the prediction error of each model on the validation set, a more robust and accurate comprehensive prediction of system load is achieved that combines the advantages of linear extrapolation and nonlinear coupling.
[0077] In some specific embodiments, obtaining and The steps are: S501. Calculate the mean square error of the linear load prediction model on the validation set and the mean square error of the pre-trained multi-dimensional time series load forecasting model on the validation set ; S502. Calculate the weight coefficient:
[0078] .
[0079] By limiting the specific formula for calculating the weight coefficient (normalized distribution based on the inverse of the mean square error Loss of each model on the validation set), it is possible to objectively and quantitatively allocate the contribution weight of the model in the comprehensive prediction based on its historical prediction performance, thereby optimizing the accuracy of the prediction results.
[0080] In some specific embodiments, the formula used for denormalization is:
[0081] For the i The load index of the computing cluster system at time point The original comprehensive forecast value of .
[0082] By explicitly denormalizing the composite forecast value using the sample mean and standard deviation of the original time window, the normalized forecast value is converted back to the original dimension and the actual system load index value, providing directly understandable forecast data for subsequent resource period identification. Step S6: Based on the comprehensive prediction sequence of the original system load indicators, the peak period, the valley period and the period of abnormal resource shortage of the system load in the ETL pipeline operation environment are identified, and the scheduling strategy of the ETL task is dynamically adjusted according to the identification results.
[0083] By identifying peak periods, off-peak periods, and periods of abnormally tight resources based on the original comprehensive forecast sequence, and dynamically adjusting the ETL task scheduling strategy according to the identification results, forward-looking and differentiated optimization configuration of task scheduling is achieved based on the predicted system load status.
[0084] In some specific embodiments, a method for identifying peak load periods, low load periods, and resource shortage periods in an ETL pipeline operating environment is as follows: The period when the average value of any cluster system load indicator within the specified time window is greater than or equal to 80% is determined as the peak period; The period when the average value of all cluster system load indicators in the specified time window is lower than 50% is determined as the valley period; The peak period is defined as the period when the highest value of any cluster system load indicator within the specified time window is greater than or equal to 90%.
[0085] By clarifying the threshold rules for determining periods of abnormal resource shortage during peak and off-peak periods, we have achieved automated and standardized identification of future periods of different load risk levels for the system based on a comprehensive forecast sequence.
[0086] In some specific embodiments, the method for dynamically adjusting the scheduling strategy of the ETL task according to the recognition result includes: During peak hours, implement the following measures: Disable retrying after ETL scheduling failures and adjust the recalculation scheduling to off-peak hours; Non-critical ETL tasks are postponed to off-peak hours; Limit the total number of concurrent users during peak hours to 80% of the original number of concurrent users, and increase the total number of concurrent users during off-peak hours to 120% of the original number of concurrent users. Provide users with new ETL scheduling time configuration suggestions. If the user's scheduling time is set during peak hours, an alarm will pop up and prompt the off-peak time period; During off-peak hours, tasks increase memory and concurrency within tasks to speed up execution. Send alarm notifications to relevant operation and maintenance personnel; During periods of exceptionally tight resource availability, perform the following actions: Immediately cease non-emergency dispatches; Send emergency alert notifications to the operations team; Record abnormal event logs for subsequent analysis; Abnormally high resource scheduling analysis.
[0087] By implementing specific scheduling policy adjustment measures (such as task retry control, task priority adjustment, concurrency limit, alarm notification, etc.) during identified peak hours and periods of abnormal resource shortage, we can proactively and meticulously intervene in ETL task scheduling based on predicted resource shortages to avoid resource bottlenecks, optimize task execution efficiency, and ensure system stability. Step S7: Execute the adjusted scheduling strategy.
[0088] By implementing the adjusted scheduling strategy, we can put prediction-based optimization decisions into practice, ultimately achieving the goals of dynamically adjusting resource allocation, optimizing task execution, and ensuring stable system operation.
[0089] In a specific embodiment, the steps of the ETL pipeline dynamic resource optimization method include: Step S1: Collect the time series data of the ETL pipeline operation environment, including the time series data of the industrial equipment indicators, ETL task scheduling indicators and computing cluster system load indicators in the ETL pipeline operation environment, where: Industrial equipment indicators include: Number of gateways operating normally; Number of production equipment operating normally; The number of quality inspection equipment in operation; Number of other equipment in operation; Number of alarms; Quantity of products produced; Quantity of products tested; Number of faulty devices; The number of unqualified products detected by equipment; ETL task scheduling indicators include: Number of offline tasks; Number of real-time tasks; Number of failed missions; Number of task failures due to external reasons; Number of task failures due to resource reasons; Number of task failures due to internal reasons; Number of parallel tasks; No ETL task execution time; Computing cluster system load indicators include: CPU utilization; Network traffic; Memory usage.
[0090] In step S2, the ETL pipeline operation environment time series data is divided into 1-hour time windows, and then normalized to obtain the standardized observation values of the ETL pipeline operation environment time series data, and organized into standardized feature vectors for each time window. The normalization formula is:
[0091] Where, The first time series data in the ETL pipeline running environment i Indicators at a time point Observed values of The first time series data in the ETL pipeline running environment i Indicators at a time point The sample mean within the time window; The first time series data in the ETL pipeline running environment i Indicators at a time point The sample standard deviation within the time window; The first time series data in the ETL pipeline running environment i Indicators at a time point The standardized observation value of .
[0092] Step S3: Input the standardized observation values of the cluster system load index time series data into the linear load forecasting model to generate the system load index linear forecast values of the specified continuous forecast step length to form the system load index linear forecast sequence. The linear load forecasting model is constructed based on the ARIMA algorithm, and the steps include: S301. Input the standardized observation values of the computing cluster system load index time series data into the linear load forecasting model and automatically select the optimal parameter combination of the ARIMA algorithm formula in the linear load forecasting model using the AIC criterion. The ARIMA algorithm formula is:
[0093] in, For the i The load index of the computing cluster system at time point The standardized observation value of For the i The load index of the computing cluster system at time point The standardized observation value of for Order difference operator; is a constant term; is the autoregressive coefficient; is the moving average coefficient; is the white noise error term; are the core parameters, where: is the number of autoregressive terms, which indicates the number of lagged observations used in the linear load forecasting model; is the difference order, which indicates the number of differences required to make the time series stationary; is the number of moving average terms, which indicates the number of lagged errors used in the linear load forecasting model; The optimal parameter combination automatically selected using the AIC criterion is expressed as ; S302. Use the optimal parameter combination as the core parameter of the ARIMA algorithm recursive formula. Use the ARIMA algorithm recursive formula to generate a linear prediction value of the system load indicator for a specified continuous prediction step. The ARIMA algorithm recursive formula is:
[0094] in, is the time point of the last known observation; is the prediction step length, =1-24, the unit prediction step is 1 hour; For the i The load index of the computing cluster system at time point The linear prediction value of The computing cluster system load index at the time point The standardized observation value of hour, = .
[0095] Step S4: Input the standardized feature vector of the continuous time window into the pre-trained multi-dimensional time series load forecasting model to generate a system load indicator coupling prediction value of a specified continuous prediction step, forming a system load indicator coupling prediction sequence. The multi-dimensional time series load forecasting model is constructed based on the LSTM algorithm, and the steps include: S401. Splitting each standardized feature vector into an input feature vector and a corresponding multidimensional label. The standardized feature vector consists of standardized observations of industrial equipment indicators and ETL task scheduling indicators, and the multidimensional label consists of standardized observations of computing cluster system load indicators. S402. Use the principal component analysis algorithm to reduce the dimension of the standardized feature vector to obtain the reduced dimension feature vector , For the After dimensionality reduction, is the dimension after dimensionality reduction; S403. Reduce the dimension of the feature vector of the past 168 continuous time windows Input the pre-trained multi-dimensional time series load forecasting model to generate the coupled prediction value of the system load indicators with a specified continuous forecast step length, with each time window being 1 hour; The training steps of the multi-dimensional time series load forecasting model include: S411. Collect historical time series data from the ETL pipeline runtime environment to construct a sample set. Retain 20% of the recent continuous historical data from the sample set as a validation set, and the remaining 80% of the samples as a training set. S412 validation set and training set and step S2, step S401, step S402 the same process, obtain the training set dimensionality reduction feature vector, training set multidimensional label, validation set dimensionality reduction feature vector, validation set multidimensional label; S413. Input the reduced-dimensional feature vector of the training set into the multidimensional time series load forecasting model to generate a preliminary system load index coupling prediction value for a specified continuous forecast step, each preliminary system load index coupling prediction value corresponding to a forecast time point; S414. Use the multi-dimensional label of the training set corresponding to each predicted time point as the true label of the predicted time point, calculate the loss using the mean square error as the loss function, and backpropagate the error to update the function; Repeat steps S413-S414 to perform training. During the training process, the validation set is used to monitor the multi-dimensional time series load prediction model, and the optimal model is retained as the pre-trained multi-dimensional time series load prediction model.
[0096] Step S5: Use the weighted average method to integrate the system load index linear prediction sequence and the system load index coupled prediction sequence into a system load index comprehensive prediction sequence. The formula is:
[0097] Where, For the i The load index of the computing cluster system at time point The comprehensive forecast value of For the i The load index of the computing cluster system at time point The linear prediction value of For the i The load index of the computing cluster system at time point The coupling prediction value of is the linear load prediction weight coefficient; Predict weight coefficients for multi-dimensional time series loads; Get and The steps are: S501. Calculate the mean square error of the linear load prediction model on the validation set and the mean square error of the pre-trained multi-dimensional time series load forecasting model on the validation set ; S502. Calculate the weight coefficient:
[0098] ; The system load index comprehensive prediction sequence is denormalized to obtain the original system load index comprehensive prediction sequence, and the formula is:
[0099] For the i The load index of the computing cluster system at time point The original comprehensive forecast value of .
[0100] Step S6: Identify the peak hours, low hours, and resource shortage hours of the system load in the ETL pipeline operation environment based on the comprehensive prediction sequence of the original system load indicators, and dynamically adjust the scheduling strategy of the ETL task according to the identification results; Among them, the method for identifying peak hours, low hours, and resource shortage periods in the ETL pipeline operation environment is as follows: The period when the average value of any cluster system load indicator within the specified time window is greater than or equal to 80% is determined as the peak period; The period when the average value of all cluster system load indicators in the specified time window is lower than 50% is determined as the valley period; The peak period is defined as the period when the highest value of any cluster system load indicator within the specified time window is greater than or equal to 90%. Methods for dynamically adjusting the scheduling strategy of ETL tasks based on recognition results include: During peak hours, implement the following measures: Disable retrying after ETL scheduling failures and adjust the recalculation scheduling to off-peak hours; Non-critical ETL tasks are postponed to off-peak hours; Limit the total number of concurrent users during peak hours to 80% of the original number of concurrent users, and increase the total number of concurrent users during off-peak hours to 120% of the original number of concurrent users. Provide users with new ETL scheduling time configuration suggestions. If the user's scheduling time is set during peak hours, an alarm will pop up and prompt the off-peak time period; During off-peak hours, tasks increase memory and concurrency within tasks to speed up execution. Send alarm notifications to relevant operation and maintenance personnel; During periods of exceptionally tight resource availability, perform the following actions: Immediately cease non-emergency dispatches; Send emergency alert notifications to the operations team; Record abnormal event logs for subsequent analysis; Analysis of abnormally high resource scheduling, including: During peak hours, especially those with severe resource shortages, we screen for high resource usage scheduling. We combine the resource consumption of each ETL job with the forecast for the next 24 hours to assess the impact each job may have on system resources during a specific time period. For each ETL job in the research period , and stipulates that its average CPU utilization is , the average memory usage is , the average network traffic is , the predicted resource utilization rate in the next t hour is , , ; Define impact factor To measure homework At the time point t Impact on system resources:
[0101] Set the weight coefficients according to the importance of the three predicted values: , , ; According to the calculated impact factor , you can determine which ETL jobs have a greater impact on system resources within a specific time period, and define the ETL job tasks that have a greater impact on system resources as high-impact factor tasks.
[0102] Step S7: Execute the adjusted scheduling policy. Based on the APIs of DolphinScheduler and Seatunnel, dynamic modification of ETL job configuration parameters is achieved to improve the system's automation and responsiveness. DolphinScheduler, as the workflow scheduling system, is responsible for scheduling and managing ETL tasks. Seatunnel, as a lightweight data synchronization tool, is responsible for data extraction, transformation, and loading (ETL) operations. DolphinScheduler calls Seatunnel's API to achieve dynamic configuration and execution of ETL tasks. The specific steps include: Convert the scheduling policy output in step S6 into an interface, call Seatunnel's modification interface through DolphinScheduler, and dynamically adjust the configuration parameters of the ETL job. The adjustment content includes parallelism, memory allocation, etc. according to the specific decision API; Based on the adjusted ETL job configuration, the execution time and order of ETL tasks during peak hours are rearranged through the DolphinScheduler task scheduling interface. Based on load forecasts, suggestions and alerts are issued for the newly added ETL scheduling execution time. Send an alert for the high-impact factor tasks identified and analyzed in step S6 and provide information such as the recommended execution time; Trigger the execution of ETL tasks and monitor their running status through DolphinScheduler; During the ETL task execution process, the real-time monitoring interface provided by DolphinScheduler and Seatunnel is used to obtain the task's running status and resource usage; If any abnormal situation or resource shortage is found, an alarm will be triggered in time and corresponding adjustment measures will be taken.
[0103] The following is an embodiment of the ETL pipeline dynamic resource optimization system provided in the embodiments of the present application. The ETL pipeline dynamic resource optimization system and the ETL pipeline dynamic resource optimization method of the above-mentioned embodiments belong to the same inventive concept. For details not fully described in the embodiments of the ETL pipeline dynamic resource optimization system, please refer to the embodiments of the above-mentioned ETL pipeline dynamic resource optimization method.
[0104] like Figure 2 As shown, the ETL pipeline dynamic resource optimization system includes: The time series data collection module is used to collect time series data of the ETL pipeline operation environment, including time series data of industrial equipment indicators, ETL task scheduling indicators and computing cluster system load indicators in the ETL pipeline operation environment; The data processing module is used to divide the time series data of the ETL pipeline operation environment into time windows, and then perform standardization processing to obtain the standardized observation values of the time series data of the ETL pipeline operation environment, and organize them into standardized feature vectors for each time window; The linear load prediction module is used to input the standardized observation values of the computing cluster system load index time series data into the linear load prediction model, generate the system load index linear prediction values with a specified continuous prediction step length, and form a system load index linear prediction sequence; A coupled load prediction module is used to input the standardized feature vectors of the continuous time window into the pre-trained multi-dimensional time series load prediction model to generate a system load indicator coupled prediction value with a specified continuous prediction step length, forming a system load indicator coupled prediction sequence; The prediction sequence integration and denormalization module is used to integrate the system load index linear prediction sequence and the system load index coupled prediction sequence into a system load index comprehensive prediction sequence, and denormalize the system load index comprehensive prediction sequence to obtain the original system load index comprehensive prediction sequence; The load period identification and scheduling strategy adjustment module is used to identify the peak period, low period and resource shortage period of the system load in the ETL pipeline operation environment based on the comprehensive prediction sequence of the original system load indicators, and dynamically adjust the scheduling strategy of the ETL task according to the identification results; The scheduling strategy execution module is used to execute the adjusted scheduling strategy.
[0105] The ETL pipeline dynamic resource optimization system of this embodiment is used to implement the ETL pipeline dynamic resource optimization method.
[0106] This application also provides an electronic device for implementing each embodiment of this application. Figure 3 A hardware structure diagram of an electronic device for implementing various embodiments of the present application is shown in FIG. Figure 3 As shown, the electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor.
[0107] Those skilled in the art will understand that the electronic device structure involved in the embodiments of the present application does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0108] In the embodiments of the present application, electronic devices include, but are not limited to, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices may also represent various forms of mobile devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described and / or claimed herein.
[0109] In the embodiment of the present application, the processor can be implemented by using at least one of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a processor, a controller, a microcontroller, a microprocessor, and an electronic unit designed to perform the functions described herein. In some cases, such an embodiment can be implemented in a controller. For software implementation, an embodiment such as a process or function can be implemented with a separate software module that allows the execution of at least one function or operation. The software code can be implemented by a software application (or program) written in any appropriate programming language, and the software code can be stored in a memory and executed by a controller.
[0110] In addition, the electronic device includes some functional modules not shown, which will not be described here.
[0111] Those skilled in the art will appreciate that various aspects of the electronic device provided herein may be implemented as a system, method, or program product. Therefore, various aspects of the present application may be specifically implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits," "modules," or "systems."
[0112] The present application also provides a storage medium storing a program product capable of implementing the method for dynamic resource optimization of an ETL pipeline. In some possible implementations, various aspects of the present application may also be implemented in the form of a program product comprising program code. When the program product is executed on a terminal device, the program code is configured to cause the terminal device to execute the steps described in the "Exemplary Methods" section above according to the various exemplary implementations of the present application.
[0113] The storage medium can be any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0114] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for dynamic resource optimization of ETL pipeline, characterized by: include: S1. Collect time series data on the ETL pipeline operating environment, including industrial equipment indicators, ETL task scheduling indicators, and computing cluster system load indicators within the ETL pipeline operating environment. S2. Divide the ETL pipeline runtime environment time series data into time windows and perform normalization to obtain standardized observations of the ETL pipeline runtime environment time series data. This data is then organized into standardized feature vectors for each time window. S3. Input the standardized observations of the computing cluster system load indicator time series data into the linear load forecasting model to generate linear forecast values of the system load indicator with a specified continuous forecast step length, forming a linear forecast sequence of the system load indicator. The linear load forecasting model is constructed based on the ARIMA algorithm. S4. Input the standardized feature vectors of the continuous time window into the pre-trained multi-dimensional time series load forecasting model to generate coupled prediction values of system load indicators with a specified continuous prediction step length, forming a coupled prediction sequence of system load indicators. The multi-dimensional time series load forecasting model is built based on the LSTM algorithm. S5. Integrate the system load index linear prediction sequence and the system load index coupled prediction sequence into a system load index comprehensive prediction sequence, and denormalize the system load index comprehensive prediction sequence to obtain the original system load index comprehensive prediction sequence; S6. Based on the comprehensive prediction sequence of raw system load indicators, identify peak and low load periods, and periods of resource shortage within the ETL pipeline operating environment, and dynamically adjust the ETL task scheduling strategy based on the identification results. S7. Execute the adjusted scheduling strategy.
2. The ETL pipeline dynamic resource optimization method according to claim 1, characterized in that: In step S1, calculating the cluster system load index includes: CPU utilization; Network traffic; Memory usage.
3. The ETL pipeline dynamic resource optimization method according to claim 1, characterized in that: In step S2, the formula for normalization is: Where, The first time series data in the ETL pipeline running environment i Indicators at a time point Observed values of The first time series data in the ETL pipeline running environment i Indicators at a time point The sample mean within the time window; The first time series data in the ETL pipeline running environment i Indicators at a time point The sample standard deviation within the time window; The first time series data in the ETL pipeline running environment i Indicators at a time point The standardized observation value of .
4. The ETL pipeline dynamic resource optimization method according to claim 1, characterized in that: In step S3, the steps of using the linear load prediction model to generate a linear prediction sequence of system load indicators with a specified prediction step size include: S301. Input the standardized observation values of the computing cluster system load index time series data into the linear load forecasting model and automatically select the optimal parameter combination of the ARIMA algorithm formula in the linear load forecasting model using the AIC criterion. The ARIMA algorithm formula is: in, For the i The load index of the computing cluster system at time point The standardized observation value of For the i The load index of the computing cluster system at time point The standardized observation value of for Order difference operator; is a constant term; is the autoregressive coefficient; is the moving average coefficient; is the white noise error term; are the core parameters, where: is the number of autoregressive terms, which indicates the number of lagged observations used in the linear load forecasting model; is the difference order, which indicates the number of differences required to make the time series stationary; is the number of moving average terms, which indicates the number of lagged errors used in the linear load forecasting model; The optimal parameter combination automatically selected using the AIC criterion is expressed as ; S302. Use the optimal parameter combination as the core parameter of the ARIMA algorithm recursive formula. Use the ARIMA algorithm recursive formula to generate a linear prediction value of the system load indicator for a specified continuous prediction step. The ARIMA algorithm recursive formula is: in, is the time point of the last known observation; is the prediction step length; For the i The load index of the computing cluster system at time point The linear prediction value of The load index of the computing cluster system at the time point The standardized observation value of hour, = .
5. The ETL pipeline dynamic resource optimization method according to claim 1, characterized in that: In step S4, the step of using the pre-trained multi-dimensional time series load forecasting model to generate a system load indicator coupled prediction value with a specified continuous prediction step size includes: S401. Splitting each standardized feature vector into an input feature vector and a corresponding multidimensional label. The standardized feature vector consists of standardized observations of industrial equipment indicators and ETL task scheduling indicators, and the multidimensional label consists of standardized observations of computing cluster system load indicators. S402. Use the principal component analysis algorithm to reduce the dimension of the standardized feature vector to obtain the reduced dimension feature vector , For the After dimensionality reduction, is the dimension after dimensionality reduction; S403. Reduce the dimension of the continuous time window feature vector Input the pre-trained multi-dimensional time series load forecasting model to generate the system load indicator coupled prediction value with the specified continuous prediction step.
6. The ETL pipeline dynamic resource optimization method according to claim 5, characterized in that: The training steps of the multi-dimensional time series load forecasting model include: S411. Collect historical time series data from the ETL pipeline runtime environment to construct a sample set. Retain 20% of the recent continuous historical data from the sample set as a validation set, and the remaining 80% of the samples as a training set. S412 validation set and training set and step S2, step S401, step S402 the same process, obtain the training set dimensionality reduction feature vector, training set multidimensional label, validation set dimensionality reduction feature vector, validation set multidimensional label; S413. Input the reduced-dimensional feature vector of the training set into the multidimensional time series load forecasting model to generate a preliminary system load index coupling prediction value for a specified continuous forecast step, each preliminary system load index coupling prediction value corresponding to a forecast time point; S414. Use the multi-dimensional label of the training set corresponding to each predicted time point as the true label of the predicted time point, calculate the loss using the mean square error as the loss function, and backpropagate the error to update the function; Repeat steps S413-S414 to perform training. During the training process, the validation set is used to monitor the multi-dimensional time series load prediction model, and the optimal model is retained as the pre-trained multi-dimensional time series load prediction model.
7. The ETL pipeline dynamic resource optimization method according to claim 1, characterized in that: In step S5, the system load index linear prediction sequence and the system load index coupled prediction sequence are integrated into a system load index comprehensive prediction sequence using a weighted average method. The formula is: Where, For the i The load index of the computing cluster system at time point The comprehensive forecast value of For the i The load index of the computing cluster system at time point The linear prediction value of For the i The load index of the computing cluster system at time point The coupling prediction value of is the linear load prediction weight coefficient; Predict weight coefficients for multi-dimensional time series loads.
8. An ETL pipeline dynamic resource optimization system, characterized by: The method for implementing the ETL pipeline dynamic resource optimization method according to any one of claims 1 to 7 comprises: The time series data collection module is used to collect time series data of the ETL pipeline operation environment, including time series data of industrial equipment indicators, ETL task scheduling indicators and computing cluster system load indicators in the ETL pipeline operation environment; The data processing module is used to divide the time series data of the ETL pipeline operation environment into time windows, and then perform standardization processing to obtain the standardized observation values of the time series data of the ETL pipeline operation environment, and organize them into standardized feature vectors for each time window; The linear load prediction module is used to input the standardized observation values of the computing cluster system load index time series data into the linear load prediction model, generate the system load index linear prediction values with a specified continuous prediction step length, and form a system load index linear prediction sequence; A coupled load prediction module is used to input the standardized feature vectors of the continuous time window into the pre-trained multi-dimensional time series load prediction model to generate a system load indicator coupled prediction value with a specified continuous prediction step length, forming a system load indicator coupled prediction sequence; The prediction sequence integration and denormalization module is used to integrate the system load index linear prediction sequence and the system load index coupled prediction sequence into a system load index comprehensive prediction sequence, and denormalize the system load index comprehensive prediction sequence to obtain the original system load index comprehensive prediction sequence; The load period identification and scheduling strategy adjustment module is used to identify the peak period, low period and resource shortage period of the system load in the ETL pipeline operation environment based on the comprehensive prediction sequence of the original system load indicators, and dynamically adjust the scheduling strategy of the ETL task according to the identification results; The scheduling strategy execution module is used to execute the adjusted scheduling strategy.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: The processor is used to implement the steps of the ETL pipeline dynamic resource optimization method according to any one of claims 1 to 7 when executing the computer program.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the ETL pipeline dynamic resource optimization method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Container cloud resource prediction method based on ARIMA-LSTM
CN117827617A
Multi-model fusion intelligent power grid cloud data center resource load prediction method
CN118363831A
Cluster node load state prediction-based job scheduling method
WO2020206705A1
Anticipatory learning method and system oriented towards short-term time series prediction
WO2021052140A1