Prediction method for batch operation time of distributed system and related device

Through the hybrid architecture of the seasonal difference autoregressive moving average model and the deep autoregressive recurrent neural network model, the linear trend and periodic laws of the batch job time of the distributed system are accurately captured. Combined with deep learning to capture nonlinear fluctuations, this solves the problem of inaccurate time prediction in existing technologies, achieves low-error time prediction, and ensures business continuity and stability.

CN120763540APending Publication Date: 2025-10-10AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511042267.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately predict the duration of batch jobs in distributed systems, resulting in wasted resources or business delays.

Method used

A hybrid architecture of the seasonal difference autoregressive sliding average model and the deep autoregressive recurrent neural network model is adopted. By extracting trend terms, seasonal terms and residual terms, a three-dimensional vector is constructed for prediction. Deep learning is combined to capture nonlinear fluctuations and generate accurate time usage prediction results.

Benefits of technology

It achieves low-error distributed system batch job duration prediction, reduces the difficulty of engineering implementation, adapts to the actual application scenarios of distributed systems, and ensures business continuity and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763540A_ABST
    Figure CN120763540A_ABST
Patent Text Reader

Abstract

The invention discloses a distributed system batch operation time prediction method and a related device, which can be used in the field of artificial intelligence, and the method comprises the steps: firstly, obtaining a batch operation time data set of a distributed system batch operation; the batch time-use data set is a daily time-use data set or a rest-settlement daily time-use data set; then, inputting the batch time data set into a seasonal difference autoregression moving average model to obtain a trend term, a seasonal term, a residual term and a baseline predicted value; then, constructing a three-dimensional vector based on the trend term, the seasonal term and the residual term; thirdly, inputting a three-dimensional vector into the deep autoregressive recurrent neural network model to obtain a future residual value; and finally, calculating the sum of the baseline prediction value and the future residual value to obtain a prediction result of the batch operation time of the distributed system. Therefore, the strong explanatory force of the statistical model to the seasonality is reserved, the fitting ability of the deep learning to the complex mode is also exerted, and a low-error prediction result can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method and related device for predicting the time taken for batch operations in a distributed system. Background Art

[0002] With the rapid development of financial technology, distributed systems have become the core architecture for data centers to handle large-scale batch jobs. For example, distributed system batch jobs such as bank account settlement, data backup, and interest calculation all rely on distributed clusters to achieve efficient parallel processing.

[0003] The accuracy of batch job duration predictions in distributed systems is directly related to business continuity. If the duration prediction is too long, it can lead to redundant resources and waste; if it is too short, task timeouts can impact normal daytime operations. Currently, batch job duration predictions in distributed systems typically rely on manual experience, making accurate predictions difficult.

[0004] Therefore, how to accurately predict the time required for batch jobs in distributed systems has become a problem that needs to be solved. Summary of the Invention

[0005] Based on the above problems, the present application provides a method and related device for predicting the time required for batch jobs in a distributed system, which can accurately predict the time required for batch jobs in a distributed system.

[0006] The embodiments of this application disclose the following technical solutions:

[0007] In a first aspect, an embodiment of the present application provides a method for predicting the duration of batch jobs in a distributed system, the method comprising:

[0008] Acquire a batch time dataset of a distributed system batch job; the batch time dataset is a weekday time dataset or an interest payment day time dataset;

[0009] Inputting the batch time usage data set into a seasonal difference autoregressive moving average model to obtain a trend term, a seasonal term, a residual term, and a baseline forecast value;

[0010] constructing a three-dimensional vector based on the trend term, the seasonal term, and the residual term;

[0011] Inputting the three-dimensional vector into a deep autoregressive recurrent neural network model to obtain a future residual value;

[0012] The sum of the baseline prediction value and the future residual value is calculated to obtain a prediction result of the batch job time of the distributed system.

[0013] Optionally, before obtaining the batch time dataset of the distributed system batch job, the method further includes:

[0014] Obtaining a batch time training set of a distributed system batch job; the batch time training set is a weekday time training set or an interest settlement day time training set;

[0015] Performing time series decomposition on the data in the batch time training set to obtain a training trend term, a training season term, and a training residual term;

[0016] A seasonal difference autoregressive sliding average model is constructed based on the training trend term, the training season term and the training residual term.

[0017] Optionally, constructing a seasonal difference autoregressive moving average model based on the training trend term, the training season term, and the training residual term includes:

[0018] Determining a cycle parameter based on the training season term; determining a difference order based on the training trend term and the training season term; and determining an autoregressive order and a moving average order based on the training residual term;

[0019] A seasonal difference autoregressive sliding average model is constructed based on the period parameter, the difference order, the autoregressive order and the moving average order.

[0020] Optionally, before performing time series decomposition on the data in the batch time training set to obtain a training trend term, a training season term, and a training residual term, the method further includes:

[0021] Data cleaning is performed on the data in the batch time training set.

[0022] Optionally, in the deep autoregressive recurrent neural network model, the encoder is a two-layer unidirectional long short-term memory network, and the decoder is an autoregressive long short-term memory network.

[0023] Optionally, the weekday time usage data set includes a timestamp, batch operation time usage data on weekdays, and emergency event occurrence information; the interest payment day time usage data set includes a timestamp, batch operation time usage data on the interest payment day, and emergency event occurrence information.

[0024] In a second aspect, an embodiment of the present application provides a device for predicting the duration of batch jobs in a distributed system, the device comprising:

[0025] The first acquisition module is used to acquire a batch time dataset of a batch job of a distributed system; the batch time dataset is a weekday time dataset or an interest payment day time dataset;

[0026] The first prediction module is configured to input the batch time dataset into a seasonal difference autoregressive moving average model to obtain a trend item, a seasonal item, a residual item, and a baseline prediction value.

[0027] The vector construction module is configured to construct a three-dimensional vector based on the trend item, the seasonal item, and the residual item.

[0028] The second prediction module is configured to input the three-dimensional vector into a deep autoregressive recurrent neural network model to obtain a future residual value.

[0029] The calculation module is configured to calculate a sum of the baseline prediction value and the future residual value to obtain a prediction result of the batch job time of the distributed system.

[0030] Optionally, the apparatus further comprises:

[0031] The second acquisition module is configured to acquire a batch time training set of the batch job of the distributed system; the batch time training set is a weekday time training set or a weekend time training set.

[0032] The decomposition module is configured to perform time series decomposition on data in the batch time training set to obtain a training trend item, a training seasonal item, and a training residual item.

[0033] The construction module is configured to construct a seasonal difference autoregressive moving average model based on the training trend item, the training seasonal item, and the training residual item.

[0034] In a third aspect, an embodiment of the present application provides a prediction device for a batch job time of a distributed system, the device comprising a memory and a processor.

[0035] The memory is configured to store program code and transmit the program code to the processor.

[0036] The processor is configured to execute steps of the prediction method for the batch job time of the distributed system according to the program code.

[0037] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program, when the computer program runs on the prediction device for the batch job time of the distributed system, the prediction device for the batch job time of the distributed system executes steps of the prediction method for the batch job time of the distributed system according to any one of the embodiments of the first aspect.

[0038] Compared with the prior art, the present application has the following beneficial effects:

[0039] The embodiment of the application provides a distributed system batch job time prediction method, and the method comprises the following steps: firstly, obtaining a batch time dataset of a distributed system batch job; the batch time dataset is a weekday time dataset or a weekend time dataset; then, inputting the batch time dataset into a seasonal difference autoregressive moving average model to obtain a trend item, a seasonal item, a residual item and a baseline prediction value; then, constructing a three-dimensional vector based on the trend item, the seasonal item and the residual item; then, inputting the three-dimensional vector into a deep autoregressive recurrent neural network model to obtain a future residual value; and finally, calculating the sum of the baseline prediction value and the future residual value to obtain a prediction result of the distributed system batch job time.

[0040] Therefore, the trend item, the seasonal item and the residual item are extracted by the seasonal difference autoregressive moving average model, the linear trend and the periodic law of the batch time are accurately captured, the deep autoregressive recurrent neural network model is used to learn the three-dimensional vector, the nonlinear fluctuation is captured, the combination of the two retains the strong explanation of the seasonal statistical model and the fitting ability of the deep learning to the complex pattern, and a prediction result of the distributed system batch job time with low error can be obtained. By using the hybrid architecture of SARIMA and DeepAR, the dilemma of high calculation cost of a single complex model is avoided, the three-dimensional vector as the input feature retains the core time sequence information and simplifies the input dimension, so that the model is easy to deploy and run in real time in the distributed system, and the engineering implementation difficulty is reduced.

[0041] In addition, the method only relies on the time sequence characteristics of the batch time dataset itself, does not need to obtain external dynamic resource indicators such as CPU utilization or static task attributes such as node priority, solves the problems of difficult external feature collection and poor feature quality in actual operation and maintenance, replaces manual feature engineering by internal decomposition features, reduces the dependence on the data collection link, and is more suitable for the actual application scene of the distributed system. BRIEF DESCRIPTION OF DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description only constitute some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0043] Figure 1 A distributed system batch job time prediction method flowchart is provided for the embodiment of the application;

[0044] Figure 2 A seasonal difference autoregressive moving average model training method flowchart is provided for the embodiment of the application;

[0045] Figure 3 A distributed system batch job time prediction device schematic diagram provided by an embodiment of the application;

[0046] Figure 4 A distributed system batch job time prediction device structure diagram provided by an embodiment of the application. DETAILED DESCRIPTION

[0047] The prediction method and related device for distributed system batch job time provided by the application can be used in the field of artificial intelligence. The above is only an example and does not limit the application of the prediction method and related device for distributed system batch job time provided by the application.

[0048] The terms "first", "second", "third", and "fourth" and the like in the specification of the application and the description of the drawings are used to distinguish different objects, and are not used to limit a specific order.

[0049] In the embodiments of the application, the words "as an example" or "for example" are used to represent an example, illustration or description. Any embodiment or design scheme described as "as an example" or "for example" in the embodiments of the application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. On the contrary, the words "as an example" or "for example" are used to present the relevant concept in a specific way.

[0050] The terms used in the embodiment part of the application are only used to explain the specific embodiments of the application, and are not intended to limit the application.

[0051] In order to enable those skilled in the art to better understand the scheme of the application, the technical solutions in the embodiments of the application will be described clearly and completely in the following with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only some of the embodiments of the application, not all. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.

[0052] Referring to Figure 1 The figure is a prediction method flow chart of a distributed system batch job time provided by an embodiment of the application, the method comprising:

[0053] S101: Obtain a batch time data set of a distributed system batch job.

[0054] Batch time usage data sets are either weekday time usage data sets or interest payment day time usage data sets. Weekday time usage data sets include timestamps, weekday batch time usage data, and emergency event information; interest payment day time usage data sets include timestamps, batch time usage data on the interest payment day, and emergency event information.

[0055] Emergency event occurrence information at least includes whether an emergency event occurred on each date. For example, if the emergency event column of a certain date in the batch time data set displays "True", it means that an emergency event occurred on that day. At this time, the batch time data of that day is marked as an outlier, and the outlier can be replaced by the daily average value of the previous month.

[0056] S102: Input the batch time usage data set into the seasonal difference autoregressive moving average model to obtain the trend term, seasonal term, residual term and baseline prediction value.

[0057] Seasonal Autoregressive Integrated Moving Average (SARIMA) is a statistical model used to process and predict time series data with obvious seasonal changes.

[0058] By fitting the batch usage data with a seasonally differencing autoregressive moving average model, a smooth trend term T can be extracted from the in-sample forecast results output by the model. This trend term T can reflect the overall long-term changes in batch usage data, such as a gradual decrease in usage due to system performance optimization or a slow increase in usage due to business volume growth.

[0059] Based on the seasonal parameters (e.g., period s) in the seasonal difference autoregressive moving average model, the cyclical fluctuation component is extracted from the model decomposition results to obtain the seasonal term S. The seasonal term S corresponds to the regular pattern of batch time data recurring with a fixed period such as monthly, such as the periodic increase in time due to the surge in data volume on the monthly interest payment date.

[0060] Subtracting the extracted trend term T and seasonal term S from the original batch duration data yields the remaining fluctuation component, the residual term R. This residual term R reflects the duration deviation caused by random factors such as temporary task insertions or instantaneous resource fluctuations after removing the trend and seasonal patterns.

[0061] Using the seasonal difference autoregressive moving average model, the forecast value for the next n steps can be generated through the iterative autoregressive process. Specifically, for the non-seasonal part, the trend change can be predicted based on the ARIMA model parameters ARIMA (p, d, q); for the seasonal part, the cyclical fluctuation can be extrapolated based on the seasonal parameters SARIMA (P, D, Q, S). Thus, the seasonal difference autoregressive moving average model is used to capture the linear trend and seasonality of the batch time data time series, and the baseline forecast value y is obtained. sarima .

[0062] In this embodiment, when the batch time usage data set is a weekday time usage data set, the seasonal difference autoregressive sliding average model obtained by training based on the weekday time usage training set is selected; when the batch time usage data set is a interest payment day time usage data set, the seasonal difference autoregressive sliding average model obtained by training based on the interest payment day time usage training set is selected.

[0063] S103: Construct a three-dimensional vector based on the trend term, the seasonal term, and the residual term.

[0064] Baseline predicted value y sarima It can predict the expected change path of batch operation time without the influence of new abnormal factors. However, the baseline prediction value y sarima It only reflects the historical trends and seasonal patterns captured by the model, and does not include random fluctuation corrections that may occur in the future.

[0065] In order to obtain more accurate prediction results, in this embodiment, a three-dimensional vector [T t ,S t ,R t ], where t represents the time step. The three-dimensional vector [Tt, St, Rt] is used as the input vector of the Deep Auto-regressive Recurrent Network (DeepAR). The Deep AR can learn the nonlinear relationship of "trend + season + residual" in the batch time data through the Deep AR model, and predict the possible deviation from the baseline prediction value y in the future. sarima The future residual value of the baseline prediction value y sarima Make corrections to obtain more accurate prediction results.

[0066] S104: Input the three-dimensional vector into the deep autoregressive recurrent neural network model to obtain the future residual value.

[0067] The Deep Auto-regressive Recurrent Network (DeepAR) model is a probabilistic generative model designed for time series forecasting. It uses an autoregressive recursive neural network (RNN) to predict time series distribution. It can effectively solve the problem of scale inconsistency between multiple time series and predict the probability distribution of time series by selecting a likelihood function based on data features.

[0068] As an example, in a deep autoregressive recurrent neural network model, a two-layer unidirectional long short-term memory network (LSTM) can be used as an encoder, where the first layer of LSTM is used to process the input three-dimensional vector [T t ,S t ,R t ], outputs the hidden state, the second layer LSTM is used to output feature representation based on the hidden state to capture temporal dependency; the decoder can be an autoregressive LATM, which is used to map the feature representation output by the second layer LSTM to Gaussian distribution parameters of future time steps through a fully connected layer; the loss function can be a negative log-likelihood loss (NLL Loss).

[0069] Input a three-dimensional vector [T t ,S t ,R t ], the deep autoregressive recurrent neural network uses the trend term and the seasonal term as conditional features, minimizes the negative log-likelihood loss, and can predict the future residual value R of "not happening in the future" DeepAR .

[0070] As an example, the three-dimensional vector [T t-k ,S t-k ,R t-k ],...,[T t ,S t ,R t ] Input the deep autoregressive recurrent neural network model. First, the model outputs the residual distribution parameter (μ t+1 ,σ t+1 ), and sample to get the residual value R t+1 Where μ represents the mean, which represents the "most likely estimate" of the future residual value; σ represents the standard deviation, which represents the uncertainty of the estimate. Then, R t+1 and the trend term T at the corresponding time step t+1 and seasonal term S t+1 Combine to form a new vector [T t+1 ,S t+1 ,Rt+1 Then, slide the window to [T t-k+1 ,S t-k+1 ,R t-k+1 ],...,[T t+1 ,S t+1 ,R t+1 ], the residual value R can be obtained by t+2 Repeat the above process to generate the future residual value R for the next m steps. t+1 , R t+2 ,……,R t+m .

[0071] S105: Calculate the sum of the baseline prediction value and the future residual value to obtain the prediction result of the batch job time of the distributed system.

[0072] As an example, to predict the batch time for a certain day in the future, SARIMA gives the baseline prediction value y sarima is 100 minutes, and DeepAR predicts the possible future residual value R DeepAR For +3 minutes, you can DeepAR y obtained by correcting SARIMA prediction sarima , and get the predicted result y=y for the batch job time of the distributed system sarima +R DeepAR =13 minutes.

[0073] Therefore, in the embodiment of the present application, the trend term, seasonal term and residual term are extracted by the seasonal difference autoregressive sliding average model to accurately capture the linear trend and periodic law of batch time. At the same time, the deep autoregressive recurrent neural network model is used to learn the three-dimensional vector to capture nonlinear fluctuations. The combination of the two not only retains the strong explanatory power of the statistical model for seasonality, but also gives full play to the fitting ability of deep learning for complex patterns, and can obtain low-error prediction results for the batch operation time of distributed systems. By adopting a hybrid architecture that integrates SARIMA and DeepAR, the dilemma of high computational cost of a single complex model is avoided. The three-dimensional vector as the input feature not only retains the core time series information, but also simplifies the input dimension, making the model easy to deploy and run in real time in a distributed system, reducing the difficulty of engineering implementation.

[0074] In addition, this method only relies on the temporal characteristics of the batch time dataset itself, without the need to obtain external dynamic resource indicators such as CPU utilization or static task attributes such as node priority. It solves the problems of difficulty in collecting external features and poor feature quality in actual operation and maintenance. By replacing manual feature engineering with internal decomposition features, it reduces dependence on the data collection link and is more suitable for the actual application scenarios of distributed systems.

[0075] The method provided by the embodiment of the present application can obtain accurate batch job time prediction results, which can help the data center determine in advance whether the task will be completed within the specified window. If the predicted time is long, resources can be allocated in advance, such as increasing node computing power; if there is a risk of timeout, emergency plans can be formulated in advance, such as splitting tasks. At key nodes such as interest payment dates, business delays caused by inaccurate time estimates can be effectively avoided, ensuring the continuity and stability of financial services.

[0076] See also Figure 2 , which is a flow chart of a training method for a seasonal difference autoregressive moving average model provided in an embodiment of the present application, the method comprising:

[0077] S201: Obtain a batch time training set of a distributed system batch job.

[0078] Among them, the batch time training set is the weekday time training set or the interest settlement day time training set.

[0079] S202: Perform data cleaning on the data in the batch time training set.

[0080] As an example, you can remove duplicate data from the batch time training set and handle missing values ​​and outliers, for example, by replacing missing values ​​and outliers with the monthly daily average.

[0081] After data cleaning, you can also perform data conversion to convert the data in the batch time training set into a numerical format that meets the requirements of the SARIMA model. For example, you can first divide the data in the batch time training set into two columns, one containing the date field and the other containing the distributed core master batch time data. Then, convert the date field to a date format, set the date as the index, and convert the distributed core master batch time data into int64 format. The distributed core master batch time data can be mission-critical data that directly impacts core business operations, such as account settlement and interest calculation.

[0082] S203: Perform time series decomposition on the data in the batch time training set to obtain a training trend term, a training season term, and a training residual term.

[0083] Specifically, we can use STL (Seasonal and Trend decomposition using LOESS, STL) to perform time series decomposition on the data in the batch time training set to obtain training trend terms, training seasonal terms, and training residual terms. STL is a time series analysis method that can represent time series data as a linear combination of trend, seasonality, and residuals.

[0084] In addition, considering that a large difference in the data value range may cause errors in the model weight distribution, all data can be further normalized and the numerical features can be Min-Max standardized to narrow the absolute numerical range of the data.

[0085] S204: Constructing a seasonal difference autoregressive sliding average model based on the training trend term, the training season term, and the training residual term.

[0086] Based on the training seasonal term, the period parameter S can be determined. Specifically, the training seasonal term can be visualized and analyzed to observe its periodic fluctuation pattern, thereby determining the seasonal cycle length of the time series and obtaining the period parameter S.

[0087] Based on the training trend term and the training seasonal term, the differencing order can be determined. Specifically, the non-seasonal differencing order d can be determined by testing whether the training trend term is stationary, and the seasonal differencing order D can be determined by testing whether the training seasonal term is stationary.

[0088] Based on the training residuals, the autoregressive order and moving average order can be determined. Specifically, the autocorrelation plot (ACF) and partial autocorrelation plot (PACF) of the training residuals can be plotted to observe the truncation (suddenly approaching 0) or tailing (slowly approaching 0) characteristics of the correlation coefficients in the graphs. For the non-seasonal part, the moving average order q and autoregressive order p are determined based on the truncation position of the residual ACF and PACF; for the seasonal part, the seasonal moving average order Q and seasonal autoregressive order P are determined based on the truncation position of the seasonally lagged ACF and PACF of the residuals.

[0089] Furthermore, by integrating the period parameter (S), difference order (d, D), autoregressive order (p, P) and moving average order (q, Q), we can obtain the ARIMA parameters (p, d, q) and seasonal parameters (P, D, Q, S), thereby constructing a complete seasonal difference autoregressive moving average model.

[0090] See also Figure 3 , which is a schematic diagram of a distributed system batch job time prediction device provided by an embodiment of the present application, the device includes:

[0091] The first acquisition module 301 is used to acquire a batch time dataset of a batch job of a distributed system; the batch time dataset is a weekday time dataset or an interest payment day time dataset;

[0092] The first prediction module 302 is used to input the batch time usage data set into the seasonal difference autoregressive moving average model to obtain the trend term, seasonal term, residual term and baseline prediction value;

[0093] A vector construction module 303 is used to construct a three-dimensional vector based on the trend term, the seasonal term, and the residual term;

[0094] A second prediction module 304 is configured to input a three-dimensional vector into the deep autoregressive recurrent neural network model to obtain a future residual value;

[0095] The calculation module 305 is used to calculate the sum of the baseline prediction value and the future residual value to obtain the prediction result of the batch job time of the distributed system.

[0096] Therefore, in the embodiment of the present application, the trend term, seasonal term and residual term are extracted by the seasonal difference autoregressive sliding average model to accurately capture the linear trend and periodic law of batch time. At the same time, the deep autoregressive recurrent neural network model is used to learn the three-dimensional vector to capture nonlinear fluctuations. The combination of the two not only retains the strong explanatory power of the statistical model for seasonality, but also gives full play to the fitting ability of deep learning for complex patterns, and can obtain low-error prediction results for the batch operation time of distributed systems. By adopting a hybrid architecture that integrates SARIMA and DeepAR, the dilemma of high computational cost of a single complex model is avoided. The three-dimensional vector as the input feature not only retains the core time series information, but also simplifies the input dimension, making the model easy to deploy and run in real time in a distributed system, reducing the difficulty of engineering implementation.

[0097] Optionally, some other distributed system batch job duration prediction devices provided by embodiments of the present application further include:

[0098] The second acquisition module is used to obtain a batch time training set of batch jobs of the distributed system; the batch time training set is a weekday time training set or an interest settlement day time training set;

[0099] The decomposition module is used to perform time series decomposition on the data in the batch time training set to obtain the training trend term, training seasonal term, and training residual term;

[0100] A construction module is used to construct a seasonal difference autoregressive moving average model based on a training trend term, a training seasonal term, and a training residual term.

[0101] Optionally, a construction module is specifically used to: determine the period parameter based on the training seasonal term; determine the difference order based on the training trend term and the training seasonal term; determine the autoregressive order and the moving average order based on the training residual term; and construct a seasonal difference autoregressive sliding average model based on the period parameter, the difference order, the autoregressive order and the moving average order.

[0102] Optionally, some other distributed system batch job duration prediction devices provided by embodiments of the present application further include: a cleaning module for performing data cleaning on the data in the batch duration training set.

[0103] Referring to Figure 4 The figure is a distributed system batch job time prediction device structure diagram provided by an embodiment of the application, the device comprising a memory 401 and a processor 402.

[0104] The memory 401 is used to store program codes and transmit the program codes to the processor.

[0105] The processor 402 is used to execute the steps of the above-mentioned distributed system batch job time prediction method according to the instructions in the program codes.

[0106] In addition, the application further provides a computer readable storage medium, the computer readable storage medium storing computer instructions, when the computer instructions run on a distributed system batch job time prediction device, the distributed system batch job time prediction device executes the steps of the above-mentioned distributed system batch job time prediction method.

[0107] It should be noted that each embodiment in the specification is described in a progressive manner, and the same and similar parts between each embodiment can be referred to each other, and each embodiment focuses on the difference from other embodiments. Especially, for the device and storage medium embodiments, since they are basically similar to the method embodiments, they are described more simply, and the related parts can be referred to the part of the description of the method embodiments. The above-described device and storage medium embodiments are only illustrative, and the units described as separate components can be or can not be physically separated, and the components indicated as units can be or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. According to the actual needs, some or all of the modules can be selected to achieve the purpose of the embodiment. Those skilled in the art can understand and implement it without creative labor.

[0108] The above is only a specific embodiment of the application, but the protection scope of the application is not limited to this, any skilled person in the art can easily think of changes or replacements within the technical range disclosed by the application, which should be covered in the protection scope of the application. Therefore, the protection scope of the application should be subject to the protection scope of the claims.

Claims

1. A method for predicting the duration of batch jobs in a distributed system, characterized in that: The method comprises: Obtaining a batch time dataset of a distributed system batch job; the batch time dataset is a weekday time dataset or an interest payment day time dataset; Inputting the batch time usage data set into a seasonal difference autoregressive moving average model to obtain a trend term, a seasonal term, a residual term, and a baseline forecast value; constructing a three-dimensional vector based on the trend term, the seasonal term, and the residual term; Inputting the three-dimensional vector into a deep autoregressive recurrent neural network model to obtain a future residual value; The sum of the baseline prediction value and the future residual value is calculated to obtain a prediction result of the batch job time of the distributed system.

2. The method according to claim 1, characterized in that Before obtaining the batch time dataset of the distributed system batch job, the method further includes: Obtaining a batch time training set of a distributed system batch job; the batch time training set is a weekday time training set or an interest settlement day time training set; Performing time series decomposition on the data in the batch time training set to obtain a training trend term, a training season term, and a training residual term; A seasonal difference autoregressive sliding average model is constructed based on the training trend term, the training season term and the training residual term.

3. The method according to claim 2, characterized in that The constructing of a seasonal difference autoregressive moving average model based on the training trend term, the training season term, and the training residual term includes: Determining a cycle parameter based on the training season term; determining a difference order based on the training trend term and the training season term; and determining an autoregressive order and a moving average order based on the training residual term; A seasonal difference autoregressive sliding average model is constructed based on the period parameter, the difference order, the autoregressive order and the moving average order.

4. The method according to claim 2, characterized in that Before performing time series decomposition on the data in the batch time training set to obtain a training trend term, a training season term, and a training residual term, the method further includes: Data cleaning is performed on the data in the batch time training set.

5. The method according to claim 1, wherein In the deep autoregressive recurrent neural network model, the encoder is a two-layer unidirectional long short-term memory network, and the decoder is an autoregressive long short-term memory network.

6. The method according to claim 1, characterized in that The daytime usage data set includes a timestamp, daytime batch operation time data, and emergency event occurrence information; the interest payment day time data set includes a timestamp, daytime batch operation time data, and emergency event occurrence information.

7. A distributed system batch job time prediction device, characterized in that: The device comprises: The first acquisition module is used to acquire a batch time dataset of a batch job of a distributed system; the batch time dataset is a weekday time dataset or an interest payment day time dataset; A first prediction module is configured to input the batch time usage data set into a seasonal difference autoregressive moving average model to obtain a trend term, a seasonal term, a residual term, and a baseline prediction value; A vector construction module, configured to construct a three-dimensional vector based on the trend term, the seasonal term, and the residual term; A second prediction module is used to input the three-dimensional vector into a deep autoregressive recurrent neural network model to obtain a future residual value; The calculation module is used to calculate the sum of the baseline prediction value and the future residual value to obtain the prediction result of the batch operation time of the distributed system.

8. The device according to claim 7, characterized in that The device further comprises: The second acquisition module is used to acquire a batch time training set of batch jobs of the distributed system; the batch time training set is a weekday time training set or an interest payment day time training set; A decomposition module, configured to perform time series decomposition on the data in the batch time training set to obtain a training trend term, a training season term, and a training residual term; A construction module is used to construct a seasonal difference autoregressive sliding average model based on the training trend term, the training season term and the training residual term.

9. A distributed system batch job time prediction device, characterized in that: The device includes: a memory and a processor; The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the steps of the method for predicting the duration of batch jobs in a distributed system according to any one of claims 1 to 6 according to the program code.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program. When the computer program is executed on a distributed system batch job time prediction device, the distributed system batch job time prediction device performs the steps of the distributed system batch job time prediction method according to any one of claims 1 to 6.