Multi-target coal product demand quantity prediction method based on sparse BiLSTM

By introducing sparse strategies and multi-objective optimization methods in BiLSTM model, the problem of overfitting and computational overhead in the prediction of coal product demand is solved, and more efficient prediction and lower computational cost are achieved.

CN120013576APending Publication Date: 2025-05-16CCTEG SHENYANG ENG CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411909177.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing BiLSTM model has problems of overfitting and weak generalization capabilities in the coal product demand forecast, and there is overconnection in large BiLSTM networks, which makes the calculation overhead too large.

Method used

The multi-objective coal product demand prediction method based on sparse BiLSTM is adopted, and the weight parameters of the BiLSTM model are optimized by improving the multi-objective differential evolution algorithm (MCMODE) and sparse strategy, and a multi-period length prediction result superposition method is designed to improve the prediction accuracy and generalization ability of the model.

Benefits of technology

It improves the prediction accuracy and generalization ability of the model, reduces the calculation cost, and avoids the phenomenon of overfitting the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013576A_ABST
    Figure CN120013576A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-target coal product demand quantity prediction method based on sparse BiLSTM, and belongs to the technical field of time sequence prediction analysis methods. According to the method, the characteristic that a BiLSTM model is good at capturing a data complex nonlinear mode is fully utilized, an MCMODE algorithm taking model precision and network complexity as targets is developed to optimize weight parameters of the BiLSTM, and a superposition method of multi-cycle length prediction results is designed to add BiLSTM prediction values, so that the prediction precision and generalization ability of the model are ensured; and meanwhile, a sparse strategy is designed to compress the parameter scale of the BiLSTM, so that the calculation cost of the model is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of time series forecasting and analysis methods, and in particular relates to a multi-objective coal product demand forecasting method based on sparse BiLSTM. Background Art

[0002] For coal mining companies, predicting the market demand for coal is extremely important for the design of the company's mining plan and the formulation of the production plan. As a popular nonlinear model, deep neural network has excellent time series learning ability and can achieve high-precision prediction of complex data. Among the many neural network models, the learning and prediction ability of the bidirectional long short-term memory network (Bi-Directional Long Short-Term Memory, referred to as BiLSTM) model is relatively outstanding, and can effectively handle the long-term dependencies in the time series. Therefore, the present invention adopts the BiLSTM model to predict the daily demand for coal products of coal mining enterprises.

[0003] In terms of the model, considering the influence of the estimation level of weight parameters and the coal product data pattern on the prediction accuracy of the model, the present invention adopts a multi-objective optimization method to optimize the weight parameters of BiLSTM, and improves the prediction accuracy of the model by superimposing multi-period prediction results. At the same time, most existing studies construct multiple optimization objectives from the perspective of model fitting accuracy, and few studies use accuracy and network complexity as optimization objectives to estimate model weight parameters, resulting in overfitting of the model and weak generalization ability. In addition, there are excessive connections in large BiLSTM networks, and the model weight parameters are redundant and the computational overhead is too high. Summary of the invention

[0004] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a multi-objective coal product demand forecasting method based on sparse BiLSTM.

[0005] The technical solution adopted by the invention is: a multi-objective coal product demand forecasting method based on sparse BiLSTM, the technical key points of which are as follows:

[0006] Step 1: Download the daily demand data of coal products from the coal mining enterprise database and convert it into a time series {y1, y2, ..., y t};

[0007] Step 2: Clean and normalize the daily demand data of coal products, and then divide it into a training set and a test set;

[0008] The cleaning method of coal product daily demand data is:

[0009] In order to achieve high-precision forecast of demand, the 3σ standard deviation method is used to remove abnormal values ​​of demand data, and the missing values ​​of demand in the middle period are filled and replaced with 1 / 2 of the demand values ​​of two adjacent periods to obtain the cleaned coal product demand data.

[0010] The normalization method for the daily demand data of cleaned coal products is:

[0011] Set y max and min is a sequence {y1, y2, ..., y t}, then we can get y based on the following method: l The normalized sequence corresponding to (l=1,2,...,t) is:

[0012]

[0013] Step 3: Plot the cleaned daily demand data of coal products into a time series graph, and determine the cycle length s of the coal product demand data based on the fixed time period in which the daily demand data of coal products repeatedly reaches the peak value and / or the valley value in the graph;

[0014] Step 4: Initially generate the coal product demand time series data cycle length multiple l = 0, the maximum multiple of the cycle length is n s ;

[0015] Step 5: Execute l=l+1, if l>n s , go to step 7; otherwise, go to step 6;

[0016] Step 6: For the period length l×s of the coal product demand time series data, the sparse BiLSTM model based on the improved multi-objective differential evolution (Multi-objective Differential Evolution with Novel Mutation and Crossover Strategies, referred to as MCMODE) algorithm is used to fit the coal product demand data on the training set, thereby obtaining the fitted value and corresponding residual of the coal product demand of the current period length;

[0017] Step 6.1: Set the number of nodes of the BiLSTM network, the value range of the weight matrix elements to be estimated, and the sparse thresholds of the corresponding parameters;

[0018] (1) The number of input layer nodes m1, the number of hidden layer nodes m2, and the number of output layer nodes m3 of the forward LSTM and reverse LSTM are set based on actual operating conditions;

[0019] (2) The weight matrix W of the forward LSTM is set based on the fitting model accuracy and network complexity.i , W o , W f , W c , U i , U o , U f , U c , W out , corresponding to the weight matrix with the same meaning as the reverse LSTM And the weighted weights W of the forward LSTM and backward LSTM prediction results egt = the range of elements in [w1,w2];

[0020] (3) Input layer vector x based on forward LSTM and reverse LSTM t The four types of input content are composed of the weight matrix W i , W o , W f , W c and Each row of is divided into four parts for sparseness, and the sparse domain of each part is set to α, α′, α″ and α″′ respectively. At the same time, the weight matrix U i , U o , U f , U c and The sparse threshold is β, W out and The sparse threshold of is λ;

[0021] Determine the input layer vector x t The four types of input content are composed of:

[0022] Based on the four-part structure of the Seasonal Autoregressive Integrated Moving Average (SARIMA) model (such as non-seasonal autoregressive part, non-seasonal moving average part, seasonal autoregressive part, seasonal moving average part), the input layer vector x tThe input of is automatically divided into four parts, namely: the part of the r previous period observations of the demand time series data in the non-seasonal component, the part of the corresponding r1 residual items in the non-seasonal component, the part of the Lc×r′ previous observations in the seasonal component and the part of the corresponding Lc1×r1′ residual items in the seasonal component; where Lc and r′ represent the number of lag periods of the previous observations on the seasonal component and the lag order of each lag period, respectively; Lc1 and r1′ represent the number of lag periods of the residual items on the seasonal component and the lag order of each lag period, respectively; where the specific form of the SARIMA model is as follows

[0023]

[0024] in They represent the non-seasonal autoregressive part, non-seasonal moving average part, seasonal autoregressive part and seasonal moving average part of the SARIMA model respectively; p, q, P and Q represent the orders of the corresponding parts respectively; s represents the cycle length;

[0025] Step 6.2: For the time series period length l×s, determine the optimal weight parameters of BiLSTM based on the MCMODE algorithm and sparse strategy;

[0026] Step 6.3: Fit the demand data based on the optimal BiLSTM model of the current cycle length on the training set and then perform denormalization to obtain a series of fitting values ​​and corresponding residuals, and go to step 5;

[0027] Step 7: Substitute the above obtained n s The demand fitting values ​​of different cycle lengths are added together to obtain the fitting results of the BiLSTM model based on the superposition of multiple cycle lengths; the specific method is:

[0028] Set the period lengths to s, 2s, ..., n respectively s ×s, the fitting values ​​obtained in the same period are added together to obtain the final fitting result;

[0029] Compared with the traditional BiLSTM model, the BiLSTM model based on period superposition can better process samples with fewer special maximum or minimum value data. In addition to being small in quantity, the time intervals between occurrences of such extreme value data are large and irregular. By superimposing the fitting values ​​obtained from different periods, the probability of such data information being used can be increased without losing the fitting accuracy of other data, thereby enhancing the overall model's prediction ability for such extreme value data.

[0030] Step 8: Check whether the fitting result of the final multi-cycle length superposition BiLSTM model meets the requirements: if yes, go to step 9; if not, go to step 3;

[0031] Step 9: Obtain the final forecast value of coal product demand based on BiLSTM with multiple cycle lengths superimposed on the test set

[0032] In the above scheme, the specific method for determining the optimal weight parameters of BiLSTM based on the MCMODE algorithm and sparse strategy described in step 6.2 is:

[0033] Step 6.2.1: Initialize the target weight parameter value scheme population, set the population size in the MCMODE algorithm to NP, the total number of parameters of the scheme individual to M, the mutation rate of the scheme individual gene to F, the crossover rate of genes between different schemes to CR, the initial generation number G = 0, and the maximum generation number G max , External Archives

[0034] The method for generating the population of initial weight parameter value solutions based on a combination of sparse and random strategies is as follows:

[0035] According to the W in the BiLSTM model i , W o , W f , W c , U i , U o , U f , U c , W out , and W egt The elements in these matrices are encoded in the order of rows first and columns later to obtain an initial weight parameter value scheme in the population. in represents the encoding value of the hth element to be optimized in the BiLSTM model, g = 1, 2, ..., NP, h = 1, 2, ..., M; The value range is from Uniform extraction, and Respectively represent The upper and lower limits of the value, x h yes The sparse threshold of The weight parameter value scheme individuals generated by the sparse strategy can reduce the algorithm's search space while ensuring the quality of individuals in the population, thereby accelerating the convergence of the MCMODE algorithm.

[0036] Step 6.2.2: Determine the individual weight parameter value scheme based on the sparse strategy The objective function is calculated and its fitness value is calculated, and then the initial external archive set Arch_Set and the best solution individual in the population are obtained;

[0037] Step 6.2.3: Implement G=G+1, if G>G max , go to step 6.2.8; otherwise, go to step 6.2.4;

[0038] Step 6.2.4: Scheme for selecting the value of the g-th weight parameter Randomly select one weight parameter value scheme in the external archive set Arch_Set and randomly select two weight parameter value schemes with different numbers and not g in the current population to combine them, and get Corresponding mutation weight parameter value scheme plan The way to obtain is as follows:

[0039]

[0040] Where F has mean μ and variance σ 2 The normal distribution N(μ,σ 2 ); Rd is a random number generated by uniform distribution [0, 1]; Arch represents the index of the solution individual randomly selected from the external archive set Arch_Set; r1 and r2 are the indexes of the solution individual randomly selected from the current population, and the indexes satisfy r1≠r2≠g; Represents the best solution individual in Arch_Set, which is determined by identifying the inverse point, that is, using all the weight parameter value solution individuals in Arch_Set to construct the Pareto frontier, determine a straight line based on the two boundary points in the frontier, calculate the distance from each point in the frontier to the straight line, and the point corresponding to the maximum distance is In addition, for invalid encoding values ​​generated after performing mutation operations Using the interval The code is repaired by randomly generating;

[0041] Step 6.2.5: Target solution and mutation scheme Performing a binary crossover operation based on the Cauchy distribution yields The corresponding crossover scheme plan The way to obtain is as follows:

[0042]

[0043] where CR follows the Cauchy distribution C(γ,x0) with scale parameter γ and location parameter x0; rand h (0,1) represents a real number randomly drawn from the interval [0, 1];

[0044] Step 6.2.6: Calculate the individual solutions and The objective function fitness value of is selected based on a greedy approach as the target solution individual in the next generation population; the method for obtaining the target solution individual is as follows:

[0045]

[0046] Step 6.2.7: Compare the target solution individuals in the new generation population with the solution individuals in the external archive set Arch_Set, and update Arch_Set based on the dominance relationship between the solutions;

[0047] The method of updating Arch_Set by using the dominant relationship between weight parameter value schemes is described as follows:

[0048] Compare each new solution in the population with the existing solutions in Arch_Set, and update Arch_Set based on the dominance relationship between the two: 1) If any new solution dominates a solution in Arch_Set, then use the new solution to replace the original dominated solution in Arch_Set; 2) If any new solution dominates a group of solutions in Arch_Set, then add the new solution to Arch_Set and delete the group of dominated solutions from Arch_Set; 3) If any new solution is dominated by the solutions in Arch_Set, then ignore this solution; After all new solutions are compared with the existing solutions in Arch_Set, the optimal non-dominated weight parameter value solution set up to the current generation is obtained, and then the update of Arch_Set is completed;

[0049] Step 6.2.8: If the current generated algebra G>G max , the algorithm terminates, outputs the optimal weight parameter value scheme of the BiLSTM model with the current cycle length l×s, and then obtains the optimal BiLSTM model; otherwise, go to step 6.2.3.

[0050] In the above scheme, the specific method of step 6.2.2 is:

[0051] For any scheme in the population Calculate the fitness value of its objective function, and then determine the non-dominated solution individuals based on the Pareto dominance principle and store them in Arch_Set. When obtaining the initial Arch_Set, the best weight parameter value scheme in the initial population is obtained. The specific method of the objective function is as follows:

[0052] To achieve high-precision prediction of the BiLSTM model, it is necessary to balance the fitting accuracy and generalization ability of the model. Therefore, the model accuracy and network complexity are considered as the objective functions for the individual weight parameter value assignment scheme, and the fitness value of the individual is evaluated by minimizing these two objective functions. The specific objective functions (6) and (7) are expressed as follows: The fitness value of the individual is evaluated by minimizing these two objective functions. The specific objective functions (6) and (7) are expressed as follows:

[0053]

[0054] where z t represents the true value of the coal product demand; n is the number of training samples; represents the predicted value of the coal product demand obtained based on the weight parameter value assignment scheme ; is a binary variable determined based on : when , otherwise,

[0055] The specific method for determining the objective function and the number of non-zero elements based on the sparse strategy is as follows:

[0056] During the process of optimizing the BiLSTM model parameters, for each weight parameter value assignment scheme generated by MCMODE (such as: W i , W o , W f , W c , U i , U o , U f , U c , W out , and W egt element combinations), perform the following operations:

[0057] 1) For the matrix If the elements of the weight matrix corresponding to the previous observed values of the input demand time series data in the non-seasonal component part in the vector x t satisfy |w u,v | < α (i.e., ) and v ≤ r, set and Otherwise, set If the elements of the weight matrix corresponding to the input residual term in the non-seasonal component part in the vector x t satisfy |w u,v | < α′ (i.e., ) and r < v ≤ r + r1, set and Otherwise, set If the vector x t the weight matrix element corresponding to the previous observation value input in the seasonal component part satisfies |w u,v | < α″ (i.e.: ) and r + r1 < v ≤ r + r1 + Lc × r′, set and Otherwise, set If the weight matrix element corresponding to the residual term input in the vector x t in the seasonal component part satisfies |w u,v | < α″′ (i.e.: ) and r + r1 + Lc × r′ < v ≤ m2, set and Here, w u,v represents the element in the u-th row and v-th column of the matrix W i , u ∈ {1, 2,..., m1}, v ∈ {1, 2,..., m2}, and the sparsity patterns of W o , W f , W c and are the same as that of W i ;

[0058] 2) For the matrix , if |u u′,v′ | < β (i.e.: ), set Otherwise, set The sparsity patterns of U o , U f , and are the same as that of U i , where u u′,v′ represents the element in the u′-th row and v′-th column of U i , u′ ∈ {1, 2,..., m2}, v′ ∈ {1, 2,..., m2}, ω u″,v″ represents the element in the u″-th row and v″-th column of W out , u″ ∈ {1, 2,..., m2}, v″ ∈ {1, 2,..., m3}. Here, to ensure that both the forward and backward LSTM information can be effectively utilized, W egt is not sparse. It should be noted that if the weight matrix becomes a zero matrix after executing the sparsity strategy, then perform the following operations to ensure the effectiveness of the BiLSTM model: In the forward LSTM model, randomly select any row index u″ and column index v″ in W out , and repeatedly and randomly generate the corresponding element ω u″,v″ based on the indices u″ and v″The value of until the condition |ω u″,v″ |≥λ is satisfied; in the reverse LSTM model, if the weight matrix is ​​executed after the sparse strategy becomes a zero matrix, then Execution and W out Same operation.

[0059] The beneficial effects of the present invention are as follows: the present invention provides a multi-objective coal product demand forecasting method based on sparse BiLSTM, which makes full use of the characteristics of the BiLSTM model that it is good at capturing complex nonlinear patterns of data, develops an MCMODE algorithm with model accuracy and network complexity as the goals to optimize the weight parameters of BiLSTM, and designs a method for superimposing multi-cycle length prediction results to add BiLSTM prediction values, thereby ensuring the prediction accuracy and generalization ability of the model, and at the same time designs a sparse strategy to compress the parameter scale of BiLSTM, thereby reducing the computational cost of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0061] Figure 1 A framework diagram of the BiLSTM prediction model provided in an embodiment of the present invention;

[0062] Figure 2 A model unit structure diagram of the forward LSTM and the reverse LSTM in the BiLSTM provided in an embodiment of the present invention;

[0063] Figure 3 A flowchart of a multi-objective coal product demand forecasting method based on sparse BiLSTM provided in an embodiment of the present invention;

[0064] Figure 4 This is a flowchart of the MCMODE algorithm provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0065] In order to make the above-mentioned objects, features and advantages of the present invention more clearly understood, the following is a brief description of the present invention in conjunction with the attached drawings. Figure 1-Figure 4 The present invention is further described in detail with reference to the accompanying drawings and specific embodiments.

[0066] This implementation takes the daily sales data of coal products of a large domestic coal enterprise as an example, and adopts a multi-objective coal product demand prediction method based on sparse BiLSTM to obtain the forecast of daily demand of coal mines, providing guidance for the design of the enterprise's coal mining plan and the formulation of production plans.

[0067] (1) BiLSTM model

[0068] In this embodiment, the BiLSTM prediction model consists of a forward LSTM model and a reverse LSTM model. Figure 1 Among them, the forward LSTM model and the reverse LSTM model are composed of an LSTM unit structure and a fully connected layer, please refer to Figure 2 In a forward LSTM unit structure, let the sigmoid function S and the hyperbolic tangent function tanh represent the activation function. Then at time t, the forget gate f t Determines what information from the previous period is retained in the cell state C t Middle; input gate i t Determine from the input layer x t What new information is learned and stored in C t In t Information stored in C t Previously, an additional layer based on the tanh activation function was required to relate x t To generate candidate states Output Gate O t Decision C t The specific information transmitted and output in O t and C t , calculate the hidden state h t And then get the final output y of the forward LSTM model t , the specific calculation process is described as follows:

[0069] f t =S(W f × t +U f ×h t-1 ) (1)

[0070] i t =S(W i × t +U i ×h t-1 ) (2)

[0071]

[0072] O t =S(W o × t +Uo ×h t-1 ) (5)

[0073] h t =O t ×tanh(C t ) (6)

[0074] y t =W out ×h t (7)

[0075] Among them, W f , W i , W c and W o They are connected to the input layer x at time t. t and forget gate f t , input gate i t , Candidate Status And the output gate O t The weight matrix between f , U i , U c and U o They are connected to the hidden state h at time t-1 respectively. t-1 and the forget gate f at time t t , input gate i t , Candidate Status And the output gate O t The weight matrix between out is the hidden state h at time t of the connection t and the final prediction y t The weight matrix of t-1 is the cell state at time t-1; the form of the reverse LSTM model is the same as the forward LSTM model, where With the same W as in the forward LSTM model i , W o , W f , W c , U i , U o , U f , U c , W out Corresponding to the same meaning.

[0076] Let the final output of the reverse LSTM model be y t ′,W egt =[w1,w2] is the weighted vector of the prediction results of the forward LSTM model and the reverse LSTM model, then the output result of a BiLSTM model is for:

[0077]

[0078] In this embodiment, a multi-objective coal product demand forecasting method based on sparse BiLSTM is provided. Figure 2 , including the following steps:

[0079] Step 1: Download the daily demand data of coal products from the coal mining enterprise database and convert it into a time series {y1, y2, ..., y t};

[0080] Step 2: Clean and normalize the daily demand data of coal products, and then divide it into a training set and a test set;

[0081] The cleaning method of coal product daily demand data is:

[0082] In order to achieve high-precision forecast of demand, the 3σ standard deviation method is used to remove outliers in the demand data. The missing values ​​of the demand in the middle period are filled and replaced with 1 / 2 of the demand values ​​in two adjacent periods to obtain the cleaned coal product demand data.

[0083] The normalization method for the daily demand data of cleaned coal products is:

[0084] Set y max and min is a sequence {y1, y2, ..., y t}, then we can get y based on the following method: l The normalized sequence corresponding to (l=1,2,...,t) is:

[0085]

[0086] Step 3: Plot the cleaned daily demand data of coal products into a time series graph, and determine the cycle length s of the coal product demand data based on the fixed time period in which the daily demand data of coal products repeatedly reaches the peak value and / or the valley value in the graph;

[0087] Step 4: Initially generate the coal product demand time series data cycle length multiple l = 0, the maximum multiple of the cycle length is n s ;

[0088] In this embodiment, n s =2.

[0089] Step 5: Execute l=l+1, if l>n s , go to step 7; otherwise, go to step 6;

[0090] Step 6: For the coal product demand time series data with a period length of l×s, the coal product demand data is fitted using the sparse BiLSTM model based on MCMODE on the training set, thereby obtaining the fitted value and corresponding residual of the coal product demand of the current period length;

[0091] Step 6.1: Set the number of nodes of the BiLSTM network, the value range of the weight matrix elements to be estimated, and the sparse thresholds of the corresponding parameters;

[0092] (1) The number of input layer nodes m1, the number of hidden layer nodes m2, and the number of output layer nodes m3 of the forward LSTM and reverse LSTM are set based on actual operating conditions;

[0093] In this embodiment, the number of nodes in the input layer, hidden layer, and output layer of the forward LSTM and the backward LSTM are set to m1=12, m2=3, and m3=1, respectively.

[0094] (2) The weight matrix W of the forward LSTM is set based on the fitting model accuracy and network complexity. i , W o , W f , W c , U i , U o , U f , U c , W out , corresponding to the weight matrix with the same meaning as the reverse LSTM And the weighted weights W of the forward LSTM and backward LSTM prediction results egt = the range of elements in [w1,w2];

[0095] In this embodiment, the weight matrix W is set i , W o , W f , W c , U i , U o , U f , U c , W out , and W egt The value range of the elements is [-1, 1].

[0096] (3) Input layer vector x based on forward LSTM and reverse LSTM t The four types of input content are composed of the weight matrix W i , W o , W f , W c and Each row of is divided into four parts for sparseness, and the sparse domain of each part is set to α, α′, α″ and α″′ respectively. At the same time, the weight matrix U i , U o , U f , U c and The sparse threshold is β, W out and The sparse threshold of is λ;

[0097] Determine the input layer vector x t The four types of input content are composed of:

[0098] Based on the four-part structure of the SARIMA model (such as non-seasonal autoregressive part, non-seasonal moving average part, seasonal autoregressive part, seasonal moving average part), the input layer vector x t The input is automatically divided into four parts, namely: the part of the r previous period observations of the demand time series data in the non-seasonal component, the part of the corresponding r1 residual items in the non-seasonal component, the part of the Lc×r′ previous observations in the seasonal component, and the part of the corresponding Lc1×r1′ residual items in the seasonal component. Among them, Lc and r′ represent the number of lag periods of the previous observations on the seasonal component and the lag order of each lag period; Lc1 and r1′ represent the number of lag periods of the residual items on the seasonal component and the lag order of each lag period. The specific form of the SARIMA model is as follows

[0099]

[0100] in They represent the non-seasonal autoregressive part, non-seasonal moving average part, seasonal autoregressive part and seasonal moving average part of the SARIMA model respectively; p, q, P and Q represent the orders of the corresponding parts respectively; s represents the cycle length.

[0101] In this embodiment, the parameters are set as r=5, r1=2, Lc=2, r′=2, Lc1=1, r1′=1; the sparse thresholds are α=0.1, α′=0.15, α″=0.1, α″′=0.1, β=0.05 and λ=0.1; and the cycle length is s=13.

[0102] Step 6.2: Determine the optimal weight parameters of BiLSTM based on the MCMODE algorithm and sparse strategy for the time series period length l×s; see Figure 3 , the specific method is:

[0103] Step 6.2.1: Initialize the target weight parameter value scheme population, set the population size in the MCMODE algorithm to NP, the total number of parameters of the scheme individual to M, the mutation rate of the scheme individual gene to F, the crossover rate of genes between different schemes to CR, the initial generation number G = 0, and the maximum generation number G max , External Archives

[0104] The method for generating the population of initial weight parameter value solutions based on a combination of sparse and random strategies is as follows:

[0105] According to the W in the BiLSTM model i , W o , W f , W c , U i , U o , U f , U c , W out , and W egt The elements in these matrices are encoded in the order of rows first and columns later to obtain an initial weight parameter value scheme in the population. in Represents the encoding value of the hth element that needs to be optimized in the BiLSTM model, g = 1, 2, ..., NP, h = 1, 2, ..., M. The value range is from Uniform extraction, and Respectively represent The upper and lower limits of the value, x h yes The sparse threshold of The weight parameter value scheme individuals generated by the sparse strategy can reduce the search space of the algorithm while ensuring the quality of individuals in the population, thereby accelerating the convergence of the MCMODE algorithm.

[0106] In this embodiment, NP=600, M=392, G max =1000, F obeys the normal distribution F~N(0,1) with mean 0 and variance 1, and CR obeys the Cauchy distribution CR~C(0.1,0) with scale parameter 0.1 and location parameter 0.

[0107] Step 6.2.2: Determine the individual weight parameter value scheme based on the sparse strategy The objective function is calculated and its fitness value is calculated, and then the initial external archive set Arch_Set and the best solution individual in the population are obtained;

[0108] For any scheme in the population Calculate the fitness value of its objective function, and then determine the non-dominated solution individuals based on the Pareto dominance principle and store them in Arch_Set. When obtaining the initial Arch_Set, the best weight parameter value solution in the initial population is obtained. The specific method of the objective function is as follows:

[0109] In order to achieve high-precision prediction of the BiLSTM model, it is necessary to balance the model's fitting accuracy and generalization ability. Therefore, considering the model accuracy (such as root mean square error ) and network complexity (e.g., the number of non-zero elements in the weight matrix of the model ) as the objective function of the individual weight parameter selection scheme, and evaluate the individual by minimizing these two objective functions. The specific objective functions (11) and (12) are expressed as follows:

[0110]

[0111] where z t represents the true value of coal product demand; n is the number of training samples; Represents a weight parameter value scheme The forecast value of coal product demand is obtained; is a Determined binary variables: When hour, otherwise,

[0112] The specific method for determining the objective function and the number of non-zero elements based on the sparse strategy is:

[0113] In the process of optimizing BiLSTM model parameters, for each weight parameter value scheme generated by MCMODE (such as: W i , W o , W f , W c , U i , U o , U f , U c , W out , and W egt ), do the following:

[0114] 1) For the matrix If the vector x t The weight matrix elements corresponding to the previous observations of the input demand time series data in the non-seasonal component satisfy |w u,v |<α(ie: ) and v ≤ r, set and Otherwise, set If the weight matrix element corresponding to the residual term input in vector x t in the non-seasonal component part satisfies |w u,v | < α′ (i.e.: ) and r < v ≤ r + r1, set and Otherwise, set If the weight matrix element corresponding to the previous observation input in vector x t in the seasonal component part satisfies |w u,v | < α″ (i.e.: ) and r + r1 < v ≤ r + r1 + Lc × r′, set and Otherwise, set If the weight matrix element corresponding to the residual term input in vector x t in the seasonal component part satisfies |w u,v | < α″′ (i.e.: ) and r + r1 + Lc × r′ < v ≤ m2, set and Here, w u,v represents the element in the u-th row and v-th column of matrix W i , u ∈ {1, 2,..., m1}, v ∈ {1, 2,..., m2}. The sparsity patterns of W o , W f , W c and are the same as that of W i .

[0115] 2) For matrix If |u u′,v′ | < β (i.e.: ), set Otherwise, set U o , U f , U c , and The sparsity patterns of are the same as that of U i . Here, u u′,v′ represents the element in the u′-th row and v′-th column of U i , u′ ∈ {1, 2,..., m2}, v′ ∈ {1, 2,..., m2}, ω u″,v″ represents W outThe u″th row and v″th column element of , u″∈{1,2,...,m2}, v″∈{1,2,...,m3}. Here, in order to ensure that both forward and reverse LSTM information can be effectively utilized, W egt Not sparse. It is worth noting that if the weight matrix is becomes a zero matrix, then perform the following operations to ensure the effectiveness of the BiLSTM model: In the forward LSTM model, randomly extract W out For any row index u″ and column index v″ in the , the corresponding element ω is repeatedly randomly generated based on the index u″ and v″ u″,v″ The value of until the condition |ω u″,v″ |≥λ is satisfied; in the reverse LSTM model, if the weight matrix is ​​executed after the sparse strategy becomes a zero matrix, then Execution and W out Same operation.

[0116] In this embodiment, n=462 is set.

[0117] Step 6.2.3: Implement G=G+1, if G>G max , go to step 6.2.8; otherwise, go to step 6.2.4;

[0118] Step 6.2.4: Scheme for selecting the value of the g-th weight parameter Randomly select one weight parameter value scheme in the external archive set Arch_Set and randomly select two weight parameter value schemes with different numbers and not g in the current population to combine them, and get Corresponding mutation weight parameter value scheme plan The way to obtain is as follows:

[0119]

[0120] Where F has mean μ and variance σ 2 The normal distribution N(μ,σ 2 ); Rd is a random number generated by uniform distribution [0, 1]; Arch represents the index of the solution individual randomly selected from the external archive set Arch_Set; r1 and r2 are the indexes of the solution individual randomly selected from the current population, and the indexes satisfy r1≠r2≠g; Represents the best solution individual in Arch_Set, which is determined by identifying the inverse point, that is, using all the weight parameter value solution individuals in Arch_Set to construct the Pareto frontier, determine a straight line based on the two boundary points in the frontier, calculate the distance from each point in the frontier to the straight line, and the point corresponding to the maximum distance is In addition, for invalid encoding values ​​generated after performing mutation operations Using the interval The code is repaired in a randomly generated manner.

[0121] Step 6.2.5: Target solution and mutation scheme Performing a binary crossover operation based on the Cauchy distribution yields The corresponding crossover scheme plan The way to obtain is as follows:

[0122]

[0123] where CR follows the Cauchy distribution C(γ,x0) with scale parameter γ and location parameter x0; rand h (0,1) represents a real number randomly drawn from the interval [0, 1].

[0124] Step 6.2.6: Calculate the individual solutions and The objective function fitness value of is selected based on a greedy approach as the target solution individual in the next generation population; the method for obtaining the target solution individual is as follows:

[0125]

[0126] Step 6.2.7: Compare the target solution individuals in the new generation population with the solution individuals in the external archive set Arch_Set, and update Arch_Set based on the dominance relationship between the solutions;

[0127] The method of updating Arch_Set by using the dominant relationship between weight parameter value schemes is described as follows:

[0128] Compare each new solution in the population with the existing solutions in Arch_Set, and update Arch_Set based on the dominance relationship between the two: 1) If any new solution dominates a solution in Arch_Set, then use the new solution to replace the original dominated solution in Arch_Set; 2) If any new solution dominates a group of solutions in Arch_Set, then add the new solution to Arch_Set and delete the group of dominated solutions from Arch_Set; 3) If any new solution is dominated by a solution in Arch_Set, then ignore this solution. After all new solutions are compared with existing solutions in Arch_Set, the optimal non-dominated weight parameter value solution set up to the current generation is obtained, and then the update of Arch_Set is completed.

[0129] Step 6.2.8: If the current generated algebra G>G max , the algorithm terminates, outputs the optimal weight parameter value scheme of the BiLSTM model with the current cycle length l×s, and then obtains the optimal BiLSTM model; otherwise, go to step 6.2.3.

[0130] Step 6.3: Fit the demand data based on the optimal BiLSTM model of the current cycle length on the training set and then perform denormalization to obtain a series of fitting values ​​and corresponding residuals, and go to step 5;

[0131] Step 7: Substitute the above obtained n s The demand fitting values ​​of different cycle lengths are added together to obtain the fitting results of the BiLSTM model based on the superposition of multiple cycle lengths; the specific method is:

[0132] Set the period lengths to s, 2s, ..., n respectively s The fitting values ​​obtained by ×s are added in the same period to obtain the final fitting result.

[0133] Compared with the traditional BiLSTM model, the BiLSTM model based on period superposition can better handle samples with fewer special maximum or minimum data. In addition to being small in quantity, such extreme value data also appear at large and irregular intervals. By superimposing the fitting values ​​obtained in different periods, the probability of such data information being used can be increased without losing the fitting accuracy of other data, thereby enhancing the overall model's ability to predict such extreme value data.

[0134] Step 8: Check whether the fitting result of the final multi-cycle length superposition BiLSTM model meets the requirements: if yes, go to step 9; if not, go to step 3;

[0135] Step 9: Obtain the final forecast value of coal product demand based on BiLSTM with multiple cycle lengths superimposed on the test set

[0136] In this implementation, based on the daily sales data of coal products of a large domestic coal enterprise, the method of the present invention is used to predict the demand for coal products in the next 60 days. The prediction accuracy comparison results of sparse BiLSTM models based on different cycle lengths are shown in Table 1. According to the table, all prediction performance indicators of the BiLSTM model superimposed by different cycle lengths are better than the performance indicators of the BiLSTM model with a single cycle length, thereby verifying the effectiveness of the method of the present invention.

[0137] Table 1 Comparison of prediction accuracy of sparse BiLSTM models based on different cycle lengths

[0138] Data cycle length Mean absolute error Mean absolute percentage error (%) Root mean square error 13 20097.7 21.9886 28922 26 21131.1 24.45 30566.2 13_26 19103.93 21.91 27838.03

[0139] Note: 13_26 indicates that the prediction results are obtained by superimposing the method of the present invention when the cycle lengths are 13 and 26 respectively; the values ​​shown in bold indicate the best results obtained based on the models with different cycle lengths.

[0140] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A multi-objective coal product demand forecasting method based on sparse BiLSTM, characterized in that: The following steps are involved: Step 1: Download the daily demand data of coal products from the coal mining enterprise database and convert it into a time series {y1, y2, ..., y t }; Step 2: Clean and normalize the daily demand data of coal products, and then divide it into a training set and a test set; Step 3: Plot the cleaned daily demand data of coal products into a time series graph, and determine the cycle length s of the coal product demand data based on the fixed time period in which the daily demand data of coal products repeatedly reaches the peak value and / or the valley value in the graph; Step 4: Initially generate the coal product demand time series data cycle length multiple l = 0, the maximum multiple of the cycle length is n s ; Step 5: Execute l=l+1, if l>n s , go to step 7; otherwise, go to step 6; Step 6: For the coal product demand time series data with a period length of l×s, the coal product demand data is fitted using the sparse BiLSTM model based on MCMODE on the training set, thereby obtaining the fitted value and corresponding residual of the coal product demand of the current period length; Step 6.1: Set the number of nodes of the BiLSTM network, the value range of the weight matrix elements to be estimated, and the sparse thresholds of the corresponding parameters; (1) The number of input layer nodes m1, the number of hidden layer nodes m2, and the number of output layer nodes m3 of the forward LSTM and reverse LSTM are set based on actual operating conditions; (2) The weight matrix W of the forward LSTM is set based on the fitting model accuracy and network complexity. i , W o , W f , W c , U i , U o , U f , U c , W out , corresponding to the weight matrix with the same meaning as the reverse LSTM And the weighted weights W of the forward LSTM and backward LSTM prediction results egt = the range of elements in [w1,w2]; (3) Input layer vector x based on forward LSTM and reverse LSTM t The four types of input content are composed of the weight matrix W i , W o , W f , W c and Each row of is divided into four parts for sparseness, and the sparse domain of each part is set to α, α′, α″ and α″′ respectively. At the same time, the weight matrix U is set i , U o , U f , U c and The sparse threshold is β, W out and The sparse threshold of is λ; Determine the input layer vector x t The four types of input content are composed of: Based on the four-part structure of the seasonal autoregressive integrated moving average model, the input layer vector x t The input is automatically divided into four parts, namely: the part of the r previous period observations of the demand time series data in the non-seasonal component, the part of the corresponding r1 residual items in the non-seasonal component, the part of the Lc×r′ previous observations in the seasonal component and the part of the corresponding Lc1×r1′ residual items in the seasonal component, where Lc and r′ represent the number of lag periods of the previous observations on the seasonal component and the lag order of each lag period, respectively; Lc1 and r1′ represent the number of lag periods of the residual items on the seasonal component and the lag order of each lag period, respectively; where the specific form of the SARIMA model is as follows in and They represent the non-seasonal autoregressive part, non-seasonal moving average part, seasonal autoregressive part and seasonal moving average part of the SARIMA model respectively; p, q, P and Q represent the orders of the corresponding parts respectively; s represents the cycle length; Step 6.2: For the time series period length l×s, determine the optimal weight parameters of BiLSTM based on the MCMODE algorithm and sparse strategy; Step 6.3: Fit the demand data based on the optimal BiLSTM model of the current cycle length on the training set and then perform denormalization to obtain a series of fitting values ​​and corresponding residuals, and go to step 5; Step 7: Substitute the above obtained n s The demand fitting values ​​of different cycle lengths are added together to obtain the fitting results of the BiLSTM model based on the superposition of multiple cycle lengths; Step 8: Check whether the fitting result of the final multi-cycle length superposition BiLSTM model meets the requirements: if yes, go to step 9; if not, go to step 3; Step 9: Obtain the final forecast value of coal product demand based on BiLSTM with multiple cycle lengths superimposed on the test set 2. The multi-objective coal product demand forecasting method based on sparse BiLSTM according to claim 1 is characterized in that: The specific method for determining the optimal weight parameters of BiLSTM based on the MCMODE algorithm and sparse strategy described in step 6.2 is: Step 6.2.1: Initialize the target weight parameter value scheme population, set the population size in the MCMODE algorithm to NP, the total number of parameters of the scheme individual to M, the mutation rate of the scheme individual gene to F, the crossover rate of genes between different schemes to CR, the initial generation number G = 0, and the maximum generation number G max , External Archives Step 6.2.2: Determine the individual weight parameter value scheme based on the sparse strategy The objective function is calculated and its fitness value is calculated, and then the initial external archive set Arch_Set and the best solution individual in the population are obtained; Step 6.2.3: Implement G=G+1, if G>G max , go to step 6.2.8; otherwise, go to step 6.2.4; Step 6.2.4: Scheme for selecting the value of the g-th weight parameter Randomly select one weight parameter value scheme in the external archive set Arch_Set and randomly select two weight parameter value schemes with different numbers and not g in the current population to combine them, and get Corresponding mutation weight parameter value scheme Step 6.2.5: Target solution and mutation scheme Performing a binary crossover operation based on the Cauchy distribution yields The corresponding crossover scheme Step 6.2.6: Calculate the individual solutions and The objective function fitness value of is used to select the better individual of the two as the target solution individual in the next generation population in a greedy way; Step 6.2.7: Compare the target solution individuals in the new generation population with the solution individuals in the external archive set Arch_Set, and update Arch_Set based on the dominance relationship between the solutions; Step 6.2.8: If the current generated algebra G>G max , the algorithm terminates, outputs the optimal weight parameter value scheme of the BiLSTM model with the current cycle length l×s, and then obtains the optimal BiLSTM model; Otherwise go to step 6.2.

3.

3. The multi-objective coal product demand forecasting method based on sparse BiLSTM according to claim 2 is characterized in that: The specific method of step 6.2.2 is: For any scheme in the population Calculate the fitness value of its objective function, and then determine the non-dominated solution individuals based on the Pareto dominance principle and store them in Arch_Set. When obtaining the initial Arch_Set, the best weight parameter value scheme in the initial population is obtained. The specific method of the objective function is as follows: In order to achieve high-precision prediction of the BiLSTM model, it is necessary to balance the model's fitting accuracy and generalization ability. Therefore, the model accuracy and network complexity are considered as the objective functions of the individual weight parameter value scheme, and the individual weight parameter value is evaluated by minimizing these two objective functions. The fitness value of , the specific objective functions (2) and (3) are expressed as follows: where z t represents the true value of coal product demand; n is the number of training samples; Represents a weight parameter value scheme The forecast value of coal product demand is obtained; is a Determined binary variables: When hour, otherwise, The specific method for determining the objective function and the number of non-zero elements based on the sparse strategy is: In the process of optimizing BiLSTM model parameters, for each weight parameter value scheme generated by MCMODE, perform the following operations: 1) For the matrix If for the vector x t the weight matrix element corresponding to the previous observation value of the input demand time series data in the non-seasonal component part satisfies |w u,v | < α, that is: and v ≤ r, set and Otherwise, set If for the vector x t the weight matrix element corresponding to the residual term of the input in the non-seasonal component part satisfies |w u,v | < α′, that is: and r < v ≤ r + r1, set and Otherwise, set If for the vector x t the weight matrix element corresponding to the previous observation value of the input in the seasonal component part satisfies |w u,v | < α″, that is: and r + r1 < v ≤ r + r1 + Lc×r′, set and Otherwise, set If for the vector x t the weight matrix element corresponding to the residual term of the input in the seasonal component part satisfies |w u,v | < α″′, that is: and r + r1 + Lc×r′ < v ≤ m2, set and where w u,v represents the element in the u-th row and v-th column of the matrix W i , u ∈ {1, 2, …, m1}, v ∈ {1, 2, …, m2}, and the sparsity patterns of W o , W f , W c and are the same as that of W i ; 2) For the matrix If |u u′,v′ |<β, that is: set up Otherwise, set U o , U f , U c , and The sparse method and U i The same; among them, u u′,v′ Represents U i The u′th row and v′th column element, u′∈{1,2,…,m2}, v′∈{1,2,…,m2}, ω u″,v″ Represents W out The u″th row and v″th column element, u″∈{1,2,…,m2}, v″∈{1,2,…,m3}, to ensure that both forward and reverse LSTM information can be effectively utilized, so W egt Not sparse; if the weight matrix is becomes a zero matrix, then perform the following operations to ensure the effectiveness of the BiLSTM model: In the forward LSTM model, randomly extract W out For any row index u″ and column index v″ in , the corresponding element ω is repeatedly randomly generated based on the index u″ and v″ u″,v″ The value of until the condition |ω u″,v″ |≥λ is satisfied; in the reverse LSTM model, if the weight matrix is ​​executed after the sparse strategy becomes a zero matrix, then Execution and W out Same operation.