A state space-based deep-cone thickener modeling method and system
By adopting a state-space-based deep cone thickener modeling method, and utilizing slicing operations and the Mamba module to process long sequence data, the problem of long-term dependency capture difficulties in neural networks in deep cone thickeners is solved, and efficient, accurate prediction and stable paste filling process are achieved.
Patent Information
- Application Number
- CN202411656449.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-19
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-11-19
AI Technical Summary
Existing neural network-based methods struggle to effectively capture long-term dependencies in deep cone thickeners, leading to problems such as low prediction accuracy, long prediction times, and wasted resources.
A state-space-based deep cone denser modeling method is adopted. By acquiring historical operating data for preprocessing, a neural network model of a sharding operation module, a Mamba module, and a prediction unit module is constructed. The linear characteristics of the state-space model and the feedforward network layer are used to process long sequence data, capture long-term dependencies, and perform training.
It improves the accuracy and efficiency of deep cone thickener prediction, enhances the model's generalization ability under different working conditions, and optimizes the stability and efficiency of paste filling process.
Smart Images

Figure CN119670535B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a deep-cone thickener modeling method and system based on state space. BACKGROUND
[0002] The deep-cone thickener is a device for solid-liquid separation, widely used in mining, metallurgy, chemical industry and other industries. Its main function is to concentrate the suspension containing solid particles into high-concentration underflow, while minimizing the liquid in the underflow. The deep-cone thickener is also a device for preparing paste. A deep-cone thickener modeling method based on state space is a technical means for mathematical modeling and prediction of deep-cone thickener, which combines the theory of state space model to describe and control the dynamic behavior of deep-cone thickener.
[0003] With the development of the times, the scale of human underground mining of minerals has also expanded, resulting in an increasing number of underground cavities. At the same time, a large amount of tailings left on the surface after mining has caused serious environmental pollution. In order to solve these problems, paste backfill technology is adopted, and the deep-cone thickener is a device for preparing paste. Since the deep-cone thickener plays a crucial role in the paste filling process in the mining field, its performance is directly related to the efficiency and stability of the entire filling system. Therefore, modeling the deep-cone thickener is of great significance to improving the stability, efficiency, and reliability of the thickener, industrial production, and environmental protection.
[0004] In the early stage, people tried to predict the underflow concentration manually, but manual prediction had great subjectivity and uncertainty, resulting in great differences and uncertainty in the prediction results. With the development of simulation technology, researchers have achieved process control and optimization through simulation technology. However, due to the unobservability of the thickener system and the complexity of its internal flow, involving multiphase flow and numerous physical and chemical actions, it is actually very challenging to rely solely on mathematical models to simulate and predict parameter changes under actual operating conditions. Although some scholars have used neural network methods based entirely on data to predict key operating parameters of the thickener, most of the methods based on neural network learning are difficult to capture long-term dependencies in industrial systems. Some transformer-like models can capture long-term dependencies, but their self-attention mechanisms often cause high time complexity when facing long time series in industrial systems, resulting in low prediction accuracy, long prediction time, and low prediction efficiency, causing resource waste. SUMMARY
[0005] In order to solve the technical problems that most of the methods based on neural network learning in the prior art are difficult to capture long-time dependence in industrial systems, some transformer model can capture long-time dependence, but the self-attention mechanism of the transformer model often causes great time complexity in the face of long-time sequence of industrial systems, resulting in low prediction accuracy, long prediction time and low prediction efficiency, and waste of resources, the present application provides a state space-based deep cone thickener modeling method and system.
[0006] The technical scheme provided by the embodiments of the present application is as follows:
[0007] The first aspect
[0008] The state space-based deep cone thickener modeling method provided by the embodiments of the present application comprises:
[0009] S1: obtaining historical operation data of the deep cone thickener;
[0010] S2: preprocessing the historical operation data;
[0011] S3: constructing a deep cone thickener prediction model based on a neural network under the framework of a state space model, wherein the deep cone thickener prediction model comprises a slicing operation module, a Mamba module and a prediction unit module;
[0012] S4: inputting the preprocessed historical operation data as a training set into the deep cone thickener prediction model to train the deep cone thickener prediction model until the loss function value of the deep cone thickener prediction model is less than a preset loss function value;
[0013] S5: outputting the trained deep cone thickener prediction model to complete the modeling of the deep cone thickener.
[0014] The second aspect
[0015] The state space-based deep cone thickener modeling system provided by the embodiments of the present application comprises:
[0016] A processor;
[0017] A memory, the memory stores computer readable instructions, and the computer readable instructions are executed by the processor to realize the state space-based deep cone thickener modeling method of the first aspect.
[0018] The third aspect
[0019] The computer readable storage medium provided by the embodiments of the present application stores a computer program, and the program is executed by the processor to realize the state space-based deep cone thickener modeling method of the first aspect.
[0020] The technical scheme provided by the embodiment of the present application has at least the following beneficial effects:
[0021] In the present application, by obtaining historical operation data and preprocessing the historical operation data, then constructing a deep cone thickener prediction model based on a neural network under the framework of a state space model, by inputting the preprocessed historical operation data into the model and using a loss function for training, the model can gradually adjust its parameters to reduce the error between the predicted value and the true value, and finally obtain a trained model, thereby completing the modeling process of the deep cone thickener and realizing the efficiency and adaptability of the state space-based deep cone thickener modeling method. The method uses a slicing technique to process long sequence data, which helps to improve the model's ability to process long sequences. Meanwhile, the Mamba module utilizes the linear characteristics of the state space model to capture long-term dependencies in sequence data through selective state spaces, while the feedforward network layer can handle more direct variable relationships. By combining the Mamba module and the feedforward network layer, the algorithm can effectively handle variable correlation and time correlation problems, improve prediction accuracy, ensure the model's generalization ability under different operating conditions, optimize operation speed, and improve the efficiency and stability of the paste filling process. BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0023] Figure 1 A flowchart of a state space-based deep cone thickener modeling method provided by the embodiment of the present application is shown in the figure.
[0024] Figure 2 A structural diagram of a deep cone thickener provided by the embodiment of the present application is shown in the figure.
[0025] Figure 3 A principle diagram of a state space equation provided by the embodiment of the present application is shown in the figure.
[0026] Figure 4 A structural diagram of a deep cone thickener prediction model provided by the embodiment of the present application is shown in the figure.
[0027] Figure 5 A structural diagram of a state space-based deep cone thickener modeling system provided by the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0028] The technical solutions in the present application will be described below with reference to the drawings.
[0029] In the embodiments of the present application, the words such as "example", "for example" are used to represent an example, illustration or description. Any embodiment or design scheme described as "example" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the word "example" is intended to present the concept in a specific manner. In addition, in the embodiments of the present application, the meaning expressed by "and / or" can be both, or can be one of the two.
[0030] In order to make the technical problems, technical solutions and advantages to be solved by the present application more clear, the following will be described in detail in combination with the drawings and specific embodiments.
[0031] Referring to the drawings accompanying the Figure 1 , a flowchart of a state space-based deep-cone thickener modeling method provided by an embodiment of the present application is shown.
[0032] The embodiment of the present application provides a state space-based deep-cone thickener modeling method, which can be realized by a state space-based deep-cone thickener modeling device. The state space-based deep-cone thickener modeling device can be a terminal or a server. The processing flow of the state space-based deep-cone thickener modeling method can include the following steps:
[0033] S1: Obtain historical operation data of the deep-cone thickener.
[0034] The historical operation data of the deep-cone thickener refers to collecting various important operation parameters of the deep-cone thickener under different working conditions and time periods. By obtaining the historical operation data of the thickener, the detailed operation of the equipment under different working conditions can be provided, which lays a data foundation for model construction and performance optimization.
[0035] In a possible implementation, the historical operation data includes underflow flow, feed flow, underflow concentration, feed concentration, sludge layer pressure, overflow flow and rake frame torque.
[0036] It should be noted that by collecting multi-dimensional parameter data, the dynamic response of the equipment can be more comprehensively understood, and the accuracy and adaptability of the prediction model can be enhanced.
[0037] Specifically, let the historical sequence segment time length be T h , and the prediction segment sequence length be T p . Then, for a given time t0, the offline prediction problem of the thickener system state can be expressed as:
[0038]
[0039] Wherein, represents the predicted system state underflow concentration, Represents the control input set {u2, u3, u4, u5, u6, u7, u8, u9, u 10 All parameters within the time range (t0, t0-T) h The union of the sequences X(t0,-T) h ) represents the historical time range (t0, t0-T) h The system state sequence of {u1} within ) Represents the control input set {u2, u3, u4, u5, u6, u7, u8, u9, u 10 All parameters within the time range (r0, t0+T) p The union of sequences.
[0040] Table 1 Input Parameters for the Prediction Model of Deep Cone Thickener
[0041] Symbol Meaning Dimension u1 Underflow flow rate m 3 / h]]> u2 Feed flow rate m 3 / h]]> u3 Underflow concentration % u4 Feed concentration % u5 Sludge blanket pressure MPa u6 Overflow flow rate m 3 / h]]> u7 Rake 1 torque N·m u8 Rake 2 torque N·m u9 Rake 3 torque N·m u 10 ]]> Rake 4 torque N·m u 11 ]]> Time stamp s
[0042] As shown in Table 1, for the established deep cone thickener prediction model, the system state is the underflow flow rate u1, and the control inputs are the feed flow rate setpoint u2, underflow concentration u3, feed concentration u4, mud pressure u5, overflow flow rate u6, and rake torque u1. i ,i=7,…,10.
[0043] Reference manual attached Figure 2 The diagram shows a structural schematic of a deep cone thickener provided in an embodiment of the present invention.
[0044] Figure 2 In China, a deep cone thickener is a device used for solid-liquid separation. Its body has a cylindrical upper part and a conical lower part with a deep cone. This device works based on the principles of centrifugal force and gravity sedimentation, and can concentrate slurry with low solids content into underflow slurry with high solids content through processes such as flocculation and gravity sedimentation.
[0045] The main components of the deep-cone thickener include a stirrer, a hopper, a deep-cone cylinder, bearings, a transmission device, an inlet water pipe, an outlet, a compressed air pipe, and the like. The deep-cone cylinder is the core component of the equipment and is composed of a cone, a cylinder, and a cone head. In the working process, the slurry to be treated is introduced into the thickening zone of the deep-cone thickener through the inlet pipe, and the forced stirring device is turned on to uniformly stir the feed suspension. Due to the centrifugal force, the solid particles begin to deposit at the bottom of the thickening barrel. The higher the slurry concentration, the greater the centrifugal force, and the faster the solid particles deposit. By adjusting the stirring speed, the height of the overflow pipe, and other parameters, the concentration and thickening effect of the slurry can be controlled. When the thickening effect meets the requirements, the stirring device and the overflow pipe are turned off, the feeding is stopped, and the solid-liquid separation is realized. The thickened slurry is discharged through the overflow pipe, and the cleaned material is discharged through the bottom outlet device. At the same time, the stability of the underflow concentration is a basic index for judging the process level. In the production process, multiple parameters affect the stability of the underflow concentration. The instability of the volume and concentration of the feed flow will cause the underflow concentration to fluctuate, and other parameters such as the feed flow rate, the underflow flow rate, and the like will also affect the underflow flow rate.
[0046] S2: Preprocessing the historical operation data.
[0047] It should be noted that preprocessing the historical data can improve the data quality and ensure its consistency, making the model training more stable and reliable.
[0048] In one possible implementation, the preprocessing specifically includes abnormal segment removal preprocessing, missing value interpolation preprocessing, and z-score standardization preprocessing.
[0049] The abnormal segment removal preprocessing refers to identifying and removing data segments in the historical data that do not conform to normal operating conditions, such as sensor failures or unexpected shutdowns. The missing value interpolation preprocessing refers to filling in missing data points using corresponding interpolation methods to ensure the continuity and integrity of the data. The z-score standardization refers to normalizing the data according to the standard score, which eliminates the dimensional differences between different data by removing the mean and scaling to the unit standard deviation, so that they conform to the same scale, facilitating model processing.
[0050] It should be noted that abnormal segment removal can eliminate biased data that is detrimental to the model, avoiding interference with the model's learning process. Missing value interpolation allows missing data to be reasonably filled in, preserving the continuity of key information. Z-score standardization balances the weights of different parameters, allowing them to play a reasonable role in training.
[0051] In one possible implementation, the missing value interpolation preprocessing specifically includes:
[0052] Determine the time i of the missing historical operation data in the historical time series.
[0053] It should be noted that in the industrial field, especially in the operation process of deep cone thickener, different sensors collect data at different frequencies according to their specific working period and monitoring tasks. For the underflow flow rate u1 and the feed flow rate u2, since they are directly affected by the upstream and downstream processes and change more sensitively, an electromagnetic flowmeter is used to collect data every 3 seconds to capture the small changes in flow rate in time. The underflow concentration u3 and the feed concentration u4 are also affected by the upstream and downstream processes, but change slowly, so a nuclear concentration meter is used to measure every 5 seconds. Such frequency can ensure the real-time of data and effectively control the measurement cost. As for the mud layer pressure u5, the overflow water flow rate u6 and the rake torque u7, u8, u9, u 10 Since their changes are relatively slow, they do not need to be measured too frequently, so data is collected every 15 seconds.
[0054] Specifically, in order to uniformly process these data, it is necessary to adjust them to the same time scale, therefore, the least common multiple of the collection frequencies of all variables, i.e. 30 seconds, is selected as the unified time interval. However, due to the lack of synchronization reference among sensors, the collected data are not aligned within the unified time interval, resulting in data missing for some variables in each 30-second interval. This non-alignment of data needs to be properly handled in subsequent data preprocessing and analysis to ensure that the model can accurately reflect the running state and trend of the deep cone thickener. Missing historical running data refers to the data that cannot be recorded at some time points under the unified time interval standard due to the lack of synchronization among sensors during the past operation of the thickener. The historical time series represents a set of running data of the thickener arranged in chronological order, including various running parameters at multiple time points, for describing the running change trend of the thickener.
[0055] Determine non-missing data around the missing historical running data:
[0056] x i-2 ,x i-1 ,x i+1
[0057] where x i-2 ,x i-1 ,x i+1 is used as the interpolation point, x i-2 represents the non-missing data at time i-2, x i-1 represents the non-missing data at time i-1, and x i+1 represents the non-missing data at time i+1. Let y be the value of each interpolation point x, and use Newton interpolation method to use the data points near the missing point x i . Let the interpolation point and the value of the point satisfy the mapping yi = f(x i ), a second-order polynomial P(x) is constructed to estimate the missing value:
[0058] P(x) = f(x i-2 ) + f[x i-2 , x i-1 ](x - x i-2 ) + f[x i-2 , x i-1 , x i+1 ](x - x i-2 )(x - x i-1 )
[0059] where the first-order difference quotient is:
[0060]
[0061] the second-order difference quotient is:
[0062]
[0063] Substitute x = x i into the polynomial p(x) to calculate p(x i ) as the estimate of the missing value y i ,
[0064] For variables with small changes, use x i-1 , x i+1 as the interpolation point, and use linear interpolation method, set y as the value of each interpolation point x, and construct the following polynomial:
[0065]
[0066] Substitute x = x i into the polynomial p(x) to calculate p(x i ) as the estimate of the missing value y i .
[0067] where, the non-missing historical running data refers to the successfully recorded, non-missing running data points in the time series, i.e. valid data, which can be used as a reference to estimate the missing value, and the missing historical running data value refers to the data points that are not recorded due to sensor failure or other reasons, i.e. missing values in the time series, which need to be estimated and completed to ensure the integrity of the data.
[0068] The z-score standardization preprocessing is specifically:
[0069]
[0070] where, represents the non-missing historical running data at time i after standardization, xi non-missing historical operation data at time i, μ i mean of non-missing historical operation data at time i, σ i standard deviation of non-missing historical operation data at time i, N represents the total number of historical operation data.
[0071] It should be noted that the missing data points are replaced by the average value of the surrounding recorded data to fill in the missing part of the data, ensuring the continuity of the data in the time series, and the z-score standardization balances the weights of different parameters to make them play a reasonable role in training, thereby enhancing the robustness and generalization ability of the model through data preprocessing, which helps to improve the accuracy and stability of the prediction.
[0072] Referring to the accompanying drawings Figure 3 , a principle schematic diagram of a state space equation provided by an embodiment of the present application is shown.
[0073] Figure 3 In the formula, x represents the input historical time series of the state space model, B represents the input matrix, h represents the state variable of the historical operation data, A represents the state transition matrix, C represents the output matrix, and y represents the predicted output of the state space model.
[0074] After the input historical time series data enters the model, the time series data is transformed or weighted by the input matrix, and the processed time series data is transmitted to the state variable. After receiving the time series data from the input matrix and the state transition matrix, the state variable calculates the state change, and the current state is fed back to the state variable through the state transition matrix to further control the change and stability of the time series data state. The updated state is output as a vector to realize the control of the deep cone thickener.
[0075] It should be noted that when the state space model is applied to a computer, the computer contacts and processes discrete data, so the zero-order hold technology is used to convert the discrete data into continuous data. First, the value of each received discrete signal is retained until a new discrete signal is received. This operation results in the creation of continuous signals that can be used by the state space model. The time for maintaining the value is represented by a new learnable parameter, called the step size, which represents the phased retention of the input. With continuous input data, continuous output can be generated, and only the time step of the input is sampled to obtain the discrete output.
[0076] In practical applications, the basic state space model performs poorly, especially when dealing with long-term dependencies, because the solution of a linear first-order ordinary differential equation is usually an exponential function, which can cause the gradient to grow exponentially over the sequence length, thereby causing the problem of gradient vanishing or explosion, so the HIPPO (High-order Polynomial Projection Operators) theory is used, the HiPPO theory specifies a special class of matrices, when these matrices are included in the equations of the SSM, the state can remember the historical information of the input, these special matrices are called HiPPO matrices, they have a specific mathematical form, which can effectively capture long-term dependencies, replace the state transition matrix with the HiPPO matrix, which attempts to compress all the input signals seen so far into a coefficient vector, and generate the optimal solution of the state matrix through function approximation.
[0077] With reference to the accompanying drawings Figure 4 , a structure diagram of a deep-cone thickener prediction model is shown.
[0078] Figure 4 In the formula, u1…u i and x i represent the historical operation data of the deep-cone thickener, represents the prediction result, t represents the time axis of the historical sequence and the prediction sequence, x represents the obtained historical operation data sequence, the historical sequence represents the historical operation data used to train the deep-cone thickener prediction model, and the prediction sequence represents the future data to be predicted by the model, in order to process missing or insufficient data, the method of filling with 0 is used in the figure, and the slice processing represents that the input historical time sequence data is processed by the slice operation module according to a fixed length, this operation is similar to cutting a long time sequence into several shorter fragments, the purpose of slicing is to facilitate subsequent feature extraction and help the model understand the data change characteristics in the local time period, and the feature extraction is used to align the sliced data into a two-dimensional matrix, one dimension of the matrix is time, d modelAnother dimension is the feature dimension, which is achieved by feature projection, mapping the original data to a high-dimensional space, so that the data of each time slice can be characterized in the high-dimensional feature space. Then, the high-dimensional feature data after feature alignment and projection is input into the Mamba module. The Mamba module can efficiently process long time series data based on the framework of the state space model, and filter and extract key information through specific mechanisms (such as a gating mechanism similar to RNN). Finally, the data processed by the Mamba module is sent to the prediction unit module. The prediction head represents the prediction unit module, which generates the prediction results of the model. The prediction unit module can be a linear layer or other deep learning structure, which is used to generate future prediction values. The final prediction results Will be generated according to the historical data sequence, which means that the model uses past data to infer future trends to achieve prediction of the operation state of the thickener.
[0079] S3: Under the framework of the state space model, a deep cone thickener prediction model based on neural network is constructed, wherein the deep cone thickener prediction model comprises a slicing operation module, a Mamba module and a prediction unit module.
[0080] Wherein, the state space model refers to a mathematical model used to represent the dynamic behavior of a system over time. It describes the internal state of the system with state variables and establishes the relationship between the state and time. It is suitable for prediction of dynamic systems. Neural network refers to a machine learning model that processes complex data by imitating the structure of human brain neurons. Neural network can automatically extract features and patterns from data for prediction and classification. The deep cone thickener prediction model is an intelligent prediction model for the operation state of the deep cone thickener. The slicing operation module refers to a module that cuts long historical data into multiple sub-slices so that the neural network model can process short time series one by one, improving the computational efficiency and stability of the model. The Mamba module refers to a special structure in the model, which is used to extract the correlation between key features and variables from historical time series, so that the model can more effectively capture the dynamic behavior of the thickener. The prediction unit module refers to receiving key features processed by the Mamba module and generating the final prediction results to help identify and predict future operation states.
[0081] It should be noted that the introduction of neural networks and their special modules under the state space model framework enables the model to flexibly cope with the complex dynamic characteristics of the deep cone thickener, the slicing operation module improves the data processing efficiency, facilitates the model to focus on local features, the Mamba module further enhances the accuracy of key feature extraction, helps to capture deep correlations in the data, and the prediction unit module converts the extracted features into prediction output, thereby improving the accuracy of prediction. Through this modular design, the model has stronger adaptability and robustness, which helps to realize accurate prediction of the operating state of the thickener.
[0082] In one possible implementation, the slicing operation module in S3 is used to cut the long time series into multiple sub-slices, i.e., divide the entire time series into multiple time period slices, specifically including:
[0083] S301: Convert the preprocessed historical operation data into a historical time series, and input the historical time series into the deep cone thickener prediction model.
[0084] The historical time series refers to the operating data of the deep cone thickener arranged in chronological order, which records the operating state of the device at different time points, such as flow, concentration, pressure, etc., and can reflect the dynamic change trend of the device.
[0085] Specifically, let x = [x1, x2, …, x n ], x represents the historical time series, x n represents the time series at the nth time point, which contains the measurement values of 11 variables (feed flow u1, underflow flow u2, underflow concentration u3, feed concentration u4, sludge layer pressure u5, overflow flow u6, rake torque u7, rake torque u8, rake torque u9, rake torque u 10 , and timestamp u 11 ).
[0086] S302: Slice the input historical time series by a fixed length:
[0087] x patch = x[:,:, stride: P + stride]
[0088] Where x patch represents the slice obtained by slicing, x represents the historical time series, P represents the slice length, and stride represents the step length.
[0089] S303: Perform value embedding operation and position embedding operation on each sub-slice:
[0090] x embed = ValueEmbedding(x patch)+ PositionEmbedding(x patch )
[0091] wherein ValueEmbedding represents a value embedding layer for mapping each slice to a fixed-dimensional vector space, which fuses time information along the slice length, and upgrades the slice length 16 to 128, PositionEmbedding represents a position embedding layer for adding position information to each slice, x embed represents the embedded historical time series.
[0092] Table 2 Deep-cone thickener prediction model hyperparameter tuning results
[0093]
[0094] As shown in Table 2, the number of Mamba modules represents the number of Mamba modules used in the model, the embedding layer dimension represents the embedding layer dimension of the model, i.e., the dimension size of the input features, when the embedding dimension is 128, a higher accuracy can be obtained, if a higher embedding dimension is selected, it will lead to too large computational complexity, the prediction length refers to the length of the time series generated by the model in the prediction task, MSE represents the average of the error squares between the prediction results and the true values, which is used to measure the accuracy of the model prediction, the smaller the value, the better, MAE represents the average of the absolute error between the prediction results and the true values, which is also an index for measuring the prediction accuracy, the smaller the value, the better.
[0095] It should be noted that through the historical time series conversion, the model can directly use the sorted running data, which is convenient for subsequent operation, then the slice processing cuts the long time series into appropriate length subsequences, which helps the model focus on the data features in a short time period, effectively reduces the computational burden and enhances the analysis of local dynamics, the value embedding and position embedding operations make each slice have a fixed-dimensional representation in the vector space while having time position information, which ensures that the model can fully utilize the numerical and time characteristics of the data, thereby improving the processing effect of time series data.
[0096] In one possible implementation, the Mamba module in S3 is used to extract key features and variable correlations in the historical time series, specifically including:
[0097] S304: processing the historical time series through a selection mechanism in the Mamba module:
[0098] g t =σ(W g x t +b g) where gtis the selection gate vector at time step t, which determines the importance of each feature or variable at the current time step, xtis the input data at time step t, and x t represents the data at time t in the historical time series, W g and b g are learnable parameters, and σ represents the sigmoid function, which is used to convert the linear output into a value between 0 and 1, representing the degree of importance of each feature.
[0099] It should be noted that the selection mechanism in the Mamba module processes the historical time series, which can adjust the processing of information according to the relevance and importance of the input data at each time step. This mechanism enables the model to identify and emphasize variables that have a greater impact on the concentration of the paste, such as feed flow and underflow flow, while reducing the weight of variables that have a smaller impact.
[0100] S305: Update the output vector according to the state space model:
[0101]
[0102] where, represents a randomly initialized learnable parameter, h k represents the state vector at the current time k, h k-1 represents the state vector at time k-1, and represents element-wise multiplication, which is used to multiply the gate vector g t with the original data x t , essentially weighting the data after selecting the importance of each variable. The model combines the state h k-1 at the previous time and the input x k at the current time to calculate the state h k at the current time.
[0103] S306: Generate a predicted output based on the updated output vector:
[0104] y k = Ch k
[0105] where y k represents the predicted output at time k, and C represents the output matrix that converts the state vector h k at the current time k to the output.
[0106] It should be noted that this model not only can handle the current input data, but also can consider the influence of past data on the current state, which is crucial for predicting the future state of the deep cone thickener. This ability enables the Mamba module to accurately simulate the dynamic interaction between variables, providing accurate concentration prediction and operation guidance for the deep cone thickener, thereby improving the prediction accuracy of the model and enhancing its adaptability and stability in complex industrial environments.
[0107] S307: Extracting key features in the prediction output through the feedforward network layer in the Mamba module:
[0108] y = SiLU(W2SiLU(W1x + b1) + b2)
[0109] Where y represents the output result after the feedforward network layer, SiLU represents the SiLU activation function, x represents the prediction output y k , W1 and W2 represent the weights of the feedforward network layer, and b1 and b2 represent the biases of the feedforward network layer.
[0110] It should be noted that the selection mechanism in the Mamba module effectively filters and selects key information from historical data, enabling the model to focus on the most important features for prediction, thereby improving the relevance and efficiency of data processing. By updating the state vector, the model can reflect the dynamic changes of the system and provide accurate state transition information, thereby enhancing the understanding of system behavior. Using the updated output vector to generate accurate prediction results, combined with the state at the previous moment and the current input, improves the coherence and reliability of the prediction. Finally, by extracting key features in the prediction output through the feedforward network layer, the model further refines information and enhances its ability to capture complex patterns.
[0111] In one possible implementation, the prediction unit module is configured to receive the key features processed by the Mamba module and output the prediction results:
[0112] y pred = W out FFN(x) + b out
[0113] Where W out represents the weights of the prediction head output layer, b out represents the biases of the prediction head output layer, y pred represents the prediction results, and FFN represents the feedforward network layer.
[0114] It should be noted that the prediction unit module effectively receives and processes the key features extracted by the Mamba module to achieve high-quality prediction results, enabling the model to achieve more accurate prediction, reduce errors, and improve the prediction quality of the deep cone thickener operating state.
[0115] S4: input the pretreated historical operation data as a training set to the deep cone thickener prediction model to train the deep cone thickener prediction model until a loss function value of the deep cone thickener prediction model is less than a preset loss function value.
[0116] The pretreated historical operation data refers to a cleaning data set obtained after removing outliers, interpolation, standardization and other pretreatments on historical operation data of the deep cone thickener, the training set refers to the pretreated data set used to train the prediction model of the deep cone thickener to help the model learn the relationship between the input data and the output prediction, the loss function refers to an evaluation index for measuring the prediction effect of the model, indicating the deviation between the model prediction value and the actual value, by optimizing the loss function, the model continuously adjusts its parameters to improve the prediction accuracy, and the preset loss function value refers to a threshold for evaluating the end of model training, when the loss function is less than this value, the model is considered to have reached the expected accuracy, and the training can be stopped.
[0117] It should be noted that using the pretreated data as the training set ensures the quality of the input data of the model, thereby improving the learning effect of the model, by introducing the loss function evaluation mechanism, the model can detect its prediction error in real time, gradually optimize the model parameters, and finally realize high-precision prediction, and setting the preset loss threshold also avoids the risk of overtraining, saves computing resources and speeds up the training process.
[0118] It should be noted that the size of the preset loss function value can be set by the person skilled in the art according to actual needs, which is not limited in the present application.
[0119] In one possible implementation, the loss function of the deep cone thickener prediction model is a mean square error loss function.
[0120] The mean square error loss function is specifically:
[0121]
[0122] Wherein, MSE represents the mean square error loss function, n represents the total number of samples in the historical operation data, y true represents the actual value of each sample in the historical operation data, y pred represents the predicted value of each sample in the historical operation data.
[0123] It should be noted that the MSE loss function provides a reliable optimization direction for the model, helping it to quickly converge to a low error state, ensuring that the prediction result of the deep cone thickener is more accurate and stable.
[0124] S5: output the trained deep cone thickener prediction model to complete the modeling of the deep cone thickener.
[0125] The trained deep-cone thickener prediction model refers to a model that has been optimized, learned and reached a target loss value through a training process, and has good prediction ability and can effectively predict the state of the deep-cone thickener in actual operation.
[0126] It should be noted that the output of the optimized deep-cone thickener prediction model ensures that the model has reached the expected prediction accuracy and can be directly applied to the production process. At the same time of completing the modeling, a verified model is provided, which not only saves the time of repeated adjustment, but also ensures that the model has good generalization ability in actual application.
[0127] Table 3 Comparison of performance indicators of different models in predicting the underflow concentration of the deep-cone thickener
[0128]
[0129] As shown in Table 3, Patch_Mamba is the model name used in the present application, AutoTransformer, Dlinear and Non-Stationary Transformer are all Transformer models for processing time series data, the output results of the present application model and the three are compared, and the lowest error is obtained by comparing the MSE, MAE and the current mainstream model, wherein the comparison indicators are MAE, MSE, RMSE, RMSE, MSPE, MAE represents the mean absolute error, MSE represents the mean square error, RMSE represents the root mean square error, MAPE represents the mean absolute percentage error, and MSPE represents the mean square percentage error.
[0130] In the present application, the IP address, port range, domain name, URL and / or TCP packet that need to be shielded are determined according to the abnormal category in an automated manner, which can quickly and timely respond to network anomalies. Without manual intervention, the response time can be reduced, and the real-time performance and efficiency of the system can be improved. The target that needs to be blocked can be accurately determined according to the abnormal category, thereby preventing the spread of malicious attacks. The source of abnormal data and related information are shielded in time, which can curb the spread of attack behavior and protect network security. According to the abnormal category of network transmission data, the corresponding bypass blocking can automatically respond to network anomalies, prevent the spread of malicious attacks, protect system and data security, and improve the reliability of the network. At the same time, it also has flexibility and accuracy, effectively coping with various network security threats.
[0131] The technical scheme provided by the embodiment of the present application has at least the following beneficial effects:
[0132] In the present application, by obtaining historical operation data, and preprocessing the historical operation data, then constructing a deep cone thickener prediction model based on neural network under the framework of state space model, by inputting the preprocessed historical operation data into the model and using the loss function for training, the model can gradually adjust its parameters, reduce the error between the predicted value and the true value, and finally obtain a trained model, thereby completing the modeling process of the deep cone thickener, and realizing the efficiency and adaptability of the state space-based deep cone thickener modeling method. The method uses a slicing technique to process long sequence data, which helps to improve the model's ability to process long sequences. At the same time, the Mamba module utilizes the linear characteristics of the state space model to capture long-term dependencies in sequence data through selective state space, while the feedforward network layer can handle more direct variable relationships. By combining the Mamba module and the feedforward network layer, the algorithm can effectively handle variable correlation and time correlation problems, improve prediction accuracy, ensure the model's generalization ability under different working conditions, optimize the operation speed, and improve the efficiency and stability of the paste filling process.
[0133] Referring to the accompanying drawings Figure 5 , a structure schematic diagram of a state space-based deep cone thickener modeling system provided by the present application is shown.
[0134] The present application also provides a state space-based deep cone thickener modeling system 20, which is applied to the above-mentioned state space-based deep cone thickener modeling method, comprising:
[0135] The processor 201.
[0136] The memory 202, the memory 202 stores computer readable instructions, and the computer readable instructions are executed by the processor 201 to realize the state space-based deep cone thickener modeling method of the method embodiment.
[0137] The state space-based deep cone thickener modeling system 20 provided by the present application can execute the above-mentioned state space-based deep cone thickener modeling method, and realize the same or similar technical effects. To avoid repetition, the present application will not be described again.
[0138] The technical scheme provided by the embodiment of the present application brings at least the following beneficial effects:
[0139] In this invention, historical operating data is acquired and preprocessed. Then, a neural network-based deep cone thickener prediction model is constructed within the framework of a state-space model. By inputting the preprocessed historical operating data into the model and training it using a loss function, the model can gradually adjust its parameters, reducing the error between predicted and actual values, ultimately resulting in a well-trained model. This completes the deep cone thickener modeling process, achieving high efficiency and adaptability of the state-space-based deep cone thickener modeling method. This method employs a sharding technique to process long-sequence data, which helps improve the model's ability to handle long sequences. Simultaneously, the Mamba module utilizes the linearity of the state-space model to capture long-term dependencies in the sequence data through selective state space, while the feedforward network layer can handle more direct inter-variable relationships. By combining the Mamba module and the feedforward network layer, the algorithm effectively handles variable correlation and temporal correlation issues, improving prediction accuracy, ensuring the model's generalization ability under different operating conditions, optimizing computational speed, and improving the efficiency and stability of the paste filling process.
[0140] It should be understood that the processor in the embodiments of the present invention can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0141] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0142] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the flow or function according to the embodiments of the present invention is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. A computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. Semiconductor media can be solid-state drives.
[0143] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0144] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.
[0145] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0146] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0147] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0148] In the embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0149] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0150] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0151] If the functionality is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0152] This invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the state-space-based deep cone thickener modeling method as described in the method embodiments.
[0153] The present invention provides a computer-readable storage medium that can implement the steps and effects of the state-space-based deep cone thickener modeling method of the above-described method embodiments. To avoid repetition, the present invention will not elaborate further.
[0154] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:
[0155] In this invention, historical operating data is acquired and preprocessed. Then, a neural network-based deep cone thickener prediction model is constructed within the framework of a state-space model. By inputting the preprocessed historical operating data into the model and training it using a loss function, the model can gradually adjust its parameters, reducing the error between predicted and actual values, ultimately resulting in a well-trained model. This completes the deep cone thickener modeling process, achieving high efficiency and adaptability of the state-space-based deep cone thickener modeling method. This method employs a sharding technique to process long-sequence data, which helps improve the model's ability to handle long sequences. Simultaneously, the Mamba module utilizes the linearity of the state-space model to capture long-term dependencies in the sequence data through selective state space, while the feedforward network layer can handle more direct inter-variable relationships. By combining the Mamba module and the feedforward network layer, the algorithm effectively handles variable correlation and temporal correlation issues, improving prediction accuracy, ensuring the model's generalization ability under different operating conditions, optimizing computational speed, and improving the efficiency and stability of the paste filling process.
[0156] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
[0157] The following points need to be explained:
[0158] (1) The accompanying drawings of the embodiments of the present invention only involve the structures involved in the embodiments of the present invention. Other structures can refer to the general design.
[0159] (2) For clarity, the thickness of layers or regions is enlarged or reduced in the drawings used to describe embodiments of the present invention; that is, these drawings are not drawn to actual scale. It is understood that when an element such as a layer, film, region, or substrate is referred to as being “above” or “below” another element, the element may be “directly” located “above” or “below” the other element, or there may be intermediate elements.
[0160] (3) Where there is no conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.
[0161] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. The scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A deep cone thickener modeling method based on state space, characterized in that, include: S1: Obtain historical operating data of the deep cone thickener; S2: Preprocess the historical operation data; S3: Under the state-space model framework, a deep cone denser machine prediction model based on neural networks is constructed, wherein the deep cone denser machine prediction model includes a slicing operation module, a Mamba module, and a prediction unit module; The slicing operation module in S3 is used to divide a long-term series into multiple sub-slices, that is, to divide the entire time series into multiple time-segment slices, specifically including: S301: Convert the preprocessed historical operating data into a historical time series, and input the historical time series into the deep cone thickener prediction model; S302: Divide the input historical time series into segments of fixed length: x patch =x[:,:,stride:P+stride] Where, x patch This represents the fragment obtained from the fragmentation process, where x represents the historical time series, P represents the fragment length, and stride represents the step size. S303: Perform value embedding and position embedding operations on each sub-segment: x embed =ValueEmbedding(x patch )+PositionEmbedding(x patch ) Here, ValueEmbedding represents the value embedding layer used to map each piece to a fixed-dimensional vector space, and PositionEmbedding represents the position embedding layer used to add position information to each piece. embed This represents the embedded historical time series; The Mamba module in S3 is used to extract key features and variable correlations from historical time series, specifically including: S304: Process historical time series using the selection mechanism in the Mamba module: g t =σ(W g x t +b g ) Where gt is the gating vector selected at time step t, which determines the importance of each feature or variable at the current time step, and x... t W represents the data at time t in a historical time series. g and b g σ represents the learnable parameters, and σ represents the sigmoid function, which is used to transform the linear output into a value between 0 and 1, representing the importance of each feature. S305: Update the output vector based on the state-space model: in, H represents a randomly initialized learnable parameter. k h represents the state vector at the current time k. k-1 This represents the state vector at time k-1, and ⊙ represents element-wise multiplication, used to transform the gating vector g. t Compared with the original data x t Multiplication, in essence, involves weighting the data after selecting the importance of each variable, and the model then applies the previous state h. k-1 and the input x at the current time k Combined, calculate the current state h. k ; S306: Generate predicted output based on the updated output vector: the k =Ch k Among them, y k Let C represent the predicted output at time k, and let C represent the state vector h at time k. k The output matrix is transformed into the output; S307: Extract key features from the predicted output using the feedforward network layer in the Mamba module: y = SiLU(W2SiLU(W1x+b1)+b2) Where y represents the output of the feedforward network layer, SiLU represents the SiLU activation function, and x represents the predicted output y. k W1 and W2 both represent the weights of the feedforward network layer, and b1 and b2 both represent the biases of the feedforward network layer. The prediction unit module is used to receive the key features obtained by the Mamba module and output the prediction result: the pred =W out FFN(x)+b out Among them, W out b represents the weights of the prediction head output layer. out This represents the bias of the prediction head output layer, y pred This represents the prediction result, and FFN represents the feedforward network layer; S4: Input the preprocessed historical running data as a training set into the deep cone thickener prediction model to train the deep cone thickener prediction model until the loss function value of the deep cone thickener prediction model is less than the preset loss function value. S5: Output the trained deep cone thickener prediction model to complete the modeling of the deep cone thickener.
2. The deep cone thickener modeling method based on state space according to claim 1, characterized in that, The historical operating data includes underflow rate, feed flow rate, underflow concentration, feed concentration, mud pressure, overflow flow rate, and rake torque.
3. The deep cone thickener modeling method based on state space according to claim 1, characterized in that, The preprocessing specifically includes outlier removal preprocessing, missing value imputation preprocessing, and z-score normalization preprocessing.
4. The deep cone thickener modeling method based on state space according to claim 3, characterized in that, The missing value imputation preprocessing specifically includes: Determine the moment i in the historical time series where the missing historical operational data is located; Identify the non-missing data surrounding the missing historical runtime data: x i-2 ,x i-1 ,x i+1 Where x is used i-2 ,x i-1 ,x i+1 As the interpolation point, x i-2 x represents the non-missing data at time i-2. i-1 x represents the non-missing data at time i-1. i+1 Let y represent the non-missing data at time i+1, and let y be the value of each interpolation point x. Using Newton's interpolation method, we utilize the missing point x... i Nearby data points, let the interpolation points and their values satisfy the mapping y i =f(x) i A second-order polynomial P(x) is constructed to estimate the missing values: P(x)=f(x i-2 )+f[x i-2 ,x i-1 ](x-x i-2 )+f[x i-2 ,x i-1 ,x i+1 ](x-x i-2 )(x-x i-1 ) One of the first-order difference quotients is: The second-order difference quotient is: Let x = x i Substitute into the polynomial p(x) and calculate p(x). i ) as missing value y i The estimate, For variables that change relatively little, use x i-1 ,x i+1 Using linear interpolation as interpolation points, let y be the value of x at each interpolation point, and construct the following polynomial: Let x = x i Substitute into the polynomial p(x) and calculate p(x). i ) as missing value y i The estimation; the z-score normalization preprocessing specifically includes: in, x represents the standardized, non-missing historical running data at time i. i μ represents the non-missing historical running data at time i. i σ represents the mean of the non-missing historical running data at time i. i Let represent the standard deviation of the non-missing historical data at time i, and N represent the total number of historical data.
5. The deep cone thickener modeling method based on state space according to claim 1, characterized in that, The loss function of the deep cone thickener prediction model is the mean square error loss function; The mean squared error loss function is specifically as follows: Where MSE represents the mean squared error loss function, n represents the total number of samples in the historical data, and y true y represents the actual value of each sample in the historical data. pred This represents the predicted value for each sample in the historical data.
6. A deep cone thickener modeling system based on state space, characterized in that, include: processor; A memory storing computer-readable instructions, which, when executed by the processor, implement the state-space-based deep cone thickener modeling method as described in any one of claims 1 to 5.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the state-space-based deep cone thickener modeling method as described in any one of claims 1 to 5.