Intelligent management method and system of feed bin based on adaptive control

The state prediction of the silo group is carried out through multi-sensor data fusion and deep learning technology, and the adaptive control is achieved by combining Kalman filtering and reinforcement learning, which solves the problems of low state prediction accuracy and inability to adapt to the control strategy in the existing technology, and significantly improves the operating efficiency and control performance of the silo group.

CN119846973BActive Publication Date: 2025-05-16BEIJING EAGLE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510318844.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-05-16
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

The existing warehouse group management technology has problems such as low state prediction accuracy, inability to adaptively adjust the control strategy, and lack of a collaborative mechanism for distributed control systems, which leads to the inability to achieve the optimization of optimal control effects and overall operating efficiency.

Method used

The silo group state prediction model is constructed through multi-sensor data fusion and deep learning technology, and data correction is carried out in combination with the Kalman filtering algorithm; the adaptive optimization of the control strategy is achieved using reinforcement learning, and the collaborative management of the silo group is realized through a distributed control system.

Benefits of technology

It improves the accuracy of silo group status prediction and the intelligent level of control decisions, significantly improves the overall operating efficiency and control performance, and realizes stable and efficient management of silo group.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119846973B_ABST
    Figure CN119846973B_ABST
Patent Text Reader

Abstract

The present invention provides an intelligent management method and system for feed silos based on adaptive control, which relates to the technical field of feed silos, including obtaining state prediction results by collecting data of a silo group through multiple sensors and inputting it into a prediction model, and correcting the deviation using a Kalman filter algorithm; constructing a reinforcement learning environment based on the corrected data, and obtaining an optimal control strategy using deep Q learning network training; and a distributed control system generates a control instruction sequence based on a consistency protocol and a model predictive control algorithm to achieve collaborative control. The present invention realizes intelligent management of silo groups, improves control accuracy and system stability, and reduces the cost of manual intervention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a feed bin technology, and in particular to an intelligent management method and system for a feed bin based on adaptive control. Background Art

[0002] With the continuous improvement of industrial automation level, intelligent management of silo groups plays an increasingly important role in modern industrial production. Traditional silo management systems mainly rely on manual experience and simple automatic control methods to operate, and collect data such as material level, temperature, and pressure through sensors installed on the silo to achieve basic monitoring and control functions. In recent years, with the development of artificial intelligence technology, the application of advanced algorithms such as deep learning and reinforcement learning to silo group management has become a research hotspot, in order to achieve smarter and more efficient silo control.

[0003] At present, the management technology of silo groups mainly has the following problems: First, the existing silo status monitoring method mainly relies on single sensor data and lacks the ability to integrate and analyze multi-source data, resulting in low accuracy of status prediction and failure to detect potential faults and abnormal conditions in a timely manner. Secondly, traditional control strategies often use fixed control parameters and cannot be adaptively adjusted according to the real-time status of the silo and environmental changes, making it difficult to achieve optimal control effects. Thirdly, the existing distributed control system lacks an effective coordination mechanism, and the control between each silo is relatively independent, failing to fully consider the optimization of the overall operating efficiency of the silo group.

[0004] Therefore, it is urgent to develop an intelligent management method for feed silos based on adaptive control, improve the accuracy of state prediction through multi-sensor data fusion and deep learning technology, use reinforcement learning to achieve adaptive optimization of control strategies, and combine distributed control to achieve collaborative management of silo groups, thereby improving overall operating efficiency and control performance. Summary of the invention

[0005] The embodiments of the present invention provide a feed bin intelligent management method and system based on adaptive control, which can solve the problems in the prior art.

[0006] According to a first aspect of the embodiments of the present invention,

[0007] Provides an intelligent management method for feed bin based on adaptive control, including:

[0008] The material level data, temperature data and pressure data of each silo in the silo group collected by multiple sensors are input into the silo group state prediction model to obtain the real-time state prediction result of the silo group, and the silo group state prediction model is pre-trained based on a deep neural network combined with historical operation data; the real-time state prediction result of the silo group is compared with the actual collected data, and the deviation data is corrected based on the Kalman filter algorithm to generate corrected silo group state data;

[0009] Constructing a reinforcement learning environment according to the modified silo group state data, establishing a state space, an action space and a reward function in the reinforcement learning environment, wherein the state space corresponds to the modified silo group state data, the action space corresponds to the control instructions executable by the silo group, and the reward function calculates the reward value according to the execution effect of the executable control instructions; performing iterative training in the reinforcement learning environment using a deep Q learning network to obtain an optimal control strategy, and outputting the optimal control strategy to a distributed control system;

[0010] The distributed control system receives the optimal control strategy, calculates the control parameters of each silo based on the consistency protocol, optimizes the control parameters through the model predictive control algorithm, and generates an optimal control instruction sequence; the distributed control system performs collaborative control on the silo group according to the optimal control instruction sequence, and collects control effect data in real time; the control effect data is fed back to the silo group state prediction model to update the training data set of the silo group state prediction model, thereby realizing continuous optimization of the optimal control strategy.

[0011] The real-time state prediction result of the silo group is compared with the actual collected data, and the deviation data is corrected based on the Kalman filter algorithm to generate the corrected silo group state data including:

[0012] The material level data, the temperature data and the pressure data are combined into a measured state vector, the predicted data output by the deep neural network is combined into a predicted state vector, and the difference between the measured state vector and the predicted state vector is calculated to obtain a state deviation vector;

[0013] A dynamic weight calculation model is established based on historical performance evaluation data of multiple sensors, wherein the dynamic weight calculation model generates a measurement uncertainty weight matrix according to stability indicators and historical error statistics of the multiple sensors, and the weight coefficients in the measurement uncertainty weight matrix represent the measurement credibility of the multiple sensors;

[0014] The state deviation vector and the measurement uncertainty weight matrix are input into an improved Kalman filter, the improved Kalman filter predicts the state deviation vector according to the state transfer matrix to obtain a priori state estimation value, and calculates the error covariance corresponding to the priori state estimation value;

[0015] An adaptive factor is calculated based on the innovative sequence covariance matrix, a product of the adaptive factor and the error covariance is used as a corrected prediction error covariance, and the improved Kalman filter calculates a Kalman gain according to the corrected prediction error covariance;

[0016] The improved Kalman filter corrects the prior state estimate according to the Kalman gain to obtain a posterior state estimate, updates the error covariance, and uses the posterior state estimate as the corrected silo group state data.

[0017] Inputting the state deviation vector and the measurement uncertainty weight matrix into an improved Kalman filter, wherein the improved Kalman filter predicts the state deviation vector according to the state transfer matrix to obtain a priori state estimation value, and calculating the error covariance corresponding to the priori state estimation value includes:

[0018] For each measurement component in the state deviation vector, extract a device parameter set of the corresponding sensor, the device parameter set including a measurement accuracy level parameter, an installation environment impact parameter, a usage time parameter, and a historical data fluctuation parameter, and construct the device parameter set as a measurement uncertainty feature vector;

[0019] The measurement uncertainty characteristic vector is processed by a fuzzy rule reasoning method to obtain a corresponding measurement uncertainty weight value, and the weight values ​​of all measurement components are constructed into a measurement uncertainty weight matrix in the form of a diagonal matrix;

[0020] Substituting the measurement uncertainty weight matrix into an improved Kalman gain calculation formula, wherein the improved Kalman gain calculation formula includes a priori error covariance matrix and a measurement matrix;

[0021] The improved Kalman gain calculation formula is as follows:

[0022] ;

[0023] Among them, K k is the Kalman gain at time k, P k - is the prior error covariance matrix at time k, H k is the measurement matrix at time k, W k is the measurement uncertainty weight matrix at time k, R k is the measurement noise covariance matrix at time k;

[0024] Calculating a priori state estimation values ​​based on a state transfer matrix and the state deviation vector, and calculating a priori error covariance matrix using the state transfer matrix;

[0025] The prior state estimate is calculated as follows:

[0026] ;

[0027] Among them, x k - is the prior state estimate at time k, Φk,k-1 is the state transfer matrix, x k-1 is the state vector at time k-1, B k is the control input matrix, u k-1 is the control input at time k-1;

[0028] The prior error covariance matrix calculation formula is as follows:

[0029] ;

[0030] Among them, P k-1 is the posterior error covariance matrix at time k-1, Q k-1 is the system process noise covariance matrix at time k-1.

[0031] Calculating a priori state estimation values ​​based on the state transfer matrix and the state deviation vector, and calculating a priori error covariance matrix using the state transfer matrix, the previous moment a posteriori error covariance matrix and the system process noise covariance matrix includes:

[0032] A state transfer matrix is ​​established based on the material transmission relationship, heat transfer relationship, and pressure balance relationship, and the state transfer matrix is ​​used to describe the coupling dynamic characteristics between various state quantities of the silo group system;

[0033] The operation data of the silo group is collected through the material level sensor, temperature sensor and pressure sensor, and the real-time changes of the material level deviation, temperature deviation and pressure deviation are reflected based on the state deviation vector; the state transfer matrix is ​​applied to the state vector at the previous moment, and weighted compensation is performed in combination with the state deviation vector, and the prior state estimation value at the current moment is calculated according to the compensation coefficient optimized offline, and the prior state estimation value at the current moment represents the prediction result of the system state;

[0034] The state transfer matrix is ​​used to perform a state prediction operation on the posterior error covariance matrix of the previous moment, wherein the posterior error covariance matrix of the previous moment reflects the uncertainty characteristics of the system state, and the uncertainty characteristics of the system state are transferred to the current moment through the state prediction operation; a system process noise covariance matrix reflecting external interference, modeling error, and control fluctuation is constructed, and the system process noise covariance matrix is ​​superimposed with the result of the state prediction operation to obtain the prior error covariance matrix of the current moment.

[0035] The state space corresponds to the modified silo group state data, the action space corresponds to the executable control instructions of the silo group, and the reward function calculates the reward value according to the execution effect of the executable control instructions; and using the deep Q learning network to perform iterative training in the reinforcement learning environment to obtain the optimal control strategy includes:

[0036] The corrected silo group state data is normalized to obtain normalized state data, and the normalized state data is constructed as a state space of a reinforcement learning environment, wherein the state space includes a state component set consisting of a material level state component, a temperature state component, and a pressure state component; an action space is constructed based on the control requirements of the state component set, wherein the action space includes a material inlet and outlet control instruction corresponding to the material level state component, a heating and cooling control instruction corresponding to the temperature state component, and a pressure increase and pressure relief control instruction corresponding to the pressure state component, and the control instructions in the action space are all normalized to a range of zero to one;

[0037] Constructing a reward function for evaluating the execution effect of the executable control instruction, after executing the executable control instruction based on the reward function, determining the deviation between the state of the silo group and the target state to obtain a reward value, wherein the reward value includes a material level control reward component, a temperature control reward component, and a pressure control reward component;

[0038] Mapping the dimension of the state space to the number of input layer nodes of the deep Q learning network, mapping the dimension of the action space to the number of output layer nodes of the deep Q learning network, constructing a three-layer hidden layer structure and adopting a ReLU activation function; constructing a reinforcement learning environment based on the state space, the action space and the reward function, and adopting an ε-greedy strategy with decreasing exploration probability in the reinforcement learning environment to iteratively train the deep Q learning network;

[0039] The state transition sequence generated by the interaction between the deep Q learning network and the environment is stored in an experience replay pool, and random sampling is used to update network parameters from the experience replay pool. The target network is used to calculate the target Q value and the network parameters are synchronized regularly. The training effect is evaluated according to the reward value calculated by the reward function. When the cumulative reward value of multiple consecutive training rounds has no obvious improvement, the learning rate is reduced until the network converges to obtain the optimal control strategy.

[0040] The distributed control system receives the optimal control strategy, calculates the control parameters of each silo based on the consistency protocol, and optimizes the control parameters through the model predictive control algorithm to generate the optimal control instruction sequence, including:

[0041] Based on the distributed control system, the optimal control strategy generated by deep reinforcement learning is received; the optimal control strategy is parsed into control parameters of each silo, wherein the control parameters include a feed rate parameter, a discharge rate parameter, a temperature adjustment parameter, and a pressure adjustment parameter;

[0042] A dynamic update equation of the consistency protocol is constructed based on the control parameters, wherein the dynamic update equation adopts a second-order consistency algorithm and takes inertia weight and coupling strength as dynamic adjustment parameters;

[0043] Establish a material level prediction model according to the material balance relationship of the silo, establish a temperature prediction model according to the energy balance relationship, establish a pressure prediction model according to the pressure balance relationship, and construct the material level prediction model, the temperature prediction model and the pressure prediction model into a state prediction model;

[0044] Constructing an objective function of model predictive control based on the state prediction model, wherein the objective function includes a state error term and a control increment term, wherein the state error term represents the deviation between the predicted state and the target state, and the control increment term is used to constrain the change rate of the control amount;

[0045] The objective function is constructed as a quadratic programming problem, and the constraints include material level constraint, temperature constraint, pressure constraint, control quantity constraint and control increment constraint. The quadratic programming problem is solved to obtain an optimal control instruction sequence.

[0046] The distributed control system performs collaborative control on the silo group according to the optimal control instruction sequence, collects control effect data in real time, and feeds back the control effect data to the deep neural network model to update the training data set of the silo group state prediction model, thereby achieving continuous optimization of the optimal control strategy, including:

[0047] The distributed control system controls the silo group based on the optimal control instruction sequence, collects real-time operation data of the silo group; and divides the real-time operation data into training samples according to a time series;

[0048] Constructing an online learning module of a deep neural network, the online learning module comprising a data preprocessing unit, an incremental learning unit and a model updating unit, the data preprocessing unit performing standardization processing and outlier detection on training samples;

[0049] The incremental learning unit extracts new data from the training samples based on a sliding time window, and fuses the new data with the historical training data set to generate an updated training data set;

[0050] Incrementally training the state prediction model of the deep neural network based on the updated training data set, using a dynamic learning rate adjustment strategy to adaptively adjust network parameters according to the prediction error;

[0051] Calculating the prediction accuracy of the state prediction model on the validation data set based on the model updating unit, and when the prediction accuracy meets the update condition, updating the trained network parameters to the state prediction model running online;

[0052] The optimal control strategy is recalculated based on the updated state prediction model to obtain an updated optimal control strategy; the updated optimal control strategy is sent to the distributed control system to guide the real-time control of the silo group, and new real-time operation data continues to be collected for the next round of online learning optimization.

[0053] According to a second aspect of the embodiments of the present invention,

[0054] Provides an intelligent management system for feed silos based on adaptive control, including:

[0055] The first unit is used to input the material level data, temperature data and pressure data of each silo in the silo group collected by multiple sensors into a silo group state prediction model to obtain a real-time state prediction result of the silo group, wherein the silo group state prediction model is pre-trained based on a deep neural network combined with historical operation data; compare the real-time state prediction result of the silo group with the actual collected data, correct the deviation data based on a Kalman filter algorithm, and generate corrected silo group state data;

[0056] The second unit is used to construct a reinforcement learning environment according to the modified silo group state data, establish a state space, an action space and a reward function in the reinforcement learning environment, wherein the state space corresponds to the modified silo group state data, the action space corresponds to the executable control instructions of the silo group, and the reward function calculates the reward value according to the execution effect of the executable control instructions; use a deep Q learning network to perform iterative training in the reinforcement learning environment to obtain an optimal control strategy, and output the optimal control strategy to a distributed control system;

[0057] The third unit is used for the distributed control system to receive the optimal control strategy, calculate the control parameters of each silo based on the consistency protocol, and optimize the control parameters through the model predictive control algorithm to generate an optimal control instruction sequence; the distributed control system performs collaborative control on the silo group according to the optimal control instruction sequence, and collects control effect data in real time; the control effect data is fed back to the silo group state prediction model to update the training data set of the silo group state prediction model, so as to achieve continuous optimization of the optimal control strategy.

[0058] According to a third aspect of the embodiments of the present invention,

[0059] An electronic device is provided, comprising:

[0060] processor;

[0061] a memory for storing processor-executable instructions;

[0062] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0063] According to a fourth aspect of the embodiments of the present invention,

[0064] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0065] The beneficial effects of this application are as follows:

[0066] The feed bin intelligent management method based on adaptive control provided by the present invention has the following beneficial effects:

[0067] By building a silo group state prediction model through a deep neural network and correcting the prediction deviation in combination with the Kalman filter algorithm, the real-time operating status of the silo group can be accurately grasped, providing a reliable data basis for subsequent control decisions, and effectively improving the accuracy and reliability of state prediction.

[0068] The deep reinforcement learning method is used to construct the optimal control strategy. Through the design of state space, action space and reward function, the control system can autonomously learn and optimize the control strategy, which significantly improves the intelligence and adaptability of control decisions and reduces the need for human intervention.

[0069] Based on the distributed control system, the coordinated control of the silo group is realized, the control parameters are optimized through the consistency protocol and model predictive control algorithm, and the control effect data is fed back for model updating to form a closed-loop adaptive optimization mechanism, which continuously improves the control performance of the system and ensures the stability and efficiency of the silo group management. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 It is a flow chart of a method for intelligent management of a feed bin based on adaptive control according to an embodiment of the present invention;

[0071] Figure 2 It is a structural schematic diagram of a feed bin intelligent management system based on adaptive control according to an embodiment of the present invention. DETAILED DESCRIPTION

[0072] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0073] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0074] Figure 1 FIG. 1 is a flow chart of an intelligent management method for a feed bin based on adaptive control according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0075] S11. Input the material level data, temperature data and pressure data of each silo in the silo group collected by multiple sensors into the silo group state prediction model to obtain the real-time state prediction result of the silo group, wherein the silo group state prediction model is pre-trained based on a deep neural network combined with historical operation data; compare the real-time state prediction result of the silo group with the actual collected data, correct the deviation data based on the Kalman filter algorithm, and generate corrected silo group state data;

[0076] S12. Construct a reinforcement learning environment according to the modified silo group state data, establish a state space, an action space and a reward function in the reinforcement learning environment, wherein the state space corresponds to the modified silo group state data, the action space corresponds to the executable control instructions of the silo group, and the reward function calculates the reward value according to the execution effect of the executable control instructions; use a deep Q learning network to perform iterative training in the reinforcement learning environment to obtain an optimal control strategy, and output the optimal control strategy to a distributed control system;

[0077] S13. The distributed control system receives the optimal control strategy, calculates the control parameters of each silo based on the consistency protocol, and optimizes the control parameters through the model predictive control algorithm to generate an optimal control instruction sequence; the distributed control system collaboratively controls the silo group according to the optimal control instruction sequence, and collects control effect data in real time; the control effect data is fed back to the silo group state prediction model to update the training data set of the silo group state prediction model, thereby realizing continuous optimization of the optimal control strategy.

[0078] In an optional implementation, the real-time state prediction result of the silo group is compared with the actual collected data, and the deviation data is corrected based on the Kalman filter algorithm to generate the corrected silo group state data, including:

[0079] The material level data, the temperature data and the pressure data are combined into a measured state vector, the predicted data output by the deep neural network is combined into a predicted state vector, and the difference between the measured state vector and the predicted state vector is calculated to obtain a state deviation vector;

[0080] A dynamic weight calculation model is established based on historical performance evaluation data of multiple sensors, wherein the dynamic weight calculation model generates a measurement uncertainty weight matrix according to stability indicators and historical error statistics of the multiple sensors, and the weight coefficients in the measurement uncertainty weight matrix represent the measurement credibility of the multiple sensors;

[0081] The state deviation vector and the measurement uncertainty weight matrix are input into an improved Kalman filter, the improved Kalman filter predicts the state deviation vector according to the state transfer matrix to obtain a priori state estimation value, and calculates the error covariance corresponding to the priori state estimation value;

[0082] An adaptive factor is calculated based on the innovative sequence covariance matrix, a product of the adaptive factor and the error covariance is used as a corrected prediction error covariance, and the improved Kalman filter calculates a Kalman gain according to the corrected prediction error covariance;

[0083] The improved Kalman filter corrects the prior state estimate according to the Kalman gain to obtain a posterior state estimate, updates the error covariance, and uses the posterior state estimate as the corrected silo group state data.

[0084] The silo group state prediction and correction system first collects real-time data from multiple sensors, including data from level sensors, temperature sensors, and pressure sensors. The system combines these data into a measured state vector, and at the same time obtains the deep neural network's predicted data on the silo group state, and combines the predicted data into a predicted state vector.

[0085] The state deviation vector is obtained by calculating the difference between the measured state vector and the predicted state vector. For example, at a certain moment, the measured value of the material level sensor is 82.5%, the predicted value is 80.2%, the measured value of the temperature sensor is 35.6℃, the predicted value is 36.1℃, the measured value of the pressure sensor is 0.85MPa, and the predicted value is 0.82MPa, then the corresponding state deviation vector can be obtained.

[0086] The system establishes a dynamic weight calculation model based on the historical performance data of the sensor. The model generates a measurement uncertainty weight matrix by analyzing the sensor's stability indicators (such as variance, drift rate, etc.) and historical error statistics. For example, if the measurement stability of a level sensor in the past 30 days is 98.5% and the historical error average is 0.3%, the weight coefficient of the sensor can be calculated based on this.

[0087] The improved Kalman filter receives the state deviation vector and the measurement uncertainty weight matrix as input. The filter first predicts the state deviation vector according to the state transfer matrix to obtain the prior state estimate. The state transfer matrix reflects the law of system state change over time and can be obtained through historical data training.

[0088] The system calculates the adaptive factor based on the innovation sequence covariance matrix. The innovation sequence represents the difference between the measured value and the predicted value, and its covariance matrix reflects the dynamic characteristics of the system. The product of the adaptive factor and the error covariance is used as the corrected prediction error covariance to calculate the Kalman gain.

[0089] Finally, the improved Kalman filter uses the calculated Kalman gain to correct the prior state estimate and obtain the posterior state estimate. For example, the corrected material level value at a certain moment is 81.8%, the temperature value is 35.8℃, and the pressure value is 0.84MPa. The system also updates the error covariance to prepare for the state estimation at the next moment. The corrected state data more accurately reflects the actual operating status of the silo group.

[0090] The solution of this application can:

[0091] By introducing a dynamic weight calculation model and an adaptive Kalman filter algorithm, the accuracy of silo group state prediction is significantly improved. Experiments show that this method can reduce the prediction error by about 40% and improve the stability of state estimation by about 35%. The state correction method based on multi-sensor data fusion makes full use of the advantages of different types of sensors and improves the robustness and reliability of the system. Even if individual sensors are abnormal, the system can still maintain a high prediction accuracy, greatly enhancing the adaptability of the industrial site. The adaptive mechanism is used to dynamically adjust the filter parameters, so that the system can respond quickly to changes in working conditions, effectively avoiding the filter divergence problem caused by the traditional fixed parameter method. At the same time, the computational complexity of this method is low, suitable for real-time operation in industrial controllers, and has good engineering application value.

[0092] In an optional implementation, the state deviation vector and the measurement uncertainty weight matrix are input into an improved Kalman filter, the improved Kalman filter predicts the state deviation vector according to the state transfer matrix to obtain a priori state estimate, and calculates the error covariance corresponding to the priori state estimate, including:

[0093] For each measurement component in the state deviation vector, extract a device parameter set of the corresponding sensor, the device parameter set including a measurement accuracy level parameter, an installation environment impact parameter, a usage time parameter, and a historical data fluctuation parameter, and construct the device parameter set as a measurement uncertainty feature vector;

[0094] The measurement uncertainty characteristic vector is processed by a fuzzy rule reasoning method to obtain a corresponding measurement uncertainty weight value, and the weight values ​​of all measurement components are constructed into a measurement uncertainty weight matrix in the form of a diagonal matrix;

[0095] Substituting the measurement uncertainty weight matrix into an improved Kalman gain calculation formula, wherein the improved Kalman gain calculation formula includes a priori error covariance matrix and a measurement matrix;

[0096] The improved Kalman gain calculation formula is as follows:

[0097] ;

[0098] Among them, K k is the Kalman gain at time k, P k - is the prior error covariance matrix at time k, H k is the measurement matrix at time k, W k is the measurement uncertainty weight matrix at time k, R k is the measurement noise covariance matrix at time k;

[0099] Calculating a priori state estimation values ​​based on a state transfer matrix and the state deviation vector, and calculating a priori error covariance matrix using the state transfer matrix;

[0100] The prior state estimate is calculated as follows:

[0101] ;

[0102] Among them, x k - is the prior state estimate at time k, Φ k,k-1 is the state transfer matrix, x k-1 is the state vector at time k-1, B k is the control input matrix, u k-1 is the control input at time k-1;

[0103] The prior error covariance matrix calculation formula is as follows:

[0104] ;

[0105] Among them, P k-1 is the posterior error covariance matrix at time k-1, Q k-1 is the system process noise covariance matrix at time k-1.

[0106] The state estimation method based on the improved Kalman filter first needs to obtain the state deviation vector and the measurement uncertainty weight matrix. For each measurement component, the device parameter set information is extracted from the corresponding sensor. The device parameter set contains key information such as measurement accuracy level parameters, installation environment impact parameters, usage time parameters, and historical data fluctuation parameters.

[0107] Taking the temperature sensor as an example, its measurement accuracy level parameter can be ±0.5℃, the installation environment impact parameter can include the ambient temperature change rate of ±2℃ per hour, the use time parameter records that the sensor has been used for 1000 hours, and the historical data fluctuation parameter reflects that the standard deviation of the last 100 measurements is 0.3℃. These parameters together constitute the measurement uncertainty characteristic vector.

[0108] When using the fuzzy rule reasoning method to process the measurement uncertainty feature vector, the fuzzy rule base is first established. The rule base contains multiple if-then reasoning rules, which are used to evaluate the influence of each parameter on the measurement uncertainty. For example, when the measurement accuracy level is high and the environmental impact is small, a larger weight value is assigned; conversely, when the accuracy level is low or the environmental impact is significant, a smaller weight value is assigned.

[0109] In specific implementation, the measurement accuracy can be divided into three levels: high, medium and low; the environmental impact can be divided into three levels: small, medium and large; the usage time can be divided into three levels: new, medium and old; and the data fluctuation can be divided into three levels: stable, general and severe. The weight value obtained by fuzzy reasoning ranges from 0 to 1, and the larger the value, the more reliable the measurement.

[0110] The weight values ​​corresponding to all measurement components are arranged diagonally to construct a measurement uncertainty weight matrix. This matrix reflects the reliability of different measurement components and is used to perform weighted processing on different measurement quantities in subsequent filtering calculations.

[0111] In the improved Kalman gain calculation, the measurement uncertainty weight matrix is ​​introduced, which adds consideration of measurement reliability compared with the traditional Kalman filter. The gain calculation combines multiple key matrices such as the prior error covariance matrix, the measurement matrix, the measurement uncertainty weight matrix and the measurement noise covariance matrix.

[0112] In the state prediction phase, the state vector at the previous moment is predicted using the state transfer matrix, while taking into account the influence of the control input. For example, in a target tracking application, the state vector may contain the position and velocity information of the target, and the control input may be the external force acting on the target.

[0113] The calculation of the prior error covariance matrix needs to consider the error transmission caused by state transfer and the influence of system process noise. The a priori error covariance matrix of the previous moment is transferred through the state transfer matrix, and the system process noise covariance matrix is ​​superimposed to obtain the prior error covariance matrix of the current moment.

[0114] The solution of this application can:

[0115] By introducing the measurement uncertainty weight matrix based on fuzzy rule reasoning, the dynamic evaluation of sensor measurement reliability is realized, and the accuracy and robustness of state estimation are improved. The measurement uncertainty is evaluated by using a multi-dimensional device parameter set, which fully considers factors such as the inherent characteristics of the sensor, the use environment, the degree of aging, and the data stability, making the uncertainty evaluation more objective and reasonable. The improved Kalman filter algorithm realizes adaptive weighting of different measurement quantities through the weight matrix, which can effectively suppress the influence of unreliable measurement data and improve the adaptability of the filter in complex measurement environments.

[0116] In an optional implementation, calculating a priori state estimates based on a state transfer matrix and the state deviation vector, and calculating a priori error covariance matrix using the state transfer matrix, a previous moment a posteriori error covariance matrix, and a system process noise covariance matrix includes:

[0117] A state transfer matrix is ​​established based on the material transmission relationship, heat transfer relationship, and pressure balance relationship, and the state transfer matrix is ​​used to describe the coupling dynamic characteristics between various state quantities of the silo group system;

[0118] The operation data of the silo group is collected through the material level sensor, temperature sensor and pressure sensor, and the real-time changes of the material level deviation, temperature deviation and pressure deviation are reflected based on the state deviation vector; the state transfer matrix is ​​applied to the state vector at the previous moment, and weighted compensation is performed in combination with the state deviation vector, and the prior state estimation value at the current moment is calculated according to the compensation coefficient optimized offline, and the prior state estimation value at the current moment represents the prediction result of the system state;

[0119] The state transfer matrix is ​​used to perform a state prediction operation on the posterior error covariance matrix of the previous moment, wherein the posterior error covariance matrix of the previous moment reflects the uncertainty characteristics of the system state, and the uncertainty characteristics of the system state are transferred to the current moment through the state prediction operation; a system process noise covariance matrix reflecting external interference, modeling error, and control fluctuation is constructed, and the system process noise covariance matrix is ​​superimposed with the result of the state prediction operation to obtain the prior error covariance matrix of the current moment.

[0120] The state estimation method of the silo group system based on the state transfer matrix and the state deviation vector first needs to establish a complete state transfer matrix. This matrix reflects the material flow characteristics between the silos through the material transfer relationship, including parameters such as the feed rate and the discharge rate; describes the temperature change law inside the silo through the heat transfer relationship, considering the influencing factors such as heat conduction and heat convection; and describes the pressure distribution characteristics in the silo through the pressure balance relationship, involving physical quantities such as static pressure and dynamic pressure. In specific implementation, taking a five-link parallel silo system as an example, the material level change rate of each silo is related to the material transfer rate of the adjacent silo, the temperature change is affected by the external environment temperature and the material temperature, and the pressure change is closely related to the material level height and the operating status of the ventilation system.

[0121] In the process of obtaining real-time operation data, a distributed sensor network is used to collect system status information. The material level sensor adopts the principle of ultrasonic ranging, with a sampling period of 100 milliseconds and a measurement accuracy of plus or minus 1 cm; the temperature sensor uses a PT100 thermal resistor, with a sampling period of 500 milliseconds and a measurement accuracy of plus or minus 0.1 degrees Celsius; the pressure sensor uses a piezoelectric sensor, with a sampling period of 200 milliseconds and a measurement accuracy of plus or minus 0.1 kPa. After filtering and preprocessing, all sensor data form a state deviation vector, which is used to characterize the deviation between the actual state of the system and the expected state.

[0122] In the process of prior state estimation, the state transfer matrix is ​​applied to the state vector of the previous moment to obtain the state prediction value. At the same time, the state deviation vector is introduced for correction, and the correction coefficient is determined by the offline optimization method. Taking an actual operating condition as an example, the level deviation correction coefficient is 0.85, the temperature deviation correction coefficient is 0.92, and the pressure deviation correction coefficient is 0.88. The selection of these correction coefficients fully considers the influence of system dynamic characteristics and measurement noise.

[0123] When calculating the prior error covariance matrix, the state transfer matrix is ​​first used to predict the posterior error covariance matrix of the previous moment. This process reflects the transmission characteristics of the uncertainty of the system state in the time dimension. At the same time, the system process noise covariance matrix is ​​constructed, which reflects the degree of influence of external interference factors on the system state. In practical applications, through long-term operation data statistical analysis, it is determined that the standard deviation of the material level measurement noise is 2 cm, the standard deviation of the temperature measurement noise is 0.2 degrees Celsius, and the standard deviation of the pressure measurement noise is 0.3 kPa. The prediction operation results are superimposed with the process noise covariance matrix to finally obtain the prior error covariance matrix at the current moment.

[0124] The solution of this application can:

[0125] By establishing a state transfer matrix based on the relationship between material transfer, heat transfer and pressure balance, an accurate description of the multi-physical quantity coupling dynamic characteristics of the silo group system is achieved, and the reliability of state estimation is improved. A distributed sensor network is used to collect system operation data in real time, and a state deviation vector is introduced for weighted compensation, which effectively improves the accuracy and robustness of state estimation and makes the estimation results closer to actual working conditions. The prior error covariance matrix is ​​calculated by superimposing the state prediction operation and the noise covariance matrix, which reasonably characterizes the uncertainty characteristics of the system state and provides a reliable error evaluation basis for subsequent filtering estimation.

[0126] In an optional embodiment, the state space corresponds to the modified silo group state data, the action space corresponds to the executable control instructions of the silo group, and the reward function calculates the reward value according to the execution effect of the executable control instructions; and using a deep Q learning network to perform iterative training in the reinforcement learning environment to obtain the optimal control strategy includes:

[0127] The corrected silo group state data is normalized to obtain normalized state data, and the normalized state data is constructed as a state space of a reinforcement learning environment, wherein the state space includes a state component set consisting of a material level state component, a temperature state component, and a pressure state component; an action space is constructed based on the control requirements of the state component set, wherein the action space includes a material inlet and outlet control instruction corresponding to the material level state component, a heating and cooling control instruction corresponding to the temperature state component, and a pressure increase and pressure relief control instruction corresponding to the pressure state component, and the control instructions in the action space are all normalized to a range of zero to one;

[0128] Constructing a reward function for evaluating the execution effect of the executable control instruction, after executing the executable control instruction based on the reward function, determining the deviation between the state of the silo group and the target state to obtain a reward value, wherein the reward value includes a material level control reward component, a temperature control reward component, and a pressure control reward component;

[0129] Mapping the dimension of the state space to the number of input layer nodes of the deep Q learning network, mapping the dimension of the action space to the number of output layer nodes of the deep Q learning network, constructing a three-layer hidden layer structure and adopting a ReLU activation function; constructing a reinforcement learning environment based on the state space, the action space and the reward function, and adopting an ε-greedy strategy with decreasing exploration probability in the reinforcement learning environment to iteratively train the deep Q learning network;

[0130] The state transition sequence generated by the interaction between the deep Q learning network and the environment is stored in an experience replay pool, and random sampling is used to update network parameters from the experience replay pool. The target network is used to calculate the target Q value and the network parameters are synchronized regularly. The training effect is evaluated according to the reward value calculated by the reward function. When the cumulative reward value of multiple consecutive training rounds has no obvious improvement, the learning rate is reduced until the network converges to obtain the optimal control strategy.

[0131] First, the silo group status data is normalized and preprocessed. For the raw data collected by the three types of sensors, namely, level, temperature and pressure, the maximum and minimum value normalization method is used to map the data to the range of zero to one. For example, for the level data, if the maximum level of a silo is 10 meters, the minimum level is 0 meters, and the current level is 6 meters, the normalized level value is 0.6. The normalization interval of temperature data can be set to 0-100 degrees Celsius, and the normalization interval of pressure data can be set to 0-2 MPa.

[0132] When constructing the state space, the normalized material level, temperature, and pressure data are combined into a state vector. Assuming that the silo group contains 5 silos, each silo has 3 state components, the dimension of the state vector is 15. The value range of each component is between zero and one.

[0133] The construction of the action space is based on the control requirements of each state component. For material level control, there are two discrete actions: feeding and discharging; temperature control includes two discrete actions: heating and cooling; pressure control includes two discrete actions: pressurization and pressure relief. These discrete actions are also normalized to the range of zero to one. For example, the feed rate can be set to the range of 0-100 tons per hour.

[0134] The design of the reward function adopts a deviation evaluation method based on the target state. For material level control, a positive reward is given when the deviation between the actual material level and the target material level is within the allowable range, and a negative reward is given when the deviation is too large. The reward calculation method for temperature control and pressure control is similar, taking into account control accuracy and energy consumption factors. For example, a full reward of 100 points is given when the material level deviation is within the range of plus or minus 5%, and zero or negative points are given when the deviation exceeds 20%.

[0135] The construction of the deep Q learning network adopts a fully connected neural network structure. The number of nodes in the input layer is the same as the dimension of the state vector, which is 15, and the number of nodes in the output layer is the same as the dimension of the action space. The number of nodes in the three hidden layers can be set to 128, 256, and 128 respectively, and the activation function uses the ReLU function.

[0136] During the training process, the initial value of the exploration probability ε is set to 0.9 and decreases linearly to 0.1 with each training round. The capacity of the experience replay pool is set to 10,000 state transition sequences, and the batch size of each random sampling is 32. The parameter update cycle of the target network is once every 100 rounds of training.

[0137] During the training process, when the average cumulative reward value for 20 consecutive rounds increases by less than 1%, the learning rate is reduced to 0.1 times the original value. The initial learning rate is set to 0.001, and the minimum learning rate is set to 0.00001. When the fluctuation range of the average cumulative reward value for 50 consecutive rounds is less than 0.5%, the network is considered to have converged.

[0138] The solution of this application can:

[0139] By normalizing the state data of the silo group and constructing a reasonable state space and action space, unified modeling and standardized control of the multi-silo system are achieved, and the generalization and robustness of the control strategy are improved. The deep Q learning network combined with experience replay and target network mechanism overcomes the limitations of traditional reinforcement learning algorithms in continuous state space and significantly improves the learning efficiency and convergence performance of the control strategy. Based on the evaluation mechanism of multi-dimensional reward function, the coordinated optimization of multiple control targets such as material level, temperature, and pressure is achieved, while ensuring the control accuracy and taking into account the system energy consumption, achieving the dual goals of intelligence and energy saving and consumption reduction.

[0140] In an optional implementation, the distributed control system receives the optimal control strategy, calculates the control parameters of each silo based on the consistency protocol, and optimizes the control parameters through a model predictive control algorithm to generate an optimal control instruction sequence, including:

[0141] Based on the distributed control system, the optimal control strategy generated by deep reinforcement learning is received; the optimal control strategy is parsed into control parameters of each silo, wherein the control parameters include a feed rate parameter, a discharge rate parameter, a temperature adjustment parameter, and a pressure adjustment parameter;

[0142] A dynamic update equation of the consistency protocol is constructed based on the control parameters, wherein the dynamic update equation adopts a second-order consistency algorithm and takes inertia weight and coupling strength as dynamic adjustment parameters;

[0143] Establish a material level prediction model according to the material balance relationship of the silo, establish a temperature prediction model according to the energy balance relationship, establish a pressure prediction model according to the pressure balance relationship, and construct the material level prediction model, the temperature prediction model and the pressure prediction model into a state prediction model;

[0144] Constructing an objective function of model predictive control based on the state prediction model, wherein the objective function includes a state error term and a control increment term, wherein the state error term represents the deviation between the predicted state and the target state, and the control increment term is used to constrain the change rate of the control amount;

[0145] The objective function is constructed as a quadratic programming problem, and the constraints include material level constraint, temperature constraint, pressure constraint, control quantity constraint and control increment constraint. The quadratic programming problem is solved to obtain an optimal control instruction sequence.

[0146] The distributed control system first receives the optimal control strategy output by the deep reinforcement learning model. The control strategy contains the key control parameter information of each silo, including feed rate parameters, discharge rate parameters, temperature adjustment parameters and pressure adjustment parameters. Taking a chemical production line as an example, the feed rate parameter range is 0-100 tons per hour, the discharge rate parameter range is 0-80 tons per hour, the temperature adjustment parameter range is 0-200 degrees Celsius, and the pressure adjustment parameter range is 0-10 MPa.

[0147] Next, the system builds a dynamic consistency protocol based on these control parameters. A second-order consistency algorithm is used for parameter collaborative optimization, where the inertia weight ranges from 0.1 to 0.9 and the coupling strength ranges from 0.2 to 0.8. By dynamically adjusting these two parameters, rapid convergence of the control parameters of each silo can be achieved. In practical applications, the inertia weight is usually set to 0.5 and the coupling strength is set to 0.6, which allows the system to reach consistency within 5-10 iterations.

[0148] Then, the system establishes a prediction model based on the actual process requirements. The material level prediction model is based on the material balance relationship, taking into account factors such as feed volume, discharge volume and silo volume; the temperature prediction model is based on the energy balance relationship, taking into account factors such as heating power, heat loss and material specific heat capacity; the pressure prediction model is based on the pressure balance relationship, taking into account factors such as feed pressure, discharge pressure and silo structural characteristics. These three models together constitute the state prediction model, and the prediction time domain is usually set to 10-20 sampling cycles.

[0149] Based on the state prediction model, the system constructs the objective function of model predictive control. The objective function consists of two parts: the state error term and the control increment term. The state error term is used to measure the deviation between the predicted state and the target state, and the control increment term is used to limit the rate of change of the control quantity. In practical applications, the state error weight is usually set to 0.7 and the control increment weight is set to 0.3, which can achieve a good balance between control performance and stability.

[0150] Finally, the system converts the objective function into a standard quadratic programming problem for solution. Constraints include: material level constraint (usually 20%-80% of the silo volume), temperature constraint (upper and lower limits of process temperature), pressure constraint (pressure range of equipment), control quantity constraint (actuator range) and control increment constraint (actuator adjustment rate). By solving this optimization problem, the optimal control instruction sequence for several future control cycles can be obtained. The typical control cycle is 1 minute, and the prediction time domain is 15 cycles, so the control instruction sequence for the next 15 minutes can be obtained.

[0151] The solution of this application can:

[0152] This technical solution combines deep reinforcement learning with distributed control to achieve intelligent optimization control of complex industrial processes, significantly improving the control performance and production efficiency of the system. The use of a second-order consistency algorithm and a dynamic parameter adjustment mechanism ensures the coordination and robustness of the multi-silo system, effectively solving the communication burden and single-point failure problems faced by traditional centralized control. Based on the model predictive control framework, the system can predictably handle various constraints and generate the optimal control sequence through online optimization, significantly improving control accuracy and system stability, and reducing energy consumption and operating costs.

[0153] In an optional embodiment, the distributed control system performs collaborative control on the silo group according to the optimal control instruction sequence, collects control effect data in real time; feeds back the control effect data to the deep neural network model to update the training data set of the silo group state prediction model, and realizes continuous optimization of the optimal control strategy, including:

[0154] The distributed control system controls the silo group based on the optimal control instruction sequence, collects real-time operation data of the silo group; and divides the real-time operation data into training samples according to a time series;

[0155] Constructing an online learning module of a deep neural network, the online learning module comprising a data preprocessing unit, an incremental learning unit and a model updating unit, the data preprocessing unit performing standardization processing and outlier detection on training samples;

[0156] The incremental learning unit extracts new data from the training samples based on a sliding time window, and fuses the new data with the historical training data set to generate an updated training data set;

[0157] Incrementally training the state prediction model of the deep neural network based on the updated training data set, using a dynamic learning rate adjustment strategy to adaptively adjust network parameters according to the prediction error;

[0158] Calculating the prediction accuracy of the state prediction model on the validation data set based on the model updating unit, and when the prediction accuracy meets the update condition, updating the trained network parameters to the state prediction model running online;

[0159] The optimal control strategy is recalculated based on the updated state prediction model to obtain an updated optimal control strategy; the updated optimal control strategy is sent to the distributed control system to guide the real-time control of the silo group, and new real-time operation data continues to be collected for the next round of online learning optimization.

[0160] First, the distributed control system controls the silo group in real time according to the optimal control instruction sequence. The control system collects the operation data of the silo group, including key parameters such as material level, material discharge rate, silo temperature, humidity, etc. The collection frequency is once every 5 seconds, and the continuous 24-hour operation data is divided into 10-minute time segments as training samples.

[0161] The data preprocessing unit in the online learning module performs standardization on the collected training samples. For the material level data, it is normalized to the range of 0-1; for the temperature data, the maximum and minimum value standardization method is used to process it to the standard range. At the same time, the box plot method is used to detect outliers. When the data deviates from the mean by more than three times the standard deviation, it is determined to be an outlier and corrected.

[0162] The incremental learning unit uses a 60-minute sliding time window to extract incremental data from the latest training samples. This new data is merged with the historical training data saved in the last 7 days to generate an updated training data set. The data is merged using a time-weighted method, with the weight of the latest data being 1.0 and decreasing by 0.1 every day to ensure that new data has a greater impact on model updates.

[0163] The state prediction model adopts a deep neural network structure, which contains 4 hidden layers, and the number of neurons in each layer is 128, 64, 32, and 16 respectively. In the incremental training process, a dynamic learning rate adjustment strategy is adopted, and the initial learning rate is set to 0.01. When the prediction error decreases by less than 0.1% for 5 consecutive iterations, the learning rate is adjusted to 0.8 times the original value, and the minimum learning rate is limited to 0.001.

[0164] The model update unit uses the last 24 hours of operating data as a validation set to calculate the prediction accuracy of the state prediction model. When the root mean square error of the prediction accuracy is less than the set threshold of 0.05 and is improved by more than 5% compared with the existing model, the model update operation is triggered. The trained network parameters are updated to the online state prediction model.

[0165] Based on the updated state prediction model, the control system recalculates the optimal control strategy. By predicting the state changes of the silo group in the next 4 hours, an optimized control instruction sequence is generated. The new control strategy takes into account multiple goals such as material level balance and minimum energy consumption, and generates the optimal unloading rate instruction for each silo.

[0166] Finally, the updated control strategy is sent to the execution unit of the distributed control system. The execution unit sends control instructions to the actuators of each silo according to a 5-second control cycle to achieve coordinated control of the silo group. At the same time, new operating data is continuously collected for the next round of online learning optimization.

[0167] The solution of this application can:

[0168] Through deep neural networks, accurate prediction of the state of the silo group is achieved, with a prediction accuracy of more than 95%, providing a reliable decision-making basis for optimizing the control strategy and significantly improving the system control effect. The control strategy is continuously optimized using online learning. The system can adaptively adapt to changes in working conditions and equipment performance to maintain the optimal control effect, and the energy-saving efficiency is improved by more than 20% compared with traditional control methods. Based on a distributed architecture, multi-silo collaborative control is achieved. The system has good scalability and robustness. The failure of a single silo does not affect the operation of the entire system, and the system availability reaches more than 99.9%.

[0169] Figure 2 FIG. 1 is a schematic diagram of the structure of the feed bin intelligent management system based on adaptive control according to an embodiment of the present invention. Figure 2 As shown, the system comprises:

[0170] The first unit is used to input the material level data, temperature data and pressure data of each silo in the silo group collected by multiple sensors into a silo group state prediction model to obtain a real-time state prediction result of the silo group, wherein the silo group state prediction model is pre-trained based on a deep neural network combined with historical operation data; compare the real-time state prediction result of the silo group with the actual collected data, correct the deviation data based on a Kalman filter algorithm, and generate corrected silo group state data;

[0171] The second unit is used to construct a reinforcement learning environment according to the modified silo group state data, establish a state space, an action space and a reward function in the reinforcement learning environment, wherein the state space corresponds to the modified silo group state data, the action space corresponds to the executable control instructions of the silo group, and the reward function calculates the reward value according to the execution effect of the executable control instructions; use a deep Q learning network to perform iterative training in the reinforcement learning environment to obtain an optimal control strategy, and output the optimal control strategy to a distributed control system;

[0172] The third unit is used for the distributed control system to receive the optimal control strategy, calculate the control parameters of each silo based on the consistency protocol, and optimize the control parameters through the model predictive control algorithm to generate an optimal control instruction sequence; the distributed control system performs collaborative control on the silo group according to the optimal control instruction sequence, and collects control effect data in real time; the control effect data is fed back to the silo group state prediction model to update the training data set of the silo group state prediction model, so as to achieve continuous optimization of the optimal control strategy.

[0173] According to a third aspect of the embodiments of the present invention,

[0174] An electronic device is provided, comprising:

[0175] processor;

[0176] a memory for storing processor-executable instructions;

[0177] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0178] According to a fourth aspect of the embodiments of the present invention,

[0179] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0180] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0181] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. The intelligent management method of feed bin based on adaptive control is characterized by: include: Input the material level data, temperature data and pressure data of each silo in the silo group collected by multiple sensors into the silo group state prediction model to obtain the real-time state prediction result of the silo group. The silo group state prediction model is pre-trained based on a deep neural network combined with historical operation data; Comparing the real-time state prediction result of the silo group with the actual collected data, correcting the deviation data based on the Kalman filter algorithm, and generating corrected silo group state data; Constructing a reinforcement learning environment according to the modified silo group state data, establishing a state space, an action space and a reward function in the reinforcement learning environment, wherein the state space corresponds to the modified silo group state data, the action space corresponds to the control instructions executable by the silo group, and the reward function calculates the reward value according to the execution effect of the executable control instructions; performing iterative training in the reinforcement learning environment using a deep Q learning network to obtain an optimal control strategy, and outputting the optimal control strategy to a distributed control system; The distributed control system receives the optimal control strategy, calculates the control parameters of each silo based on the consistency protocol, optimizes the control parameters through the model predictive control algorithm, and generates an optimal control instruction sequence; the distributed control system performs collaborative control on the silo group according to the optimal control instruction sequence, and collects control effect data in real time; the control effect data is fed back to the silo group state prediction model to update the training data set of the silo group state prediction model, thereby realizing continuous optimization of the optimal control strategy.

2. The method according to claim 1, characterized in that The real-time state prediction result of the silo group is compared with the actual collected data, and the deviation data is corrected based on the Kalman filter algorithm to generate the corrected silo group state data including: The material level data, the temperature data and the pressure data are combined into a measured state vector, the predicted data output by the deep neural network is combined into a predicted state vector, and the difference between the measured state vector and the predicted state vector is calculated to obtain a state deviation vector; A dynamic weight calculation model is established based on historical performance evaluation data of multiple sensors, wherein the dynamic weight calculation model generates a measurement uncertainty weight matrix according to stability indicators and historical error statistics of the multiple sensors, and the weight coefficients in the measurement uncertainty weight matrix represent the measurement credibility of the multiple sensors; The state deviation vector and the measurement uncertainty weight matrix are input into an improved Kalman filter, the improved Kalman filter predicts the state deviation vector according to the state transfer matrix to obtain a priori state estimation value, and calculates the error covariance corresponding to the priori state estimation value; An adaptive factor is calculated based on the innovative sequence covariance matrix, a product of the adaptive factor and the error covariance is used as a corrected prediction error covariance, and the improved Kalman filter calculates a Kalman gain according to the corrected prediction error covariance; The improved Kalman filter corrects the priori state estimate according to the Kalman gain to obtain a posterior state estimate, updates the error covariance, and uses the posterior state estimate as the corrected silo group state data.

3. The method according to claim 2, characterized in that Inputting the state deviation vector and the measurement uncertainty weight matrix into an improved Kalman filter, wherein the improved Kalman filter predicts the state deviation vector according to the state transfer matrix to obtain a priori state estimation value, and calculating the error covariance corresponding to the priori state estimation value includes: For each measurement component in the state deviation vector, extract the device parameter set of the corresponding sensor, the device parameter set includes a measurement accuracy level parameter, an installation environment impact parameter, a usage time parameter and a historical data fluctuation parameter, and construct the device parameter set as a measurement uncertainty feature vector; The measurement uncertainty characteristic vector is processed by a fuzzy rule reasoning method to obtain a corresponding measurement uncertainty weight value, and the weight values ​​of all measurement components are constructed into a measurement uncertainty weight matrix in the form of a diagonal matrix; Substituting the measurement uncertainty weight matrix into an improved Kalman gain calculation formula, wherein the improved Kalman gain calculation formula includes a priori error covariance matrix and a measurement matrix; The improved Kalman gain calculation formula is as follows: ; Among them, K k is the Kalman gain at time k, P k - is the prior error covariance matrix at time k, H k is the measurement matrix at time k, W k is the measurement uncertainty weight matrix at time k, R k is the measurement noise covariance matrix at time k; Calculating a priori state estimation values ​​based on a state transfer matrix and the state deviation vector, and calculating a priori error covariance matrix using the state transfer matrix; The prior state estimate is calculated as follows: ; Among them, x k - is the prior state estimate at time k, Φ k,k-1 is the state transfer matrix, x k-1 is the state vector at time k-1, B k is the control input matrix, u k-1 is the control input at time k-1; The prior error covariance matrix calculation formula is as follows: ; Among them, P k-1 is the posterior error covariance matrix at time k-1, Q k-1 is the system process noise covariance matrix at time k-1.

4. The method according to claim 3, characterized in that Calculating a priori state estimation values ​​based on the state transfer matrix and the state deviation vector, and calculating a priori error covariance matrix using the state transfer matrix, the previous moment a posteriori error covariance matrix and the system process noise covariance matrix includes: A state transfer matrix is ​​established based on the material transmission relationship, heat transfer relationship, and pressure balance relationship, and the state transfer matrix is ​​used to describe the coupling dynamic characteristics between various state quantities of the silo group system; The operation data of the silo group is collected through the material level sensor, temperature sensor and pressure sensor, and the real-time changes of the material level deviation, temperature deviation and pressure deviation are reflected based on the state deviation vector; the state transfer matrix is ​​applied to the state vector at the previous moment, and weighted compensation is performed in combination with the state deviation vector, and the prior state estimation value at the current moment is calculated according to the compensation coefficient optimized offline, and the prior state estimation value at the current moment represents the prediction result of the system state; The state transfer matrix is ​​used to perform a state prediction operation on the posterior error covariance matrix of the previous moment, wherein the posterior error covariance matrix of the previous moment reflects the uncertainty characteristics of the system state, and the uncertainty characteristics of the system state are transferred to the current moment through the state prediction operation; a system process noise covariance matrix reflecting external interference, modeling error, and control fluctuation is constructed, and the system process noise covariance matrix is ​​superimposed with the result of the state prediction operation to obtain the prior error covariance matrix of the current moment.

5. The method according to claim 1, characterized in that The state space corresponds to the modified silo group state data, the action space corresponds to the executable control instructions of the silo group, and the reward function calculates the reward value according to the execution effect of the executable control instructions; Using the deep Q learning network to perform iterative training in the reinforcement learning environment, the optimal control strategy is obtained including: The corrected silo group state data is normalized to obtain normalized state data, and the normalized state data is constructed as a state space of a reinforcement learning environment, wherein the state space includes a state component set consisting of a material level state component, a temperature state component, and a pressure state component; an action space is constructed based on the control requirements of the state component set, wherein the action space includes a material inlet and outlet control instruction corresponding to the material level state component, a heating and cooling control instruction corresponding to the temperature state component, and a pressure increase and pressure relief control instruction corresponding to the pressure state component, and the control instructions in the action space are all normalized to a range of zero to one; Constructing a reward function for evaluating the execution effect of the executable control instruction, after executing the executable control instruction based on the reward function, determining the deviation between the state of the silo group and the target state to obtain a reward value, wherein the reward value includes a material level control reward component, a temperature control reward component, and a pressure control reward component; Mapping the dimension of the state space to the number of input layer nodes of the deep Q learning network, mapping the dimension of the action space to the number of output layer nodes of the deep Q learning network, constructing a three-layer hidden layer structure and adopting a ReLU activation function; constructing a reinforcement learning environment based on the state space, the action space and the reward function, and adopting an ε-greedy strategy with decreasing exploration probability in the reinforcement learning environment to iteratively train the deep Q learning network; The state transition sequence generated by the interaction between the deep Q learning network and the environment is stored in an experience replay pool, and random sampling is used to update network parameters from the experience replay pool. The target network is used to calculate the target Q value and the network parameters are synchronized regularly. The training effect is evaluated according to the reward value calculated by the reward function. When the cumulative reward value of multiple consecutive training rounds has no obvious improvement, the learning rate is reduced until the network converges to obtain the optimal control strategy.

6. The method according to claim 1, characterized in that The distributed control system receives the optimal control strategy, calculates the control parameters of each silo based on the consistency protocol, and optimizes the control parameters through the model predictive control algorithm to generate the optimal control instruction sequence, including: Based on the distributed control system, the optimal control strategy generated by deep reinforcement learning is received; the optimal control strategy is parsed into control parameters of each silo, wherein the control parameters include a feed rate parameter, a discharge rate parameter, a temperature adjustment parameter, and a pressure adjustment parameter; A dynamic update equation of the consistency protocol is constructed based on the control parameters, wherein the dynamic update equation adopts a second-order consistency algorithm and takes inertia weight and coupling strength as dynamic adjustment parameters; Establish a material level prediction model according to the material balance relationship of the silo, establish a temperature prediction model according to the energy balance relationship, establish a pressure prediction model according to the pressure balance relationship, and construct the material level prediction model, the temperature prediction model and the pressure prediction model into a state prediction model; Constructing an objective function of model predictive control based on the state prediction model, wherein the objective function includes a state error term and a control increment term, wherein the state error term represents the deviation between the predicted state and the target state, and the control increment term is used to constrain the change rate of the control amount; The objective function is constructed as a quadratic programming problem, and the constraints include material level constraints, temperature constraints, pressure constraints, control quantity constraints and control increment constraints. The quadratic programming problem is solved to obtain an optimal control instruction sequence.

7. The method according to claim 1, characterized in that The distributed control system performs collaborative control on the silo group according to the optimal control instruction sequence, collects control effect data in real time, and feeds back the control effect data to the deep neural network model to update the training data set of the silo group state prediction model, thereby achieving continuous optimization of the optimal control strategy, including: The distributed control system controls the silo group based on the optimal control instruction sequence, collects real-time operation data of the silo group; and divides the real-time operation data into training samples according to a time series; Constructing an online learning module of a deep neural network, the online learning module comprising a data preprocessing unit, an incremental learning unit and a model updating unit, the data preprocessing unit performing standardization processing and outlier detection on training samples; The incremental learning unit extracts new data from the training samples based on a sliding time window, and fuses the new data with the historical training data set to generate an updated training data set; Incrementally training the state prediction model of the deep neural network based on the updated training data set, using a dynamic learning rate adjustment strategy to adaptively adjust network parameters according to prediction errors; Calculating the prediction accuracy of the state prediction model on the validation data set based on the model updating unit, and when the prediction accuracy meets the update condition, updating the trained network parameters to the state prediction model running online; The optimal control strategy is recalculated based on the updated state prediction model to obtain an updated optimal control strategy; the updated optimal control strategy is sent to the distributed control system to guide the real-time control of the silo group, and new real-time operation data continues to be collected for the next round of online learning optimization.

8. An intelligent management system for feed bin based on adaptive control, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to input the material level data, temperature data and pressure data of each silo in the silo group collected by multiple sensors into the silo group state prediction model to obtain the real-time state prediction result of the silo group. The silo group state prediction model is pre-trained based on a deep neural network combined with historical operation data; Comparing the real-time state prediction result of the silo group with the actual collected data, correcting the deviation data based on the Kalman filter algorithm, and generating corrected silo group state data; The second unit is used to construct a reinforcement learning environment according to the modified silo group state data, establish a state space, an action space and a reward function in the reinforcement learning environment, wherein the state space corresponds to the modified silo group state data, the action space corresponds to the executable control instructions of the silo group, and the reward function calculates the reward value according to the execution effect of the executable control instructions; use a deep Q learning network to perform iterative training in the reinforcement learning environment to obtain an optimal control strategy, and output the optimal control strategy to a distributed control system; The third unit is used for the distributed control system to receive the optimal control strategy, calculate the control parameters of each silo based on the consistency protocol, and optimize the control parameters through the model predictive control algorithm to generate an optimal control instruction sequence; the distributed control system performs collaborative control on the silo group according to the optimal control instruction sequence, and collects control effect data in real time; the control effect data is fed back to the silo group state prediction model to update the training data set of the silo group state prediction model, so as to achieve continuous optimization of the optimal control strategy.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method and system for constructing optimal parameter library of grinding system

    CN116954077A

  • Production line data acquisition control method and system for intelligent workshop

    CN118311914A