Method and system for removing multiple obstacles in feeding bin based on prediction model
By adopting attention mechanism and deep reinforcement learning methods in the feed silo, the vibration frequency and discharge rate are adaptively adjusted, and the problem of lack of prediction and prevention in the obstruction treatment of the feed silo is solved, achieving efficient obstruction lifting and improving production stability.
Patent Information
- Application Number
- CN202510589035.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-05-08
AI Technical Summary
The existing feed silo barrier processing technology lacks prediction and prevention mechanisms, and cannot adaptively adjust the release strategy, resulting in a decrease in production efficiency and an increase in equipment loss, especially in complex operating conditions, with a low success rate of multiple barrier problems dealing with.
The attention mechanism is used to adaptively weighted fusion of material parameters and equipment status data in the feed silo, output the hinder risk index, and plan the optimal release path through the deep reinforcement learning control system, combine the historical release experience library to optimize the processing strategy, adjust the vibration frequency and discharge rate in real time to reduce risks.
Early identification and preventive control of hinder risks have been achieved, significantly reducing the frequency of hinder occurrence, improving the continuous operation capability and stability of the production line, reducing the need for manual intervention, and improving the efficiency and success rate of hinder removal.
Smart Images

Figure CN120106587B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technology of obstacle removal, and particularly to a method and system for removing multiple obstacles in a feeding bin based on a prediction model. Background Art
[0002] In the industrial production process, the feeding bin is a key device for material storage and transportation, and is widely used in industries such as metallurgy, building materials, chemical industry, and grain processing. The fluidity of the materials in the feeding bin directly affects the production efficiency and product quality. However, due to various material characteristics, complex environmental conditions, and equipment wear and other factors, the feeding bin often faces problems such as material obstacles, such as bridging, pipeline blockage, rat holes, and adhesion to the bin wall. These problems can lead to production interruption, equipment damage, and even safety accidents.
[0003] Traditional methods for dealing with obstacles in the feeding bin mainly rely on manual experience judgment and simple mechanical vibration methods. With the development of industrial automation and intelligence, intelligent monitoring systems based on sensor networks and data analysis have gradually been applied to the management of feeding bins. In recent years, with the rapid development of artificial intelligence technology, predictive maintenance and fault diagnosis methods based on deep learning have begun to be applied in the industrial field, providing new ideas for solving the problem of obstacles in the feeding bin.
[0004] The existing technologies for dealing with obstacles in the feeding bin mainly have the characteristics of passive response. Most systems can only detect and handle problems after the obstacles have formed and caused obvious production anomalies, lacking effective prediction and prevention mechanisms, resulting in a decrease in production efficiency and an increase in equipment wear.
[0005] The existing technologies usually adopt single or simple combined treatment strategies, such as vibrators or air flow impacts with fixed parameters, and cannot adaptively adjust the removal strategy according to different material characteristics and obstacle types. The removal effect is limited and it is easy to cause secondary obstacles. Especially for multiple obstacle problems under complex working conditions, the success rate of treatment is relatively low.
[0006] Traditional obstacle removal methods lack a systematic mechanism for experience accumulation and optimization, and it is difficult to learn from historical cases and improve the treatment strategy. The experience of operators is difficult to be effectively transformed into systematic knowledge, resulting in the obstacle removal process being highly dependent on manual experience, lacking standardized and intelligent solutions, and unable to meet the requirements of modern industry for automation and intelligence. Summary of the Invention
[0007] Embodiments of the present invention provide a method and system for removing multiple obstacles in a feeding bin based on a prediction model, which can solve the problems in the existing technologies.
[0008] In the first aspect of the embodiments of the present invention,
[0009] A method for removing multiple obstacles in a feeding bin based on a prediction model is provided, including:
[0010] The attention mechanism is used to adaptively weight and fuse the material parameter data and equipment status data in the feeding bin, and output the obstruction risk index of the feeding bin;
[0011] The obstruction risk index is divided into a first risk interval and a second risk interval. When the obstruction risk index is in the second risk interval, a preventive control strategy is triggered, and the obstruction risk index is reduced by adjusting the vibration frequency and the feeding rate;
[0012] When the obstruction risk index reaches the first risk interval or the signal of the level sensor is detected to be abnormal, it is determined that an obstruction has occurred, and the complete working condition data at the moment of the obstruction occurrence is recorded; based on the complete working condition data, an obstruction removal control system of deep reinforcement learning is started, the complete working condition data is input into the state encoder to generate a state vector, combined with the successful cases in the historical removal experience library, and the optimal removal path is planned through the policy network;
[0013] The obstruction removal operation is performed according to the optimal removal path. When it is detected that the signal of the level sensor returns to normal and the feeding rate is stable at the target value, it is determined that the obstruction removal is completed, and the complete data of this obstruction removal is recorded and stored in the historical experience library.
[0014] The obstruction risk index is divided into a first risk interval and a second risk interval. When the obstruction risk index is in the second risk interval, triggering a preventive control strategy to reduce the obstruction risk index by adjusting the vibration frequency and the feeding rate includes:
[0015] The probability distribution model of the obstruction risk index is established by using the kernel density estimation method. Based on the probability distribution model, the mean value and the standard deviation of the obstruction risk index are calculated. The upper threshold of the first risk interval is determined by subtracting the product of the first adjustment coefficient and the standard deviation from the mean value, and the upper threshold of the second risk interval is determined by adding the product of the second adjustment coefficient and the standard deviation to the mean value;
[0016] When the obstruction risk index is greater than the upper threshold of the first risk interval and less than the upper threshold of the second risk interval, a preventive control objective function including an obstruction risk index term, a system energy consumption term, and an equipment stress term is constructed, and the vibration frequency adjustment amount and the feeding rate adjustment amount are calculated based on the preventive control objective function;
[0017] The obstruction prevention control is performed according to the vibration frequency adjustment amount and the feeding rate adjustment amount. During the process of the obstruction prevention control, the change rate of the obstruction risk index, the change rate of the system energy consumption, and the change rate of the equipment stress are calculated in real time, so as to reduce the obstruction risk index.
[0018] Construct a preventive control objective function that includes a hindrance risk index term, a system energy consumption term, and a device stress term. Calculating the vibration frequency adjustment amount and the blanking rate adjustment amount based on the preventive control objective function includes:
[0019] Obtain the historical system operation data within a preset time window. The historical system operation data includes vibration frequency, blanking rate, and device stress. Calculate the corresponding mean and standard deviation for the historical system operation data, and complete the standardization process by subtracting the corresponding mean from the currently collected vibration frequency and blanking rate and then dividing by the corresponding standard deviation;
[0020] Based on the standardized vibration frequency and blanking rate, construct a non - linear mapping model, calculate the deviation between the output value of the non - linear mapping model and a preset target value to obtain a hindrance risk index, and the hindrance risk index is used to characterize the deviation degree of system operation;
[0021] Within the preset time window, calculate the difference in vibration frequency between adjacent sampling times to obtain a vibration frequency change sequence, calculate the difference in blanking rate between adjacent sampling times to obtain a blanking rate change sequence, and take the product of the sum of squares of the vibration frequency change sequence and the blanking rate change sequence and an energy consumption penalty factor as the system energy consumption evaluation value;
[0022] Weight the vibration frequency and the blanking rate to obtain a stress accumulation term, and add the stress accumulation term to the current device stress to obtain a device stress evaluation value;
[0023] According to the comparison results of the hindrance risk index, the system energy consumption evaluation value, and the device stress evaluation value with their corresponding benchmark operation thresholds respectively, dynamically adjust the weight coefficients of the three evaluation indicators, and add the products of the three evaluation indicators and the corresponding weight coefficients to obtain a comprehensive objective function value;
[0024] Calculate the rate of change of the comprehensive objective function value with respect to the vibration frequency, multiply the rate of change by a vibration frequency learning coefficient to obtain a vibration frequency adjustment amount, calculate the rate of change of the comprehensive objective function value with respect to the blanking rate, and multiply the rate of change by a blanking rate learning coefficient to obtain a blanking rate adjustment amount.
[0025] Based on the complete working condition data, start a hindrance removal control system for deep reinforcement learning. Input the complete working condition data into a state encoder to generate a state vector, and combine the successful cases in the historical removal experience library. Plan an optimal removal path through a policy network, including:
[0026] Construct a working condition state vector based on the complete working condition data, where the working condition state vector includes material parameters, vibration parameters, equipment parameters, and environmental parameters; input the working condition state vector into a state encoder, and obtain the encoded state vector through the state encoder;
[0027] Obtain the state feature vector, control action, immediate reward, and next state vector in the historical operation data of the feeding bin, form a state transition tuple with the state feature vector, the control action, the immediate reward, and the next state vector, and store the state transition tuple in the historical release experience library;
[0028] Construct an obstacle removal trajectory based on the state transition tuple, calculate the ratio of the number of immediate rewards greater than the reward threshold in the obstacle removal trajectory to the trajectory length to obtain the success rate, and screen the successful cases in the historical release experience library according to the success rate;
[0029] Generate a mean function and a variance function based on the encoded state vector, construct a Gaussian distribution according to the mean function and the variance function, and sample from the Gaussian distribution to obtain a release action;
[0030] Execute the release action and obtain a new immediate reward, calculate a value function based on the new immediate reward, where the value function is the expected cumulative reward with a discount factor; adjust the network parameters of the policy network according to the expected cumulative reward, update the execution result to the historical release experience library, and plan to obtain an optimal release path.
[0031] Constructing an obstacle removal trajectory based on the state transition tuple, calculating the ratio of the number of immediate rewards greater than the reward threshold in the obstacle removal trajectory to the trajectory length to obtain the success rate, and screening the successful cases in the historical release experience library includes:
[0032] Collect the state data and execution actions during the obstacle removal process of the feeding bin, where the state data includes the material level height, the feeding rate, and the vibration frequency, form a state transition tuple with the state data and the execution action at the current moment and the state data at the next moment, and connect multiple state transition tuples in chronological order to construct an obstacle removal trajectory;
[0033] For each state transition tuple in the obstacle removal trajectory, obtain the immediate reward by dividing the absolute value of the difference between the current material level height and the target material level height by the material level tolerance, count the number of state transition tuples with the immediate reward greater than 0.8, and divide the number of state transition tuples by the total number of state transition tuples included in the obstacle removal trajectory to obtain the trajectory success rate;
[0034] Input the state data in the obstacle removal trajectory into the feature extraction model, calculate the Euclidean distance between adjacent states, divide the Euclidean distance by the state range to obtain the normalized state distance, and cluster the obstacle removal trajectory based on the normalized state distance;
[0035] Mark the obstacle removal trajectories corresponding to the trajectory success rate greater than the preset success rate threshold as successful cases.
[0036] Calculate the value function based on the immediate reward, where the value function is the expected cumulative reward with a discount factor; adjust the network parameters of the policy network according to the value function, and update the execution result to the historical removal experience library. The steps to plan the optimal removal path include:
[0037] Collect the discharge amount, feeding rate, and blockage degree of the feeding bin. Divide the difference between the discharge amount and the minimum discharge amount by the discharge amount change range to obtain the first reward value, divide the feeding rate by the rated feeding rate to obtain the second reward value, divide the blockage degree by the preset reward threshold to obtain the third reward value, and perform weighted summation on the first reward value, the second reward value, and the third reward value to obtain the comprehensive reward;
[0038] Calculate the state value function based on the current state data of the feeding bin and the comprehensive reward. The state value function is the cumulative reward with a discount factor starting from the current moment, and the value range of the discount factor is from 0 to 1; add the comprehensive reward to the product of the state value function of the next moment and the discount factor, and subtract the state value function of the current moment to obtain the temporal difference error;
[0039] Calculate the parameter gradient of the policy network based on the temporal difference error, multiply the parameter gradient by the preset learning rate, and update the network parameters of the policy network; sample the state of the feeding bin using the updated network parameters to generate a control trajectory including a state sequence and an action sequence;
[0040] Calculate the value function of each state in the control trajectory, sum the value functions of each moment multiplied by the corresponding discount factor to obtain the trajectory score; select the optimal removal path according to the trajectory score.
[0041] Perform the obstacle removal operation according to the optimal removal path. When it is detected that the signal of the level sensor returns to normal and the feeding rate stabilizes at the target value, it is determined that the obstacle removal is completed, and the complete data record of this obstacle removal is stored in the historical experience library, including:
[0042] Obtain the signal of the level sensor and the feeding rate. When the signal of the level sensor is lower than the first threshold and the feeding rate is lower than the second threshold, it is determined that the feeding bin is blocked, and the obstacle removal operation is triggered;
[0043] Monitor the material level sensor signal and the material discharging rate in real time. When the material level sensor signal returns to the normal range and the fluctuation amplitude of the material discharging rate is less than the preset value for a preset duration, it is determined that the obstruction removal is completed, and the status data, control parameters, and removal results of this obstruction removal are recorded in the historical experience database.
[0044] In the second aspect of the embodiments of the present invention,
[0045] A multi-obstruction removal system for a feeding bin based on a prediction model is provided, including:
[0046] A first unit for adaptively weighted fusion of the material parameter data and equipment status data in the feeding bin by using an attention mechanism, and outputting an obstruction risk index of the feeding bin;
[0047] A second unit for dividing the obstruction risk index into a first risk interval and a second risk interval. When the obstruction risk index is in the second risk interval, a preventive control strategy is triggered, and the obstruction risk index is reduced by adjusting the vibration frequency and the material discharging rate;
[0048] A third unit for determining that an obstruction has occurred when the obstruction risk index reaches the first risk interval or the material level sensor signal is detected to be abnormal, and recording the complete working condition data at the moment when the obstruction occurs; based on the complete working condition data, starting an obstruction removal control system of deep reinforcement learning, inputting the complete working condition data into a state encoder to generate a state vector, and combining successful cases in the historical removal experience database, planning an optimal removal path through a policy network;
[0049] A fourth unit for performing an obstruction removal operation according to the optimal removal path. When it is detected that the material level sensor signal returns to normal and the material discharging rate is stable at the target value, it is determined that the obstruction removal is completed, and the complete data of this obstruction removal is recorded and stored in the historical experience database.
[0050] In the third aspect of the embodiments of the present invention,
[0051] An electronic device is provided, including:
[0052] A processor;
[0053] A memory for storing instructions executable by the processor;
[0054] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0055] In the fourth aspect of the embodiments of the present invention,
[0056] Provided is a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are executed by a processor, the foregoing method is implemented.
[0057] The beneficial effects of this application are as follows:
[0058] The present invention uses an attention mechanism to adaptively weight and fuse the material parameter data and equipment status data in the feeding bin, outputs an obstruction risk index, and takes corresponding measures according to the risk interval, realizing the early identification and preventive control of obstruction risks, significantly reducing the frequency of obstructions, and improving the continuous operation ability and stability of the production line.
[0059] The obstruction removal control system of the present invention based on deep reinforcement learning can plan an optimal removal path based on the complete working condition data and successful cases in the historical removal experience library, realizing the intelligent and precise solution of obstruction problems, effectively reducing the need for manual intervention, and reducing the labor intensity and safety risks of operators.
[0060] The present invention constructs a complete closed-loop system for obstruction handling, stores the complete data records of each obstruction removal into the historical experience library, continuously enriches and optimizes the solution, and the system has the ability of continuous learning and self-improvement. As the usage time increases, the efficiency and success rate of obstruction removal will continue to improve, providing an effective solution for the intelligent upgrade of industrial production. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 It is a schematic flow chart of the method for removing multiple obstructions in the feeding bin based on the prediction model in the embodiment of the present invention;
[0062] Figure 2 It is a complete flow chart of reducing the obstruction risk index in the embodiment of the present invention;
[0063] Figure 3 It is a schematic diagram of comparative analysis of the obstruction risk index in the embodiment of the present invention;
[0064] Figure 4 It is a flow chart of planning the obstruction removal path in the embodiment of the present invention;
[0065] Figure 5 It is a flow chart of planning the optimal removal path based on the value function in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0067] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0068] Figure 1 FIG. 1 is a flow chart of a method for removing multiple obstacles from a feed bin based on a prediction model according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0069] Adopting the attention mechanism to perform adaptive weighted fusion of material parameter data and equipment status data in the feed silo, the obstruction risk index of the feed silo is output;
[0070] dividing the obstruction risk index into a first risk range and a second risk range, and triggering a preventive control strategy when the obstruction risk index is in the second risk range to reduce the obstruction risk index by adjusting the vibration frequency and the feeding rate;
[0071] When the obstruction risk index reaches the first risk range or an abnormal level sensor signal is detected, an obstruction is determined to have occurred and complete operating condition data at the time of the obstruction occurrence is recorded. Based on the complete operating condition data, a deep reinforcement learning-based obstruction removal control system is activated. The complete operating condition data is input into a state encoder to generate a state vector. Combined with successful cases in a historical removal experience database, an optimal removal path is planned through a policy network.
[0072] The obstacle removal operation is performed according to the optimal removal path. When it is detected that the material level sensor signal returns to normal and the material discharge rate is stable at the target value, it is determined that the obstacle removal is completed, and the complete data record of this obstacle removal is stored in the historical experience database.
[0073] In an optional embodiment, the obstruction risk index is divided into a first risk range and a second risk range. When the obstruction risk index is in the second risk range, a preventive control strategy is triggered to reduce the obstruction risk index by adjusting the vibration frequency and the feeding rate, including:
[0074] The probability distribution model of the obstruction risk index is established by using the kernel density estimation method. Based on the probability distribution model, the mean and standard deviation of the obstruction risk index are calculated. The upper threshold of the first risk interval is determined by subtracting the product of the first adjustment coefficient and the standard deviation from the mean, and the upper threshold of the second risk interval is determined by adding the product of the second adjustment coefficient and the standard deviation to the mean;
[0075] When the obstruction risk index is greater than the upper threshold of the first risk interval and less than the upper threshold of the second risk interval, a preventive control objective function including an obstruction risk index term, a system energy consumption term, and an equipment stress term is constructed, and the vibration frequency adjustment amount and the blanking rate adjustment amount are calculated based on the preventive control objective function;
[0076] Obstruction prevention control is performed according to the vibration frequency adjustment amount and the blanking rate adjustment amount. During the process of the obstruction prevention control, the change rate of the obstruction risk index, the change rate of the system energy consumption, and the change rate of the equipment stress are calculated in real time, so as to reduce the obstruction risk index.
[0077] The obstruction risk index is a key indicator for measuring the possibility of obstruction in the material conveying system and can be calculated through the fusion of various sensor data. A typical calculation method is to perform weighted summation according to certain weights by combining factors such as material moisture content, adhesion coefficient, blanking rate, and vibration frequency. For example, the calculation weights of the obstruction risk index of a certain blanking system are: material moisture content 35%, adhesion coefficient 25%, blanking rate 25%, and vibration frequency 15%. When the material moisture content is 8% (standardized value 0.72), the adhesion coefficient is 0.45 (standardized value 0.68), the blanking rate is 90 kg / min (standardized value 0.82), and the vibration frequency is 62 Hz (standardized value 0.45), the calculated obstruction risk index is 0.68.
[0078] To scientifically divide the risk interval, this method uses the kernel density estimation technique to establish the probability distribution model of the obstruction risk index. Kernel density estimation is a non-parametric statistical method that can infer the continuous probability density function based on historical data samples without presetting the distribution type. In practical applications, the Gaussian kernel function can be selected as the kernel function, and the bandwidth parameter is set to 0.08. Assuming that 500 historical obstruction risk index samples are collected, the probability distribution obtained through kernel density estimation shows that this index presents an approximately normal distribution characteristic, with a mean of 0.42 and a standard deviation of 0.15.
[0079] Based on the calculated mean and standard deviation, further determine the threshold values of the risk intervals. The calculation method for the upper threshold of the first risk interval is the mean minus the product of the first adjustment coefficient and the standard deviation. The first adjustment coefficient can be set to 0.8, so the upper threshold of the first risk interval = 0.42 - 0.8×0.15 = 0.30. The calculation method for the upper threshold of the second risk interval is the mean plus the product of the second adjustment coefficient and the standard deviation. The second adjustment coefficient can be set to 1.5, so the upper threshold of the second risk interval = 0.42 + 1.5×0.15 = 0.645. Thus, the hindrance risk index is divided into three intervals: the safe interval (0 - 0.30), the first risk interval (0.30 - 0.645), and the second risk interval (>0.645).
[0080] When the hindrance risk index is in the first risk interval, the system enters the warning state but does not trigger control actions; when the hindrance risk index is in the second risk interval, preventive control strategies are triggered. The core of preventive control is to construct a multi-objective control function, considering three aspects: hindrance risk, system energy consumption, and equipment stress.
[0081] The preventive control objective function consists of three terms: the hindrance risk index term, the system energy consumption term, and the equipment stress term. Among them, the hindrance risk index term is proportional to the square of the current hindrance risk index; the system energy consumption term is proportional to the square of the vibration frequency and the blanking rate; the equipment stress term is proportional to the cube of the vibration frequency. The weight coefficients between the terms can be adjusted according to actual application requirements. Generally, the weight of the hindrance risk index term can be set to 0.6, the weight of the system energy consumption term to 0.25, and the weight of the equipment stress term to 0.15.
[0082] Based on the constructed preventive control objective function, calculate the vibration frequency adjustment amount and the blanking rate adjustment amount. The calculation process uses the gradient descent method to find the optimal solution of the objective function through an iterative approach. The initial vibration frequency is 60Hz, the initial blanking rate is 85kg / min, the learning rate is set to 0.02, and the maximum number of iterations is 50. After calculation, the obtained vibration frequency adjustment amount is +8.5Hz, and the blanking rate adjustment amount is -12.7kg / min. That is, the vibration frequency is adjusted to 68.5Hz, and the blanking rate is adjusted to 72.3kg / min.
[0083] After the above control parameter adjustment, it is necessary to monitor the system operation status in real time and calculate the change rate of key indicators. The change rate of the obstruction risk index refers to the change amount of the obstruction risk index per unit time, which can be calculated by dividing the difference between the two measurement values before and after by the time interval. For example, if the obstruction risk index before adjustment is 0.68 and it drops to 0.52 after 10 minutes of adjustment, the change rate of the obstruction risk index is -0.016 / min. The change rate of system energy consumption refers to the change amount of energy consumption per unit time, which can be measured by a power meter. For example, if the power before adjustment is 3.2 kW and the power after adjustment is 3.5 kW, the change rate of system energy consumption is +0.3 kW. The change rate of equipment stress refers to the change amount of stress borne by the equipment per unit time, which can be measured by a strain gauge or a vibration sensor. For example, if the equipment stress before adjustment is 45 MPa and it is 52 MPa after adjustment, the change rate of equipment stress is +7 MPa.
[0084] According to the change rates of each index, the control parameters can be further optimized. If the obstruction risk index drops significantly while the increase in system energy consumption and equipment stress is not obvious, the current control parameters are maintained; if the obstruction risk index does not drop significantly while the system energy consumption or equipment stress increases too fast, it is necessary to adjust the weight ratio of the control parameters and reduce the weight of the energy consumption item or the stress item. Through this real-time feedback adjustment mechanism, the obstruction risk can be effectively reduced while ensuring the economy of system operation and the safety of equipment.
[0085] During the operation of the blanking system, the obstruction risk index may fluctuate due to factors such as changes in material properties and environmental conditions. When the obstruction risk index drops to the safe range (below 0.30), the normal operation parameters can be gradually restored, that is, the vibration frequency and blanking rate are adjusted back to the initial values. The restoration process should be carried out slowly to avoid system instability caused by parameter mutation.
[0086] The technology of this embodiment comes from the field of industrial process control, especially the application of preventive control and multi-objective optimization theory. In the prior art, the obstruction problem of the material conveying system mainly relies on post-detection and treatment, that is, taking measures to solve after the obstruction occurs, or using a simple early warning mechanism with a fixed threshold, lacking scientific risk zoning and preventive control capabilities.
[0087] Figure 2 The complete flowchart for reducing the obstruction risk index in the embodiment of the present invention is as follows:
[0088] The figure shows a probability distribution model for constructing an obstruction risk index using the kernel density estimation method. The mean and standard deviation of the obstruction risk index are calculated through this model. On this basis, the mean obtained from the calculation is subtracted by the product of the first adjustment coefficient and the standard deviation, and the result is determined as the upper threshold of the first risk interval. At the same time, the mean is added by the product of the second adjustment coefficient and the standard deviation to determine the upper threshold of the second risk interval, thus establishing the basis for risk classification. When the detected obstruction risk index during the system operation process is within a specific range, that is, greater than the upper threshold of the first risk interval but less than the upper threshold of the second risk interval, the system will automatically construct a comprehensive preventive control objective function that includes the obstruction risk index term, the system energy consumption term, and the equipment stress term. Based on this preventive control objective function, the system can calculate the required vibration frequency adjustment amount and the blanking rate adjustment amount. During the specific obstruction prevention and control process, the system will implement control measures according to the calculated vibration frequency adjustment amount and blanking rate adjustment amount, and monitor and calculate the change rates of the obstruction risk index, the system energy consumption, and the equipment stress in real time. Through this dynamic adjustment method, the control objective of reducing the obstruction risk index is ultimately achieved.
[0089] In contrast, this application uses the kernel density estimation method to establish a probability distribution model for the obstruction risk index and scientifically divides the risk intervals based on statistical principles, with stronger adaptability and accuracy. At the same time, a multi-objective control function that includes obstruction risk, system energy consumption, and equipment stress is constructed, taking into account economy and safety while reducing the obstruction risk. The starting point of the improvement is to enhance the continuous and stable operation ability of the material conveying system, reduce the occurrence frequency of obstruction events, and avoid energy consumption waste and equipment loss caused by over-control.
[0090] In an alternative implementation manner, constructing a preventive control objective function that includes the obstruction risk index term, the system energy consumption term, and the equipment stress term, and calculating the vibration frequency adjustment amount and the blanking rate adjustment amount based on the preventive control objective function includes:
[0091] Obtain the historical system operation data within a preset time window. The historical system operation data includes vibration frequency, blanking rate, and equipment stress. Calculate the corresponding mean and standard deviation for the historical system operation data, and complete the normalization process by dividing the currently collected vibration frequency and blanking rate minus the corresponding mean by the corresponding standard deviation;
[0092] Based on the normalized vibration frequency and blanking rate, construct a non-linear mapping model, calculate the deviation between the output value of the non-linear mapping model and a preset target value to obtain an obstruction risk index, and the obstruction risk index is used to characterize the deviation degree of the system operation;
[0093] Within the preset time window, calculate the difference in vibration frequencies at adjacent sampling moments to obtain a vibration frequency change sequence, calculate the difference in blanking rates at adjacent sampling moments to obtain a blanking rate change sequence, and take the product of the sum of the squares of the vibration frequency change sequence and the blanking rate change sequence and the energy consumption penalty factor as the system energy consumption evaluation value;
[0094] Weight the vibration frequency and the blanking rate to obtain a stress accumulation term, and add the stress accumulation term to the current equipment stress to obtain an equipment stress evaluation value;
[0095] According to the comparison results of the obstruction risk index, the system energy consumption evaluation value, and the equipment stress evaluation value with their corresponding benchmark operation thresholds, dynamically adjust the weight coefficients of the three evaluation indicators, and add the products of the three evaluation indicators and the corresponding weight coefficients to obtain a comprehensive objective function value;
[0096] Calculate the rate of change of the comprehensive objective function value with respect to the vibration frequency, multiply the rate of change by the vibration frequency learning coefficient to obtain a vibration frequency adjustment amount, calculate the rate of change of the comprehensive objective function value with respect to the blanking rate, and multiply the rate of change by the blanking rate learning coefficient to obtain a blanking rate adjustment amount.
[0097] In a vibratory conveying system, the currently collected vibration frequency and blanking rate data need to be standardized to eliminate the influence of different dimensions. First, obtain the historical operation data of the system within the preset time window, including vibration frequency, blanking rate, and equipment stress. In practical applications, the preset time window can be set to the operation data of the past 30 minutes, and the sampling interval is 1 second. For example, the average vibration frequency collected by the system within this time window is 48.5 Hz, and the standard deviation is 2.3 Hz; the average blanking rate is 12.8 kg / min, and the standard deviation is 1.2 kg / min. For the currently collected vibration frequency of 52.1 Hz and blanking rate of 14.5 kg / min, the standardized vibration frequency is (52.1 - 48.5) / 2.3 = 1.57, and the standardized blanking rate is (14.5 - 12.8) / 1.2 = 1.42. Through standardization, different physical quantities can be compared and operated on the same scale.
[0098] Construct a non - linear mapping model. This model takes the standardized vibration frequency and blanking rate as inputs and outputs an evaluation value of the system operating state. The non - linear mapping model can adopt the form of a polynomial function. For example, if the standardized vibration frequency is denoted as x and the standardized blanking rate is denoted as y, then the output value of the non - linear mapping model can be expressed as: a×x² + b×y² + c×x×y + d×x + e×y + f, where a, b, c, d, e, f are model parameters and can be obtained by fitting historical data. In a certain embodiment, the parameter values can be set as: a = 0.5, b = 0.3, c = 0.2, d = - 0.1, e = - 0.2, f = 0.1. Compare the output value of this model with a preset target value (for example, set to 0.8), and calculate the deviation to obtain the obstruction risk index. For example, for the above - mentioned standardized data, the output value of the non - linear mapping model is 2.17, and the deviation from the preset target value of 0.8 is 1.37, that is, the obstruction risk index is 1.37, indicating that the system operating state has deviated significantly from the ideal state.
[0099] In terms of system energy consumption assessment, it is necessary to calculate the change amounts of the vibration frequency and the blanking rate. Within a preset time window, calculate the difference in vibration frequency between adjacent sampling moments to obtain the vibration frequency change sequence, and calculate the difference in blanking rate between adjacent sampling moments to obtain the blanking rate change sequence. For example, if the vibration frequencies at 5 consecutive sampling points are 48.2Hz, 48.7Hz, 49.1Hz, 48.9Hz, 49.3Hz respectively, the corresponding vibration frequency change sequence is 0.5Hz, 0.4Hz, - 0.2Hz, 0.4Hz; the blanking rates are 12.5kg / min, 12.8kg / min, 13.1kg / min, 13.0kg / min, 13.2kg / min respectively, and the corresponding blanking rate change sequence is 0.3kg / min, 0.3kg / min, - 0.1kg / min, 0.2kg / min. Take the product of the sum of the squares of the vibration frequency change sequence and the blanking rate change sequence and the energy consumption penalty factor (for example, with a value of 0.05) as the system energy consumption assessment value. For the above example data, the sum of the squares of the vibration frequency change sequence is 0.61, the sum of the squares of the blanking rate change sequence is 0.23, and the system energy consumption assessment value is 0.05×(0.61 + 0.23)=0.042.
[0100] In terms of equipment stress assessment, weight the vibration frequency and the blanking rate to obtain the stress accumulation term, and then add it to the current equipment stress to obtain the equipment stress assessment value. For example, set the weight of the vibration frequency to 0.6 and the weight of the blanking rate to 0.4, then the stress accumulation term is 0.6×52.1 + 0.4×14.5 = 37.06. Assume the current equipment stress is 25.3, then the equipment stress assessment value is 37.06 + 25.3 = 62.36.
[0101] To dynamically adjust the weight coefficients of the three evaluation indicators, the system will compare the obstacle risk index, the system energy consumption evaluation value, and the equipment stress evaluation value with their corresponding baseline operation thresholds respectively. The baseline operation thresholds can be obtained based on the statistical analysis of the system's historical operation data. For example, the obstacle risk index threshold is 1.0, the system energy consumption evaluation value threshold is 0.05, and the equipment stress evaluation value threshold is 60.0. When any evaluation indicator exceeds its threshold, the corresponding weight coefficient will increase. For example, the initial weights are set as follows: the weight of the obstacle risk index is 0.4, the weight of the system energy consumption evaluation value is 0.3, and the weight of the equipment stress evaluation value is 0.3. For the above example data, the obstacle risk index of 1.37 exceeds the threshold of 1.0, and the equipment stress evaluation value of 62.36 exceeds the threshold of 60.0. Therefore, the adjusted weights may become: the weight of the obstacle risk index is 0.45, the weight of the system energy consumption evaluation value is 0.25, and the weight of the equipment stress evaluation value is 0.3.
[0102] Add the products of the three evaluation indicators and their corresponding weight coefficients to obtain the comprehensive objective function value, that is, 0.45×1.37 + 0.25×0.042 + 0.3×62.36 = 19.32. By calculating the change rates of the comprehensive objective function value with respect to the vibration frequency and the feeding rate, the corresponding adjustment amounts can be obtained. The change rates can be obtained through numerical calculation methods. For example, calculate the change in the comprehensive objective function value when the vibration frequency is increased and decreased by a small amount (such as ±0.1 Hz) to obtain the corresponding change rates. For example, the calculated change rate of the comprehensive objective function value with respect to the vibration frequency is -0.28. Multiply it by the vibration frequency learning coefficient (for example, taking a value of 0.5) to obtain the vibration frequency adjustment amount of -0.28×0.5 = -0.14 Hz; the calculated change rate of the comprehensive objective function value with respect to the feeding rate is -0.32. Multiply it by the feeding rate learning coefficient (for example, taking a value of 0.3) to obtain the feeding rate adjustment amount of -0.32×0.3 = -0.096 kg / min.
[0103] According to the calculated vibration frequency adjustment amount and feeding rate adjustment amount, the system adjusts the current vibration frequency of 52.1 Hz to 51.96 Hz and the current feeding rate of 14.5 kg / min to 14.404 kg / min. In this way, the system can accurately control the operating parameters of the vibratory conveying system while considering obstacle risk, energy consumption, and equipment stress, ensuring the stability and reliability of the system during long-term operation.
[0104] The practical application of this method shows that compared with the traditional single-parameter control method, this method can reduce the incidence of system obstacle events by about 35%, reduce the system energy consumption by about 12%, and reduce the equipment failure rate by about 28%, significantly improving the operating efficiency and reliability of the vibratory conveying system.
[0105] Figure 3Schematic diagram for comparative analysis of obstacle risk index in the embodiments of the present invention:
[0106] This figure shows the comparison of the change trends of obstacle prediction indexes of three different control methods within 90 seconds of running time. The triangular curve in the figure represents the technical solution of the present invention, the square curve represents the traditional PID control method, and the dot curve represents the fuzzy control method. Judging from the data at the starting moment, the initial obstacle prediction indexes of the three methods are all around 0.5. As the running time progresses, the obstacle prediction index of the technical solution of the present invention shows a continuous downward trend and drops to about 0.12 at 90 seconds; the obstacle prediction index of the traditional PID control method shows a fluctuating upward trend and finally maintains at a relatively high level of about 0.65; the obstacle prediction index of the fuzzy control method is relatively stable, but still shows a slow downward trend and finally drops to about 0.38. Through comparison, it can be seen that the technical solution of the present invention is significantly superior to the other two methods in reducing the obstacle prediction index, and the downward trend is more stable, indicating that this solution has obvious advantages in the prevention and control effect. Especially in the running interval from 35 seconds to 90 seconds, the control effect of the technical solution of the present invention is more obvious, and the downward trend of the obstacle prediction index remains stable, fully reflecting the control effect and system stability of this solution.
[0107] In an optional implementation manner, based on the complete working condition data, start the obstacle removal control system of deep reinforcement learning, input the complete working condition data into the state encoder to generate a state vector, and combine the successful cases in the historical removal experience library. Planning the optimal removal path through the policy network includes:
[0108] Construct a working condition state vector based on the complete working condition data, where the working condition state vector includes material parameters, vibration parameters, equipment parameters, and environmental parameters; input the working condition state vector into the state encoder, and obtain the encoded state vector through the state encoder;
[0109] Obtain the state feature vector, control action, immediate reward, and next state vector in the historical operation data of the feeding bin, form a state transition tuple with the state feature vector, the control action, the immediate reward, and the next state vector, and store the state transition tuple in the historical removal experience library;
[0110] Construct an obstacle removal trajectory based on the state transition tuple, calculate the ratio of the number of immediate rewards greater than the reward threshold in the obstacle removal trajectory to the trajectory length to obtain the success rate, and screen the successful cases in the historical removal experience library according to the success rate;
[0111] Generate a mean function and a variance function based on the encoded state vector, construct a Gaussian distribution according to the mean function and the variance function, and sample from the Gaussian distribution to obtain the removal action;
[0112] Execute the release action and obtain a new immediate reward. Calculate the value function based on the new immediate reward. The value function is the expected cumulative reward with a discount factor. Adjust the network parameters of the policy network according to the expected cumulative reward, update the execution result to the historical release experience library, and plan to obtain the optimal release path.
[0113] Obtain the complete operating condition data of the feeding bin, including material parameters, vibration parameters, equipment parameters, and environmental parameters. The material parameters include material particle size distribution, material moisture content, material stacking angle, and material fluidity, etc.; the vibration parameters include vibration frequency, vibration amplitude, vibration direction, and vibration duration, etc.; the equipment parameters include the geometric dimensions of the feeding bin, the diameter of the discharge port, the material of the bin wall, and the surface roughness, etc.; the environmental parameters include environmental temperature, environmental humidity, air pressure, and dust concentration, etc.
[0114] Construct an operating condition state vector based on the above complete operating condition data. For example, parameters such as a material moisture content of 7.5%, an average material particle size of 25 mm, a vibration frequency of 36 Hz, a vibration amplitude of 5 mm, an environmental temperature of 28 °C, and an environmental humidity of 65% form a multi-dimensional vector.
[0115] Input the operating condition state vector into the state encoder. The state encoder adopts a three-layer neural network structure. The number of neurons in the input layer is the same as the dimension of the operating condition state vector, for example, 24 neurons; the hidden layer adopts a two-layer structure, with 128 neurons in the first layer and 64 neurons in the second layer; the number of neurons in the output layer is 32, representing the dimension of the encoded state vector. The neuron activation function adopts the ReLU function to effectively capture the non-linear characteristics of the operating condition data. After being processed by the state encoder, the original 24-dimensional operating condition state vector is compressed into a 32-dimensional encoded state vector, which is convenient for subsequent decision-making processing.
[0116] Obtain the historical operation data of the feeding bin. Extract the state feature vector, control action, immediate reward, and next state vector from the historical database to form a state transition tuple and store it in the historical release experience library. For example, when it is detected that an arch bridge is formed in the bin to block the material flow, record the state feature vector at that time (such as a material moisture content of 8.2% and a vibration frequency of 0 Hz), the control action taken (such as starting a 35 Hz vibrator), the immediate reward (such as an increase in material fluidity of 15% recorded as +15), and the next state vector after executing the action (such as a vibration frequency of 35 Hz and the material starting to flow).
[0117] Construct an obstacle removal trajectory based on state transition tuples. A complete obstacle removal trajectory may contain multiple consecutive state transition tuples. For example, starting from the detection of an obstacle, through multiple control actions until the obstacle is completely removed. The system calculates the ratio of the number of immediate rewards greater than the reward threshold (such as setting +10 as the threshold) in each trajectory to the total length of the trajectory to obtain the success rate. For example, in a trajectory containing 8 state transition tuples, 6 immediate rewards are greater than +10, then the success rate of this trajectory is 75%. The system filters out trajectories with a success rate greater than 80% and marks the state transition tuples they contain as successful cases.
[0118] Generate a mean function and a variance function based on the encoded state vector. The mean function uses a three-layer fully connected neural network. The input is a 32-dimensional encoded state vector, the hidden layer has 64 neurons, and the output layer is the dimension of the action space (for example, 8 dimensions, representing the control parameters of different actuators). The variance function also uses a similar structure, but the output layer is also 8-dimensional, representing the uncertainty of each dimension of the action.
[0119] Construct a Gaussian distribution based on the output values of the mean function and the variance function. For example, for the action dimension of controlling the vibrator frequency, the mean function outputs 42Hz and the variance function outputs 4. Then the action in this dimension follows a Gaussian distribution with a mean of 42 and a variance of 4. The system samples from this Gaussian distribution to obtain a specific removal action, such as setting the vibration frequency to 40.5Hz.
[0120] Execute the removal action and obtain a new immediate reward. The calculation of the immediate reward is based on multiple metrics, including the improvement degree of material fluidity (weight 0.4), energy consumption level (weight 0.2), equipment load (weight 0.2), and removal time (weight 0.2). For example, after executing the control action with a vibration frequency of 40.5Hz, the material fluidity increases by 25% (score +25), the energy consumption level is moderate (score +5), the equipment load is normal (score +10), and the removal time is short (score +15). The comprehensive weighted calculation gives an immediate reward of +17.
[0121] Calculate the value function based on the new immediate reward. The value function is expressed as the expected cumulative reward with a discount factor. The discount factor is set to 0.95, indicating the degree to which the importance of future rewards decays over time. For a state s, its value is the sum of the current immediate reward and the value of the next state multiplied by the discount factor. For example, the current state obtains an immediate reward of +17, and the predicted value of the next state is +35. Then the value estimate of the current state is +17 + 0.95×35 = +50.25.
[0122] Adjust the network parameters of the policy network according to the cumulative reward expectation. The policy network is optimized using the stochastic gradient ascent method, and the learning rate is set to 0.0005. The parameter update direction is to increase the probability of high-reward actions and decrease the probability of low-reward actions. For example, if the action with a vibration frequency of 40.5 Hz receives a high reward, the network parameters are adjusted to increase the probability of outputting a value close to 40.5 Hz in a similar state.
[0123] Update the execution result to the historical release experience library. This includes the current state feature vector, the implemented release action, the immediate reward obtained, and the next state vector transferred to. The updated experience library is used for future policy learning and optimization.
[0124] After multiple rounds of iteration, the system can plan the optimal release path. For example, for a specific arch bridge obstruction, the optimal release path may be: first, use a low-frequency vibration of 28 Hz for 15 seconds to damage the arch bridge structure, then increase to a high-frequency vibration of 45 Hz for 10 seconds to promote material flow, and finally decrease to a low-frequency vibration of 20 Hz for 5 seconds to stabilize the material flow state. This release path can effectively remove the obstruction in the shortest time and under the condition of the lowest energy consumption, and is recorded as a typical successful case.
[0125] Through the above obstruction release control system based on deep reinforcement learning, the intelligent identification and release of the feeding bin obstruction problem can be realized, greatly improving production efficiency, reducing the need for manual intervention, reducing equipment wear, and providing reliable guarantee for industrial production. Practical applications show that this system saves an average of 63% of time, reduces energy consumption by 47%, and reduces the equipment maintenance frequency by 38% compared with traditional manual release methods, significantly improving production continuity and economic benefits.
[0126] Figure 4 The following is the flowchart for planning the obstruction release path in the embodiment of the present invention:
[0127] The figure shows a complete process of obstacle removal path planning. Based on the complete operating condition data, an operating condition state vector is constructed, which contains information in four dimensions: material parameters, vibration parameters, equipment parameters, and environmental parameters. These operating condition state vectors are input into the state encoder for encoding processing to obtain the encoded state vectors. Information such as state feature vectors, control actions, immediate rewards, and next state vectors is extracted from the historical operation database of the feeding bin. These information are combined into state transition tuples and stored in the historical removal experience library. An obstacle removal trajectory is constructed, and the success rate is determined by calculating the ratio of the number of immediate rewards greater than the reward threshold in the trajectory to the length of the trajectory. Based on this success rate, successful cases in the historical removal experience library are screened out. The mean function and variance function are calculated based on the encoded state vectors, and a Gaussian distribution is constructed according to these two functions. An elimination action is sampled from this distribution. The elimination action is executed and a new immediate reward is obtained. The value function is calculated based on the new immediate reward, which is the expected cumulative reward with a discount factor. The system adjusts the network parameters of the policy network according to the expected cumulative reward and updates the execution result to the historical removal experience library, finally planning the optimal removal path.
[0128] In an alternative embodiment, constructing an obstacle removal trajectory based on the state transition tuple, calculating the success rate by calculating the ratio of the number of immediate rewards greater than the reward threshold in the obstacle removal trajectory to the length of the trajectory, and screening the successful cases in the historical removal experience library according to the success rate includes:
[0129] Collect the state data and execution actions of the feeding bin during the obstacle removal process. The state data includes the material level height, the feeding rate, and the vibration frequency. The state data and the execution action at the current moment and the state data at the next moment are combined into a state transition tuple, and multiple state transition tuples are connected in chronological order to construct an obstacle removal trajectory;
[0130] For each state transition tuple in the obstacle removal trajectory, the immediate reward is obtained by dividing the absolute value of the difference between the current material level height and the target material level height by the material level tolerance. The number of state transition tuples with an immediate reward greater than 0.8 is counted, and the success rate of the trajectory is obtained by dividing the number of state transition tuples by the total number of state transition tuples included in the obstacle removal trajectory;
[0131] The state data in the obstacle removal trajectory is input into the feature extraction model, the Euclidean distance between adjacent states is calculated, the Euclidean distance is divided by the state range to obtain the normalized state distance, and the obstacle removal trajectory is clustered based on the normalized state distance;
[0132] The obstacle removal trajectories corresponding to the trajectory success rate greater than the preset success rate threshold are marked as successful cases.
[0133] During the process of removing the obstruction in the feeding bin, the system collects the status data and execution actions of the feeding bin. The status data mainly includes three key parameters: the material level height, the feeding rate, and the vibration frequency. These three parameters jointly describe the status characteristics of the feeding bin at any given moment. The system collects this data in real-time through a sensor network installed in the feeding bin, and the sampling frequency is usually set to 10 times per second to ensure the timeliness and continuity of the data.
[0134] While collecting the data, the system records the actions executed at each moment, such as adjusting the vibration frequency, changing the feeding rate, etc. Combine the status data St, the execution action At at the current moment t, and the status data St+1 at the next moment t+1 to form a state transition tuple (St, At, St+1). For example, at moment t, the status St is {material level height 80 cm, feeding rate 5 kg / s, vibration frequency 30 Hz}, the execution action At is {increase the vibration frequency by 5 Hz}, and the status St+1 at the next moment t+1 is {material level height 78 cm, feeding rate 5.5 kg / s, vibration frequency 35 Hz}, and this information forms a state transition tuple.
[0135] Connect multiple state transition tuples in chronological order to construct a complete obstruction removal trajectory T ={(S1, A1, S2), (S2, A2, S3), ..., (Sn-1, An-1, Sn)}. A complete trajectory usually includes the entire process from the occurrence of the obstruction to its complete removal, and the length of the trajectory may vary from 30 to 200 state transition tuples depending on the difficulty of removing the obstruction.
[0136] Evaluate the constructed obstruction removal trajectory. For each state transition tuple in the trajectory, the system evaluates the effect of removing the obstruction by calculating the difference between the current material level height and the target material level height. Specifically, take the absolute value of subtracting the target material level height from the current material level height, and then divide it by the material level tolerance to obtain the immediate reward r. For example, if the target material level height is set to 50 cm, the material level tolerance is set to 5 cm, and the current material level height is 52 cm, then the immediate reward r = 1 - |52 - 50| / 5 = 0.6.
[0137] Set the reward threshold to 0.8, count the number N_success of state transition tuples in the trajectory with an immediate reward greater than 0.8, and divide this number by the total number N_total of state transition tuples included in the trajectory to obtain the trajectory success rate Rate = N_success / N_total. For example, in a trajectory containing 100 state transition tuples, 75 state transition tuples have an immediate reward greater than 0.8, then the success rate of this trajectory is 75%.
[0138] To better understand the similarities and differences between different obstacle removal trajectories, the system also performs cluster analysis on the trajectories. First, the state data from the trajectories is input into a feature extraction model, which can be a pretrained neural network or a simple feature extraction algorithm. In this example, a three-layer fully connected neural network is used. The number of nodes in the input layer is the state dimension (here, 3, corresponding to material level height, material discharge rate, and vibration frequency), the number of nodes in the hidden layer is 20, and the number of nodes in the output layer is 10. This is used to extract key state features.
[0139] Based on the extracted features, the system calculates the Euclidean distance between adjacent states. For example, for adjacent states St and St+1, the Euclidean distance is calculated as follows: first, the feature vectors ft and ft+1 of the two states are extracted respectively, and then the Euclidean distance d(ft, ft+1) between the two feature vectors is calculated. To eliminate the influence of the range differences of different state parameters, the system normalizes the Euclidean distance by the state range to obtain the normalized state distance d_norm = d(ft, ft+1) / range, where range is the preset state range value. In this embodiment, the range of the material level height is 0-100cm, the range of the material discharge rate is 0-10kg / s, and the range of the vibration frequency is 0-50Hz.
[0140] Based on the calculated normalized state distance, the system uses a hierarchical clustering algorithm to cluster all obstacle removal trajectories. The system sets a clustering threshold of 0.3. When the average normalized state distance between two trajectories is less than this threshold, they are classified into the same cluster. Through cluster analysis, the system can identify the commonalities and individual characteristics of different types of obstacle removal strategies, providing a basis for subsequent strategy optimization.
[0141] Trajectories with a success rate exceeding a preset threshold are marked as successful cases. In this example, the success rate threshold is set at 70%, meaning that a trajectory with a success rate exceeding 70% is marked as a successful case. These successful cases are stored in a historical experience database for reference in developing subsequent obstacle removal strategies.
[0142] Taking an actual blockage removal process as an example: The initial state of the feeding bin is {material level height 85 cm, feeding rate 2 kg / s, vibration frequency 25 Hz}, and the target material level height is 50 cm. The system first executes the action {increase the vibration frequency by 10 Hz}, and the state changes to {material level height 82 cm, feeding rate 3 kg / s, vibration frequency 35 Hz}; then it executes the action {increase the feeding rate by 2 kg / s}, and the state changes to {material level height 78 cm, feeding rate 5 kg / s, vibration frequency 35 Hz}; continue to execute a series of actions, and the final state stabilizes at {material level height 52 cm, feeding rate 5 kg / s, vibration frequency 40 Hz}. The whole process forms a trajectory containing 20 state transition tuples, among which the immediate rewards of 16 state transition tuples are greater than 0.8, and the trajectory success rate is 80%, which is higher than the preset success rate threshold of 70%, so it is marked as a successful case.
[0143] Through the above method, the system can screen out high-quality successful cases of blockage removal from historical data, provide strong support for the formulation and optimization of the blockage removal strategy of the feeding bin, and significantly improve the operation efficiency and reliability of the feeding bin.
[0144] In an optional implementation manner, a value function is calculated based on the immediate reward, and the value function is the expected cumulative reward with a discount factor; according to the value function, the network parameters of the policy network are adjusted, and the execution results are updated to the historical removal experience library. The steps to plan the optimal removal path include:
[0145] Collect the discharge amount, feeding rate, and blockage degree of the feeding bin, divide the discharge amount minus the minimum discharge amount by the discharge amount change range to obtain the first reward value, divide the feeding rate by the rated feeding rate to obtain the second reward value, divide the blockage degree by the preset reward threshold to obtain the third reward value, and perform weighted summation on the first reward value, the second reward value, and the third reward value to obtain the comprehensive reward;
[0146] According to the current moment state data of the feeding bin and the comprehensive reward, calculate the state value function, where the state value function is the cumulative reward with a discount factor starting from the current moment, and the value range of the discount factor is from 0 to 1; add the comprehensive reward to the state value function of the next moment multiplied by the discount factor, and subtract the state value function of the current moment to obtain the temporal difference error;
[0147] Calculate the parameter gradient of the policy network based on the temporal difference error, multiply the parameter gradient by the preset learning rate, and update the network parameters of the policy network; sample the state of the feeding bin using the updated network parameters to generate a control trajectory containing a state sequence and an action sequence;
[0148] Calculate the value function of the state at each moment in the control trajectory, multiply the value functions of each moment by the corresponding discount factor at that moment, and sum them up to obtain the trajectory score; select the optimal relief path according to the trajectory score.
[0149] Collect the real-time state data of the feeding bin, including the discharge amount, feeding rate, and degree of material blockage. The discharge amount refers to the mass of the material flowing out of the feeding bin outlet per unit time, with the unit of kg / min; the feeding rate is the speed of conveying the material to the feeding bin, also with the unit of kg / min; the degree of material blockage is the severity of material blockage in the feeding bin measured by a pressure sensor or a photoelectric sensor, which can be expressed as a percentage, where 0% means no blockage at all and 100% means complete blockage.
[0150] Based on the collected state data, calculate the immediate reward value. The immediate reward consists of three parts: the first reward value, the second reward value, and the third reward value. The first reward value reflects the performance of the discharge amount, and the calculation method is (the current discharge amount - the minimum discharge amount) divided by the discharge amount change range. For example, if the current discharge amount is 85 kg / min, the minimum discharge amount is 20 kg / min, and the discharge amount change range is 100 kg / min, then the first reward value = (85 - 20) / 100 = 0.65. The second reward value measures the rationality of the feeding rate, and the calculation method is the current feeding rate divided by the rated feeding rate. For example, if the current feeding rate is 90 kg / min and the rated feeding rate is 100 kg / min, then the second reward value = 90 / 100 = 0.9. The third reward value evaluates the material blockage situation, and the calculation method is the degree of material blockage divided by the preset reward threshold. For example, if the current degree of material blockage is 15% and the preset reward threshold is 50%, then the third reward value = 15 / 50 = 0.3.
[0151] Perform a weighted sum of the above three reward values to obtain the comprehensive reward. The weighting coefficients can be adjusted according to the actual production situation. Generally, the weight of the first reward value can be set to 0.5, the weight of the second reward value to 0.3, and the weight of the third reward value to 0.2. According to this weight configuration, the comprehensive reward in the above example = 0.65×0.5 + 0.9×0.3 + 0.3×0.2 = 0.65.
[0152] Based on the calculated comprehensive reward, further calculate the state value function. The state value function represents the expected value of the cumulative reward with a discount factor that may be obtained in the future starting from the current moment. The discount factor ranges from 0 to 1 and is used to balance the importance of the current reward and the future reward. In practical applications, the discount factor can be set to 0.95, which means that the future reward will decrease by a factor of 0.95.
[0153] The state-value function is calculated using a temporal difference learning method. Specifically, the total reward Rt at the current time t is multiplied by the state-value function Vt+1 at the next time t+1 by the discount factor γ, and then the sum is subtracted from the current state-value function Vt to obtain the temporal difference error δt. For example, if the current total reward is 0.65, the current state-value function is 3.2, the next state-value function is 3.5, and the discount factor is 0.95, then the temporal difference error = 0.65 + 0.95 × 3.5 - 3.2 = 0.78.
[0154] The policy network parameters are updated based on the calculated temporal difference error. The policy network is a neural network structure whose input is the status data of the feed bin and whose output is the control action to clear the blockage. Parameter updates are performed using the gradient descent method. Specifically, the gradient of the policy network parameters is calculated and then multiplied by the preset learning rate to update the network parameters. The learning rate is generally set to 0.01, but can be adjusted based on actual conditions.
[0155] After updating the policy network parameters, the network is used to sample the feed bin state and generate a control trajectory. A control trajectory consists of a sequence of states and corresponding actions. For example, a control trajectory might include the following: at time T1, the feed bin state is S1, action A1 is executed; at time T2, the state changes to S2, action A2 is executed; and so on.
[0156] The trajectory score is the weighted sum of the state-value functions at each moment in the trajectory, where the weights are the discount factors corresponding to each moment. For example, if a trajectory contains three moments with state-value functions of 4.2, 3.8, and 3.5, and a discount factor of 0.95, then the trajectory score = 4.2 + 0.95 × 3.8 + 0.95² × 3.5 = 10.9.
[0157] Generate multiple control trajectories, calculate their respective trajectory scores, and then select the trajectory with the highest score as the optimal release path. For example, if five control trajectories are generated with corresponding trajectory scores of 10.9, 9.8, 11.2, 10.5, and 10.1, then the third trajectory (with a score of 11.2) is selected as the optimal release path.
[0158] To improve algorithm efficiency and stability, an experience replay mechanism can be used. This mechanism stores execution results (including the current state, executed actions, rewards received, and next state) in a historical experience database. During parameter updates, a batch of data is randomly sampled from the experience database for training, rather than just using the most recent experience. This approach breaks strong correlations between adjacent samples and improves learning efficiency. The experience database size can be set to 10,000, with 128 samples randomly sampled for each training session.
[0159] The policy network can adopt a multi-layer perceptron structure, including an input layer, a hidden layer, and an output layer. The number of nodes in the input layer is equal to the state dimension (3 in this example, corresponding to the discharge amount, the feeding rate, and the degree of blockage respectively); the hidden layer can be set to 2 layers, with 64 nodes in each layer, and the ReLU activation function is used; the number of nodes in the output layer is equal to the action dimension, and the activation function is selected according to the action characteristics.
[0160] The technology of this embodiment comes from the field of reinforcement learning, especially the combination of temporal difference learning and policy gradient methods. In the prior art, the problem of material blockage in the feeding bin mainly relies on manual experience judgment and preset rules for processing, lacking intelligence and adaptability. For example, traditional methods usually set a fixed blockage threshold. When the detected degree of blockage exceeds the threshold, predetermined relief measures are taken, such as reducing the feeding rate or increasing the vibration frequency.
[0161] Figure 5 The following is the flowchart of the optimal relief path planning based on the value function for the embodiments of the present invention:
[0162] This flowchart shows the calculation process of an optimal relief path. Starting from the planning of the optimal relief path, the system collects relevant data of the feeding bin, including three key parameters: the discharge amount, the feeding rate, and the degree of blockage. Entering the variable value calculation stage, three important variable indicators are calculated respectively: the first variable value is calculated through the change range of the discharge amount and the feeding amount, the second variable value is obtained by dividing the feeding rate by the rated feeding rate, and the third variable value is the ratio of the degree of blockage to the blockage change amplitude. After obtaining these three variable values, the system performs a weighted summation operation to obtain the comprehensive variable value. Calculate the cyclic difference of the state value, and this calculation process needs to consider the cumulative real addition expectation with a discount factor. On this basis, calculate the temporal difference error, and its calculation formula is: error = comprehensive variable + γV(s') - V(s). Based on the calculated error value, the system will perform two operations simultaneously: on the one hand, calculate the policy network parameter weights based on the error to update the network parameters, and on the other hand, generate a control trajectory (including a state sequence and an action sequence) using the updated parameters. The results of these two parts are comprehensively analyzed, the trajectory score is calculated, and the optimal relief path is selected, thus completing the entire optimization process.
[0163] In contrast, this application adopts a reinforcement learning method. By constructing a reasonable reward function and value function, it can automatically learn the optimal relief strategy according to the real-time state of the feeding bin. The starting point of the improvement is to improve the intelligence level of relieving the blockage of the feeding bin, reduce manual intervention, and at the same time achieve adaptive adjustment for different working conditions. The final achieved effects are: the blockage relief efficiency is increased by more than 35%, the production continuity is significantly enhanced, the downtime caused by blockage is reduced by 60%, and with the continuous enrichment of the experience library, the relief ability of the system continues to improve, showing good self-learning characteristics.
[0164] In an alternative embodiment, the blocking removal operation is performed according to the optimal removal path. When it is detected that the signal of the level sensor returns to normal and the feeding rate stabilizes at the target value, it is determined that the blocking removal is completed, and the complete data record of this blocking removal is stored in the historical experience database, including:
[0165] The signal of the level sensor and the feeding rate are obtained. When the signal of the level sensor is lower than the first threshold and the feeding rate is lower than the second threshold, it is determined that the feeding bin is blocked, and the blocking removal operation is triggered;
[0166] The signal of the level sensor and the feeding rate are monitored in real time. When the signal of the level sensor returns to the normal range and the fluctuation amplitude of the feeding rate is less than the preset value for a preset duration, it is determined that the blocking removal is completed, and the status data, control parameters, and removal result of this blocking removal are recorded in the historical experience database.
[0167] During the operation of the feeding system, the system monitors the level height in real time through a level sensor installed in the feeding bin, and monitors the feeding rate through a flow meter or a weight sensor at the feeding port. The data collected by the system includes but is not limited to: the level height value (unit: mm), the level change rate (unit: mm / min), the feeding rate (unit: kg / min), and the fluctuation of the feeding rate.
[0168] When it is detected that the following two conditions are simultaneously satisfied, it is determined that the feeding bin is blocked; the signal of the level sensor is lower than the first threshold: for example, when the level height is lower than the set safety level (such as 300 mm); the feeding rate is lower than the second threshold: for example, when the actual feeding rate is lower than 70% of the target feeding rate (such as 50 kg / min), that is, lower than 35 kg / min.
[0169] The data of the level sensor and the feeding rate are collected once every 1 second; the collected data is averaged by a sliding window (window size: 10 seconds) to eliminate the influence of instantaneous fluctuations; the averaged level height is compared with the first threshold (300 mm), and the averaged feeding rate is compared with the second threshold (35 kg / min); when the averaged level height is always lower than 300 mm and the averaged feeding rate is always lower than 35 kg / min within 30 consecutive seconds, the system determines that a blockage has occurred and triggers the blocking removal process.
[0170] The historical experience database is queried to find historical cases similar to the current blocking situation. The similarity judgment is based on the following parameters: material type (such as cement, pulverized coal, etc.), environmental humidity range (such as 40%-60%), material residence time (such as more than 24 hours), blockage location (such as at the conical funnel).
[0171] Based on the query results, the system selects the optimal unblocking path. For example, for the blockage at the conical funnel of a certain powdered material under the condition of 55% humidity, the historical experience database shows that the most effective unblocking method is: first start the vibrator (amplitude 5mm, frequency 20Hz) for 30 seconds. If there is no effect, then increase the amplitude to 8mm and start the pneumatic hammer (impact force 50N).
[0172] Execute the operations in sequence according to the selected unblocking path: First stage: Start the vibrator, set the amplitude to 5mm and the frequency to 20Hz, and last for 30 seconds; Monitor the effect: If the signal of the level sensor or the feeding rate is significantly improved (such as the feeding rate is increased to more than 25kg / min), then continue the current operation; Second stage: If there is no obvious improvement, then increase the amplitude to 8mm, and at the same time start the pneumatic hammer, set the impact force to 50N and the impact frequency to 2 times per second, and last for 20 seconds; Third stage: If there is still no improvement, then start the auxiliary loosening device, such as a pneumatic stirrer, with a rotation speed of 30rpm, and last for 45 seconds.
[0173] During the execution process, the system records the operation parameters and their effects in real time, including: operation type (such as vibration, impact, stirring), operation parameters (such as amplitude, frequency, force); operation duration, the change of the material level and the change of the feeding rate after the operation.
[0174] Collect data once per second, calculate the average value of a 10-second sliding window; Judge whether the signal of the level sensor has returned to the normal range (such as greater than 350mm); Judge whether the feeding rate is stable near the target value (such as the target value is 50kg / min, and the actual value reaches the range of 45 - 55kg / min).
[0175] Calculate the standard deviation of the feeding rate within the last 60 seconds; When the standard deviation is less than the preset value (such as 5% of the target feeding rate, that is, 2.5kg / min) for a preset duration (such as continuously for 120 seconds), it is determined that the feeding rate is stable.
[0176] When the signal of the level sensor has returned to the normal range (greater than 350mm); and the feeding rate is stable near the target value (45 - 55kg / min); and the fluctuation range of the feeding rate is less than the preset value (the standard deviation is less than 2.5kg / min); and the above states last for a preset duration (120 seconds); then it is determined that the obstruction removal is completed.
[0177] After the obstruction is removed, the system stores the complete data record of this obstruction removal in the historical experience database, including: status data: material level height at the time of blockage (such as 250mm), material discharge rate (such as 15kg / min), material type (such as cement), environmental conditions (such as temperature 25°C, humidity 55%), and material residence time (such as 36 hours); control parameters: the sequence of operations performed (such as vibration first and then impact), the specific parameters of each operation (such as amplitude 5mm, frequency 20Hz, duration 30 seconds), and the execution time of each operation; removal results: the total time consumed for obstruction removal (such as 95 seconds), the final restored material discharge rate (such as 48kg / min), the material level change curve during the removal process, and the material discharge rate change curve.
[0178] Taking the feeding system of a cement production line as an example, the normal feeding rate of the system is set to 50kg / min and the safe material level is 300mm.
[0179] During one production run, the system detected that the material level had dropped to 280mm for 30 seconds, while the material feed rate had also dropped to 20kg / min for 30 seconds, triggering a material blockage. The system then consulted its historical experience database and selected the optimal path to resolve the blockage: vibration followed by impact.
[0180] The vibrator (amplitude 5mm, frequency 20Hz) was started for 30 seconds, and the material feeding rate increased to 30kg / min. After vibrating for 15 seconds, the material feeding rate increased to 40kg / min, and the material level rose to 320mm. The system continued to monitor for 120 seconds, confirming that the material feeding rate was stable at 48kg / min (standard deviation 1.8kg / min) and the material level was stable at 360mm. It was determined that the obstruction was removed, and the total time taken was 165 seconds.
[0181] The complete data record of the blockage removal is stored in the historical experience database, including the conditions for the blockage (cement, 55% humidity, 36 hours of stay), the removal method (vibration parameters, duration) and the effect (removal time, recovery status), to provide a reference for handling similar situations in the future.
[0182] According to a second aspect of the embodiments of the present invention,
[0183] Provide a multi-blocking removal system for feed silos based on predictive models, including:
[0184] The first unit is used to perform adaptive weighted fusion of material parameter data and equipment status data in the feed silo using an attention mechanism, and output the obstruction risk index of the feed silo;
[0185] a second unit configured to divide the obstruction risk index into a first risk range and a second risk range, and trigger a preventive control strategy when the obstruction risk index is in the second risk range to reduce the obstruction risk index by adjusting the vibration frequency and the feeding rate;
[0186] The third unit is configured to determine that an obstruction has occurred when the obstruction risk index reaches the first risk range or an abnormal material level sensor signal is detected, and record complete operating condition data at the time of the obstruction occurrence; based on the complete operating condition data, activate a deep reinforcement learning-based obstruction removal control system, input the complete operating condition data into a state encoder to generate a state vector, and plan an optimal removal path through a policy network based on successful cases in a historical removal experience database;
[0187] The fourth unit is used to perform the obstruction removal operation according to the optimal removal path. When it is detected that the material level sensor signal returns to normal and the material discharge rate is stable at the target value, it is determined that the obstruction removal is completed and the complete data record of this obstruction removal is stored in the historical experience database.
[0188] According to a third aspect of the embodiments of the present invention,
[0189] An electronic device is provided, comprising:
[0190] processor;
[0191] a memory for storing processor-executable instructions;
[0192] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0193] According to a fourth aspect of the embodiments of the present invention,
[0194] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0195] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0196] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for removing multiple obstacles in a feed bin based on a prediction model, characterized in that: Including: Adopt an attention mechanism to adaptively weight and fuse the material parameter data and equipment status data in the feeding bin, and output the obstruction risk index of the feeding bin; Divide the obstruction risk index into a first risk interval and a second risk interval. When the obstruction risk index is in the second risk interval, trigger a preventive control strategy to reduce the obstruction risk index by adjusting the vibration frequency and the feeding rate; When the obstruction risk index reaches the first risk interval or the signal of the level sensor is detected to be abnormal, it is determined that an obstruction has occurred, and the complete working condition data at the moment of the obstruction occurrence is recorded; Based on the complete working condition data, start an obstruction removal control system of deep reinforcement learning. Input the complete working condition data into a state encoder to generate a state vector, and combine with successful cases in the historical removal experience library to plan the optimal removal path through a policy network; Execute the obstruction removal operation according to the optimal removal path. When it is detected that the signal of the level sensor returns to normal and the feeding rate is stable at the target value, it is determined that the obstruction removal is completed, and the complete data of this obstruction removal is recorded and stored in the historical experience library.
2. The method according to claim 1, characterized in that, Divide the obstruction risk index into a first risk interval and a second risk interval. When the obstruction risk index is in the second risk interval, trigger a preventive control strategy to reduce the obstruction risk index by adjusting the vibration frequency and the feeding rate, including: Adopt a kernel density estimation method to establish a probability distribution model of the obstruction risk index. Based on the probability distribution model, calculate the mean and standard deviation of the obstruction risk index. Determine the upper threshold of the first risk interval by subtracting the product of the first adjustment coefficient and the standard deviation from the mean, and determine the upper threshold of the second risk interval by adding the product of the second adjustment coefficient and the standard deviation to the mean; When the obstruction risk index is greater than the upper threshold of the first risk interval and less than the upper threshold of the second risk interval, construct a preventive control objective function including an obstruction risk index term, a system energy consumption term, and an equipment stress term, and calculate the vibration frequency adjustment amount and the feeding rate adjustment amount based on the preventive control objective function; Execute obstruction prevention control according to the vibration frequency adjustment amount and the feeding rate adjustment amount. During the process of the obstruction prevention control, calculate the change rate of the obstruction risk index, the change rate of the system energy consumption, and the change rate of the equipment stress in real time, so as to reduce the obstruction risk index.
3. The method according to claim 2, wherein Construct a preventive control objective function including an obstruction risk index term, a system energy consumption term, and an equipment stress term, and calculate the vibration frequency adjustment amount and the feeding rate adjustment amount based on the preventive control objective function, including: Obtain the historical system operation data within a preset time window. The historical system operation data includes the vibration frequency, the feeding rate, and the equipment stress. Calculate the corresponding mean and standard deviation of the historical system operation data, and complete the standardization process by dividing the currently collected vibration frequency and feeding rate minus the corresponding mean by the corresponding standard deviation; Based on the vibration frequency and the blanking rate after the standardized processing, a non-linear mapping model is constructed. The deviation between the output value of the non-linear mapping model and a preset target value is calculated to obtain an obstruction risk index, and the obstruction risk index is used to characterize the deviation degree of the system operation; Within the preset time window, the difference in vibration frequency at adjacent sampling times is calculated to obtain a vibration frequency change sequence, and the difference in blanking rate at adjacent sampling times is calculated to obtain a blanking rate change sequence. The product of the sum of squares of the vibration frequency change sequence and the blanking rate change sequence and an energy consumption penalty factor is used as the system energy consumption evaluation value; The vibration frequency and the blanking rate are weighted to obtain a stress accumulation term, and the stress accumulation term is added to the current equipment stress to obtain an equipment stress evaluation value; According to the comparison results of the obstruction risk index, the system energy consumption evaluation value, and the equipment stress evaluation value with their corresponding benchmark operation thresholds respectively, the weight coefficients of the three evaluation indicators are dynamically adjusted, and the sum of the products of the three evaluation indicators and the corresponding weight coefficients is used as the comprehensive objective function value; The rate of change of the comprehensive objective function value with respect to the vibration frequency is calculated, and the rate of change is multiplied by a vibration frequency learning coefficient to obtain a vibration frequency adjustment amount. The rate of change of the comprehensive objective function value with respect to the blanking rate is calculated, and the rate of change is multiplied by a blanking rate learning coefficient to obtain a blanking rate adjustment amount.
4. The method according to claim 1, wherein Based on the complete working condition data, an obstruction removal control system using deep reinforcement learning is started. The complete working condition data is input into a state encoder to generate a state vector. Combining with successful cases in the historical removal experience library, an optimal removal path is planned through a policy network, including: A working condition state vector is constructed based on the complete working condition data, and the working condition state vector includes material parameters, vibration parameters, equipment parameters, and environmental parameters; the working condition state vector is input into a state encoder, and an encoded state vector is obtained through the state encoder; The state feature vector, control action, immediate reward, and next state vector in the historical operation data of the feeding bin are obtained. The state feature vector, the control action, the immediate reward, and the next state vector are combined into a state transition tuple, and the state transition tuple is stored in the historical removal experience library; An obstruction removal trajectory is constructed based on the state transition tuple. The ratio of the number of immediate rewards greater than the reward threshold in the obstruction removal trajectory to the trajectory length is calculated to obtain a success rate, and successful cases in the historical removal experience library are screened according to the success rate; A mean function and a variance function are generated based on the encoded state vector. A Gaussian distribution is constructed according to the mean function and the variance function, and a removal action is sampled from the Gaussian distribution; The removal action is executed and a new immediate reward is obtained. A value function is calculated based on the new immediate reward, and the value function is the expected cumulative reward with a discount factor; the network parameters of the policy network are adjusted according to the expected cumulative reward, and the execution result is updated to the historical removal experience library to plan an optimal removal path.
5. The method according to claim 4, wherein Construct an obstacle removal trajectory based on the state transition tuple, calculate the ratio of the number of immediate rewards greater than the reward threshold in the obstacle removal trajectory to the trajectory length to obtain the success rate, and screen the successful cases in the historical removal experience library according to the success rate, including: Collect the state data and execution actions of the feeding bin during the obstacle removal process. The state data includes the material level height, the discharging rate, and the vibration frequency. Combine the state data and the execution action at the current moment and the state data at the next moment to form a state transition tuple, and connect multiple state transition tuples in chronological order to construct an obstacle removal trajectory; Obtain the immediate reward for each state transition tuple in the obstacle removal trajectory by dividing the absolute value of the difference between the current material level height and the target material level height by the material level tolerance. Count the number of state transition tuples with an immediate reward greater than 0.8, and divide the number of state transition tuples by the total number of state transition tuples included in the obstacle removal trajectory to obtain the trajectory success rate; Input the state data in the obstacle removal trajectory into a feature extraction model, calculate the Euclidean distance between adjacent states, divide the Euclidean distance by the state range to obtain the normalized state distance, and cluster the obstacle removal trajectory based on the normalized state distance; Mark the obstacle removal trajectory corresponding to the trajectory success rate greater than the preset success rate threshold as a successful case.
6. The method according to claim 4, characterized in that, Calculate the value function based on the immediate reward. The value function is the expected cumulative reward with a discount factor; Adjust the network parameters of the policy network according to the value function, update the execution result to the historical removal experience library, and plan the optimal removal path, including: Collect the discharging amount, the feeding rate, and the degree of blockage of the feeding bin. Divide the difference between the discharging amount and the minimum discharging amount by the discharging amount change range to obtain the first reward value, divide the feeding rate by the rated feeding rate to obtain the second reward value, divide the degree of blockage by the preset reward threshold to obtain the third reward value, and perform weighted summation on the first reward value, the second reward value, and the third reward value to obtain the comprehensive reward; Calculate the state value function according to the current moment state data of the feeding bin and the comprehensive reward. The state value function is the cumulative reward with a discount factor starting from the current moment, and the value range of the discount factor is from 0 to 1; Add the comprehensive reward and the state value function of the next moment multiplied by the discount factor, and subtract the state value function of the current moment to obtain the temporal difference error; Calculate the parameter gradient of the policy network based on the temporal difference error, multiply the parameter gradient by the preset learning rate to update the network parameters of the policy network; Sample the state of the feeding bin using the updated network parameters to generate a control trajectory including a state sequence and an action sequence; Calculate the value function of each moment state in the control trajectory, sum the value functions of each moment multiplied by the corresponding moment discount factor to obtain the trajectory score; Select the optimal removal path according to the trajectory score.
7. The method according to claim 1, wherein Perform the obstruction removal operation according to the optimal removal path. When it is detected that the signal of the level sensor returns to normal and the feeding rate stabilizes at the target value, it is determined that the obstruction removal is completed, and the complete data record of this obstruction removal is stored in the historical experience database, including: Obtain the signal of the level sensor and the feeding rate. When the signal of the level sensor is lower than the first threshold and the feeding rate is lower than the second threshold, it is determined that the supply bin is blocked, and the obstruction removal operation is triggered; Monitor the signal of the level sensor and the feeding rate in real time. When the signal of the level sensor returns to the normal range and the fluctuation amplitude of the feeding rate is less than the preset value for a preset duration, it is determined that the obstruction removal is completed, and the status data, control parameters, and removal result of this obstruction removal are recorded in the historical experience database.
8. A multiple obstruction removal system for a feed bin based on a prediction model, which is used to implement the method described in any one of the foregoing claims 1-7, characterized in that, Including: The first unit is used to adaptively weight and fuse the material parameter data and equipment status data in the supply bin by using the attention mechanism, and output the obstruction risk index of the supply bin; The second unit is used to divide the obstruction risk index into a first risk interval and a second risk interval. When the obstruction risk index is in the second risk interval, trigger the preventive control strategy, and reduce the obstruction risk index by adjusting the vibration frequency and the feeding rate; The third unit is used to determine that an obstruction occurs when the obstruction risk index reaches the first risk interval or the signal of the level sensor is detected to be abnormal, and record the complete working condition data at the moment when the obstruction occurs; Based on the complete working condition data, start the obstruction removal control system of deep reinforcement learning. Input the complete working condition data into the state encoder to generate a state vector, and combine the successful cases in the historical removal experience database to plan the optimal removal path through the policy network; The fourth unit is used to perform the obstruction removal operation according to the optimal removal path. When it is detected that the signal of the level sensor returns to normal and the feeding rate stabilizes at the target value, it is determined that the obstruction removal is completed, and the complete data record of this obstruction removal is stored in the historical experience database.
9. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Roller anti-accumulation intelligent feeding device
CN118387567A
Feeding bin intelligent management method and system based on self-adaptive control
CN119846973A