Method and system for removing multiple obstacles of feeding bin based on prediction model

By adopting multiple obstacle removal methods based on prediction models in the feed silo, combining attention mechanisms and deep reinforcement learning technology, the problem of lack of prediction and prevention mechanisms in the existing technology is solved, and early identification and intelligent removal of obstacle risks is achieved, which significantly improves production efficiency and equipment stability.

CN120106587AActive Publication Date: 2025-06-06BEIJING EAGLE TECH CO LTD

Patent Information

Application Number
CN202510589035.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The existing feed silo barrier processing technology lacks prediction and prevention mechanisms, resulting in a decrease in production efficiency and an increase in equipment loss, especially in complex operating conditions, with a low success rate of multiple barrier problems being handled.

Method used

The multiple obstacle removal method of the feed silo based on the prediction model is adopted, and the material parameter data and equipment status data are adaptively weighted and fused through the attention mechanism to output the obstacle risk index, and a preventive control strategy is triggered according to the risk range. When obstacles occur, use the obstacle relief control system of deep reinforcement learning, and combine the successful cases in the historical relief experience library to plan the optimal relief path.

Benefits of technology

Early identification and preventive control of hinder risks have been achieved, the frequency of hinder occurrence has been significantly reduced, the continuous operation capability and stability of the production line have been improved, and through intelligent path planning, the demand for manual intervention has been reduced, and the labor intensity and safety risks of operators have been reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106587A_ABST
    Figure CN120106587A_ABST
Patent Text Reader

Abstract

The invention provides a method and system for removing multiple obstacles of a feeding bin based on a prediction model, and relates to the technical field of obstacle removal, and the method comprises the steps: carrying out the weighted fusion of material parameters and equipment state data in the feeding bin through an attention mechanism, and outputting an obstacle risk index. Dividing a risk interval according to the index, and triggering a preventive control strategy to reduce the risk. And when the index reaches a first risk interval or the sensor signal is abnormal, recording working condition data and starting a deep reinforcement learning control system, generating a state vector and planning an optimal release path, and finally executing operation and recording empirical data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to an obstruction removal technology, and in particular to a method and system for removing multiple obstructions in a feed bin based on a prediction model. Background Art

[0002] In the industrial production process, the feed silo is a key equipment for material storage and transportation, and is widely used in metallurgy, building materials, chemicals, grain processing and other industries. The fluidity of materials in the feed silo directly affects production efficiency and product quality. However, due to factors such as diverse material properties, complex environmental conditions and equipment wear, the feed silo often faces material obstruction problems, such as bridging, pipe blockage, rat holes and silo wall attachment, which can lead to production interruptions, equipment damage and even safety accidents.

[0003] Traditionally, the handling of feed silo obstruction mainly relies on manual experience and simple mechanical vibration methods. With the development of industrial automation and intelligence, intelligent monitoring systems based on sensor networks and data analysis are gradually applied to feed silo management. In recent years, with the rapid development of artificial intelligence technology, predictive maintenance and fault diagnosis methods based on deep learning have begun to be applied in the industrial field, providing new ideas for solving the problem of feed silo obstruction.

[0004] The existing feed silo obstruction handling technology mainly has passive response characteristics. Most systems can only detect and handle problems after the obstruction has been formed and caused obvious production abnormalities. The lack of effective prediction and prevention mechanisms leads to reduced production efficiency and increased equipment losses.

[0005] Existing technologies usually adopt a single or simple combination of processing strategies, such as vibrators or airflow impacts with fixed parameters. They are unable to adaptively adjust the removal strategy according to different material characteristics and obstruction types. The removal effect is limited and it is easy to cause secondary obstruction. Especially for multiple obstruction problems under complex working conditions, the processing success rate is low.

[0006] Traditional obstruction removal methods lack systematic experience accumulation and optimization mechanisms, making it difficult to learn from historical cases and improve processing strategies. Operators' experience is difficult to effectively transform into system knowledge, resulting in the obstruction removal process being highly dependent on manual experience and lacking standardized and intelligent solutions, which cannot adapt to the modern industry's demand for automation and intelligence. Summary of the invention

[0007] The embodiments of the present invention provide a method and system for removing multiple obstacles in a feed bin based on a prediction model, which can solve the problems in the prior art.

[0008] According to a first aspect of the embodiments of the present invention, Provides a method to remove multiple obstacles in the feed silo based on a predictive model, including: The attention mechanism is used to perform adaptive weighted fusion of material parameter data and equipment status data in the feed bin, and output the obstruction risk index of the feed bin; The obstruction risk index is divided into a first risk interval and a second risk interval, and when the obstruction risk index is in the second risk interval, a preventive control strategy is triggered to reduce the obstruction risk index by adjusting the vibration frequency and the material feeding rate; When the obstruction risk index reaches the first risk interval or an abnormal signal of the material level sensor is detected, it is determined that an obstruction occurs, and the complete operating condition data at the time when the obstruction occurs is recorded; based on the complete operating condition data, the obstruction release control system of deep reinforcement learning is started, the complete operating condition data is input into the state encoder to generate a state vector, and the optimal release path is planned through the strategy network in combination with the successful cases in the historical release experience library; The obstacle removal operation is performed according to the optimal removal path. When it is detected that the material level sensor signal returns to normal and the material discharge rate is stabilized at the target value, it is determined that the obstacle removal is completed, and the complete data record of this obstacle removal is stored in the historical experience library.

[0009] The obstruction risk index is divided into a first risk interval and a second risk interval. When the obstruction risk index is in the second risk interval, a preventive control strategy is triggered to reduce the obstruction risk index by adjusting the vibration frequency and the material feeding rate. The strategy includes: A probability distribution model of the obstacle risk index is established by using a kernel density estimation method, and a mean and a standard deviation of the obstacle risk index are calculated based on the probability distribution model, and the upper limit threshold of the first risk interval is determined by subtracting the product of the first adjustment coefficient and the standard deviation from the mean, and the upper limit threshold of the second risk interval is determined by adding the product of the second adjustment coefficient and the standard deviation to the mean; When the obstruction risk index is greater than the upper threshold of the first risk interval and less than the upper threshold of the second risk interval, constructing a preventive control objective function including an obstruction risk index term, a system energy consumption term, and an equipment stress term, and calculating a vibration frequency adjustment amount and a material feeding rate adjustment amount based on the preventive control objective function; Obstruction prevention control is performed according to the vibration frequency adjustment amount and the material feeding rate adjustment amount. During the obstruction prevention control process, the obstruction risk index change rate, the system energy consumption change rate and the equipment stress change rate are calculated in real time to reduce the obstruction risk index.

[0010] Constructing a preventive control objective function including an obstacle risk index item, a system energy consumption item, and an equipment stress item, and calculating the vibration frequency adjustment amount and the material feeding rate adjustment amount based on the preventive control objective function includes: Obtaining system operation history data within a preset time window, the system operation history data including vibration frequency, material feeding rate and equipment stress, calculating the corresponding mean and standard deviation of the system operation history data, subtracting the corresponding mean from the currently collected vibration frequency and material feeding rate and dividing the result by the corresponding standard deviation to complete standardization processing; Based on the standardized vibration frequency and feeding rate, a nonlinear mapping model is constructed, and the deviation between the output value of the nonlinear mapping model and the preset target value is calculated to obtain an obstacle risk index, wherein the obstacle risk index is used to characterize the degree of deviation of the system operation; In the preset time window, the vibration frequency difference between adjacent sampling moments is calculated to obtain a vibration frequency change sequence, the material feeding rate difference between adjacent sampling moments is calculated to obtain a material feeding rate change sequence, and the product of the square sum of the vibration frequency change sequence and the material feeding rate change sequence and the energy consumption penalty factor is used as the system energy consumption evaluation value; The vibration frequency and the material feeding rate are weighted to obtain a stress accumulation term, and the stress accumulation term is added to the current equipment stress to obtain an equipment stress assessment value; According to the comparison results of the obstacle risk index, the system energy consumption evaluation value and the equipment stress evaluation value with their corresponding benchmark operation thresholds, the weight coefficients of the three evaluation indicators are dynamically adjusted, and the products of the three evaluation indicators and the corresponding weight coefficients are added to obtain a comprehensive objective function value; Calculate the frequency change rate of the comprehensive objective function value to the vibration frequency, multiply the frequency change rate by the vibration frequency learning coefficient to obtain the vibration frequency adjustment amount, calculate the rate change rate of the comprehensive objective function value to the feeding rate, multiply the rate change rate by the feeding rate learning coefficient to obtain the feeding rate adjustment amount.

[0011] Based on the complete working condition data, the deep reinforcement learning obstacle removal control system is started, the complete working condition data is input into the state encoder to generate a state vector, and the optimal removal path is planned through the strategy network in combination with the successful cases in the historical removal experience library, including: Constructing a working condition state vector based on the complete working condition data, wherein the working condition state vector includes material parameters, vibration parameters, equipment parameters and environmental parameters; inputting the working condition state vector into a state encoder, and obtaining an encoded state vector through the state encoder; Obtain the state feature vector, control action, immediate reward and next state vector in the historical operation data of the feed silo, form a state transition tuple with the state feature vector, the control action, the immediate reward and the next state vector, and store the state transition tuple in a historical release experience database; Constructing an obstacle removal trajectory based on the state transition tuple, calculating the ratio of the number of instant rewards greater than the reward threshold in the obstacle removal trajectory to the trajectory length to obtain a success rate, and screening successful cases in the historical removal experience library according to the success rate; generating a mean function and a variance function based on the encoded state vector, constructing a Gaussian distribution according to the mean function and the variance function, and sampling from the Gaussian distribution to obtain a release action; Execute the release action and obtain a new immediate reward, calculate the value function based on the new immediate reward, and the value function is the cumulative reward expectation with a discount factor; adjust the network parameters of the strategy network according to the cumulative reward expectation, update the execution result to the historical release experience library, and plan the optimal release path.

[0012] Constructing an obstacle removal trajectory based on the state transition tuple, calculating the ratio of the number of instant rewards greater than the reward threshold in the obstacle removal trajectory to the trajectory length to obtain a success rate, and screening the successful cases in the historical removal experience library according to the success rate includes: Collecting the state data and execution actions of the feed bin during the process of removing the obstruction, wherein the state data includes the material level, the material discharge rate and the vibration frequency, and forming a state transition tuple with the state data and the execution action at the current moment and the state data at the next moment, and connecting multiple state transition tuples in chronological order to construct an obstruction removal trajectory; For each state transition tuple in the obstacle removal trajectory, an immediate reward is obtained by subtracting the absolute value of the target material level from the current material level and dividing it by the material level tolerance, and the number of state transition tuples whose immediate rewards are greater than 0.8 is counted, and the number of state transition tuples is divided by the total number of state transition tuples included in the obstacle removal trajectory to obtain a trajectory success rate; Inputting the state data in the obstacle removal trajectory into a feature extraction model, calculating the Euclidean distance between adjacent states, dividing the Euclidean distance by the state range to obtain a normalized state distance, and clustering the obstacle removal trajectory based on the normalized state distance; The obstacle removal trajectory corresponding to the trajectory success rate greater than the preset success rate threshold is marked as a success case.

[0013] Calculating a value function based on the instant reward, the value function is a cumulative reward expectation with a discount factor; adjusting the network parameters of the strategy network according to the value function, updating the execution result to the historical release experience library, and planning the optimal release path include: Collect the discharge volume, feed rate and blockage degree of the feed bin, subtract the minimum discharge volume from the discharge volume and divide the result by the discharge volume variation range to obtain a first reward value, divide the feed rate by the rated feed rate to obtain a second reward value, divide the blockage degree by a preset reward threshold to obtain a third reward value, and perform weighted summation of the first reward value, the second reward value and the third reward value to obtain a comprehensive reward; According to the current state data of the feed bin and the comprehensive reward, a state value function is calculated, where the state value function is a cumulative reward with a discount factor starting from the current moment, and the value range of the discount factor is 0 to 1; the comprehensive reward and the state value function of the next moment are multiplied by the discount factor and then added, and the state value function of the current moment is subtracted to obtain a time series difference error; Calculating the parameter gradient of the policy network based on the temporal difference error, and updating the network parameters of the policy network after multiplying the parameter gradient with a preset learning rate; sampling the state of the feed bin using the updated network parameters to generate a control trajectory including a state sequence and an action sequence; Calculate the value function of the state at each moment in the control trajectory, multiply the value function at each moment by the discount factor at the corresponding moment and then sum them to obtain a trajectory score; and select the optimal release path according to the trajectory score.

[0014] The obstacle removal operation is performed according to the optimal removal path. When it is detected that the material level sensor signal returns to normal and the material discharge rate is stable at the target value, the obstacle removal is determined to be completed, and the complete data record of this obstacle removal is stored in the historical experience database, including: Acquire a material level sensor signal and a material discharge rate, and when the material level sensor signal is lower than a first threshold and the material discharge rate is lower than a second threshold, determine that a material blockage occurs in the feed bin, and trigger an obstruction removal operation; The material level sensor signal and the material discharge rate are monitored in real time. When the material level sensor signal returns to a normal range and the fluctuation amplitude of the material discharge rate is less than a preset value for a predetermined period of time, it is determined that the obstacle removal is completed, and the status data, control parameters and removal results of this obstacle removal are recorded in the historical experience database.

[0015] According to a second aspect of the embodiments of the present invention, Provide a multi-blocking removal system for feed silos based on predictive models, including: The first unit is used to use the attention mechanism to perform adaptive weighted fusion on the material parameter data and equipment status data in the feed bin, and output the obstruction risk index of the feed bin; The second unit is used to divide the obstruction risk index into a first risk interval and a second risk interval, and when the obstruction risk index is in the second risk interval, trigger a preventive control strategy to reduce the obstruction risk index by adjusting the vibration frequency and the material feeding rate; The third unit is used to determine that an obstruction occurs when the obstruction risk index reaches the first risk interval or an abnormal signal of the material level sensor is detected, and record the complete working condition data at the time when the obstruction occurs; based on the complete working condition data, start the obstruction release control system of deep reinforcement learning, input the complete working condition data into the state encoder to generate a state vector, and plan the optimal release path through the strategy network in combination with the successful cases in the historical release experience library; The fourth unit is used to perform the obstacle removal operation according to the optimal removal path. When it is detected that the material level sensor signal returns to normal and the material discharge rate is stabilized at the target value, it is determined that the obstacle removal is completed, and the complete data record of this obstacle removal is stored in the historical experience database.

[0016] According to a third aspect of the embodiments of the present invention, An electronic device is provided, comprising: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0017] According to a fourth aspect of the embodiments of the present invention, A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0018] The beneficial effects of this application are as follows: The present invention adopts the attention mechanism to perform adaptive weighted fusion of material parameter data and equipment status data in the feed silo, outputs an obstruction risk index, and takes corresponding measures according to the risk interval, thereby realizing early identification and preventive control of obstruction risks, significantly reducing the frequency of obstruction occurrence, and improving the continuous operation capability and stability of the production line.

[0019] The present invention uses an obstruction removal control system based on deep reinforcement learning to plan the optimal removal path based on complete working condition data and successful cases in the historical removal experience library, thereby achieving intelligent and precise solutions to obstruction problems, effectively reducing the need for manual intervention, and reducing the labor intensity and safety risks of operators.

[0020] The present invention constructs a complete closed-loop system for obstruction processing, stores the complete data records of each obstruction removal in the historical experience library, continuously enriches and optimizes the solution, and the system has the ability of continuous learning and self-improvement. With the increase of usage time, the efficiency and success rate of obstruction removal will continue to improve, providing an effective solution for the intelligent upgrade of industrial production. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a flow chart of a method for removing multiple obstacles in a feed bin based on a prediction model according to an embodiment of the present invention; Figure 2 A complete flow chart of reducing the obstacle risk index according to an embodiment of the present invention; Figure 3 This is a schematic diagram of comparative analysis of obstacle risk indexes according to an embodiment of the present invention; Figure 4 This is a flow chart of obstacle removal path planning according to an embodiment of the present invention; Figure 5 This is a flow chart of optimal release path planning based on a value function according to an embodiment of the present invention. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0023] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0024] Figure 1 FIG. 1 is a flow chart of a method for removing multiple obstacles in a feed bin based on a prediction model according to an embodiment of the present invention. Figure 1 As shown, the method includes: The attention mechanism is used to perform adaptive weighted fusion of material parameter data and equipment status data in the feed bin, and output the obstruction risk index of the feed bin; The obstruction risk index is divided into a first risk interval and a second risk interval, and when the obstruction risk index is in the second risk interval, a preventive control strategy is triggered to reduce the obstruction risk index by adjusting the vibration frequency and the material feeding rate; When the obstruction risk index reaches the first risk interval or an abnormal signal of the material level sensor is detected, it is determined that an obstruction occurs, and the complete operating condition data at the time when the obstruction occurs is recorded; based on the complete operating condition data, the obstruction release control system of deep reinforcement learning is started, the complete operating condition data is input into the state encoder to generate a state vector, and the optimal release path is planned through the strategy network in combination with the successful cases in the historical release experience library; The obstacle removal operation is performed according to the optimal removal path. When it is detected that the material level sensor signal returns to normal and the material discharge rate is stabilized at the target value, it is determined that the obstacle removal is completed, and the complete data record of this obstacle removal is stored in the historical experience library.

[0025] In an optional implementation, the obstruction risk index is divided into a first risk interval and a second risk interval. When the obstruction risk index is in the second risk interval, a preventive control strategy is triggered to reduce the obstruction risk index by adjusting the vibration frequency and the material feeding rate, including: A probability distribution model of the obstacle risk index is established by using a kernel density estimation method, and a mean and a standard deviation of the obstacle risk index are calculated based on the probability distribution model, and the upper limit threshold of the first risk interval is determined by subtracting the product of the first adjustment coefficient and the standard deviation from the mean, and the upper limit threshold of the second risk interval is determined by adding the product of the second adjustment coefficient and the standard deviation to the mean; When the obstruction risk index is greater than the upper threshold of the first risk interval and less than the upper threshold of the second risk interval, constructing a preventive control objective function including an obstruction risk index term, a system energy consumption term, and an equipment stress term, and calculating a vibration frequency adjustment amount and a material feeding rate adjustment amount based on the preventive control objective function; Obstruction prevention control is performed according to the vibration frequency adjustment amount and the material feeding rate adjustment amount. During the obstruction prevention control process, the obstruction risk index change rate, the system energy consumption change rate and the equipment stress change rate are calculated in real time to reduce the obstruction risk index.

[0026] The obstruction risk index is a key indicator to measure the possibility of obstruction in the material conveying system, which can be calculated by fusion of multiple sensor data. The typical calculation method is to combine factors such as material moisture content, adhesion coefficient, feeding rate and vibration frequency, and perform weighted summation according to certain weights. For example, the calculation weights of the obstruction risk index of a certain feeding system are: material moisture content 35%, adhesion coefficient 25%, feeding rate 25% and vibration frequency 15%. When the material moisture content is 8% (standardized value 0.72), the adhesion coefficient is 0.45 (standardized value 0.68), the feeding rate is 90kg / min (standardized value 0.82), and the vibration frequency is 62Hz (standardized value 0.45), the calculated obstruction risk index is 0.68.

[0027] In order to scientifically divide the risk interval, this method uses kernel density estimation technology to establish a probability distribution model of the obstacle risk index. Kernel density estimation is a non-parametric statistical method that can infer a continuous probability density function based on historical data samples without presetting the distribution type. In practical applications, the Gaussian kernel function can be selected as the kernel function, and the bandwidth parameter is set to 0.08. Assuming that 500 historical obstacle risk index samples are collected, the probability distribution obtained by kernel density estimation shows that the index presents an approximate normal distribution characteristic with a mean of 0.42 and a standard deviation of 0.15.

[0028] Based on the calculated mean and standard deviation, the threshold of the risk interval is further determined. The upper threshold of the first risk interval is calculated by subtracting the product of the first adjustment coefficient and the standard deviation from the mean. The first adjustment coefficient can be set to 0.8, then the upper threshold of the first risk interval = 0.42-0.8×0.15=0.30. The upper threshold of the second risk interval is calculated by adding the mean to the product of the second adjustment coefficient and the standard deviation. The second adjustment coefficient can be set to 1.5, then the upper threshold of the second risk interval = 0.42+1.5×0.15=0.645. Thus, the obstacle risk index is divided into three intervals: the safe interval (00.645) and the second risk interval (>0.645).

[0029] When the obstruction risk index is in the first risk interval, the system enters the warning state but does not trigger the control action; when the obstruction risk index is in the second risk interval, the preventive control strategy is triggered. The core of preventive control is to build a multi-objective control function, taking into account the three aspects of obstruction risk, system energy consumption and equipment stress.

[0030] The preventive control objective function contains three items: the obstacle risk index item, the system energy consumption item, and the equipment stress item. Among them, the obstacle risk index item is proportional to the square of the current obstacle risk index; the system energy consumption item is proportional to the square of the vibration frequency and the material feeding rate; the equipment stress item is proportional to the cube of the vibration frequency. The weight coefficients between each item can be adjusted according to the actual application requirements. In general, the weight of the obstacle risk index item can be set to 0.6, the weight of the system energy consumption item can be set to 0.25, and the weight of the equipment stress item can be set to 0.15.

[0031] Based on the constructed preventive control objective function, the vibration frequency adjustment amount and the feeding rate adjustment amount are calculated. The calculation process adopts the gradient descent method to find the optimal solution of the objective function through iteration. The initial vibration frequency is 60Hz, the initial feeding rate is 85kg / min, the learning rate is set to 0.02, and the maximum number of iterations is 50. After calculation, the vibration frequency adjustment amount is +8.5Hz, and the feeding rate adjustment amount is -12.7kg / min. That is, the vibration frequency is adjusted to 68.5Hz, and the feeding rate is adjusted to 72.3kg / min.

[0032] After executing the above control parameter adjustment, it is necessary to monitor the system operation status in real time and calculate the change rate of key indicators. The change rate of the obstacle risk index refers to the change in the obstacle risk index per unit time, which can be calculated by dividing the difference between the two measured values ​​by the time interval. For example, the obstacle risk index was 0.68 before adjustment and dropped to 0.52 after 10 minutes of adjustment, so the change rate of the obstacle risk index is -0.016 / min. The change rate of system energy consumption refers to the change in energy consumption per unit time, which can be measured by a power meter. For example, the power before adjustment is 3.2kW and the power after adjustment is 3.5kW, so the change rate of system energy consumption is +0.3kW. The change rate of equipment stress refers to the change in stress borne by the equipment per unit time, which can be measured by strain gauges or vibration sensors. For example, the equipment stress before adjustment is 45MPa and after adjustment is 52MPa, so the equipment stress change rate is +7MPa.

[0033] According to the change rate of each indicator, the control parameters can be further optimized. If the obstacle risk index decreases significantly but the system energy consumption and equipment stress increase insignificantly, the current control parameters are maintained; if the obstacle risk index decreases insignificantly but the system energy consumption or equipment stress increases too fast, the weight ratio of the control parameters needs to be adjusted to reduce the weight of the energy consumption item or the stress item. This real-time feedback adjustment mechanism can effectively reduce the obstacle risk and ensure the economy of system operation and the safety of equipment.

[0034] During the operation of the feeding system, the obstruction risk index may fluctuate due to changes in material properties, environmental conditions, etc. When the obstruction risk index drops to a safe range (below 0.30), the normal operating parameters can be gradually restored, that is, the vibration frequency and feeding rate can be adjusted back to the initial values. The recovery process should be carried out slowly to avoid sudden changes in parameters that may cause system instability.

[0035] The technology of this embodiment comes from the field of industrial process control, especially the application of preventive control and multi-objective optimization theory. In the prior art, the obstruction problem of the material conveying system mainly relies on post-detection and processing, that is, taking measures to solve it after the obstruction occurs, or adopting a simple early warning mechanism with a fixed threshold, which lacks scientific risk zoning and preventive control capabilities.

[0036] Figure 2 This is a complete flow chart of reducing the obstacle risk index according to an embodiment of the present invention: The figure shows a probability distribution model for constructing an obstacle risk index using the kernel density estimation method, through which the mean and standard deviation of the obstacle risk index are calculated. On this basis, the calculated mean is subtracted from the product of the first adjustment coefficient and the standard deviation, and it is determined as the upper threshold of the first risk interval; at the same time, the mean is added to the product of the second adjustment coefficient and the standard deviation, and it is determined as the upper threshold of the second risk interval, thereby establishing the basis for risk classification. When the obstacle risk index detected during system operation is within a specific range, that is, greater than the upper threshold of the first risk interval but less than the upper threshold of the second risk interval, the system will automatically construct a comprehensive prevention and control objective function including the obstacle risk index term, the system energy consumption term, and the equipment stress term. Based on this prevention and control objective function, the system can calculate the required vibration frequency adjustment amount and the material feeding rate adjustment amount. In the specific obstacle prevention and control process, the system will implement control measures based on the calculated vibration frequency adjustment amount and material feeding rate adjustment amount, and monitor and calculate the change rate of the obstacle risk index, the change rate of the system energy consumption, and the change rate of the equipment stress in real time. Through this dynamic adjustment method, the control goal of reducing the obstacle risk index is finally achieved.

[0037] In contrast, this application uses the kernel density estimation method to establish a probability distribution model of the obstruction risk index, and scientifically divides the risk interval based on statistical principles, which has stronger adaptability and accuracy. At the same time, a multi-objective control function including obstruction risk, system energy consumption and equipment stress is constructed to take into account both economy and safety while reducing obstruction risk. The starting point of the improvement is to improve the continuous and stable operation capability of the material conveying system, reduce the frequency of obstruction events, and avoid energy waste and equipment loss caused by excessive control.

[0038] In an optional implementation, constructing a preventive control objective function including an obstacle risk index term, a system energy consumption term, and an equipment stress term, and calculating the vibration frequency adjustment amount and the material feeding rate adjustment amount based on the preventive control objective function includes: Obtaining system operation history data within a preset time window, the system operation history data including vibration frequency, material feeding rate and equipment stress, calculating the corresponding mean and standard deviation of the system operation history data, subtracting the corresponding mean from the currently collected vibration frequency and material feeding rate and dividing the result by the corresponding standard deviation to complete standardization processing; Based on the standardized vibration frequency and feeding rate, a nonlinear mapping model is constructed, and the deviation between the output value of the nonlinear mapping model and the preset target value is calculated to obtain an obstacle risk index, wherein the obstacle risk index is used to characterize the degree of deviation of the system operation; In the preset time window, the vibration frequency difference between adjacent sampling moments is calculated to obtain a vibration frequency change sequence, the material feeding rate difference between adjacent sampling moments is calculated to obtain a material feeding rate change sequence, and the product of the square sum of the vibration frequency change sequence and the material feeding rate change sequence and the energy consumption penalty factor is used as the system energy consumption evaluation value; The vibration frequency and the material feeding rate are weighted to obtain a stress accumulation term, and the stress accumulation term is added to the current equipment stress to obtain an equipment stress assessment value; According to the comparison results of the obstacle risk index, the system energy consumption evaluation value and the equipment stress evaluation value with their corresponding benchmark operation thresholds, the weight coefficients of the three evaluation indicators are dynamically adjusted, and the products of the three evaluation indicators and the corresponding weight coefficients are added to obtain a comprehensive objective function value; Calculate the frequency change rate of the comprehensive objective function value to the vibration frequency, multiply the frequency change rate by the vibration frequency learning coefficient to obtain the vibration frequency adjustment amount, calculate the rate change rate of the comprehensive objective function value to the feeding rate, multiply the rate change rate by the feeding rate learning coefficient to obtain the feeding rate adjustment amount.

[0039] In the vibration conveying system, the currently collected vibration frequency and feeding rate data need to be standardized to eliminate the impact of different dimensions. First, obtain the system operation history data within the preset time window, including vibration frequency, feeding rate and equipment stress. In practical applications, the preset time window can be set to the operation data of nearly 30 minutes, and the sampling interval is 1 second. For example, the mean vibration frequency collected by the system within this time window is 48.5Hz, and the standard deviation is 2.3Hz; the mean feeding rate is 12.8kg / min, and the standard deviation is 1.2kg / min. For the currently collected vibration frequency of 52.1Hz and feeding rate of 14.5kg / min, the standardized vibration frequency is (52.1-48.5) / 2.3=1.57, and the standardized feeding rate is (14.5-12.8) / 1.2=1.42. Through standardization, different physical quantities can be compared and calculated on the same scale.

[0040] A nonlinear mapping model is constructed, which takes the standardized vibration frequency and feeding rate as input and outputs the evaluation value of the system operation status. The nonlinear mapping model can be in the form of a polynomial function. For example, the standardized vibration frequency is recorded as x, and the standardized feeding rate is recorded as y. Then the output value of the nonlinear mapping model can be expressed as: a×x²+b×y²+c×x×y+d×x+e×y+f, where a, b, c, d, e, and f are model parameters, which can be obtained by fitting historical data. In a certain embodiment, the parameter values ​​can be set as: a=0.5, b=0.3, c=0.2, d=-0.1, e=-0.2, and f=0.1. The output value of the model is compared with the preset target value (for example, set to 0.8), and the deviation is calculated to obtain the obstacle risk index. For example, for the above-mentioned standardized data, the output value of the nonlinear mapping model is 2.17, and the deviation from the preset target value of 0.8 is 1.37, that is, the obstacle risk index is 1.37, indicating that the system operation state has significantly deviated from the ideal state.

[0041] In terms of system energy consumption evaluation, it is necessary to calculate the changes in vibration frequency and feeding rate. Within the preset time window, the difference in vibration frequency at adjacent sampling moments is calculated to obtain the vibration frequency change sequence, and the difference in feeding rate at adjacent sampling moments is calculated to obtain the feeding rate change sequence. For example, the vibration frequencies at 5 consecutive sampling points are 48.2Hz, 48.7Hz, 49.1Hz, 48.9Hz, and 49.3Hz, respectively, and the corresponding vibration frequency change sequence is 0.5Hz, 0.4Hz, -0.2Hz, and 0.4Hz; the feeding rates are 12.5kg / min, 12.8kg / min, 13.1kg / min, 13.0kg / min, and 13.2kg / min, respectively, and the corresponding feeding rate change sequence is 0.3kg / min, 0.3kg / min, -0.1kg / min, and 0.2kg / min. The product of the sum of the squares of the vibration frequency change sequence and the feeding rate change sequence and the energy consumption penalty factor (for example, a value of 0.05) is used as the system energy consumption evaluation value. For the above example data, the sum of squares of the vibration frequency change sequence is 0.61, the sum of squares of the feeding rate change sequence is 0.23, and the system energy consumption assessment value is 0.05×(0.61+0.23)=0.042.

[0042] In terms of equipment stress assessment, the vibration frequency and the material feeding rate are weighted to obtain the stress accumulation term, which is then added to the current equipment stress to obtain the equipment stress assessment value. For example, if the weight of the vibration frequency is set to 0.6 and the weight of the material feeding rate is set to 0.4, the stress accumulation term is 0.6×52.1+0.4×14.5=37.06. Assuming that the current equipment stress is 25.3, the equipment stress assessment value is 37.06+25.3=62.36.

[0043] In order to dynamically adjust the weight coefficients of the three evaluation indicators, the system will compare the obstruction risk index, system energy consumption evaluation value, and equipment stress evaluation value with their corresponding benchmark operation thresholds. The benchmark operation threshold can be obtained based on the statistics of the system's historical operation data. For example, the obstruction risk index threshold is 1.0, the system energy consumption evaluation value threshold is 0.05, and the equipment stress evaluation value threshold is 60.0. When any evaluation indicator exceeds its threshold, the corresponding weight coefficient will increase. For example, the initial weight is set as follows: the obstruction risk index weight is 0.4, the system energy consumption evaluation value weight is 0.3, and the equipment stress evaluation value weight is 0.3. For the above example data, the obstruction risk index of 1.37 exceeds the threshold of 1.0, and the equipment stress evaluation value of 62.36 exceeds the threshold of 60.0, so the adjusted weights may become: the obstruction risk index weight is 0.45, the system energy consumption evaluation value weight is 0.25, and the equipment stress evaluation value weight is 0.3.

[0044] The product of the three evaluation indicators and the corresponding weight coefficients is added to obtain the comprehensive objective function value, that is, 0.45×1.37+0.25×0.042+0.3×62.36=19.32. The corresponding adjustment amount can be obtained by calculating the change rate of the comprehensive objective function value to the vibration frequency and the feeding rate. The change rate can be obtained by numerical calculation methods, such as calculating the change of the comprehensive objective function value when the vibration frequency increases and decreases by a small amount (such as ±0.1Hz) to obtain the corresponding change rate. For example, the change rate of the comprehensive objective function value to the vibration frequency is calculated to be -0.28, which is multiplied by the vibration frequency learning coefficient (for example, the value is 0.5) to obtain the vibration frequency adjustment amount of -0.28×0.5=-0.14Hz; the change rate of the comprehensive objective function value to the feeding rate is calculated to be -0.32, which is multiplied by the feeding rate learning coefficient (for example, the value is 0.3) to obtain the feeding rate adjustment amount of -0.32×0.3=-0.096kg / min.

[0045] According to the calculated vibration frequency adjustment and material feeding rate adjustment, the system adjusts the current vibration frequency of 52.1Hz to 51.96Hz and the current material feeding rate of 14.5kg / min to 14.404kg / min. In this way, the system can accurately control the operating parameters of the vibration conveying system while considering the risk of obstruction, energy consumption and equipment stress, ensuring the stability and reliability of the system during long-term operation.

[0046] The practical application of this method shows that compared with the traditional single parameter control method, this method can reduce the occurrence rate of system obstruction events by about 35%, the system energy consumption by about 12%, and the equipment failure rate by about 28%, significantly improving the operating efficiency and reliability of the vibration conveying system.

[0047] Figure 3This is a schematic diagram of comparative analysis of the obstacle risk index according to an embodiment of the present invention: The figure shows the comparison of the change trend of the obstruction prediction index of three different control methods within 90 seconds of operation. The triangle curve in the figure represents the present technical solution, the box curve represents the traditional PID control method, and the dot curve represents the fuzzy control method. From the data at the starting time, the initial obstruction prediction index of the three methods is around 0.5. As the running time goes by, the obstruction prediction index of the present technical solution shows a continuous downward trend, which drops to about 0.12 at 90 seconds; the obstruction prediction index of the traditional PID control method shows a fluctuating upward trend, and finally maintains a high level of about 0.65; the obstruction prediction index of the fuzzy control method is relatively stable, but still shows a slow downward trend, and finally drops to about 0.38. By comparison, it can be seen that the present technical solution is significantly better than the other two methods in reducing the obstruction prediction index, and the downward trend is more stable, indicating that the scheme has obvious advantages in the preventive control effect. Especially in the operating range of 35 seconds to 90 seconds, the control effect of the present technical solution becomes more and more obvious, and the downward trend of the obstruction prediction index remains stable, which fully reflects the control effect and system stability of the scheme.

[0048] In an optional implementation, based on the complete working condition data, a deep reinforcement learning obstacle release control system is started, the complete working condition data is input into a state encoder to generate a state vector, and combined with successful cases in a historical release experience library, an optimal release path is planned through a strategy network, including: Constructing a working condition state vector based on the complete working condition data, wherein the working condition state vector includes material parameters, vibration parameters, equipment parameters and environmental parameters; inputting the working condition state vector into a state encoder, and obtaining an encoded state vector through the state encoder; Obtain the state feature vector, control action, immediate reward and next state vector in the historical operation data of the feed silo, form a state transition tuple with the state feature vector, the control action, the immediate reward and the next state vector, and store the state transition tuple in a historical release experience database; Constructing an obstacle removal trajectory based on the state transition tuple, calculating the ratio of the number of instant rewards greater than the reward threshold in the obstacle removal trajectory to the trajectory length to obtain a success rate, and screening successful cases in the historical removal experience library according to the success rate; generating a mean function and a variance function based on the encoded state vector, constructing a Gaussian distribution according to the mean function and the variance function, and sampling from the Gaussian distribution to obtain a release action; Execute the release action and obtain a new immediate reward, calculate the value function based on the new immediate reward, and the value function is the cumulative reward expectation with a discount factor; adjust the network parameters of the strategy network according to the cumulative reward expectation, update the execution result to the historical release experience library, and plan the optimal release path.

[0049] Obtain complete working data of the feeding silo, including material parameters, vibration parameters, equipment parameters and environmental parameters. Material parameters include material particle size distribution, material moisture content, material stacking angle and material fluidity; vibration parameters include vibration frequency, vibration amplitude, vibration direction and vibration duration; equipment parameters include feeding silo geometry, discharge port diameter, silo wall material and surface roughness; environmental parameters include ambient temperature, ambient humidity, air pressure and dust concentration.

[0050] The working state vector is constructed based on the above complete working condition data. For example, the material moisture content is 7.5%, the average material particle size is 25mm, the vibration frequency is 36Hz, the vibration amplitude is 5mm, the ambient temperature is 28℃, the ambient humidity is 65%, and other parameters form a multidimensional vector.

[0051] The operating state vector is input into the state encoder, which adopts a three-layer neural network structure. The number of neurons in the input layer is the same as the dimension of the operating state vector, for example, 24 neurons; the hidden layer adopts a two-layer structure, with 128 neurons in the first layer and 64 neurons in the second layer; the number of neurons in the output layer is 32, representing the dimension of the encoded state vector. The neuron activation function adopts the ReLU function to effectively capture the nonlinear characteristics of the operating condition data. After being processed by the state encoder, the original 24-dimensional operating state vector is compressed into a 32-dimensional encoded state vector, which is convenient for subsequent decision-making processing.

[0052] Obtain the historical operation data of the feed silo. Extract the state feature vector, control action, immediate reward and next state vector from the historical database, form a state transition tuple, and store it in the historical release experience database. For example, when an arch bridge is detected in the silo, record the state feature vector at that time (such as material moisture content 8.2%, vibration frequency 0Hz, etc.), the control action taken (such as starting a 35Hz vibrator), immediate reward (such as a 15% increase in material fluidity is recorded as +15), and the next state vector after the action is executed (such as a vibration frequency of 35Hz, the material begins to flow, etc.).

[0053] Construct an obstacle removal trajectory based on state transition tuples. A complete obstacle removal trajectory may contain multiple continuous state transition tuples, such as the entire process from detecting an obstacle to completely removing the obstacle after multiple control actions. The system calculates the ratio of the number of immediate rewards greater than the reward threshold (such as setting +10 as the threshold) in each trajectory to the total length of the trajectory to obtain the success rate. For example, in a trajectory containing 8 state transition tuples, 6 immediate rewards are greater than +10, then the success rate of the trajectory is 75%. The system selects trajectories with a success rate greater than 80% and marks the state transition tuples contained in them as successful cases.

[0054] The mean function and variance function are generated based on the encoded state vector. The mean function uses a three-layer fully connected neural network with a 32-dimensional encoded state vector as input, 64 neurons in the hidden layer, and an action space dimension (e.g., 8 dimensions, representing the control parameters of different actuators). The variance function also uses a similar structure, but the output layer is also 8 dimensions, representing the uncertainty of each dimension of the action.

[0055] A Gaussian distribution is constructed based on the output values ​​of the mean function and the variance function. For example, for the action dimension that controls the frequency of the vibrator, the mean function outputs 42Hz and the variance function outputs 4. The action of this dimension obeys a Gaussian distribution with a mean of 42 and a variance of 4. The system samples from this Gaussian distribution to obtain a specific release action, such as setting the vibration frequency to 40.5Hz.

[0056] Execute the release action and get a new instant reward. The calculation of the instant reward is based on multiple indicators, including the degree of improvement in material flow (weight 0.4), energy consumption level (weight 0.2), equipment load (weight 0.2) and release time (weight 0.2). For example, after executing the control action with a vibration frequency of 40.5Hz, the material flow is improved by 25% (score +25), the energy consumption level is moderate (score +5), the equipment load is normal (score +10), and the release time is short (score +15). The comprehensive weighted calculation results in an instant reward of +17.

[0057] The value function is calculated based on the new immediate reward. The value function is expressed as the cumulative reward expectation with a discount factor. The discount factor is set to 0.95, which indicates the degree to which the importance of future rewards decays over time. For a state s, its value is the sum of the current immediate reward and the value of the next state multiplied by the discount factor. For example, the current state receives an immediate reward of +17, and the predicted value of the next state is +35, then the value of the current state is estimated to be +17+0.95×35=+50.25.

[0058] The network parameters of the policy network are adjusted according to the expected cumulative reward. The policy network is optimized using stochastic gradient ascent, and the learning rate is set to 0.0005. The parameter update direction is to increase the probability of high-reward actions and reduce the probability of low-reward actions. For example, if an action with a vibration frequency of 40.5Hz receives a high reward, the network parameters are adjusted so that the probability of outputting close to 40.5Hz in similar states increases.

[0059] Update the execution results to the historical release experience library. This includes the current state feature vector, the release action implemented, the immediate reward obtained, and the next state vector transferred to. The updated experience library is used for future strategy learning and optimization.

[0060] After multiple rounds of iterations, the system can plan the optimal removal path. For example, for a specific arch bridge obstruction, the optimal removal path may be: first use 28Hz low-frequency vibration for 15 seconds to destroy the arch bridge structure, then increase to 45Hz high-frequency vibration for 10 seconds to promote material flow, and finally reduce to 20Hz low-frequency vibration for 5 seconds to stabilize the material flow state. This removal path can effectively remove the obstruction in the shortest time and with the lowest energy consumption, and is recorded as a typical success case.

[0061] The above-mentioned deep reinforcement learning-based obstruction removal control system can realize intelligent identification and removal of obstruction problems in the feed bin, greatly improve production efficiency, reduce the need for manual intervention, reduce equipment loss, and provide reliable protection for industrial production. Practical applications have shown that the system saves an average of 63% time, 47% energy consumption, and 38% equipment maintenance frequency compared to traditional manual removal methods, significantly improving production continuity and economic benefits.

[0062] Figure 4 This is a flowchart of the obstacle removal path planning according to an embodiment of the present invention: The figure shows a complete obstacle removal path planning process. Based on the complete working condition data, the working condition state vector is constructed. This vector contains information in four dimensions: material parameters, vibration parameters, equipment parameters, and environmental parameters. These working condition state vectors are input into the state encoder for encoding processing to obtain the encoded state vector. The state feature vector, control action, immediate reward, and next state vector are extracted from the historical operation database of the feed silo, and this information is combined into a state transfer tuple, which is stored in the historical removal experience library. The obstacle removal trajectory is constructed, and the success rate is determined by calculating the ratio of the number of immediate rewards greater than the reward threshold in the trajectory to the trajectory length, and the successful cases in the historical removal experience library are screened out based on this success rate. The mean function and variance function are calculated based on the encoded state vector, and a Gaussian distribution is constructed based on these two functions. The removal action is sampled from the distribution. The removal action is executed and a new immediate reward is obtained. The value function is calculated based on the new immediate reward, which is the cumulative reward expectation with a discount factor. The system adjusts the network parameters of the policy network according to the cumulative reward expectation, updates the execution results to the historical removal experience library, and finally plans the optimal removal path.

[0063] In an optional implementation, constructing an obstacle removal trajectory based on the state transition tuple, calculating the ratio of the number of instant rewards greater than the reward threshold in the obstacle removal trajectory to the trajectory length to obtain a success rate, and screening the successful cases in the historical removal experience library according to the success rate includes: Collecting the state data and execution actions of the feed bin during the process of removing the obstruction, wherein the state data includes the material level, the material discharge rate and the vibration frequency, and forming a state transition tuple with the state data and the execution action at the current moment and the state data at the next moment, and connecting multiple state transition tuples in chronological order to construct an obstruction removal trajectory; For each state transition tuple in the obstacle removal trajectory, an immediate reward is obtained by subtracting the absolute value of the target material level from the current material level and dividing it by the material level tolerance, and the number of state transition tuples whose immediate rewards are greater than 0.8 is counted, and the number of state transition tuples is divided by the total number of state transition tuples included in the obstacle removal trajectory to obtain a trajectory success rate; Inputting the state data in the obstacle removal trajectory into a feature extraction model, calculating the Euclidean distance between adjacent states, dividing the Euclidean distance by the state range to obtain a normalized state distance, and clustering the obstacle removal trajectory based on the normalized state distance; The obstacle removal trajectory corresponding to the trajectory success rate greater than the preset success rate threshold is marked as a success case.

[0064] During the process of removing the obstruction of the feed bin, the system collects the status data of the feed bin and performs actions. The status data mainly includes three key parameters: material level height, material discharge rate and vibration frequency. These three parameters together describe the status characteristics of the feed bin at any time. The system collects this data in real time through the sensor network set up in the feed bin. The sampling frequency is usually set to 10 times per second to ensure the timeliness and continuity of the data.

[0065] While collecting data, the system records the actions performed at each moment, such as adjusting the vibration frequency, changing the feeding rate, etc. The state data St at the current moment t, the execution action At, and the state data St+1 at the next moment t+1 are combined together to form a state transition tuple (St, At, St+1). For example, at moment t, the state St is {material level height 80cm, feeding rate 5kg / s, vibration frequency 30Hz}, the execution action At is {increase vibration frequency 5Hz}, and the state St+1 at the next moment t+1 is {material level height 78cm, feeding rate 5.5kg / s, vibration frequency 35Hz}. This information constitutes a state transition tuple.

[0066] Connect multiple state transition tuples in chronological order to construct a complete obstacle removal trajectory T ={(S1, A1, S2), (S2, A2, S3), ..., (Sn-1, An-1, Sn)}. A complete trajectory usually includes the entire process from the occurrence of the obstacle to its complete removal. The length of the trajectory may range from 30 to 200 state transition tuples depending on the difficulty of removing the obstacle.

[0067] Evaluate the constructed obstacle removal trajectory. For each state transition tuple in the trajectory, the system evaluates the effect of obstacle removal by calculating the difference between the current material level height and the target material level height. Specifically, the absolute value of the target material level height is subtracted from the current material level height, and then divided by the material level tolerance to obtain the immediate reward r. For example, if the target material level height is set to 50cm, the material level tolerance is set to 5cm, and the current material level height is 52cm, then the immediate reward r = 1 - |52-50| / 5 = 0.6.

[0068] Set the reward threshold to 0.8, count the number of state transition tuples in the trajectory with an immediate reward greater than 0.8, and divide this number by the total number of state transition tuples in the trajectory, N_total, to get the trajectory success rate Rate = N_success / N_total. For example, in a trajectory containing 100 state transition tuples, if 75 state transition tuples have an immediate reward greater than 0.8, then the success rate of the trajectory is 75%.

[0069] In order to better understand the similarities and differences between different obstacle removal trajectories, the system also performs cluster analysis on the trajectories. First, the state data in the trajectory is input into a feature extraction model, which can be a pre-trained neural network or a simple feature extraction algorithm. In this embodiment, a three-layer fully connected neural network is used, with the number of input layer nodes being the state dimension (here 3, corresponding to material level height, material discharge rate, and vibration frequency), the number of hidden layer nodes being 20, and the number of output layer nodes being 10, to extract the key features of the state.

[0070] Based on the extracted features, the system calculates the Euclidean distance between adjacent states. For example, for adjacent states St and St+1, the Euclidean distance is calculated as follows: first, the feature vectors ft and ft+1 of the two states are extracted respectively, and then the Euclidean distance d(ft, ft+1) between the two feature vectors is calculated. In order to eliminate the influence of the difference in the range of different state parameters, the system normalizes the Euclidean distance by the state range to obtain the normalized state distance d_norm = d(ft, ft+1) / range, where range is the preset state range value. In this embodiment, the range of the material level height is 0-100cm, the range of the material discharge rate is 0-10kg / s, and the range of the vibration frequency is 0-50Hz.

[0071] Based on the calculated normalized state distance, the system uses a hierarchical clustering algorithm to cluster all obstacle removal trajectories. The system sets the clustering threshold to 0.3. When the average normalized state distance between two trajectories is less than this threshold, they are classified into the same category. Through cluster analysis, the system can discover the commonalities and individualities between different types of obstacle removal strategies, providing a basis for subsequent strategy optimization.

[0072] The obstacle removal trajectory corresponding to the trajectory success rate greater than the preset success rate threshold is marked as a successful case. In this embodiment, the success rate threshold is set to 70%, that is, when the trajectory success rate is greater than 70%, the trajectory is marked as a successful case. These successful cases are saved in the historical removal experience library for reference in the formulation of subsequent obstacle removal strategies.

[0073] Take an actual obstacle removal process as an example: the initial state of the feed bin is {material level height 85cm, material discharge rate 2kg / s, vibration frequency 25Hz}, and the target material level height is 50cm. The system first executes the action {increase vibration frequency 10Hz}, and the state changes to {material level height 82cm, material discharge rate 3kg / s, vibration frequency 35Hz}; then executes the action {increase discharge rate 2kg / s}, and the state changes to {material level height 78cm, material discharge rate 5kg / s, vibration frequency 35Hz}; continue to execute a series of actions, and the final state stabilizes at {material level height 52cm, material discharge rate 5kg / s, vibration frequency 40Hz}. The whole process forms a trajectory containing 20 state transition tuples, of which 16 state transition tuples have an immediate reward greater than 0.8, and the trajectory success rate is 80%, which is higher than the preset success rate threshold of 70%, so it is marked as a successful case.

[0074] Through the above method, the system can filter out high-quality successful cases of obstruction removal from historical data, provide strong support for the formulation and optimization of feed silo obstruction removal strategies, and significantly improve the operating efficiency and reliability of the feed silo.

[0075] In an optional implementation, a value function is calculated based on the instant reward, the value function being the cumulative reward expectation with a discount factor; the network parameters of the strategy network are adjusted according to the value function, the execution result is updated to the historical release experience library, and the optimal release path is planned, including: Collect the discharge volume, feed rate and blockage degree of the feed bin, subtract the minimum discharge volume from the discharge volume and divide the result by the discharge volume variation range to obtain a first reward value, divide the feed rate by the rated feed rate to obtain a second reward value, divide the blockage degree by a preset reward threshold to obtain a third reward value, and perform weighted summation of the first reward value, the second reward value and the third reward value to obtain a comprehensive reward; According to the current state data of the feed bin and the comprehensive reward, a state value function is calculated, where the state value function is a cumulative reward with a discount factor starting from the current moment, and the value range of the discount factor is 0 to 1; the comprehensive reward and the state value function of the next moment are multiplied by the discount factor and then added, and the state value function of the current moment is subtracted to obtain a time series difference error; Calculating the parameter gradient of the policy network based on the temporal difference error, and updating the network parameters of the policy network after multiplying the parameter gradient with a preset learning rate; sampling the state of the feed bin using the updated network parameters to generate a control trajectory including a state sequence and an action sequence; Calculate the value function of the state at each moment in the control trajectory, multiply the value function at each moment by the discount factor at the corresponding moment and then sum them to obtain a trajectory score; and select the optimal release path according to the trajectory score.

[0076] Collect real-time status data of the feed bin, including discharge volume, feed rate and blockage degree. The discharge volume refers to the mass of material flowing out of the feed bin outlet per unit time, in kg / min; the feed rate refers to the speed at which the material is transported to the feed bin, also in kg / min; the blockage degree is the severity of the material blockage in the feed bin measured by a pressure sensor or photoelectric sensor, which can be expressed as a percentage, 0% means no blockage at all, and 100% means complete blockage.

[0077] Based on the collected status data, the instant reward value is calculated. The instant reward consists of three parts: the first reward value, the second reward value and the third reward value. The first reward value reflects the performance of the discharge amount, which is calculated by subtracting the minimum discharge amount from the current discharge amount and dividing it by the discharge amount change range. For example, if the current discharge amount is 85 kg / min, the minimum discharge amount is 20 kg / min, and the discharge amount change range is 100 kg / min, then the first reward value = (85-20) / 100 = 0.65. The second reward value measures the rationality of the feeding rate, which is calculated by dividing the current feeding rate by the rated feeding rate. For example, if the current feeding rate is 90 kg / min and the rated feeding rate is 100 kg / min, then the second reward value = 90 / 100 = 0.9. The third reward value evaluates the blocking situation, which is calculated by dividing the blocking degree by the preset reward threshold. For example, if the current blocking degree is 15% and the preset reward threshold is 50%, then the third reward value = 15 / 50 = 0.3.

[0078] The above three reward values ​​are weighted and summed to get the comprehensive reward. The weighting coefficient can be adjusted according to the actual production situation. Generally, it can be set to 0.5 for the first reward value, 0.3 for the second reward value, and 0.2 for the third reward value. According to this weight configuration, the comprehensive reward of the above example = 0.65×0.5+0.9×0.3+0.3×0.2=0.65.

[0079] Based on the calculated comprehensive reward, the state value function is further calculated. The state value function represents the expected value of the cumulative reward with a discount factor that may be obtained in the future starting from the current moment. The discount factor ranges from 0 to 1 and is used to balance the importance of current rewards and future rewards. In practical applications, the discount factor can be set to 0.95, which means that future rewards will decrease by multiples of 0.95.

[0080] The calculation of the state value function uses the temporal difference learning method. Specifically, the comprehensive reward Rt at the current moment t and the state value function Vt+1 at the next moment t+1 are multiplied by the discount factor γ and then added, and then the state value function Vt at the current moment is subtracted to obtain the temporal difference error δt. For example, if the comprehensive reward at the current moment is 0.65, the current state value function is 3.2, the state value function at the next moment is 3.5, and the discount factor is 0.95, then the temporal difference error = 0.65 + 0.95 × 3.5-3.2 = 0.78.

[0081] The parameters of the policy network are updated based on the calculated time difference error. The policy network is a neural network structure, with the input being the state data of the feed bin and the output being the control action to remove the blockage. The parameter update adopts the gradient descent method, which is to calculate the gradient of the policy network parameters, and then update the network parameters by multiplying the gradient with the preset learning rate. The learning rate is generally set to 0.01 and can be adjusted according to the actual situation.

[0082] After updating the policy network parameters, the network is used to sample the state of the feed bin and generate a control trajectory. The control trajectory contains a series of states and corresponding action sequences. For example, the control trajectory may contain the following content: at time T1, the state of the feed bin is S1, and action A1 is executed; at time T2, the state changes to S2, and action A2 is executed; and so on.

[0083] The trajectory score is the weighted sum of the state value function at each moment in the trajectory, with the weight being the discount factor corresponding to each moment. For example, assuming that the trajectory contains 3 moments, the corresponding state value functions are 4.2, 3.8, and 3.5, and the discount factor is 0.95, then the trajectory score = 4.2 + 0.95 × 3.8 + 0.95² × 3.5 = 10.9.

[0084] Generate multiple control trajectories, calculate their respective trajectory scores, and then select the trajectory with the highest score as the optimal release path. For example, if 5 control trajectories are generated, and the corresponding trajectory scores are 10.9, 9.8, 11.2, 10.5, and 10.1, then select the third trajectory (with a score of 11.2) as the optimal release path.

[0085] To improve the efficiency and stability of the algorithm, an experience replay mechanism can be used. This mechanism stores the execution results (including the current state, the executed actions, the rewards obtained, and the next state) in the historical experience library. When updating the parameters, a batch of data is randomly extracted from the experience library for training, rather than just using the most recent experience. This approach can break the strong correlation between adjacent samples and improve learning efficiency. The experience library size can be set to 10,000, and 128 samples are randomly extracted for each training.

[0086] The policy network can adopt a multi-layer perceptron structure, including an input layer, a hidden layer, and an output layer. The number of nodes in the input layer is equal to the state dimension (3 in this case, corresponding to the discharge volume, feeding rate, and blocking degree respectively); the hidden layer can be set to 2 layers, each with 64 nodes, and the activation function uses ReLU; the number of nodes in the output layer is equal to the action dimension, and the activation function is selected according to the action characteristics.

[0087] The technology of this embodiment comes from the field of reinforcement learning, especially the combination of temporal difference learning and policy gradient method. In the prior art, the problem of feed bin blockage is mainly handled by manual experience judgment and preset rules, lacking intelligence and adaptive capabilities. For example, the traditional method usually sets a fixed blockage threshold. When the blockage degree exceeds the threshold, a predetermined relief measure is taken, such as reducing the feed rate or increasing the vibration frequency.

[0088] Figure 5 The following is a flow chart of the optimal release path planning based on the value function according to an embodiment of the present invention: The flowchart shows the calculation process of an optimal release path. From the beginning of planning the optimal release path, the system collects relevant data of the feed bin, including three key parameters: discharge volume, feed rate, and blockage degree. Entering the change value calculation stage, three important change indicators are calculated respectively: the first change value is calculated by the change range of discharge volume and feed volume, the second change value is obtained by dividing the feed rate by the rated feed rate, and the third change value is the ratio of the blockage degree to the blockage change amplitude. After obtaining these three change values, the system performs a weighted summation operation to obtain a comprehensive change value. Calculate the state value cycle difference. This calculation process needs to consider the cumulative real added expectation with a discount factor. On this basis, calculate the time difference error, and its calculation formula is: error = comprehensive change + γV(s')-V(s). Based on the calculated error value, the system will perform two operations at the same time: on the one hand, update the network parameters based on the error calculation strategy network parameter weight, and on the other hand, use the updated parameters to generate the control trajectory (including state sequence and action sequence). The results of these two parts are comprehensively analyzed, the trajectory score is calculated, and the optimal release path is selected, thereby completing the entire optimization process.

[0089] In contrast, this application adopts a reinforcement learning method, which can automatically learn the optimal release strategy according to the real-time status of the feed bin by constructing a reasonable reward function and value function. The starting point of the improvement is to improve the intelligence level of the release of the feed bin blockage, reduce manual intervention, and achieve adaptive adjustment to different working conditions. The final effect achieved is: the efficiency of blockage release is increased by more than 35%, production continuity is significantly enhanced, and the downtime caused by blockage is reduced by 60%. With the continuous enrichment of the experience base, the system's release ability continues to improve, showing good self-learning characteristics.

[0090] In an optional implementation, the obstacle removal operation is performed according to the optimal removal path, and when it is detected that the material level sensor signal returns to normal and the material discharge rate is stable at the target value, it is determined that the obstacle removal is completed, and the complete data record of this obstacle removal is stored in the historical experience library, including: Acquire a material level sensor signal and a material discharge rate, and when the material level sensor signal is lower than a first threshold and the material discharge rate is lower than a second threshold, determine that a material blockage occurs in the feed bin, and trigger an obstruction removal operation; The material level sensor signal and the material discharge rate are monitored in real time. When the material level sensor signal returns to a normal range and the fluctuation amplitude of the material discharge rate is less than a preset value for a predetermined period of time, it is determined that the obstacle removal is completed, and the status data, control parameters and removal results of this obstacle removal are recorded in the historical experience database.

[0091] During the operation of the feeding system, the system monitors the material level in real time through the material level sensor installed in the feeding bin, and monitors the material discharge rate through the flow meter or weight sensor at the discharge port. The data collected by the system include but are not limited to: material level value (unit: mm), material level change rate (unit: mm / min), discharge rate (unit: kg / min) and fluctuation of discharge rate.

[0092] When the following two conditions are detected to be met at the same time, it is determined that the feed bin is blocked; the material level sensor signal is lower than the first threshold: for example, when the material level is lower than the set safety level (such as 300mm); the feeding rate is lower than the second threshold: for example, when the actual feeding rate is lower than 70% of the target feeding rate (such as 50kg / min), that is, lower than 35kg / min.

[0093] The material level sensor data and material discharge rate data are collected every 1 second. The collected data are averaged in a sliding window (the window size is 10 seconds) to eliminate the impact of instantaneous fluctuations. The averaged material level height is compared with the first threshold (300mm), and the averaged material discharge rate is compared with the second threshold (35kg / min). When the average material level height is always lower than 300mm and the average material discharge rate is always lower than 35kg / min for 30 consecutive seconds, the system determines that a blockage has occurred and triggers the obstruction removal process.

[0094] Query the historical experience database to find historical cases similar to the current blockage situation. The similarity judgment is based on the following parameters: material type (such as cement, coal powder, etc.), ambient humidity range (such as 40%-60%), material residence time (such as more than 24 hours), and blockage location (such as the conical funnel).

[0095] Based on the query results, the system selects the optimal removal path. For example, for a certain powder material that is blocked at the cone funnel under 55% humidity, the historical experience database shows that the most effective removal method is to start the vibrator (amplitude 5mm, frequency 20Hz) for 30 seconds. If there is no effect, increase the amplitude to 8mm and start the pneumatic hammer (impact force 50N).

[0096] Perform the operations in sequence according to the selected release path: Stage 1: Start the vibrator, set the amplitude to 5mm, the frequency to 20Hz, and continue for 30 seconds; Monitor the effect: If there is a significant improvement in the material level sensor signal or the material discharge rate (such as the material discharge rate is increased to above 25kg / min), continue the current operation; Stage 2: If there is no obvious improvement, increase the amplitude to 8mm, and start the pneumatic hammer at the same time, set the impact force to 50N, the impact frequency to 2 times / second, and continue for 20 seconds; Stage 3: If there is still no improvement, start the auxiliary loosening device, such as the pneumatic agitator, at a speed of 30rpm, and continue for 45 seconds.

[0097] During the execution process, the system records various operating parameters and their effects in real time, including: operation type (such as vibration, impact, stirring), operation parameters (such as amplitude, frequency, force); operation duration, material level changes after the operation and changes in material discharge rate.

[0098] Collect data once a second and calculate the average value of the 10-second sliding window; determine whether the material level sensor signal has returned to the normal range (such as greater than 350mm); determine whether the material discharge rate is stable near the target value (such as the target value is 50kg / min, and the actual value reaches 45-55kg / min).

[0099] Calculate the standard deviation of the feeding rate in the last 60 seconds; when the standard deviation is less than a preset value (such as 5% of the target feeding rate, i.e. 2.5 kg / min) for a predetermined period of time (such as 120 consecutive seconds), the feeding rate is determined to be stable.

[0100] When the material level sensor signal returns to the normal range (greater than 350mm); and the material discharge rate stabilizes near the target value (45-55kg / min); and the fluctuation range of the material discharge rate is less than the preset value (standard deviation is less than 2.5kg / min); and the above state lasts for a predetermined period of time (120 seconds); it is determined that the obstacle removal is complete.

[0101] After the obstacle is removed, the system will store the complete data record of this obstacle removal into the historical experience database, including: status data: material level height when the blockage occurs (such as 250mm), material discharge rate (such as 15kg / min), material type (such as cement), environmental conditions (such as temperature 25℃, humidity 55%), material residence time (such as 36 hours); control parameters: the sequence of operations to be performed (such as vibration first and then impact), the specific parameters of each operation (such as amplitude 5mm, frequency 20Hz, duration 30 seconds), the execution time point of each operation; removal results: total time consumed for obstacle removal (such as 95 seconds), the final restored material discharge rate (such as 48kg / min), the material level change curve during the removal process, and the material discharge rate change curve.

[0102] Taking the feeding system of a cement production line as an example, the normal feeding rate of the system is set to 50kg / min and the safe material level is 300mm.

[0103] During a production process, the system detected that the material level dropped to 280mm and lasted for 30 seconds, and the material feeding rate dropped to 20kg / min and lasted for 30 seconds, triggering a material blockage determination. The system queries the historical experience database and selects the optimal release path: vibration first, then impact.

[0104] The vibrator was started (amplitude 5mm, frequency 20Hz) for 30 seconds, and the material feeding rate increased to 30kg / min. After vibrating for 15 seconds, the material feeding rate increased to 40kg / min, and the material level rose to 320mm. The system continued to monitor for 120 seconds, confirming that the material feeding rate was stable at 48kg / min (standard deviation 1.8kg / min), and the material level was stable at 360mm. It was determined that the obstacle removal was completed, and the total time was 165 seconds.

[0105] The complete data record of the blockage removal is stored in the historical experience database, including the conditions for the blockage (cement, humidity 55%, stay for 36 hours), the removal method (vibration parameters, duration) and the effect (removal time, recovery status), to provide a reference for the handling of similar situations in the future.

[0106] According to a second aspect of the embodiments of the present invention, Provide a multi-blocking removal system for feed silos based on predictive models, including: The first unit is used to use the attention mechanism to perform adaptive weighted fusion on the material parameter data and equipment status data in the feed bin, and output the obstruction risk index of the feed bin; The second unit is used to divide the obstruction risk index into a first risk interval and a second risk interval, and when the obstruction risk index is in the second risk interval, trigger a preventive control strategy to reduce the obstruction risk index by adjusting the vibration frequency and the material feeding rate; The third unit is used to determine that an obstruction occurs when the obstruction risk index reaches the first risk interval or an abnormal signal of the material level sensor is detected, and record the complete working condition data at the time when the obstruction occurs; based on the complete working condition data, start the obstruction release control system of deep reinforcement learning, input the complete working condition data into the state encoder to generate a state vector, and plan the optimal release path through the strategy network in combination with the successful cases in the historical release experience library; The fourth unit is used to perform the obstacle removal operation according to the optimal removal path. When it is detected that the material level sensor signal returns to normal and the material discharge rate is stabilized at the target value, it is determined that the obstacle removal is completed, and the complete data record of this obstacle removal is stored in the historical experience database.

[0107] According to a third aspect of the embodiments of the present invention, An electronic device is provided, comprising: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0108] According to a fourth aspect of the embodiments of the present invention, A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0109] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for removing multiple obstacles in a feed bin based on a prediction model, characterized in that: include: The attention mechanism is used to perform adaptive weighted fusion of material parameter data and equipment status data in the feed bin, and output the obstruction risk index of the feed bin; The obstruction risk index is divided into a first risk interval and a second risk interval, and when the obstruction risk index is in the second risk interval, a preventive control strategy is triggered to reduce the obstruction risk index by adjusting the vibration frequency and the material feeding rate; When the obstruction risk index reaches the first risk interval or an abnormal signal of the material level sensor is detected, it is determined that an obstruction occurs, and complete operating condition data at the time when the obstruction occurs is recorded; Based on the complete working condition data, a deep reinforcement learning obstacle removal control system is started, the complete working condition data is input into a state encoder to generate a state vector, and the optimal removal path is planned through a strategy network in combination with successful cases in a historical removal experience library; The obstacle removal operation is performed according to the optimal removal path. When it is detected that the material level sensor signal returns to normal and the material discharge rate is stabilized at the target value, it is determined that the obstacle removal is completed, and the complete data record of this obstacle removal is stored in the historical experience library.

2. The method according to claim 1, characterized in that The obstruction risk index is divided into a first risk interval and a second risk interval. When the obstruction risk index is in the second risk interval, a preventive control strategy is triggered to reduce the obstruction risk index by adjusting the vibration frequency and the material feeding rate. The strategy includes: A probability distribution model of the obstacle risk index is established by using a kernel density estimation method, and a mean and a standard deviation of the obstacle risk index are calculated based on the probability distribution model, and the upper limit threshold of the first risk interval is determined by subtracting the product of the first adjustment coefficient and the standard deviation from the mean, and the upper limit threshold of the second risk interval is determined by adding the product of the second adjustment coefficient and the standard deviation to the mean; When the obstruction risk index is greater than the upper threshold of the first risk interval and less than the upper threshold of the second risk interval, constructing a preventive control objective function including an obstruction risk index term, a system energy consumption term, and an equipment stress term, and calculating a vibration frequency adjustment amount and a material feeding rate adjustment amount based on the preventive control objective function; Obstruction prevention control is performed according to the vibration frequency adjustment amount and the material feeding rate adjustment amount. During the obstruction prevention control process, the obstruction risk index change rate, the system energy consumption change rate and the equipment stress change rate are calculated in real time to reduce the obstruction risk index.

3. The method according to claim 2, characterized in that Constructing a preventive control objective function including an obstacle risk index item, a system energy consumption item, and an equipment stress item, and calculating the vibration frequency adjustment amount and the material feeding rate adjustment amount based on the preventive control objective function includes: Obtaining system operation history data within a preset time window, the system operation history data including vibration frequency, material feeding rate and equipment stress, calculating the corresponding mean and standard deviation of the system operation history data, subtracting the corresponding mean from the currently collected vibration frequency and material feeding rate and dividing the result by the corresponding standard deviation to complete standardization processing; Based on the standardized vibration frequency and feeding rate, a nonlinear mapping model is constructed, and the deviation between the output value of the nonlinear mapping model and the preset target value is calculated to obtain an obstacle risk index, wherein the obstacle risk index is used to characterize the degree of deviation of the system operation; In the preset time window, the vibration frequency difference between adjacent sampling moments is calculated to obtain a vibration frequency change sequence, the material feeding rate difference between adjacent sampling moments is calculated to obtain a material feeding rate change sequence, and the product of the square sum of the vibration frequency change sequence and the material feeding rate change sequence and the energy consumption penalty factor is used as the system energy consumption evaluation value; The vibration frequency and the material feeding rate are weighted to obtain a stress accumulation term, and the stress accumulation term is added to the current equipment stress to obtain an equipment stress assessment value; According to the comparison results of the obstacle risk index, the system energy consumption evaluation value and the equipment stress evaluation value with their corresponding benchmark operation thresholds, the weight coefficients of the three evaluation indicators are dynamically adjusted, and the products of the three evaluation indicators and the corresponding weight coefficients are added to obtain a comprehensive objective function value; Calculate the frequency change rate of the comprehensive objective function value to the vibration frequency, multiply the frequency change rate by the vibration frequency learning coefficient to obtain the vibration frequency adjustment amount, calculate the rate change rate of the comprehensive objective function value to the feeding rate, multiply the rate change rate by the feeding rate learning coefficient to obtain the feeding rate adjustment amount.

4. The method according to claim 1, characterized in that: Based on the complete working condition data, the deep reinforcement learning obstacle removal control system is started, the complete working condition data is input into the state encoder to generate a state vector, and the optimal removal path is planned through the strategy network in combination with the successful cases in the historical removal experience library, including: Constructing a working condition state vector based on the complete working condition data, wherein the working condition state vector includes material parameters, vibration parameters, equipment parameters and environmental parameters; inputting the working condition state vector into a state encoder, and obtaining an encoded state vector through the state encoder; Obtain the state feature vector, control action, immediate reward and next state vector in the historical operation data of the feed silo, form a state transition tuple with the state feature vector, the control action, the immediate reward and the next state vector, and store the state transition tuple in a historical release experience database; Constructing an obstacle removal trajectory based on the state transition tuple, calculating the ratio of the number of instant rewards greater than the reward threshold in the obstacle removal trajectory to the trajectory length to obtain a success rate, and screening successful cases in the historical removal experience library according to the success rate; generating a mean function and a variance function based on the encoded state vector, constructing a Gaussian distribution according to the mean function and the variance function, and sampling from the Gaussian distribution to obtain a release action; Execute the release action and obtain a new immediate reward, calculate the value function based on the new immediate reward, and the value function is the cumulative reward expectation with a discount factor; adjust the network parameters of the strategy network according to the cumulative reward expectation, update the execution result to the historical release experience library, and plan the optimal release path.

5. The method according to claim 4, characterized in that Constructing an obstacle removal trajectory based on the state transition tuple, calculating the ratio of the number of instant rewards greater than the reward threshold in the obstacle removal trajectory to the trajectory length to obtain a success rate, and screening the successful cases in the historical removal experience library according to the success rate includes: Collecting the state data and execution actions of the feed bin during the process of removing the obstruction, wherein the state data includes the material level, the material discharge rate and the vibration frequency, and forming a state transition tuple with the state data and the execution action at the current moment and the state data at the next moment, and connecting multiple state transition tuples in chronological order to construct an obstruction removal trajectory; For each state transition tuple in the obstacle removal trajectory, an immediate reward is obtained by subtracting the absolute value of the target material level from the current material level and dividing it by the material level tolerance, and the number of state transition tuples whose immediate rewards are greater than 0.8 is counted, and the number of state transition tuples is divided by the total number of state transition tuples included in the obstacle removal trajectory to obtain a trajectory success rate; Inputting the state data in the obstacle removal trajectory into a feature extraction model, calculating the Euclidean distance between adjacent states, dividing the Euclidean distance by the state range to obtain a normalized state distance, and clustering the obstacle removal trajectory based on the normalized state distance; The obstacle removal trajectory corresponding to the trajectory success rate greater than the preset success rate threshold is marked as a success case.

6. The method according to claim 4, characterized in that Calculating a value function based on the instant reward, wherein the value function is a cumulative reward expectation with a discount factor; Adjusting the network parameters of the strategy network according to the value function, updating the execution results to the historical release experience library, and planning the optimal release path include: Collect the discharge volume, feed rate and blockage degree of the feed bin, subtract the minimum discharge volume from the discharge volume and divide the result by the discharge volume variation range to obtain a first reward value, divide the feed rate by the rated feed rate to obtain a second reward value, divide the blockage degree by a preset reward threshold to obtain a third reward value, and perform weighted summation of the first reward value, the second reward value and the third reward value to obtain a comprehensive reward; According to the current state data of the feed bin and the comprehensive reward, a state value function is calculated, where the state value function is a cumulative reward with a discount factor starting from the current moment, and the value range of the discount factor is 0 to 1; the comprehensive reward and the state value function of the next moment are multiplied by the discount factor and then added, and the state value function of the current moment is subtracted to obtain a time series difference error; Calculating the parameter gradient of the policy network based on the temporal difference error, and updating the network parameters of the policy network after multiplying the parameter gradient with a preset learning rate; sampling the state of the feed bin using the updated network parameters to generate a control trajectory including a state sequence and an action sequence; Calculate the value function of the state at each moment in the control trajectory, multiply the value function at each moment by the discount factor at the corresponding moment and then sum them to obtain a trajectory score; and select the optimal release path according to the trajectory score.

7. The method according to claim 1, characterized in that The obstacle removal operation is performed according to the optimal removal path. When it is detected that the material level sensor signal returns to normal and the material discharge rate is stable at the target value, the obstacle removal is determined to be completed, and the complete data record of this obstacle removal is stored in the historical experience database, including: Acquire a material level sensor signal and a material discharge rate, and when the material level sensor signal is lower than a first threshold and the material discharge rate is lower than a second threshold, determine that a material blockage occurs in the feed bin, and trigger an obstruction removal operation; The material level sensor signal and the material discharge rate are monitored in real time. When the material level sensor signal returns to a normal range and the fluctuation amplitude of the material discharge rate is less than a preset value for a predetermined period of time, it is determined that the obstacle removal is completed, and the status data, control parameters and removal results of this obstacle removal are recorded in the historical experience database.

8. A feed bin multiple obstruction removal system based on a prediction model, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used to use the attention mechanism to perform adaptive weighted fusion on the material parameter data and equipment status data in the feed bin, and output the obstruction risk index of the feed bin; The second unit is used to divide the obstruction risk index into a first risk interval and a second risk interval, and when the obstruction risk index is in the second risk interval, trigger a preventive control strategy to reduce the obstruction risk index by adjusting the vibration frequency and the material feeding rate; The third unit is used to determine that an obstruction occurs when the obstruction risk index reaches the first risk interval or an abnormal signal of the material level sensor is detected, and record complete working condition data at the time when the obstruction occurs; Based on the complete working condition data, a deep reinforcement learning obstacle removal control system is started, the complete working condition data is input into a state encoder to generate a state vector, and the optimal removal path is planned through a strategy network in combination with successful cases in a historical removal experience library; The fourth unit is used to perform the obstacle removal operation according to the optimal removal path. When it is detected that the material level sensor signal returns to normal and the material discharge rate is stabilized at the target value, it is determined that the obstacle removal is completed, and the complete data record of this obstacle removal is stored in the historical experience library.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Precise discharging and distributing machine for strip-shaped and block-shaped materials suitable for clamping and bridging

    CN111547281A

  • Roller anti-accumulation intelligent feeding device

    CN118387567A

  • Multi-bin intelligent feeding platform

    CN118655806A

  • Intelligent warehousing method and device based on inventory prediction and electronic equipment

    CN118710179A

  • Collecting and subpackaging equipment for zinc powder

    CN119660095A

Cited By

  • Feed production ingredient automatic control system and control method thereof

    CN122499702A