An artificial intelligence-based rice air-drying parameter optimization method and system

CN122734491APending Publication Date: 2026-09-11LIANGHE COUNTY YIKUN GRAIN & OIL IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611088961.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-22
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0005]针对现有技术的不足,本发明提供了一种基于人工智能的稻谷风干参数优化方法及系统,本发明解决的技术问题在于,现有稻谷风干控制依赖滞后的宏观含水率数据,微观应力与含水率存在时空检测错位;且常规控制策略难以平衡脱水效率、能耗与应力损伤,缺乏安全校验机制易导致稻谷爆腰;此外,在线检测的光学窗口在高温粉尘环境下易受污染失效

Benefits of technology

1、本发明通过计算物理时延数值,将偏振透射图像提取的微观应力表征指数与电容水分传感器获取的未补偿平均含水率参数进行时延配准,消除了烘干塔内因物料下落流动引起的传感器空间位置差异所带来的时间错位。该方式确保了输入强化学习模型的状态向量在时间维度上对应同一批次稻谷,提高了多参数闭环控制的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122734491A_ABST
    Figure CN122734491A_ABST
Patent Text Reader

Abstract

The present application relates to the field of agricultural product processing and intelligent control technology, and discloses a rice air-drying parameter optimization method and system based on artificial intelligence, which comprises the following steps: extracting the micro stress characterization index and the change rate data of the polarized transmission image of rice; obtaining the temperature and humidity and moisture content parameters of the drying tower, calculating the physical time delay and generating the state vector; inputting the reinforcement learning model to output the air volume and heat source adjustment parameters, and issuing the action instruction after safety check; and calculating the multi-objective reward function value based on the feedback state and dynamic stress tolerance to update the network node weight. The system comprises a drying execution system, a polarized optical bypass configured with a dehumidification purge gas path, and an edge computing node. The present application eliminates the time dislocation of multi-source detection data, realizes the dynamic balance of dehydration efficiency, energy consumption and stress damage, and guarantees the equipment safety and image acquisition stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of agricultural product processing and intelligent control technology, specifically to a method and system for optimizing rice drying parameters based on artificial intelligence. Background Technology

[0002] In the rice drying process, the control system needs to dynamically adjust the airflow and heat source temperature of the drying tower according to the dehydration status of the rice. Existing control methods mainly rely on moisture sensors to obtain macroscopic moisture content data of the material. However, during the drying process, microscopic moisture gradients and thermal stresses are generated inside the rice. When the stress exceeds its physical tolerance limit, it can cause the rice to burst and break, thus reducing the quality of the processed grain. Macroscopic average moisture content data cannot directly reflect these internal microscopic stress changes. Even if visual sensors are introduced into existing systems for appearance inspection of the rice, the large volume of the drying tower and the continuous downward flow of the material within the tower result in data from optical sensors and moisture sensors installed at different heights belonging to different time points. Existing methods typically do not compensate for the physical time delay caused by the spatial flow of the material when processing such multi-source sensor data, leading to misalignment of the data input to the control system in the time dimension and reducing the accuracy of subsequent parameter optimization.

[0003] Furthermore, rice air-drying is a multivariate coupled nonlinear thermodynamic process. Traditional control strategies struggle to achieve a dynamic balance among multiple objectives, including dehydration efficiency, energy consumption, and stress damage control. Even when some systems attempt to introduce artificial intelligence algorithms to generate control strategies, they often lack hard verification mechanisms based on equipment safety boundaries and material physical properties. When the model explores optimization in the parameter space, it may output extreme temperature or airflow adjustment commands. Directly executing these commands can easily lead to mass scorching of the grain or runaway overload of the heat source equipment.

[0004] Meanwhile, when using optical equipment for online inspection, the interior of the drying tower is a high-temperature, high-dust environment, making the optical inspection windows easily obstructed by dust, affecting image transmittance and the effectiveness of feature extraction. Adding an external high-pressure cleaning air source to the system would increase overall system complexity and equipment cost; conversely, directly channeling the drying tower's own airflow to purge the windows would result in rapid condensation of water vapor upon contact with the relatively cool optical windows, due to the high temperature and humidity of this airflow, which would also disrupt visual inspection conditions. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides an artificial intelligence-based method and system for optimizing rice drying parameters. The technical problem solved by this invention is that existing rice drying control relies on lagging macroscopic moisture content data, and there is a spatiotemporal misalignment between microscopic stress and moisture content detection; moreover, conventional control strategies are difficult to balance dehydration efficiency, energy consumption, and stress damage, and the lack of a safety verification mechanism can easily lead to rice cracking; in addition, the optical window for online detection is easily contaminated and malfunctions in high-temperature and dusty environments.

[0006] To address the above problems, the present invention provides the following technical solution: The first aspect of this invention provides a method for optimizing rice air-drying parameters based on artificial intelligence, comprising the following steps: Polarized transmission images of rice grains were acquired, features were extracted using a convolutional neural network, and after smoothing and filtering, the micro-stress characterization index and stress change rate data were output. The system acquires the inlet temperature, inlet humidity, outlet temperature, outlet humidity, uncompensated average moisture content, and discharge speed data of the drying tower. It calculates the physical time delay value and retrieves the corresponding period data from the first-in-first-out queue containing historical micro-stress characterization index and historical stress change rate data based on the physical time delay value. It then aligns the current uncompensated average moisture content parameter with the retrieved historical micro-stress characterization index and historical stress change rate data in the same batch to generate a state vector. The state vector is input into a Markov decision process model constructed based on reinforcement learning Actor network and Critic network, and the output air volume candidate adjustment parameters and heat source candidate adjustment parameters are evaluated. After verification by built-in safety rules, the final thermodynamic action command is issued. The multi-objective reward function value is calculated based on the physical feedback state after the execution of the final thermodynamic action command and the dynamic stress tolerance calculated based on the current moisture content, so as to update the node weights in the reinforcement learning Actor network and Critic network.

[0007] This invention extracts the micro-stress characterization index from polarized transmission images and combines it with physical time delay calculations to achieve time delay registration with the uncompensated average moisture content parameter, thus solving the time misalignment problem caused by the spatial distribution of parameters. By inputting the registered state vector into a Markov decision process model to generate thermodynamic action commands, and using physical feedback states and dynamic stress tolerances to update the network node weights, dynamic optimization and control of airflow and heat source parameters under multi-parameter conditions are achieved.

[0008] Furthermore, before acquiring the polarization transmission image of the rice, the process also includes: In the state of polarization optical bypass emptying, an empty window transmission image is captured, the gray level probability distribution function of the empty window transmission image is extracted to calculate the global information entropy, and at least one parameter among the average gray level value of the central region of the empty window transmission image, gray level variance, and similarity with a clean reference image is combined to calculate the optical window cleanliness evaluation value, and the optical window cleanliness evaluation value is compared with a preset cleanliness threshold. When the cleanliness evaluation value of the optical window is not lower than the preset cleanliness threshold, the average gray value of the central region of the transmitted image through the empty window is calculated, and the digital PID control algorithm is run to adjust the driving current of the backlight array in order to maintain the preset reference target light intensity. When the cleanliness evaluation value of the optical window is lower than the preset cleanliness threshold, the purge valve is opened, and the airflow from the drying tower's main air supply pipe is sequentially intercepted, separated, purified, and stabilized by the dust removal filter, condensation and dehumidification component, and pressure reducing valve. Then, the airflow is purged through the purge nozzle in the form of a targeted jet to perform a surface purging action on the optical detection window until the cleanliness evaluation value of the optical window for multiple consecutive detection cycles recovers to not lower than the preset cleanliness threshold.

[0009] This invention utilizes global information entropy to quantitatively evaluate the cleanliness of the optical detection window. When the cleanliness meets the standard, a PID control algorithm is used to maintain the reference light intensity of the backlight array. When the cleanliness is low, the airflow from the drying tower is purified and stabilized before performing surface purging on the optical detection window, thereby achieving automatic cleaning and maintaining the light transmittance of the optical detection window.

[0010] Furthermore, the step of extracting features through a convolutional neural network and then smoothing and filtering them to output micro-stress characterization index and stress change rate data specifically includes: The polarization transmission image is a near-infrared polarization transmission image, which includes a first polarization direction transmission image and a second polarization direction transmission image that are orthogonal to each other. After normalizing the two, they are concatenated into an input data tensor, which is then passed to a convolutional neural network based on a residual network architecture. The local contrast features of the image are extracted by a convolutional neural network, and the orthogonal polarization variance parameter based on the gray level difference between the first polarization direction transmission image and the second polarization direction transmission image is used to output the original proxy feature data characterizing the average micro-stress damage degree. The original proxy feature data is used as the observation input value and substituted into the discrete Kalman filter algorithm composed of the time update equation and the observation update equation for iterative calculation. High-frequency detection noise is filtered out, and smoothed micro-stress characterization index and stress change rate data are separated and extracted.

[0011] This invention utilizes the grayscale difference features of orthogonal polarization transmission images to extract original proxy feature data, reflecting the micro-stress damage state of rice, and uses a discrete Kalman filter algorithm to filter out high-frequency noise data, thereby improving the accuracy of micro-stress characterization index and stress change rate data.

[0012] Furthermore, the calculation of the physical time delay value involves time delay registration of the uncompensated average moisture content parameter with the micro-stress characterization index and stress change rate data to generate a state vector, specifically including: Based on the vertical distance difference between the inlet of the collected polarized transmission image and the moisture sensor inside the drying tower, the material flow coefficient, and the discharge speed data, the physical time delay value corresponding to the descent of the rice is calculated. Divide the physical delay value by the system sampling control period to obtain the discrete queue index steps; Using the queue index step number as the backtracking offset, backtracking is performed from the first-in-first-out queue storing historical micro-stress characterization index and historical stress change rate data in the historical direction. The corresponding historical micro-stress characterization index and historical stress change rate data are extracted and aligned with the uncompensated average moisture content parameter in the time dimension. The state vector is generated by concatenating the inlet temperature parameter, inlet humidity parameter, exhaust temperature parameter, and exhaust humidity parameter.

[0013] This invention calculates the physical time delay value based on the vertical distance difference and material flow characteristics within the drying tower, converts the physical time delay into queue index steps, and extracts historical data by backtracking the first-in-first-out queue, thereby achieving temporal alignment between micro-stress data and moisture content data caused by different spatial locations.

[0014] Furthermore, the issuance of the final thermodynamic action command after verification by the built-in safety rules specifically includes: It is determined whether the micro-stress characterization index after alignment of the same batch is not less than the preset stress limit threshold, and whether the air inlet temperature parameter is not less than the preset upper limit threshold of air inlet temperature; when the air inlet temperature parameter is not less than the preset upper limit threshold of air inlet temperature, it is further determined whether the opening control coefficient corresponding to the candidate heat source adjustment parameter is greater than the preset safe heat source opening coefficient, or whether the air volume control coefficient corresponding to the candidate air volume adjustment parameter is less than the preset safe air volume control coefficient. When any of the above judgment conditions that violate the physical safety boundary are met, the safety action overwrite mechanism is triggered, the candidate action vector containing the candidate air volume adjustment parameters and the candidate heat source adjustment parameters is blocked, and the built-in safety action vector that outputs the maximum rated air volume of the system and maintains the minimum non-extinguishing combustion state of the heat source is forcibly generated as the final thermodynamic action command to be issued and executed.

[0015] This invention compares the micro-stress characterization index, inlet temperature parameters with corresponding thresholds, and verifies the control opening coefficient to trigger a safety action overwrite mechanism when parameters violate safety boundary conditions. By outputting the maximum rated airflow and maintaining a minimum non-extinguishing combustion state, it avoids equipment failure or excessive heat damage to materials caused by abnormal thermodynamic action commands.

[0016] Furthermore, the step of calculating the multi-objective reward function value based on the physical feedback state after the execution of the final thermodynamic action command and the dynamic stress tolerance calculated based on the current water content, in order to update the node weights in the reinforcement learning Actor network and Critic network, specifically includes: Based on the difference between the registered moisture content data fed back after the action is executed and the target outlet moisture content, combined with the set conservative safety stress limit and stress-moisture coupling slope constant, the dynamic stress tolerance at the current moisture content is calculated. Based on the comparison between the registered micro-stress characterization index of the next cycle and the dynamic stress tolerance, a stress penalty term is generated; The dehydration reward item before and after feedback, the energy consumption penalty item based on system operating frequency and valve opening, the stress penalty item, the stress change rate penalty item based on the stress change rate data after registration in the next cycle, and the parameter overwrite trigger status flag are linearly weighted to calculate the value of the multi-objective reward function. The temporal difference error is calculated using an empirical data stream containing the values ​​of the multi-objective reward function to update the parameters of the Critic network, and the weight parameters of the reinforcement learning Actor network are updated using the action value evaluation gradient provided by the updated Critic network.

[0017] This invention determines the dynamic stress tolerance based on the moisture content difference and the stress-moisture coupling slope constant. It constructs a multi-objective reward function value by linearly weighting parameters such as dehydration reward, energy consumption penalty, and stress penalty, and updates the parameters of the Actor network and Critic network accordingly, so that the action instructions output by the Markov decision process model can adapt to the rice dehydration requirements and stress constraints at different stages.

[0018] A second aspect of the present invention provides an artificial intelligence-based rice air-drying parameter optimization system for implementing the above-mentioned method, comprising: The drying execution system is based on the main body of the drying tower. The main body of the drying tower is equipped with an inlet air temperature and humidity sensor, an exhaust air temperature and humidity sensor, a capacitive moisture sensor, a bottom discharge roller speed encoder, a variable frequency fan that supplies drying airflow, and a proportional regulating valve that adjusts the heat source supply. A polarization optical bypass is arranged in parallel outside the main body of the drying tower, and an optical detection chamber is set inside. A backlight array and a CMOS sensor equipped with a micro-polarization array polarization filter are respectively installed on both sides of the optical detection chamber. Edge computing nodes establish communication and control connections with the drying execution system and the polarization optical bypass, respectively, for the following purposes: The convolutional neural network is called to process the polarization transmission image of rice obtained by the CMOS sensor to extract features, and after smoothing and filtering, the micro-stress characterization index and stress change rate data are output. Acquire inlet temperature parameters, inlet humidity parameters, outlet temperature parameters, outlet humidity parameters, uncompensated average moisture content parameters, and discharge speed data. Perform physical time delay registration between the uncompensated average moisture content parameters and the micro-stress characterization index and stress change rate data to generate a state vector. The Markov decision process model, constructed based on reinforcement learning Actor and Critic networks, outputs candidate air volume adjustment parameters and candidate heat source adjustment parameters. After safety verification, action commands are issued, and the node weights in the reinforcement learning Actor and Critic networks are updated based on the multi-objective reward function value calculated from the feedback state after the action is executed.

[0019] Furthermore, the top of the polarization optical bypass is provided with a flow-guiding valve driven by a stepper motor, and the interior is provided with a row of flow channels from top to bottom. The micro-polarization array polarization filter includes at least two mutually orthogonal polarization directions; A purge nozzle is installed at an angle on the outside of the optical detection window of the optical detection chamber. The air inlet of the purge nozzle is connected in sequence to a purge valve, a pressure reducing valve, a condensation dehumidification component and a dust removal filter through a pneumatic pipeline, and finally connected to the air supply main pipe corresponding to the air outlet of the variable frequency fan to form a physical air passage.

[0020] This invention incorporates a polarization optical bypass on the exterior of the drying tower, allowing rice sampling and testing to proceed independently of the main flow path, thus avoiding interference with the flow field within the main tower. Airflow is introduced into the main air supply pipe via a physical air passage for purging, improving the system's structural integration.

[0021] Furthermore, the edge computing node has a built-in graphics processor; The inlet air temperature and humidity sensor, the exhaust air temperature and humidity sensor, the capacitive moisture sensor, and the bottom discharge roller speed encoder in the drying execution system are all connected to the programmable logic controller. The output of the programmable logic controller is connected to the centrifugal fan inverter that drives the variable frequency fan, and the proportional regulating valve, respectively. The edge computing node establishes action execution connections with the stepper motor, the purge valve, and the power drive module of the backlight array through a programmable logic controller.

[0022] This invention uses edge computing nodes in conjunction with programmable logic controllers to handle the computational needs of neural networks and the underlying logic control of electromechanical equipment, thereby improving the synchronization of data flow and equipment response.

[0023] Furthermore, the condensation dehumidification component is internally equipped with a semiconductor cooling chip, a condensate collection chamber, and a gas-liquid separation structure. This is used to forcibly cool the small-flow bypass airflow introduced from the main air supply pipe to below the dew point temperature to condense water. After passing through a reheating section or heat exchange section, the output airflow after liquid drainage is restored to a low-humidity state higher than its current dew point temperature before being input into the pressure reducing valve for dynamic pressure regulation. The flow rate of the small-flow bypass airflow is less than a preset proportion of the main airflow flow rate of the main air supply pipe, and is limited within a preset purging flow range by a throttling orifice, a flow limiting valve, or a pressure reducing valve, so that the condensation dehumidification component can maintain a stable cooling and dehumidification capacity under continuous purging conditions.

[0024] This invention uses a condensation and dehumidification component to cool, dehumidify, and raise the temperature of the introduced high-temperature airflow, thereby reducing the relative humidity of the output airflow and preventing condensation on the surface of the optical detection window, thus ensuring the optical stability of the detection environment.

[0025] This invention provides a method and system for optimizing rice air-drying parameters based on artificial intelligence. It has the following beneficial effects: 1. This invention calculates the physical time delay value and performs time delay registration between the micro-stress characterization index extracted from the polarization transmission image and the uncompensated average moisture content parameter obtained from the capacitive moisture sensor, eliminating the time misalignment caused by the spatial position difference of the sensor due to the material falling and flowing within the drying tower. This method ensures that the state vector input to the reinforcement learning model corresponds to the same batch of rice in the time dimension, improving the accuracy of multi-parameter closed-loop control.

[0026] 2. This invention utilizes a reinforcement learning Actor and Critic network architecture, combined with a multi-objective reward function based on dynamic stress tolerance calculation, to optimize airflow and heat source adjustment parameters, achieving a multi-objective dynamic balance between dehydration efficiency, energy consumption, and stress damage. Simultaneously, the built-in safety rule verification mechanism can directly overwrite and issue an action vector that outputs the system's maximum rated airflow and maintains the heat source in a minimum non-extinguishing combustion state when stress or temperature violates physical boundaries, eliminating the risk of equipment overheating and mass grain bursting during the algorithm exploration process.

[0027] 3. This invention introduces an optical window cleanliness assessment mechanism based on at least one parameter among global information entropy, average gray level, gray level variance, and reference image similarity in polarization optics bypass. When the window is contaminated with dust, the high-pressure airflow from the original main air supply pipe of the drying tower is diverted, cooled, dehumidified, purified, and stabilized before surface purging of the detection window. This structure reuses the drying tower's own air path to complete window cleaning and prevents surface condensation by removing moisture from the purging airflow, ensuring continuous and stable acquisition of polarization transmission images under high-temperature and dusty conditions. Attached Figure Description

[0028] Figure 1 This is a system architecture diagram of the rice air-drying parameter optimization system of the present invention; Figure 2 This is a flowchart of the rice air-drying parameter optimization method of the present invention; Figure 3 This is a schematic diagram of the physical delay compensation mechanism of the present invention; Figure 4 This is a schematic diagram of the closed-loop control principle for optical reference calibration of the present invention; Figure 5 This is a schematic diagram of the physical closed-loop principle of the targeted self-cleaning method based on the information entropy of the empty window in this invention. Figure 6 This is a schematic diagram illustrating the principle of micro-stress proxy feature extraction and smoothing processing in this invention. Figure 7 This is a schematic diagram of the restricted MDP state assessment and safety action overwriting principle of the present invention; Figure 8 This is a schematic diagram illustrating the dynamic tolerance coupling and multi-objective reward function iteration principle of the present invention. Figure 9 The following is a time-series diagram of the micro-stress inside the rice grain and the action overwrite response of the present invention, wherein (A) is the air inlet temperature curve and the safety threshold line, (B) is the original feature data scatter and the micro-stress continuous curve after Kalman filtering, and (C) is the original output signal of the Actor network and the overwrite action signal after the safety rule is triggered. Figure 10 The figures show the comparison curves of drying performance and micro-damage evolution under different control strategies of the present invention, where (A) is the dynamic stress evolution tracking curve and (B) is the biaxial bar chart of comprehensive energy efficiency and quality. Detailed Implementation

[0029] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0030] See attached document Figure 1 The present invention provides a rice drying parameter optimization system, which may include: a drying execution system, a polarization optical bypass, and an edge computing node.

[0031] The drying system is built upon the main body of the drying tower. Inlet and outlet air temperature and humidity sensors are installed at the air inlet and outlet of the drying tower, respectively. A capacitive moisture sensor is located inside the drying tower. A speed encoder is connected to the discharge roller at the bottom of the drying tower. The inlet and outlet air temperature and humidity sensors, the capacitive moisture sensor, and the speed encoder are all connected to a programmable logic controller (PLC).

[0032] The outputs of the programmable logic controller (PLC) are connected to the variable frequency fan and the proportional control valve, respectively. The variable frequency fan operates at a frequency regulated by a centrifugal fan inverter, and the proportional control valve is a heat source proportional valve located on the heat source supply pipeline. Control commands output from the edge computing node are converted by the PLC and then applied to the centrifugal fan inverter and the heat source proportional valve, respectively. The outlet of the variable frequency fan is connected to the main air supply duct to provide drying airflow to the main drying tower; the proportional control valve is used to regulate the amount of heat source supplied into the main air supply duct.

[0033] The polarization optical bypass is arranged in parallel on the outside of the drying tower body. A flow-guiding valve is installed at the top of the polarization optical bypass, driven by a stepper motor. Inside the polarization optical bypass, from top to bottom, are a series of flow channels and an optical detection chamber. On both sides of the optical detection chamber are a backlight array and a CMOS sensor equipped with a micro-polarization array polarization filter. The micro-polarization array polarization filter includes at least two mutually orthogonal polarization directions, enabling the CMOS sensor to acquire transmission image data corresponding to different polarization directions within the same sampling period.

[0034] As one specific implementation, the micro-polarization array polarization filter includes 0° polarization units and 90° polarization units, or includes 0°, 45°, 90° and 135° polarization units.

[0035] The optical inspection window of the optical inspection chamber is inclined to one side facing the material flow channel and is equipped with a purge nozzle. The air inlet of the purge nozzle is connected in sequence to the purge valve, pressure reducing valve, condensation dehumidification component and dust removal filter through a pneumatic pipeline, and finally connected to the air supply main pipe corresponding to the air outlet of the variable frequency fan to form a physical air path channel. The purge nozzle is used to output low-humidity clean airflow to the material side surface of the optical inspection window, so that the airflow forms a targeted shear jet along the window surface.

[0036] Edge computing nodes have built-in graphics processors. They establish data communication connections with both programmable logic controllers (PLCs) and CMOS sensors. Through the PLC or an industrial input / output module connected to it, the edge computing node establishes control connections with stepper motors, purge valves, and the power drive module of the backlight array. These connections are used to send sampling frequency control commands to the stepper motors, opening and closing control commands to the purge valves, and drive current adjustment commands to the power drive module of the backlight array.

[0037] See attached document Figure 2 This invention provides a method for optimizing rice air-drying parameters, comprising the following steps: S11, the edge computing node synchronously acquires the inlet temperature parameter, inlet humidity parameter, exhaust temperature parameter, exhaust humidity parameter, and uncompensated average moisture content parameter through the programmable logic controller. It obtains the discharge speed data output by the speed encoder and substitutes it into the system dynamic transfer function to calculate the physical delay value. It calls the first-in-first-out queue to perform time delay registration on the historical micro-feature data acquired by the polarization optical bypass, and extracts the matching micro-feature data of the same physical batch as the current moisture content data. Among them, the historical micro-feature data is written into the first-in-first-out queue after the edge computing node performs near-infrared polarization transmission image acquisition, convolutional neural network feature extraction, and discrete Kalman filtering in the previous control cycle, for physical delay backtracking in subsequent control cycles. S12, the edge computing node sends control pulses to the stepper motor to adjust the sampling frequency of bypass materials. When the calibration cycle is reached, it sends a stop command to the stepper motor to empty the entire flow channel. The CMOS sensor captures the transmission image of the empty window. The edge computing node first calculates the global information entropy of the transmission image of the empty window and judges the cleanliness status of the optical detection window. When the global information entropy is not lower than the cleanliness threshold, the edge computing node then calculates the average gray value of the central area of ​​the transmission image of the empty window and runs a digital PID control algorithm to adjust the driving current of the backlight array so that the optical background reference light intensity of the optical detection chamber remains constant. S13, edge computing nodes process the windowed transmission image in parallel to calculate the global information entropy. When the global information entropy satisfies... Under certain conditions, the edge computing node determines that there are contaminants attached to the optical detection window and sends a command to the purge valve first. The airflow output from the main air supply pipe is purified and stabilized after passing through the dust filter, condensation and dehumidification components and pressure reducing valve. Then, it performs a surface purging action on the optical detection window of the optical detection chamber through the purge nozzle. When the global information entropy recovers to no less than the cleanliness threshold after purging, the edge computing node then performs the backlight array light intensity PID calibration. S14, the stepper motor receives the instruction to resume material sampling operation, the CMOS sensor acquires the near-infrared polarization transmission image of the single-layer material, the edge computing node inputs the near-infrared polarization transmission image into the preset convolutional neural network to extract the local contrast and orthogonal polarization variance parameters of the image and outputs the original proxy feature data, and calls the discrete Kalman filter algorithm to smooth the original proxy feature data and output the micro-stress characterization index and stress change rate data. S15, the edge computing node combines temperature and humidity parameters, moisture content data registered in the same batch, micro-stress characterization index matched in the same batch, and stress change rate data matched in the same batch into a state vector. The Markov decision process model is run, and the state vector is used as input conditions for forward network reasoning to output candidate air volume adjustment parameters and candidate heat source adjustment parameters. The built-in safety rule library is called to compare and evaluate the candidate parameters. When the micro-stress characterization index matched in the same batch is not less than the stress limit threshold, or when the air inlet temperature parameter is not less than the upper limit threshold of the air inlet temperature, and the heat source candidate opening control coefficient is greater than the preset safe heat source opening coefficient or the air volume candidate control coefficient is less than the preset safe air volume control coefficient, the candidate parameter is determined to violate the safety threshold rule. The candidate adjustment parameter is overwritten as a cooling and air-increasing value and converted into the final thermodynamic action command and sent to the programmable logic controller for execution. S16, the physical feedback state after the edge computing node reads the action is completed, calculates the dynamic stress tolerance benchmark value based on the registered moisture content data of the same batch, integrates the dehydration efficiency index, energy consumption penalty index, stress deviation value, stress change rate penalty term and parameter overwrite trigger penalty term to generate a multi-objective reward function value, and uses the reward function value to update the node weights in the reinforcement learning Actor network and Critic network.

[0038] To enable those skilled in the art to better understand the technical solution of the present invention, the specific technical details of each step above will be explained in detail below in conjunction with relevant formulas and principles.

[0039] See attached document Figure 3 The specific process for physical delay compensation is explained below: S111, the edge computing node interacts with the programmable logic controller (PLC) via a preset communication cycle. The edge computing node acquires the inlet temperature parameters collected by the inlet temperature and humidity sensor. and air inlet humidity parameters Obtain the exhaust outlet temperature parameters collected by the exhaust temperature and humidity sensor. Humidity parameters of exhaust vent Edge computing nodes synchronously read the large-volume uncompensated average moisture content parameter output from the capacitive moisture sensor. Regarding the analog-to-digital conversion of sensor signals and the reading of underlying registers, those skilled in the art can configure the hardware according to conventional industrial control specifications. The acquisition of underlying signals is a well-known technology in this field and will not be elaborated here.

[0040] S112, during the grain drying process, the material flows downwards within the drying tower body due to gravity. Typically, the inlet of the polarization optical bypass is physically higher than the internal capacitive moisture sensor. This results in the microscopic features collected by the polarization optical bypass and the average moisture content collected by the capacitive moisture sensor belonging to different batches of falling material at the same time. To reduce data registration misalignment during subsequent state variable fusion, the edge computing node acquires the discharge speed data output by the speed encoder. Edge computing nodes are based on discharge speed data. By combining the height parameters of the internal mechanical structure of the drying tower body with the system dynamics transfer function, the physical time delay value is calculated. .

[0041] When the elevation is based on the bottom of the drying tower body, and the top feed inlet of the polarization optical bypass is higher than the capacitive moisture sensor, the edge computing node will calculate the vertical distance difference between the two. Defined as: In the formula, This represents the installation height elevation of the top feed inlet of the polarization optics bypass. This represents the installation height elevation of the capacitive moisture sensor within the main body of the drying tower.

[0042] When the discharge speed is within a stable operating range, the physical time delay value The calculation formula is: In the formula, The material flow coefficient is used to characterize the proportional mapping relationship between the rotational speed of the discharge roller and the vertical descent linear velocity of the material inside. This represents the discharge speed data output by the speed encoder; This represents the equivalent vertical descent speed of the material, calculated from the rotational speed of the discharge roller.

[0043] When the discharge speed data satisfies the following formula: ; If the edge computing node determines that the current material flow is stagnant or in a low-speed and unstable state, it will suspend the execution of physical delay calculation based on the current rotation speed and maintain the previous effective delay registration result, or execute a preset safety action vector to avoid queue index abnormalities due to an excessively small divisor.

[0044] As a specific implementation method, the material flow coefficient This can be determined through offline calibration experiments. The specific calibration process is as follows: Under full load conditions on the drying tower body, labeled material is introduced and the actual time it takes for it to fall from the height of the inlet to the height of the sensor is recorded. This is then combined with the corresponding discharge speed to obtain the result. In this embodiment, based on the cross-sectional area of ​​the drying tower body and the volume parameters of the discharge roller, the material flow coefficient... The value is set between 0.5 and 1.5.

[0045] S113, the edge computing node establishes a first-in-first-out queue in memory. During the system control cycle, the edge computing node will process the raw proxy feature data obtained by polarization optical bypass acquisition. Micro-stress characterization index Stress change rate data The corresponding system timestamp is stored in the first-in-first-out queue.

[0046] Because the system data acquisition is based on a fixed system sampling control cycle The edge computing nodes will calculate the continuous physical delay values. The conversion to discrete queue index steps n follows the following rules: Where `round` is the rounding function. When there are invalid flag data packets in the FIFO queue due to window calibration, purging, or sampling interruption, the edge computing node prioritizes retrieving and matching the data packet timestamp. The closest historical valid micro-feature data; if no valid micro-feature data exists within the corresponding time range, the period is marked as a failed registration period in the same batch and is not used for reinforcement learning experience updates.

[0047] Under conditions where the discharge speed fluctuates significantly, the edge computing node can also use the historical speed integration method to determine the discrete queue index step number n, that is, to accumulate the material descent distance from the current moment to the historical direction until the accumulated descent distance is not less than the vertical distance difference between the bypass inlet and the capacitive moisture sensor. The determination relationship is as follows: ,in, This represents the equivalent vertical descent speed of the material corresponding to the qth historical control cycle.

[0048] The smallest integer n that satisfies the above conditions is used as the queue backtracking index step number for the current period. Since the drainage port of the polarization optics bypass is located higher than the capacitive moisture sensor, the uncompensated average moisture content parameter currently collected by the capacitive moisture sensor... This corresponds to the same batch of materials that previously underwent polarization optical bypass. The edge computing node uses the queue index step number *n* as the backtracking offset to backtrack the corresponding number of periodic data packets from the FIFO queue storing historical polarization optical characteristic data, extracting the historical micro-stress characterization index from these periodic data packets. and historical stress change rate data and the current uncompensated average moisture content parameter As a batch of registered moisture content data belonging to the same physical batch as historical polarization optical characteristic data .

[0049] The aforementioned data backtracking operation based on physical latency enables edge computing nodes to acquire the current batch of registered moisture content data. Compared with historical polarization optical characteristic data , Maintaining consistency across physical batches. This mechanism reduces data registration bias caused by lags in material spatial flow, providing time-aligned foundational data for subsequent state vector generation.

[0050] See attached document Figure 4 The specific process of closed-loop control for optical reference calibration is explained below: S121, the edge computing node sends pulse control signals to the stepper motor according to the set operating rhythm, driving the diversion valve to control the material flow rate entering the entire flow channel. The stepper motor sampling frequency is defined as follows: The calibration period constant is set inside the edge computing node. As an example, calibrating the period constant. The value is determined based on the thermodynamic equilibrium time of the backlight array; for example, it can be set to perform calibration once every 3000 control cycles. The edge computing node also sets the maximum storage length of the first-in-first-out queue, ensuring that this maximum storage length is greater than the number of discrete control cycles corresponding to the maximum physical delay value. When the system starts up and the queue data is insufficient, the edge computing node temporarily suspends the execution of closed-loop control based on delay registration, or uses a preset safety action vector as a transitional control action.

[0051] When the system cycle count reaches the calibrated cycle constant When the sampling frequency is an integer multiple of the set value, the system enters calibration mode. The edge computing node sends a stop command to the stepper motor, adjusting the stepper motor sampling frequency. Set to zero to cut off the feed to the drain valve.

[0052] During the system's empty window calibration mode or purge cleaning mode, the edge computing node marks the corresponding control cycle as an invalid material optical sampling cycle and does not write the data corresponding to the empty window transmission image as valid material microscopic feature data into the first-in-first-out queue. The edge computing node retains the timestamp and invalid flag of the corresponding control cycle in the time delay registration buffer. When subsequent moisture content data is traced back to this invalid sampling cycle, it is determined that the physical batch lacks valid microscopic feature data, and the state fusion and experience data writing of this cycle are suspended. The closed-loop control module maintains the previous valid execution action vector, or executes a preset safety action vector when the amount of valid data in the queue is insufficient.

[0053] Under the influence of gravity, the remaining material in the entire flow channel is naturally discharged downwards. The edge computing node maintains the shutdown command state for a first preset time to empty the entire flow channel and create an unobstructed window period for optical detection. The value of this first preset time is greater than the time required for the material to freely fall through the entire flow channel and the optical detection chamber. It can be set according to the vertical height from the diversion valve to the discharge port, the free fall velocity of the material, and the material viscosity margin, and the first preset time is made greater than the longest emptying time required for the material to pass through the entire flow channel and the optical detection chamber.

[0054] In actual engineering configurations, the first preset time can be supplemented with a redundancy of 0.5 to 1.0 seconds based on this calculation result to cope with the falling delay caused by material stickiness.

[0055] In step S122, during the window period after the entire array of channels is emptied, the edge computing node sends a trigger signal to the CMOS sensor. The CMOS sensor captures the transmitted image through the window in the optical detection chamber. During the window calibration process, the exposure time, analog gain, and digital gain of the CMOS sensor are kept at preset fixed values ​​so that the average gray value of the central region of the transmitted image can characterize the actual transmitted light intensity of the backlight array. For the low-level hardware operations of the CMOS sensor, such as exposure time configuration and image format output, those skilled in the art can perform conventional configurations based on the sensor datasheet; these are well-known techniques in the field and will not be elaborated upon here.

[0056] The edge computing node extracts a preset pixel matrix region at the center of the transmissive image from the empty window as the target region of interest. Since the transmitted light at the edge of the optical inspection chamber is easily affected by diffuse reflection from the mechanical inner walls and lens edge distortion, selecting the central region more accurately reflects the true direct light intensity of the backlight array. The size of this preset pixel matrix region can be set according to the image resolution, for example, selecting a 100×100 pixel rectangular array at the center of the image. The edge computing node accumulates the grayscale values ​​of all pixels within this target region of interest and divides the sum by the total number of pixels in the target region of interest to calculate the average grayscale value of the central region of the transmissive image from the empty window. .

[0057] S123, the backlight array exhibits thermal decay under long-term operating conditions, which can cause a drift in the reference luminous intensity of the light source. A reference target light intensity value is pre-set within the edge computing node. The reference target light intensity value The value is set based on the factory-calibrated grayscale value when the optical testing room is free of pollution and the backlight array is at its rated power.

[0058] Edge computing nodes will calculate the global entropy of the current window transmission image. Compare with the cleanliness threshold Eth; When satisfied When the optical detection window is determined to be clean, the edge computing node will calculate the reference target light intensity value. Compared with the measured average gray value Perform a difference operation to calculate the light intensity error of the current calibration period. ; When satisfied If contaminants are found on the optical detection window, the edge computing node will suspend the current light intensity PID calibration and prioritize the purging and cleaning procedure. Only when the global information entropy is not lower than the cleanliness threshold Under these conditions, edge computing nodes will only reduce light intensity errors. The input is a preset discrete digital PID control algorithm program. Based on the set control coefficients, the algorithm program calculates the update drive current of the output backlight array. The calculation formula for this algorithm is: In the formula, This is the rated base drive current of the backlight array; This is the proportional control coefficient; These are integral control coefficients; These are the differential control coefficients; This is the cumulative sum of light intensity errors from system startup to the current calibration period; This represents the light intensity error from the previous calibration period; This is the rate of change of light intensity error. Those skilled in the art can determine the initial values ​​of the above PID control coefficients using conventional Ziegler-Nichols tuning methods or the critical proportional gain method.

[0059] To protect the hardware security of the backlight array and prevent the control algorithm from outputting extreme values ​​due to excessive errors, the edge computing nodes update the calculated drive current. Perform threshold limiting operation. Set the maximum safe current limit value. and minimum sustaining current limit .when When, overwrite it as ;when When, overwrite it as Edge computing nodes will update the drive current after it has been clipped. The signal is converted into an analog control electrical signal and sent to the power drive module of the backlight array. This control closed loop dynamically compensates the drive current in each calibration cycle, reducing detection interference caused by the physical attenuation of the light source brightness in subsequent transmission image feature extraction.

[0060] See attached document Figure 5 The specific process of the targeted self-cleaning physical closed loop based on the information entropy of the empty window is explained as follows: S131, during the acquisition of the aforementioned windowed transmission image, the edge computing node simultaneously initiates the contamination assessment process. The edge computing node converts the windowed transmission image into a single-channel grayscale image, counts the frequency of each grayscale pixel in the single-channel grayscale image, calculates its proportion in the total number of pixels in the entire image, and generates a grayscale probability distribution function. .

[0061] Edge computing nodes calculate the global information entropy of the current empty window polarization image using the following formula. In the formula, g represents the gray level of a single-channel grayscale image, and its value ranges from 0 to 255. This represents the probability of a pixel with gray level g appearing in a windowed transmission image with time period t.

[0062] As the physical basis of this quantitative characterization mechanism, when the optical detection window of the optical detection chamber is clean and unobstructed, the image usually contains a large amount of backlight array direct light distribution and internal mechanical texture features, with a large grayscale distribution range, corresponding to a high global information entropy value.

[0063] Conversely, when dust adheres to the surface of the optical detection window or water condensation occurs, diffuse reflection of light and physical obstruction usually cause the image to exhibit a blurred and degraded characteristic with uniform grayscale, thus reducing the global information entropy value. The system utilizes this physical-optical correspondence to establish a quantitative characterization mechanism for the contamination state. For basic image processing operations such as image conversion to grayscale and pixel traversal statistics, those skilled in the art can use conventional computer vision function libraries to implement them; these are well-known techniques in the field and will not be elaborated upon here.

[0064] S132, the edge computing node reads the preset cleanliness threshold from the system memory. Cleanliness threshold The global information entropy is determined by adjusting the average value of multiple empty window information entropies continuously acquired under contamination-free conditions during the factory calibration phase. The adjustment ratio can be set to, for example, 85% to 90% of the original average value. The edge computing node will then calculate the global information entropy. With cleanliness threshold Perform numerical comparison. When At this time, the edge computing node determines that there are contaminants or deposits in the optical detection window that affect feature extraction. The edge computing node then outputs a drive level signal and sends it to the relay control terminal of the purge valve, driving the purge valve to open the physical air passage.

[0065] S133, In conventional drying systems, the high-temperature airflow transported inside the main air supply duct of the continuous drying tower typically contains suspended dust and has a high moisture content. If this high-temperature airflow is directly introduced to purge optical components, there is a risk of secondary dust adhesion or condensation on the lens surface, which could exacerbate optical contamination or cause thermal damage to the sensor. Therefore, after the purge valve is opened, the airflow from the main air supply duct undergoes multi-stage purification and conditioning treatment sequentially along the physical air path.

[0066] Airflow enters the dust collector filter for physical interception to separate solid dust particles. The airflow then enters the condensation dehumidification unit. This unit uses a semiconductor cooling chip to rapidly cool the airflow, lowering its temperature below the dew point to precipitate gaseous moisture. The condensation dehumidification unit is equipped with a condensate collection chamber and a drain outlet to discharge the precipitated liquid moisture, preventing it from entering the purge nozzles with the airflow.

[0067] After cooling and dehumidifying the airflow, the condensation and dehumidification component discharges the condensate through a gas-liquid separation structure. The output airflow then passes through a reheating section or heat exchange section to rise to a temperature higher than its current dew point, thus creating a low-humidity, clean airflow that is less prone to condensation on the optical detection window surface. Simultaneously, the temperature of this output airflow is set below the upper limit of the CMOS sensor's operating temperature. The temperature- and humidity-controlled airflow then enters a pressure reducing valve for dynamic pressure regulation, outputting a constant-pressure airflow within the preset purge pressure range to prevent excessively high pressure from causing mechanical damage to the optical detection window.

[0068] S134, the constant pressure clean airflow, after the above purification and pressure stabilization treatment, is accelerated by the purging nozzle and impacts the outer surface of the optical inspection window of the optical inspection chamber in the form of a targeted jet, using the airflow shearing force to peel off the attached water mist and dust.

[0069] During the purging process, the edge computing node keeps the stepper motor stopped, and the drainage valve closed or at a low flow rate to ensure the optical inspection chamber is in an unobstructed, empty state. Simultaneously, the edge computing node re-acquires the transmission image through the empty window at set time intervals and calculates the real-time global information entropy. The edge computing node keeps the purge valve open until the real-time global information entropy recovers and remains at the cleanliness threshold. above.

[0070] Specifically, when the following conditions are met for m consecutive detection cycles, the edge computing node determines that the optical detection window has returned to a clean state: In the formula, r is the continuous cleanliness judgment cycle index; m is the preset number of continuous judgments, which is used to avoid the purging action from being mistakenly terminated due to a single image fluctuation.

[0071] To avoid the system getting stuck in a continuous purging loop when dealing with solidified and stubborn deposits, a maximum purging time constant is preset within the edge computing node. When the continuous opening time of the purging valve reaches this maximum purging time constant, and the real-time global information entropy is still below the cleanliness threshold... At this time, the edge computing node forcibly cancels the drive level signal to close the purge valve and generates a manual maintenance alarm log, which is then uploaded to the host computer terminal.

[0072] After the purging and cleaning is completed, the edge computing node controls the stepper motor to resume material sampling. Only after obtaining a valid near-infrared polarization transmission image of the material again will the corresponding microscopic feature data be written into the first-in-first-out queue for time delay registration in subsequent control cycles.

[0073] See attached document Figure 6 The specific process for extracting and smoothing micro-stress proxy features is explained below: S141, after completing the self-cleaning procedure of the optical detection window or after a preset window period, the edge computing node re-sends the set operating frequency pulse signal to the stepper motor. The stepper motor drives the diversion valve to resume material sampling of the drying tower body. The diversion valve continuously extracts dispersed material samples from the drying tower body according to the preset sampling frequency, enabling the polarization optical bypass to obtain bypass sample data within multiple continuous control cycles. The material enters the entire flow channel under the action of gravity.

[0074] The physical space constraints of the flow channel cause the material to form a single-layer arrangement during its descent. Specifically, the channel thickness is set to be greater than the average thickness of a single grain of rice but less than the average thickness of two grains stacked together. The channel width is set to allow single or a few rows of rice to pass through sequentially, thus restricting the rice from forming a single-layer arrangement within the optical detection chamber. This arrangement helps reduce beam transmittance attenuation and optical path distortion interference caused by multi-layer particle stacking. The near-infrared beam emitted by the backlight array penetrates the single-layer material and is polarized filtered by a polarization filter.

[0075] The backlight array can be a near-infrared LED array with a wavelength of 850nm or 940nm, or a near-infrared light source with a wavelength range of 800nm ​​to 1000nm. The CMOS sensor performs synchronous exposure to obtain a transmission image in a first polarization direction and a transmission image in a second polarization direction, respectively. The first polarization direction and the second polarization direction are orthogonal to each other, thereby acquiring a set of near-infrared polarized transmission images containing information on the optical anisotropy inside the grain.

[0076] S142, the edge computing node performs preprocessing operations on the acquired near-infrared polarized transmission image. The edge computing node extracts the effective light-receiving area at the center of the image and scales it to a set pixel size, such as 224×224 pixels. The edge computing node normalizes the pixel grayscale values ​​of the image, mapping them to a floating-point range of 0 to 1. Normalization is performed on the first polarization direction transmission image and the second polarization direction transmission image respectively, and the two are concatenated along the channel dimension to construct an input data tensor with dimensions of 224×224×2. This input data tensor carries the original optical mapping information characterizing the micro-cracks and internal stress distribution within the grain caused by moisture and thermodynamic gradients.

[0077] S143, the edge computing node invokes a pre-built convolutional neural network in system memory. This convolutional neural network is configured as a lightweight feature extraction model based on a residual network architecture. In a specific implementation, the network's internal hierarchical structure includes an input layer, multiple cascaded residual convolutional blocks, and a fully connected output layer. Each residual convolutional block contains a sequentially connected two-dimensional convolutional layer, a batch normalization layer, and a ReLU activation function layer. The number of residual convolutional blocks is 3 to 6; the kernel size of the two-dimensional convolutional layer is 3×3 or 5×5; the number of channels in each residual convolutional block increases in a stepwise manner of 16, 32, 64, or 128; the fully connected output layer includes at least one hidden fully connected layer and a scalar output node, which outputs the original proxy feature data in the range of 0 to 1 through a sigmoid function or a linear normalization function. Edge computing nodes pass the input data tensor to the input layer of the convolutional neural network. The shallow convolutional kernels of the network extract local contrast features of the image, while the deep convolutional kernels extract the orthogonal polarization variance parameter corresponding to the birefringence changes of light caused by defects inside the grain, based on the gray-level difference between the transmitted images in the first polarization direction and the transmitted images in the second polarization direction.

[0078] As a specific calculation method, the orthogonal polarization variance parameter can be expressed as: In the formula, This represents the transmission image in the first polarization direction. This represents the transmission image along the second polarization direction, which is orthogonal to it. This indicates variance calculation. Before calculating the orthogonal polarization variance parameter, both the transmission images in the first polarization direction and the transmission images in the second polarization direction are normalized.

[0079] After the feature data flows through the fully connected output layer, the network's forward inference outputs a one-dimensional scalar result. This scalar result is the original surrogate feature data. Its physical meaning characterizes the average micro-stress damage level within the currently observed batch of materials. (Original proxy feature data) Normalized to a dimensionless range of 0 to 1, with larger values ​​indicating higher levels of micro-stress damage within the current batch of materials being observed.

[0080] S144, the aforementioned convolutional neural network needs to undergo offline supervised training before deployment to establish the mapping relationship between optical features and actual stress damage. Sample acquisition stage: Under different drying temperatures and initial moisture contents, near-infrared polarization transmission images of single-layer grains in various states are collected as training sample data.

[0081] Simultaneously, corresponding sample grains were extracted, and the actual microcrack volume ratio within the sample was obtained using X-ray microtomography, serving as the true label data for micro-stress. The actual microcrack volume ratio was mapped to the 0-1 interval according to preset minimum and maximum values ​​in the training sample set to obtain the original proxy feature data output by the convolutional neural network. The corresponding monitoring label; the higher the proportion of microcrack volume, the larger the corresponding monitoring label value.

[0082] Training phase: Construct a training set containing the above sample data and their corresponding ground truth labels. Set the mean squared error function as the loss function for model training. Its mean squared error loss function... The calculation formula is: In the formula, This represents the number of sample batches in a single iteration. This represents the true label data for the j-th sample; This represents the predicted data output by the convolutional neural network for the j-th sample. An adaptive moment estimation optimization algorithm (such as the Adam optimization algorithm) is called to calculate the gradient of the loss function, and the layer weight parameters within the network are iteratively updated through backpropagation.

[0083] Training stops when the loss function value converges to the set allowable error range, the network node weights are fixed, and the network is deployed to edge computing nodes. The underlying code implementation for model training and the routine operations of microscopic tomography can be performed by those skilled in the art based on existing deep learning frameworks and laboratory specifications; these are well-known techniques in the field and will not be elaborated upon here.

[0084] S145, because the material is in a dynamic falling and random tumbling state in the entire flow channel, the attenuation of light varies with different poses, resulting in different outputs of raw proxy feature data over continuous time. High-frequency detection noise is present. Directly introducing feature data containing random noise into the underlying control system can easily cause frequent actions and control oscillations in downstream actuators. Therefore, the edge computing node calls the Discrete Kalman Filter algorithm to process the original proxy feature data. Perform smoothing processing.

[0085] Edge computing node construction includes micro-stress characterization indices and stress change rate data Two-dimensional state vector The time update equation and the observation update equation are set as follows: ; ; ; ; In the formula, This is the state prediction vector; This is the optimal state estimation vector for the previous period; This is the optimal state estimation vector for the current period; Let the state transition matrix be set. ,in This is the system sampling control cycle; To predict the covariance matrix; This is the estimated covariance matrix for the previous period; This is the estimated covariance matrix for the current period; The process excitation noise covariance matrix; The Kalman gain matrix; For the observation matrix, since only stress values ​​are observed, it is set as follows: ; To measure the noise covariance matrix; It is a two-dimensional identity matrix; Set the observed input value for the current period of the system. .

[0086] The above process excitation noise covariance matrix and measurement noise covariance matrix The specific value can be obtained by collecting sensor test data of static standard materials and calculating their statistical variance during the equipment calibration phase. During the filter startup phase, the edge computing node will initially acquire the raw proxy feature data. As a micro-stress characterization index The initial value will be the stress change rate data. Initialize to zero and initialize the estimated covariance matrix to a preset diagonal matrix.

[0087] Edge computing nodes iterate through the five equations along the time axis to filter out high-frequency interference from noisy observation sequences. The edge computing nodes estimate the optimal state vector for the current period. The smoothed micro-stress characterization index was extracted from the middle. and stress change rate data This provides stable micro-state variable inputs for subsequent reinforcement learning control.

[0088] To improve the representativeness of the polarization optics bypass for the same physical batch of material within the drying tower, the edge computing node also performs a moving average processing on the micro-stress characterization index obtained within multiple consecutive effective material sampling periods. The moving averaged micro-stress characterization index is then used as a micro-stress parameter for writing into a first-in-first-out queue or constructing a state vector. When the moving averaged micro-stress characterization index is used to construct a state vector, the edge computing node simultaneously performs a moving average processing on the corresponding stress change rate data, or recalculates the stress change rate data based on the moving averaged micro-stress characterization index.

[0089] The micro-stress characterization index after moving average can be expressed as: ,in, The micro-stress characterization index is obtained after moving average; The number of effective material sampling periods participating in the moving average; The micro-stress characterization index is the index corresponding to the p-th sampling period back from the current moment; t represents the system sampling control period; t represents the time point corresponding to the current control period.

[0090] If the rate of change of stress is recalculated based on the micro-stress characterization index after moving average, it can be expressed as: ,in, The stress change rate data is calculated based on the micro-stress characterization index after moving average. The micro-stress characterization index is the moving average of the current control cycle. It is the micro-stress characterization index after the moving average of the previous control cycle.

[0091] See attached document Figure 7 The specific process for restricted MDP status assessment and safety action overriding is explained below: S151, the edge computing node will acquire the inlet temperature parameters. Inlet humidity parameters Exhaust vent temperature parameters Exhaust vent humidity parameters Moisture content data of the same batch Micro-stress characterization index after moving average and stress change rate data based on moving average The environment state vector of the Markov decision process in the current control cycle is constructed by concatenating the elements in a fixed dimensional order. Among them, the micro-stress characterization index after moving average and stress change rate data based on moving average This refers to the moisture content data extracted from the first-in-first-out queue after physical time delay registration in step S11, and registered with the same batch. The vector expression for the microscopic feature data corresponding to the same physical batch is as follows: ; This environment state vector By integrating the thermodynamic boundary conditions outside the system, the macroscopic dehydration process, and the microscopic physical damage rate inside the material, a Markov observation space is constructed that can comprehensively characterize the current state of the drying system. To eliminate numerical calculation biases caused by different physical dimensions and accelerate the convergence speed of the neural network, the edge computing nodes, based on the pre-calibrated upper and lower limits of the sensor ranges for each physical quantity, analyze the environmental state vector. Each element within the tensor undergoes min-max normalization, mapping it to a numerical range of 0 to 1, thus generating the input state tensor. ,in, Represents the environment state vector The micro-stress characterization index scalar in [the text] Represents the environment state vector The input state tensor formed after normalization differs from the input state tensor in terms of data dimension and physical meaning.

[0092] S152, the edge computing node invokes a reinforcement learning Actor network model deployed in memory. This Actor network is configured as a multi-layer perceptron architecture with multiple fully connected layers. Internally, the network consists of an input layer, two hidden layers, and an output layer. The number of nodes in the input layer corresponds to the input state tensor. The dimensions are kept consistent, and in this embodiment, they are set to 7 nodes. The number of nodes in the two hidden layers are set to 64 and 32, or 128 and 64, respectively; the hidden layers use Tanh or ReLU activation functions to ensure the non-linear fitting ability of the feature mapping. The output layer has two nodes, corresponding to the airflow control coefficient and the heat source opening control coefficient, respectively, and the output values ​​are mapped to the floating-point range of 0 to 1 using the Sigmoid activation function. The output layer has two neuron nodes, and the output values ​​are mapped to the floating-point range of 0 to 1 using the Sigmoid activation function.

[0093] Edge computing nodes will input state tensors The vector is passed to the Actor network for forward inference computation. The network outputs a candidate action vector containing two elements. The candidate action vector includes candidate airflow control coefficients. and heat source candidate opening control coefficient Its expression is In the formula, , which is the candidate control coefficient for air volume of the centrifugal fan inverter in the variable frequency fan, used to characterize the expected intensity of air volume regulation; This is the heat source candidate opening control coefficient of the heat source proportional valve, used to characterize the expected heat source energy supply intensity.

[0094] S153. In the initial exploration phase or when facing uncharted system conditions, the decision actions output by the Actor network in reinforcement learning algorithms may exceed the capacity of the underlying physical devices. Edge computing nodes establish a safety rule base with priority execution levels within their storage area. This safety rule base includes stress limit thresholds and upper limits for inlet air temperature. and upper limit threshold of intake air temperature Stress limit threshold The specific values ​​are determined based on the critical data from compressive and tensile failure tests of specific grain types under laboratory conditions. Upper limit threshold for inlet air temperature. The temperature is set based on the biochemical critical temperature to avoid starch gelatinization and protein denaturation inside the grain. For example, it can be set to 60 degrees Celsius for rice.

[0095] Edge computing nodes combined with environment state vectors With candidate action vectors Perform boundary assessment. When the current microstress characterization index... When the following conditions are met: ; Or when the air inlet temperature parameter Meets the upper limit threshold of the inlet air temperature. When the conditions are met, the edge computing nodes further determine whether the candidate action has the ability to dissipate heat. When the heat source candidate opening control coefficient Greater than the preset safe heat source opening coefficient, or the candidate air volume control coefficient When the airflow is less than the preset safe airflow control coefficient, the edge computing node determines that the current network decision violates the physical security boundary, immediately intercepts the instruction, and triggers the safety action overwrite mechanism. The preset safe heat source opening coefficient is used to limit the maximum heat source supply intensity that can be maintained under high temperature conditions, and the preset safe airflow control coefficient is used to limit the minimum heat dissipation airflow required to be met under high temperature conditions.

[0096] S154, Under normal operating conditions where the security action overwriting mechanism is not triggered, the edge computing node directly transmits the candidate action vector. Confirmed as an action vector to be executed .

[0097] When the security action overriding mechanism is triggered, the edge computing node blocks the candidate action parameters output by the Actor network and forces the generation of the built-in security action vector. As a specific example of security parameter configuration, the expression for this security action vector is: Its underlying control logic is fixed as follows: the fan control coefficient is overwritten to 1.0 to output the system's maximum rated air volume, and the heat source opening coefficient is overwritten to 0.1 to maintain the hot blast furnace in a minimum non-quench combustion state. This overwriting action forces the system into a physical dissipation mode to vent the thermodynamic accumulation inside the drying tower and suppress the further propagation of microcrystalline cracks.

[0098] The edge computing node will determine the execution action vector The vector is sent to the programmable logic controller (PLC), which then executes the action vector. This is converted into corresponding underlying industrial analog control electrical signals and sent to the centrifugal fan inverter in the variable frequency fan and the heat source proportional valve in the proportional control valve, respectively, to complete the physical action closed loop of the current control cycle. Among these, the execution action vector... Candidate control coefficients for air volume The target operating frequency of the centrifugal fan inverter is converted using the following linear mapping: In the formula, This refers to the minimum permissible operating frequency of the variable frequency fan. This is the maximum permissible operating frequency of the variable frequency fan.

[0099] Execution action vector Candidate control coefficients for air volume The target opening degree of the heat source proportional valve is converted through the following linear mapping: In the formula, This is the minimum maintaining opening degree of the heat source proportional valve. This represents the maximum permissible opening degree of the proportional valve for the heat source.

[0100] For the digital-to-analog conversion and communication configuration of the underlying digital signals to 4-20mA or 0-10V industrial standard electrical signals, those skilled in the art can call conventional PLC communication module instruction libraries to achieve this, which is a well-known technology in the field and will not be elaborated here.

[0101] See attached document Figure 8 The specific process of dynamic tolerance coupling and multi-objective reward function iteration is explained below: S161, in the underlying control logic, the overall optimization goal of the drying system is to achieve the maximum moisture reduction while consuming reasonable energy, and at the same time, to minimize grain breakage. The underlying physical equipment executes action vectors... Run a system sampling control cycle Afterwards, the thermodynamic state and the moisture state of the material inside the drying system change.

[0102] Because there is a process response delay in the control actions of the continuous drying tower after they act on the material and reach the moisture sensor and polarization optics bypass detection position, the edge computing node establishes an action timestamp queue to execute the action vector. It is bound to its corresponding material batch identifier or timestamp; after the batch of materials arrives at the feedback detection location and forms a valid same-batch registration feedback status, a new environmental state vector for reward calculation is then constructed. .

[0103] Edge computing nodes continue to perform physical delay registration on the feedback data after the action is executed. When both the registered moisture content data and the matched micro-feature data of the same batch are valid, the dehydration reward and subsequent multi-objective reward function values ​​are calculated. Edge computing nodes extract the registered moisture content data of the physical batch corresponding to the current action and its next valid feedback batch, and calculate the dehydration reward. The formula is: ,in, Register moisture content data for the same batch in the current control cycle. This is to register the moisture content data for the same batch in the next effective feedback cycle.

[0104] Dehydration Bonus A positive value indicates that the system has achieved effective physical dehumidification. Simultaneously, edge computing nodes execute action vectors... Energy consumption penalty item for control coefficient calculation equipment According to the basic principles of fluid mechanics and thermodynamics, the shaft power of a centrifugal fan is approximately in a cubic relationship with its operating frequency, while the energy consumption of the heat source is usually linearly positively correlated with the opening degree of the proportional valve.

[0105] Therefore, energy consumption penalty item The calculation formula is configured as follows: In the formula, The rated power conversion factor for the centrifugal fan frequency converter; The full-open reference heat loss coefficient of the heat source proportional valve; The target operating frequency for the centrifugal fan inverter. This refers to the maximum permissible operating frequency of the variable frequency fan. This is the heat source candidate opening control coefficient for the heat source proportional valve.

[0106] S162. In the physical process of grain drying, the stress resistance to microcracks within the grain is not a constant value. When the moisture content is high, the grain exhibits strong plasticity and can withstand higher thermal stress; as the moisture content decreases, the grain tends to become hard and brittle, making it prone to bursting damage. Edge computing nodes introduce a dynamic tolerance coupling mechanism based on the current moisture content. Calculate dynamic stress tolerance In the formula, The conservative safety stress limit of grains obtained during the calibration phase at the end of low moisture drying; The target moisture content at the outlet of the system; The stress-moisture coupling slope constant is obtained by linear fitting of rheological compression experimental data of the same batch of materials; This is a function that takes the maximum value and is used to limit the tolerance value from decreasing further when the moisture content is below the target value.

[0107] Specifically, compression, bending, or thermal stress loading tests were conducted on rice samples at multiple moisture content levels. The critical micro-stress characterization index corresponding to the occurrence of a preset bursting rate or preset micro-crack volume percentage was recorded. The relationship between moisture content and the critical micro-stress characterization index was then linearly fitted or piecewise linearly fitted to obtain... and .

[0108] Based on this, the edge computing nodes are evaluated according to the micro-stress characterization index after the moving average of the next period. With dynamic stress tolerance The comparison relationship is used to calculate the stress penalty term. : when hour, ; when hour, ; In the formula, The nonlinear stress penalty weight constant can be calibrated through trial and error during the initial trial operation of the system. This quadratic exponential penalty term causes a gradually amplifying negative reward feedback once the system state exceeds the dynamic tolerance boundary.

[0109] S163, the edge computing node linearly weights the reward and penalty values ​​of each of the above individual objectives to calculate the overall multi-objective reward value of the current action under the current environmental state. Before entering the overall reward function, the values ​​of each individual target item are normalized to a dimensionless reward item according to a preset calibration range, or converted to a uniform reward value scale through the corresponding weight coefficient.

[0110] To ensure that the reward function simultaneously reflects stress change rate data and safe overwrite trigger state, the overall multi-objective reward value is... Configured as follows: In the formula, For efficiency weighting coefficients, This is the energy-saving weighting coefficient. The stress change rate penalty weighting coefficient, The weighting coefficient for the penalty of safe overwriting; For the stress change rate data of the next control cycle; Overwrite the trigger status flag for the parameter.

[0111] When the current control cycle triggers the safety action overwrite mechanism When the safety action overwrite mechanism is not triggered in the current control cycle, .

[0112] When the feedback data after the action is executed has not yet met the physical time delay registration conditions, the edge computing node temporarily suspends the generation of the corresponding experience data stream; after the next effective batch of registered moisture content data and batch of matched micro-feature data are formed, the reward function value is calculated and written into the experience playback pool.

[0113] Edge computing nodes will use the current environment state vector Execution action vector Overall multi-objective reward value and the normalized input state tensor of the next effective feedback cycle Packaged into a Markov empirical data stream In this process, the environmental state data in the experience replay pool is converted into normalized input state tensors before being input into the Actor network and the Critic network. The edge computing node stores the experience data stream into an allocated experience replay pool in memory. The allocation and overwrite / elimination mechanism of the experience replay pool's data structure can be implemented using a conventional circular queue data structure, which is well-known in the field and will not be elaborated upon here.

[0114] Before being deployed to edge computing nodes, the S164 Actor and Critic networks are pre-trained offline based on historical drying operation data or simulated drying environments. The initial network weights are then solidified after offline verification confirms that the safety boundary constraints are met. During the initial deployment phase, edge computing nodes employ conservative control strategies or preset safety action vectors as a control transition.

[0115] Once the amount of data accumulated in the experience replay pool reaches the set batch processing threshold, the edge computing node triggers online iterative updates of the reinforcement learning model in the background process. The experience replay pool capacity is set to 10. 4 Up to 10 6 A number of empirical data points, and the number of small batch samples updated online iteratively. The learning rates were set to 32 to 256, with the Actor and Critic networks each having a learning rate of 10. -5 Up to 10 -3 Within the range, time discount factor The target network soft update rate coefficient is set to a range of 0.001 to 0.01, within the range of 0.90 to 0.99.

[0116] The online iterative update process runs in isolation from the real-time control process, and the action overriding priority of the security rule base is always higher than the Actor network output and the online updated network parameters. Edge computing nodes randomly draw from the experience replay pool... A small batch of empirical data, in which, This represents the number of mini-batch samples during the online update process of reinforcement learning.

[0117] Edge computing nodes invoke the Critic network, which is deployed in memory in conjunction with the Actor network. As the evaluation module of this reinforcement learning framework, the hierarchical topology of the Critic network is similar to that of the Actor network. The difference is that the input layer simultaneously receives state tensors and action vectors, while the output layer outputs a single continuous action state evaluation value, namely the Q-value.

[0118] Specifically, the Critic network normalizes the input state tensor. The input is a concatenation of the action vector along its dimensions, with the number of input nodes equal to the sum of the state and action dimensions. The number of hidden layer nodes is set to 64 and 32, or 128 and 64. The output layer consists of a single linear node, used to output the action state value. For the i-th empirical sample in the mini-batch sample set, the edge computing node calculates the loss function of the Critic network using the following temporal difference error formula. : ; ; In the formula, This is the target action state evaluation value corresponding to the i-th experience sample; This represents the overall multi-objective reward value in the sampled sample. The time discount factor is set, and its value ranges from 0 to 1. It is usually set to 0.9 to 0.99 to take into account long-term returns. For the target Critic network; For the target Actor network; For the node weight parameters of the target Actor network; represents the node weight parameters of the target Critic network; Q represents the current Critic network. These are the weight parameters of the current Critic network; This represents the number of samples in a small batch.

[0119] Edge computing nodes call the optimizer to calculate the loss function Compared to The gradient of the action value evaluation gradient is used to update the parameters of the current Critic network using backpropagation. Then, the edge computing nodes use the updated Critic network's action value evaluation gradient to update the weight parameters of the current Actor network using the policy gradient ascent method, making the Actor network more likely to output action instructions with higher overall multi-objective rewards. Finally, the edge computing nodes update the gradient at a set soft update rate coefficient. The value range can be set to, for example, 0.001 to 0.01, so that the parameters of the current network are smoothly passed to the corresponding target network to complete one cycle of adaptive optimization iteration.

[0120] Specific application examples: The system is currently processing a batch of early indica rice with an initial moisture content of 24%, and the target output moisture content is set as follows: .

[0121] 1. Physical delay registration (corresponding to step S11) The vertical distance difference between the polarization optical bypass drain and the capacitive moisture sensor is known. Meters, material flow coefficient obtained from offline calibration meters per revolution per minute (rpm). Current speed encoder feedback of the discharge roller speed. Revolutions per minute.

[0122] The equivalent vertical descent speed of the material is 1.2 × 0.5 = 0.6 meters per minute.

[0123] System calculates physical delay minute.

[0124] If the system control cycle If the edge computing node backtracks n=5 cycles to the first-in-first-out queue, it matches the micro-features from 5 minutes ago with the current uncompensated moisture content data to ensure that the registered data corresponds to the same physical batch.

[0125] 2. Self-cleaning and light intensity calibration (corresponding steps S12-S13) During the calibration cycle, the entire flow channel is emptied. The CMOS sensor captures a transmission image through the empty window and calculates the global information entropy.

[0126] Preset cleanliness threshold The current calculation yields... .

[0127] Since 5.2 < 6.5, the system determines that there is contamination on the optical detection window and prioritizes triggering a purging and cleaning action. After purging for 20 seconds, a second detection is performed. The pH level rose to 7.2 (greater than 6.5), confirming that the cleanliness standard has been met.

[0128] Then calculate the center gray value. Preset benchmark Light intensity error Based on the PID formula (assuming...) ),calculate need Increased drive current The system updates the backlight array drive signal to compensate for the light source attenuation error.

[0129] 3. Proxy feature extraction and smoothing (corresponding to step S14) To restore material sampling, the convolutional neural network performs forward inference on three consecutively acquired polarization images and outputs the original surrogate feature data. The values ​​are [0.68, 0.85, 0.62].

[0130] Input the sequence into a discrete Kalman filter. Assume the optimal estimate of the previous period... After Kalman gain matrix After calculation correction, the influence of transient high-frequency data is reduced, and the micro-stress characterization index of the current period is output. Stress change rate .

[0131] 4. Status assessment and security overwrite mechanism (corresponding step S15) Preset upper limit of air intake temperature stress limit Edge computing nodes construct their current state vectors and input them into the Actor network. The Actor network outputs the corresponding candidate action vectors: candidate airflow control coefficients. Heat source candidate opening control coefficient .

[0132] At this time, read the sensor data, current intake air temperature. (Exceeding the upper limit of 60℃), and the candidate heat source opening degree (0.95) is greater than the preset safe heat source opening degree coefficient, for example, 0.80.

[0133] If the boundary assessment and interception conditions are met, the security action overwrite mechanism is triggered.

[0134] Generate an appendix based on the system state evolution data of this drying cycle. Figure 9 This refers to the time-series diagram of microscopic stress and motion-written response within rice grains. (Attached) Figure 9 Subplot (A) shows the inlet temperature curve and the safety threshold line; subplot (B) includes the original feature data scatter plot and the micro-stress continuous curve after Kalman filtering; subplot (C) compares the original output signal of the Actor network with the overwrite action signal after the safety rule is triggered. After triggering the overwrite, the action vector is updated to full fan operation (1.0) and heat source lower limit maintenance opening (0.1), and the system switches to heat dissipation mode to suppress micro-crack propagation.

[0135] 5. Reward Calculation (corresponding to step S16) After one cycle of the operation, the moisture content decreased from 24.0% to 23.8%. ).

[0136] Calculate the dynamic stress tolerance at the current moisture content (23.8%): Set a safety baseline. Coupling slope .

[0137] .

[0138] Predicted microstress for the next cycle The value is 0.65.

[0139] Since 0.65 ≤ 0.694, the stress penalty term .

[0140] Because the overwrite mechanism was triggered in the current cycle, the overwrite flag was updated, and the overall reward function was updated. Overall reward function A corresponding weighted penalty term is superimposed. After this value is written into the experience replay pool, it will reduce the Critic network's value assessment of the state-action pair, causing the Actor network to tend to output a smaller heat source opening when approaching the temperature critical point in subsequent iterations.

[0141] Experimental verification section: Three sets of operational experiments were conducted using a continuous drying tower with a loading capacity of 15 tons. The material used was the same batch of early indica rice (initial moisture content approximately 26%, target moisture content 14.5%). The comparative control conditions are as follows: Control group A (traditional PID control): The control uses a set inlet air temperature (55℃) and constant air volume, and adjusts the discharge speed based on the moisture feedback at the outlet.

[0142] Control group B (basic reinforcement learning control): The DDPG algorithm is used to optimize air volume and heat source, without introducing physical time delay registration mechanism, optical micro-stress feedback and security overwrite logic.

[0143] Experimental Group C (Scheme of this invention): Polarization optics bypass is enabled to acquire features, and self-cleaning closed loop, Kalman filtering, time delay registration and multi-objective reinforcement learning optimization control based on restricted Markov decision process are run.

[0144] The relevant performance indicators are statistically summarized in Table 1:

[0145] The full-cycle micro-stress evolution trajectory and comprehensive index data of the three sets of experiments are as follows: Figure 10 As shown, Figure 10 The curves show the comparison between drying performance and micro-damage evolution under different control strategies.

[0146] exist Figure 10 In subplot (A) (dynamic stress evolution tracking curve), the horizontal axis represents the average moisture content of the grains, and the vertical axis represents the micro-stress characterization index. The micro-stress curve of control group B crosses the dynamic safety tolerance zone when the moisture content drops to around 18%, corresponding to a measured burst rate of 5.20%. The micro-stress curve of experimental group C lies below the dynamic stress tolerance boundary. Approaching the boundary region, due to the system processing the rate of change of micro-stress and implementing safety action overwrite intervention, the rise of its micro-stress characterization index slows down and remains within the preset tolerance range.

[0147] Figure 10Subplot (B) (a biaxial bar chart of comprehensive energy efficiency and quality) shows that, compared to control groups A and B, the final breakage rate of experimental group C was controlled at 0.95%. Compared to control group A, the comprehensive energy consumption of experimental group C was reduced by approximately 21.1%. Simultaneously, combined with physical time-delay registration processing, the outlet moisture uniformity of experimental group C was narrowed to ±0.4%. The above experimental data demonstrate that the present invention can effectively balance dehydration efficiency and control breakage damage during the material drying process.

Claims

1. A method for optimizing rice air-drying parameters based on artificial intelligence, characterized in that, Includes the following steps: Polarized transmission images of rice grains were collected, features were extracted using a convolutional neural network, and after smoothing and filtering, the micro-stress characterization index and stress change rate data were output. The inlet temperature parameter, inlet humidity parameter, outlet temperature parameter, outlet humidity parameter, uncompensated average moisture content parameter, and discharge speed data of the drying tower are obtained. The physical time delay value is calculated. The uncompensated average moisture content parameter is time-delay registered with the micro-stress characterization index and stress change rate data to generate a state vector. The state vector is input into a Markov decision process model constructed based on reinforcement learning Actor network and Critic network, and the output air volume candidate adjustment parameters and heat source candidate adjustment parameters are evaluated. After verification by built-in safety rules, the final thermodynamic action command is issued. The multi-objective reward function value is calculated based on the physical feedback state after the execution of the final thermodynamic action command and the dynamic stress tolerance calculated based on the current moisture content, so as to update the node weights in the reinforcement learning Actor network and Critic network.

2. The method according to claim 1, characterized in that, Before acquiring the polarization transmission image of the rice, the following is also included: In the state of polarized optical bypass emptying, a window transmission image is captured, the gray-level probability distribution function of the window transmission image is extracted to calculate the global information entropy, and the global information entropy is compared with a preset cleanliness threshold. When the global information entropy is not lower than the cleanliness threshold, the average gray value of the central region of the transparent image is calculated, and the digital PID control algorithm is run to adjust the driving current of the backlight array in order to maintain the preset reference target light intensity. When the global information entropy is lower than the cleanliness threshold, the purge valve is opened, and the airflow from the drying tower's main air supply pipe is sequentially intercepted, separated, purified, and stabilized by the dust removal filter, condensation and dehumidification component, and pressure reducing valve. Then, the airflow is purged through the purge nozzle in the form of a targeted jet to perform a surface purging action on the optical detection window until the global information entropy of multiple consecutive detection cycles is restored to not lower than the cleanliness threshold.

3. The method according to claim 1, characterized in that, The process of extracting features through a convolutional neural network and then smoothing and filtering them to output micro-stress characterization index and stress change rate data specifically includes: The polarization transmission image is a near-infrared polarization transmission image, which includes a first polarization direction transmission image and a second polarization direction transmission image that are orthogonal to each other. After normalizing the two, they are concatenated into an input data tensor, which is then passed to a convolutional neural network based on a residual network architecture. The local contrast features of the image are extracted by a convolutional neural network, and the orthogonal polarization variance parameter based on the gray level difference between the first polarization direction transmission image and the second polarization direction transmission image is used to output the original proxy feature data characterizing the average micro-stress damage degree. The original proxy feature data is used as the observation input value and substituted into the discrete Kalman filter algorithm composed of the time update equation and the observation update equation for iterative calculation. High-frequency detection noise is filtered out, and smoothed micro-stress characterization index and stress change rate data are separated and extracted.

4. The method according to claim 1, characterized in that, The calculation of the physical time delay involves time-delay registration of the uncompensated average moisture content parameter with the micro-stress characterization index and stress change rate data to generate a state vector, specifically including: Based on the vertical distance difference between the inlet of the collected polarized transmission image and the moisture sensor inside the drying tower, the material flow coefficient, and the discharge speed data, the physical time delay value corresponding to the descent of the rice is calculated. Divide the physical delay value by the system sampling control period to obtain the discrete queue index steps; Using the queue index step number as the backtracking offset, backtracking is performed from the first-in-first-out queue storing historical micro-stress characterization index and historical stress change rate data in the historical direction. The corresponding historical micro-stress characterization index and historical stress change rate data are extracted and aligned with the uncompensated average moisture content parameter in the time dimension. The state vector is generated by concatenating the inlet temperature parameter, inlet humidity parameter, exhaust temperature parameter, and exhaust humidity parameter.

5. The method according to claim 1, characterized in that, The final thermodynamic action command, issued after verification by built-in safety rules, specifically includes: Determine whether the micro-stress characterization index after time delay registration is not less than the preset stress limit threshold, and determine whether the air inlet temperature parameter is not less than the preset upper limit threshold of air inlet temperature and whether the opening control coefficient corresponding to the candidate heat source adjustment parameter is greater than the preset safe heat source opening coefficient or whether the control coefficient corresponding to the candidate air volume adjustment parameter is less than the preset safe air volume control coefficient. When any of the above judgment conditions that violate the physical safety boundary are met, the safety action overwrite mechanism is triggered, the candidate action vector containing the candidate air volume adjustment parameters and the candidate heat source adjustment parameters is blocked, and the built-in safety action vector that outputs the maximum rated air volume of the system and maintains the minimum non-extinguishing combustion state of the heat source is forcibly generated as the final thermodynamic action command to be issued and executed.

6. The method according to claim 1, characterized in that, The step of calculating the multi-objective reward function value based on the physical feedback state after the execution of the final thermodynamic action command and the dynamic stress tolerance calculated based on the current water content, in order to update the node weights in the reinforcement learning Actor network and Critic network, specifically includes: Based on the difference between the registered moisture content data fed back after the action is executed and the target outlet moisture content, combined with the set conservative safety stress limit and stress-moisture coupling slope constant, the dynamic stress tolerance at the current moisture content is calculated. Based on the comparison between the registered micro-stress characterization index of the next cycle and the dynamic stress tolerance, a stress penalty term is generated; The dehydration reward item before and after feedback, the energy consumption penalty item based on system operating frequency and valve opening, the stress penalty item, the stress change rate penalty item based on the stress change rate data after registration in the next cycle, and the parameter overwrite trigger status flag are linearly weighted to calculate the value of the multi-objective reward function. The temporal difference error is calculated using an empirical data stream containing the values ​​of the multi-objective reward function to update the parameters of the Critic network, and the weight parameters of the reinforcement learning Actor network are updated using the action value evaluation gradient provided by the updated Critic network.

7. An artificial intelligence-based rice air-drying parameter optimization system, used to implement the method described in any one of claims 1 to 6, characterized in that, include: The drying execution system is based on the main body of the drying tower. The main body of the drying tower is equipped with an inlet air temperature and humidity sensor, an exhaust air temperature and humidity sensor, a capacitive moisture sensor, a bottom discharge roller speed encoder, a variable frequency fan that supplies drying airflow, and a proportional regulating valve that adjusts the heat source supply. A polarization optical bypass is arranged in parallel outside the main body of the drying tower, and an optical detection chamber is set inside. A backlight array and a CMOS sensor equipped with a micro-polarization array polarization filter are respectively installed on both sides of the optical detection chamber. Edge computing nodes establish communication and control connections with the drying execution system and the polarization optical bypass, respectively, for the following purposes: The convolutional neural network is called to process the polarization transmission image of rice obtained by the CMOS sensor to extract features, and after smoothing and filtering, the micro-stress characterization index and stress change rate data are output. Acquire inlet temperature parameters, inlet humidity parameters, outlet temperature parameters, outlet humidity parameters, uncompensated average moisture content parameters, and discharge speed data. Perform physical time delay registration between the uncompensated average moisture content parameters and the micro-stress characterization index and stress change rate data to generate a state vector. The Markov decision process model, constructed based on reinforcement learning Actor and Critic networks, outputs candidate air volume adjustment parameters and candidate heat source adjustment parameters. After safety verification, action commands are issued, and the node weights in the reinforcement learning Actor and Critic networks are updated based on the multi-objective reward function value calculated from the feedback state after the action is executed.

8. The system according to claim 7, characterized in that, The top of the polarization optical bypass is equipped with a flow-guiding valve driven by a stepper motor, and the interior is equipped with a row of flow channels from top to bottom. The micro-polarization array polarization filter includes at least two mutually orthogonal polarization directions; The optical inspection window of the optical inspection chamber is inclined to one side facing the material flow channel and is equipped with a purging nozzle. The air inlet of the purging nozzle is connected in sequence to the purging valve, pressure reducing valve, condensation dehumidification component and dust removal filter through a pneumatic pipeline, and finally connected to the air supply main pipe corresponding to the air outlet of the variable frequency fan to form a physical air path channel.

9. The system according to claim 8, characterized in that, The edge computing node has a built-in graphics processor; The inlet air temperature and humidity sensor, the exhaust air temperature and humidity sensor, the capacitive moisture sensor, and the bottom discharge roller speed encoder in the drying execution system are all connected to the programmable logic controller. The output of the programmable logic controller is connected to the centrifugal fan inverter that drives the variable frequency fan, and the proportional regulating valve, respectively. The edge computing node establishes action execution connections with the stepper motor, the purge valve, and the power drive module of the backlight array through a programmable logic controller.

10. The system according to claim 8, characterized in that, The condensation dehumidification component is equipped with a semiconductor cooling chip, a condensate collection chamber, and a gas-liquid separation structure. It is used to force the high-temperature dusty airflow introduced by the main air supply pipe to cool down to below the dew point temperature to condense water. After passing through the reheat section or heat exchange section, the output airflow after liquid drainage is raised to a low-humidity state higher than its current dew point temperature before being input into the pressure reducing valve for dynamic pressure regulation.