A method and system for controlling the concentration of a raw slurry in a soy product processing process
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-19
- Publication Date
- 2026-08-11
AI Technical Summary
[0005]本发明提供一种豆制品加工过程中的原浆浓度控制方法及系统,旨在解决相关技术中在复杂多变的实际应用场景中,极易导致电量告急的车辆在前往拥堵路段的途中因额外耗电骤增而直接半路抛锚,或者到达目的地后因无可用空桩而面临扑空的窘境,影响了车主的出行安全与补能体验的问题
[0006]在第一方面中,本发明提供了豆制品加工过程中的原浆浓度控制方法,包括:获取当前生产批次下的大豆原料特征以及煮浆腔内浆液的浓度与温度,并将其构建为多维状态向量;基于历史多维状态向量构建状态空间和动作空间,基于所述状态空间输入强化学习网络,得到补水调节阀的开度调节指令和蒸汽调节阀的开度调节指令;构建下开度发调节指令后的总奖励,并将其反馈至强化学习网络,其中,所述总奖励为目标偏差奖励值、浓度波动惩罚值和资源协同惩罚值之和,所述目标偏差奖励值反映了浓度与目标值之间的接近程度;所述浓度波动惩罚值反映了生产过程中浓度的平稳性;检测补水调节阀与蒸汽调节阀的开度,若两阀门开度在同一时刻均大于设定的开度阈值,则施加资源协同惩罚值。相较于现有技术中依赖目标浓度单一反馈的固定调节机制所导致的响应滞后与水热抵消问题,本发明在豆制品连续化生产场景下,通过融合大豆原料特征及浆液理化状态构建强化学习多维状态向量,并引入涵盖目标偏差、浓度波动和资源协同的多维度复合奖励机制来训练网络,不仅能够自适应不同批次原料的理化差异,精准输出补水与蒸汽阀门的协同调节指令,还有效避免了因边加冷水边灌蒸汽引起的高能耗与浓度剧烈震荡,从根本上提升了煮浆工序的良品率和系统的运行经济性。
Smart Images

Figure CN122547110A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of raw material production control technology. More specifically, this invention relates to a method and system for controlling the concentration of raw pulp in the processing of soybean products. Background Technology
[0002] Soy product processing is an important part of the modern food industry. Among them, the boiling process is the core production link that determines the gel network structure, texture, and overall yield of the final soy products (such as tofu and dried bean curd). In this process, the core purpose of the quality control system is to dynamically adjust the flow rate of clean water and heating steam entering the boiling chamber to stabilize the physicochemical properties of the raw soy milk, especially its concentration and temperature, within the standard range required by specific processes.
[0003] With the widespread adoption of automated processing technology, modern soybean product factories generally use continuous high-speed cooking lines. In actual large-scale continuous production scenarios, the source and storage conditions of soybean raw materials vary greatly, resulting in significant objective differences in the initial moisture content, water absorption swelling rate, and protein dissolution characteristics of different batches of soybeans. In addition, there is a strong physical coupling effect between the water replenishment operation and steam heating in the closed cooking chamber. That is, injecting clean water will dilute the concentration and instantly lower the temperature of the soy milk, while adding high-temperature steam will not only raise the temperature, but the condensate it produces will also have a reverse effect on the system concentration.
[0004] However, existing control schemes have significant shortcomings in dealing with such complex production scenarios. In actual production, existing control technologies mostly rely on fixed adjustment mechanisms based on a single feedback of target concentration. When using such conventional mechanisms to handle multiple batches of continuous boiling tasks, because their control logic completely ignores the differences in the underlying physicochemical properties of soybean raw materials and lacks a global assessment of the synergistic relationship between water and heat, the production line is prone to severe control oscillations when encountering new batches of raw materials. Specifically, when the protein colloidal rate of a new batch of soybeans changes, causing the actual concentration to deviate from the expected value, existing systems often fall into overcompensation due to response lag, blindly and simultaneously opening the water and steam regulating valves significantly. This crude approach not only causes drastic fluctuations in the concentration of the slurry in the chamber within a short period of time, resulting in serious defects such as loose texture and poor shaping of soybean products produced during this period due to uneven protein denaturation; at the same time, the mutual cancellation effect of water and heat due to the simultaneous addition of large amounts of cold water and the vigorous injection of steam causes significant energy waste under real working conditions, severely restricting the overall yield and operational economy of the production line. Summary of the Invention
[0005] This invention provides a method and system for controlling the concentration of raw soy pulp during the processing of soybean products. It aims to solve the problem in related technologies that, in complex and ever-changing practical application scenarios, vehicles with low battery levels are prone to break down halfway due to a sudden increase in additional power consumption on their way to congested sections of the road, or face the predicament of arriving at their destination but finding no available charging stations, thus affecting the travel safety and charging experience of car owners.
[0006] In a first aspect, the present invention provides a method for controlling the concentration of raw slurry during soybean product processing, comprising: acquiring the characteristics of soybean raw materials in the current production batch and the concentration and temperature of the slurry in the cooking chamber, and constructing them as a multi-dimensional state vector; constructing a state space and an action space based on the historical multi-dimensional state vector, inputting the state space into a reinforcement learning network to obtain the opening adjustment command of the water supply regulating valve and the steam regulating valve; constructing the total reward after issuing the opening adjustment command and feeding it back to the reinforcement learning network, wherein the total reward is the sum of the target deviation reward value, the concentration fluctuation penalty value, and the resource coordination penalty value, the target deviation reward value reflecting the closeness between the concentration and the target value; the concentration fluctuation penalty value reflecting the stability of the concentration during the production process; detecting the opening of the water supply regulating valve and the steam regulating valve, and if the opening of both valves is greater than a set opening threshold at the same time, then applying a resource coordination penalty value. Compared to the response lag and hydrothermal offsetting problems caused by the fixed adjustment mechanism that relies on a single feedback of target concentration in existing technologies, this invention, in the continuous production scenario of soybean products, constructs a reinforcement learning multidimensional state vector by integrating the characteristics of soybean raw materials and the physicochemical state of the slurry, and introduces a multidimensional composite reward mechanism covering target deviation, concentration fluctuation and resource coordination to train the network. This not only adapts to the physicochemical differences of different batches of raw materials and accurately outputs coordinated adjustment commands for water replenishment and steam valves, but also effectively avoids the high energy consumption and violent concentration fluctuations caused by adding cold water and steam at the same time, fundamentally improving the yield of the slurry cooking process and the economic efficiency of the system operation.
[0007] Furthermore, a multi-dimensional state vector is constructed, including: soybean raw material features such as soybean-to-water ratio and protein thermal effect features; the multi-dimensional state vector also includes the concentration change rate of the current sampling period. This invention deeply integrates the soybean-to-water ratio feature reflecting the true liquid-to-solid ratio, the protein thermal effect feature quantifying the protein's activity potential, and the concentration change rate reflecting the intensity of dynamics into the state vector. This enables the reinforcement learning model to proactively perceive and understand the physicochemical evolution trajectory of the soybean raw material's underlying structure under complex thermal conditions, further improving the model's decision-making accuracy and robustness under complex operating conditions.
[0008] Furthermore, the method for obtaining the target deviation reward value includes: calculating the absolute difference between the real-time slurry concentration and the standard target concentration at the current sampling time, correcting the absolute difference based on a first weighting coefficient, and using the negative value of the correction result as the target deviation reward value. By calculating the absolute difference between the real-time slurry concentration and the standard target concentration and combining it with the weighting coefficient for negative incentive, the reinforcement learning model is guided to always prioritize approaching the target concentration as the primary optimization direction during action exploration, ensuring that the finally generated adjustment strategy can strictly meet the baseline requirements of the specific process standards for soybean product processing.
[0009] Furthermore, the method for obtaining the concentration fluctuation penalty value includes: comparing the maximum absolute value of the concentration change rate within the current sampling time and a preset number of sampling periods with a preset fluctuation threshold to obtain the concentration fluctuation penalty value. By extracting the extreme values of the concentration change rate over multiple periods and comparing them with historical stability thresholds to apply the penalty, the overcompensation behavior that the model may produce when pursuing the target concentration is effectively suppressed, avoiding the risk of uneven protein denaturation and poor bean product formation caused by drastic fluctuations in slurry concentration within a short period of time.
[0010] Furthermore, the method also includes an adaptive processing mechanism for new raw materials. Specifically, it involves: acquiring the protein and moisture content of a new batch of soybean raw materials to form a current raw material feature vector; retrieving and extracting a set of historical raw material feature vectors with the highest cosine similarity to the current raw material feature vector from a historical raw material feature vector database; and when the highest cosine similarity is lower than a set threshold, entering an exploration mode and superimposing random perturbations into the action commands. By accurately quantifying the degree of physicochemical deviation between new and old materials through cosine similarity, and actively introducing random perturbations when identifying new raw materials with high deviations, the reinforcement learning network is driven to adaptively evolve in actual production, greatly enhancing the system's generalization ability and response speed to raw materials with unknown characteristics.
[0011] Furthermore, the system enters an exploration mode, which includes: continuously monitoring the deviation between the slurry concentration and the standard target concentration; when the deviation exceeds a set threshold, pausing the reinforcement learning model control and switching to PID mode; after collecting a set number of new samples and updating the model in PID mode, automatically switching back to reinforcement learning control mode. During the exploration and evolution of new raw materials, by setting a safety monitoring line for concentration deviation, and decisively temporarily handing over control to a conservative single-loop proportional-integral control mode when oscillations exceed limits, the system resolves the risk of production runaway that may arise in the early stages of reinforcement learning, and smoothly completes the smooth iteration and upgrade of the model while ensuring the basic operational safety of the production line.
[0012] Furthermore, the process of issuing the opening adjustment command also includes: adding the incremental adjustment command to the current valve opening to obtain the target absolute opening; calculating the absolute value of the difference between the target absolute opening and the current valve opening; if the absolute value of this difference is less than a preset physical dead zone threshold, then the current opening adjustment command is discarded; otherwise, the target absolute opening is converted into an analog electrical signal and sent to the valve positioner. Introducing a filtering and judgment mechanism based on a hardware dead zone threshold at the physical execution end directly discards invalid, small opening adjustment commands, effectively avoiding mechanical wear and lifespan reduction caused by frequent high-frequency fine-tuning of pneumatic valves and positioners, and improving the operational stability of the underlying actuator in harsh industrial environments.
[0013] Furthermore, random perturbations are superimposed on the action commands, including: superimposing small-amplitude opening and closing test signals on the initial predicted incremental commands for the water supply valve and the steam valve output by the reinforcement learning network. In exploratory mode, the superposition of these small-amplitude opening and closing test signals on the initial predicted commands is precisely constrained. This ensures that the system can safely and effectively acquire real dynamic response data of new materials to action commands to expand the experience pool, while preventing excessive perturbations from disrupting the basic hydrothermal balance of the current boiling process.
[0014] Furthermore, a resource coordination penalty value is applied, wherein the resource coordination penalty value is a fixed value.
[0015] In a second aspect, a raw pulp concentration control system for soybean product processing is also provided, comprising a processor and a memory, the memory storing a computer program, the processor executing the computer program to implement the raw pulp concentration control method for soybean product processing as described in any of the above embodiments.
[0016] Beneficial Effects: To address the issue of concentration fluctuations and energy waste easily caused by batch differences in raw materials and the coupling of hydrothermal physics in the continuous cooking process of soybean products, a production quality control scheme based on reinforcement learning is proposed. Firstly, a multi-dimensional state space is constructed by integrating the underlying physicochemical characteristics of soybeans with the dynamic changes in the slurry. A composite reward function is designed with the aim of suppressing concentration fluctuations and penalizing hydrothermal offsetting, guiding the model to output a precise coordinated control strategy for water replenishment and steam valves. Secondly, by combining adaptive exploration of new materials and a safety retreat mechanism, dynamic adaptation to complex thermal environments is achieved while ensuring stable production line operation, significantly improving the forming yield and processing economy of soybean products. Attached Figure Description
[0017] Figure 1 This is a schematic flowchart illustrating a power regulation method according to an embodiment of the present invention. Detailed Implementation
[0018] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0019] like Figure 1 As shown, S101: Data acquisition and preprocessing.
[0020] In this embodiment, the moisture content and crude protein content of the current batch of soybeans are obtained by a near-infrared spectrometer; the temperature of the slurry and the concentration of the raw slurry in the cooking chamber are obtained by a temperature sensor and a high-frequency microwave concentration meter at a preset sampling period; and the cumulative water injection volume of the water supply pipe from the start of processing to the current sampling time is obtained by a flow meter. The parameters obtained by the above means are used as the basic parameter set.
[0021] Following the above, a fixed-length historical data sequence is extracted using a sliding time window approach. Low-pass filtering is then applied to this historical data sequence to eliminate transient high-frequency noise interference from the sensor. By calculating the first-order difference of the pulp concentration and the first-order difference of the temperature within the time window, the concentration change rate feature and temperature change rate feature are extracted, respectively. These features directly reflect the dynamic intensity of pulp moisture evaporation and gelatinization during the current processing period. This step eliminates transient high-frequency noise interference from industrial field sensors and accurately reconstructs the underlying physicochemical evolution trajectory of the pulp. It should be noted that the first-order difference method used in this step can also be replaced by a Kalman filter state estimation method, which can also achieve the effect of smoothing dynamic feature extraction.
[0022] S102: Construct the motion space and process state space.
[0023] The specific process involves: calculating the ratio of the cumulative water injection volume to the soybean input mass at the current sampling time; then weighting and correcting this ratio based on the soybean moisture content to obtain the actual soybean-to-water ratio at the current sampling time. This actual soybean-to-water ratio characterizes the true liquid-to-solid ratio after deducting the material's own moisture content. Next, the estimated crude soybean protein content at the current sampling time is multiplied by the real-time slurry temperature to obtain the protein thermal effect characteristic at the current sampling time. This protein thermal effect characteristic quantifies the protein activity and dissolution potential under the current thermal environment. Finally, these derived characteristics are combined with the original slurry concentration at the current sampling time and the concentration change rate during the current sampling period to generate a multidimensional state vector characterizing the working conditions at the current sampling time.
[0024] Then, a historical multidimensional state vector sample set is extracted. Deviation standardization is used to eliminate differences in the dimensions of each feature, and an unsupervised clustering algorithm is used to spatially partition the sample set, mapping continuous processing states to a finite number of working condition regions, with each working condition region corresponding to a process state space. By calculating the spatial distance between the state vector at the current sampling moment and the center of each working condition region, the process state space to which the current working condition belongs is determined. This step enables the reinforcement learning network to proactively understand the evolution of the physical environment of the processing system. As an alternative implementation, the feature concatenation operation in this step can also be replaced by a feature weighted fusion network based on a self-attention mechanism to directly obtain a dimensionality-reduced state representation with key features highlighted.
[0025] After obtaining the multidimensional state vector, the agent outputs adjustment commands in the action space, which contains two-dimensional incremental adjustment commands, namely the incremental opening of the water replenishment valve and the incremental opening of the heating steam valve. When the raw pulp concentration is too high, the agent outputs the water replenishment command and simultaneously fine-tunes the heat input using the incremental steam valve command, thereby maintaining the heat balance inside the entire pulping tank and realizing the water-thermal coordinated action space.
[0026] S103: Construct the reward function.
[0027] This step is responsible for generating composite reward signals for the reinforcement learning model. After the action is issued, it ensures that the reward calculation is performed after the physical effective period of the action. Specifically, after the physical effective period ends, the actual raw soy milk concentration is collected, and the deviation between it and the preset standard target concentration is calculated to generate the target deviation reward value. The standard target concentration is derived by the laboratory through back-calculation of the dry matter of historically excellent batches of soy milk. The physical effective period refers to the inherent lag time of the system required from the issuance of the control command to the sensor capturing the corresponding physical quantity change.
[0028] After the physical activation period ends, the system constructs a quantitative reward standard based on the following three dimensions:
[0029] The first is the target deviation reward value, obtained by collecting the real-time raw soy milk concentration at the current sampling moment and calculating the absolute deviation between the real-time raw soy milk concentration and the standard target concentration. The standard target concentration is determined by a laboratory calibration experiment using dry matter drying and weighing of historically high-quality batches of soy milk. The reward calculation uses a negative absolute error model, and the formula is as follows: In the formula, The target deviation reward value at the current sampling time t reflects the effectiveness of the water replenishment and steam regulation commands issued at the current sampling time t in approaching the target. This is the first weighting coefficient; The real-time pulp concentration after the control command is issued at the current sampling time t and after the physical effective period has elapsed; The standard target concentration is defined as . This term aims to force the model to output actions that bring the concentration closer to the target.
[0030] The second is the concentration fluctuation penalty value, which is obtained as follows: Using the end time of the physical effective period as a benchmark, the concentration change rate within that time and the three consecutive sampling periods preceding it is extracted. The end time of the physical effective period refers to the time one physical effective period later than the current sampling time. If the absolute value of any concentration change rate exceeds a fluctuation threshold set based on historical stable data, a fluctuation penalty is applied; if it does not exceed the threshold, the penalty is zero. This fluctuation threshold is set by extracting the concentration change rate data from the stable operation phase of historical qualified batches and statistically calculating it according to the three-times-standard-deviation principle. It should be noted that the formula for calculating the concentration fluctuation penalty value is: In the formula, This is the concentration fluctuation penalty value at the current sampling time t; This is the second weighting coefficient; This represents the highest rate of concentration change across the three sampling periods; The fluctuation threshold; This is a function that maximizes concentration fluctuations; the concentration fluctuation penalty value is designed to constrain the stability of the production process and suppress concentration oscillations caused by over-adjustment.
[0031] The third is the resource coordination penalty value, which is obtained by real-time monitoring of the feedback signals from the opening of the water supply regulating valve and the steam regulating valve. The system presets an 80% upper limit for valve opening as a threshold. If it detects that the opening of both the water supply and steam valves exceeds this threshold at the same time, a negative penalty is triggered, directly assigning a fixed penalty value. The absolute value of the fixed penalty value should be significantly greater than the average expected value of the target deviation reward value within a single sampling period, typically set to 10 to 50 times the normal feedback level. For example, when the concentration deviation reward is in the range of -1 to -5, the fixed penalty value can be set to -200. This logic aims to force the model to recognize the water-thermal offsetting relationship and avoid unnecessary compensation actions that consume high energy.
[0032] Finally, the three rewards are summed to form a total reward, which is then fed back into the reinforcement learning network. This step, through a clear quantification standard, forces the reinforcement learning model to spontaneously converge to the optimal policy path that smoothly approaches the target concentration while minimizing the waste of hydrothermal resources.
[0033] To address the risk of model mismatch caused by batch changes in soybean raw materials, this step quantifies the degree of physicochemical deviation between new and old materials. Specifically, it involves extracting the moisture content and crude protein estimate of the current batch of soybeans to construct a feature vector for the current raw material. A global search is performed in the historical process database to identify the set of historical data with the highest cosine similarity, and this set is extracted as the historical raw material feature vector. To quantify the severity of the new raw material's deviation from the known process model, the similarity calculation must satisfy the following relationship: In the formula, This represents the numerical similarity between the current raw material feature vector and the historical raw material feature vectors. This is the current raw material feature vector. This is a historical raw material feature vector matched against the database. The lower the similarity value, the more the current soybean's water absorption and swelling characteristics and protein colloidal rate deviate from historical experience. This step accurately identifies the potential source of disturbances causing drastic fluctuations in pulp concentration.
[0034] S104: Apply random perturbations to new raw material inputs to drive adaptive evolution of the strategy.
[0035] The system receives the similarity score from step S103. Specifically, the difference between 1 and the highest cosine similarity is taken as the shortest cosine distance. When the calculated shortest cosine distance is greater than a preset raw material difference threshold, the system actively increases the exploration rate parameter of the action output space. In this embodiment, the raw material difference threshold is 0.15. Based on the initial prediction increment commands for the water supply valve and the initial prediction increment commands for the steam valve output by the reinforcement learning network, small amplitude opening and closing test signals are superimposed respectively. The amplitude range of the small amplitude opening and closing test signals is constrained to within 1% to 3% of the full range of the regulating valve. In this embodiment, 2% can be selected. The actual response of the concentration change rate of the new raw material under the test signal is collected and stored in the experience playback pool. The data stream containing the new characteristics in the pool is called concurrently to perform gradient updates of the neural network weights. During the exploration phase, the system continuously monitors the process safety boundaries: if the real-time pulp concentration deviation exceeds 15% of the standard target concentration, or if the concentration change rate triggers an oscillation alarm for 3 consecutive seconds, the system immediately forces the reinforcement learning network to suspend, switching the control of the actuators to a single-loop proportional-integral (PI) conservative control mode. In this mode, water replenishment and steam regulation are decoupled, and the system maintains basic operation using preset low-gain PI parameters. Once the number of new material samples accumulated in the experience playback pool reaches the network update threshold (e.g., 1000 sets) and the training loss function converges, the system automatically switches back from PI mode to reinforcement learning control mode.
[0036] S105: Perform dead-zone filtering on the opening increment command to convert it into a target electrical signal, and apply an analog signal to the control valve positioner to complete the physical execution closed loop.
[0037] After receiving the incremental command, the controller adds it to the current absolute opening of the pneumatic valve to obtain the target absolute opening. The absolute value of the difference between the target opening and the current absolute opening is compared using the internal physical dead zone range. If the absolute value of the difference is less than the physical dead zone threshold, the update is discarded and the existing command remains unchanged. If the absolute value of the difference is greater than or equal to the physical dead zone threshold, the target absolute opening is converted from digital to analog signal. In this embodiment, the physical dead zone threshold is 1.0% of the full stroke of the control valve.
[0038] For example, the current valve opening is 50.0%. The physical dead zone is set to 1.0%. In the first case, the reinforcement learning network calculates an incremental instruction of +0.3%. The target opening (50.3%) minus the current opening (50.0%) equals 0.3%. Since 0.3% < 1.0%, which is less than the physical dead zone threshold, the instruction is discarded, and the valve remains stationary. In the second case, the reinforcement learning network calculates an incremental instruction of +1.5%. Since 1.5% > 1.0%, which is greater than the physical dead zone threshold, the instruction is executed, adjusting the opening to 51.5%.
[0039] The system receives analog electrical signals, which are then fed directly into the positioner of the pneumatic diaphragm regulating valve via an industrial cable. The positioner drives the air supply circuit to change the valve core position, thereby synchronously altering the flow rate of clean water and heating steam entering the cooking system. The resulting new round of hydrothermal mixing and concentration is captured again by the multi-dimensional sensor array in the next sampling cycle and returned to the data access stage, thus achieving a closed-loop process for soybean product production quality.
[0040] The present invention also provides a raw pulp concentration control system in the processing of soy products. The system includes a processor and a memory. The memory stores computer program instructions. When the computer program instructions are executed by the processor, a raw pulp concentration control method in the processing of soy products according to the first aspect of the present invention is implemented.
[0041] The system also includes other components well known to those skilled in the art, such as communication buses and communication interfaces, the settings and functions of which are known in the art and therefore will not be described in detail here.
[0042] In this invention, the aforementioned memory can be any tangible medium containing or storing a program that can be used or combined with an instruction execution system, apparatus, or device. For example, a computer-readable storage medium can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc., or any other medium that can be used to store desired information and can be accessed by an application, module, or both. Any such computer storage medium can be part of a device or accessible to or connected to a device. Any application or module described in this invention can be implemented using computer-readable / executable instructions stored or otherwise maintained on such a computer-readable medium.
Claims
1. A method for controlling the concentration of raw slurry in a soy product processing process, characterized by, include: The characteristics of soybean raw materials in the current production batch and the concentration and temperature of the slurry in the boiling chamber are obtained and constructed into a multi-dimensional state vector. A state space and an action space are constructed based on historical multidimensional state vectors. The state space is then input into a reinforcement learning network to obtain the opening adjustment commands for the water supply regulating valve and the steam regulating valve. The total reward after issuing the opening adjustment command is constructed and fed back to the reinforcement learning network. The total reward is the sum of the target deviation reward value, the concentration fluctuation penalty value, and the resource coordination penalty value. The target deviation reward value reflects the closeness between the concentration and the target value. The concentration fluctuation penalty value reflects the stability of the concentration during the production process. The opening of the water supply regulating valve and the steam regulating valve is detected. If the opening of both valves is greater than the set opening threshold at the same time, the resource coordination penalty value is applied.
2. The method for controlling the concentration of raw slurry in a soy product processing process according to Claim 1, characterized by, Construct a multi-dimensional state vector, including: The soybean raw material characteristics in the multidimensional state vector include soybean-to-water ratio characteristics and protein thermal effect characteristics; the multidimensional state vector also includes the concentration change rate in the current sampling period.
3. The method for controlling the concentration of raw pulp in the processing of soybean products according to claim 1, characterized in that, The method for obtaining the target deviation reward value includes: The absolute difference between the real-time slurry concentration at the current sampling time and the standard target concentration is calculated, and the absolute difference is corrected based on the first weighting coefficient. The negative value of the correction result is used as the target deviation reward value.
4. The method for controlling the concentration of raw soy pulp during soybean product processing according to claim 1, characterized in that, The method for obtaining the concentration fluctuation penalty value includes: The concentration fluctuation penalty value is obtained by comparing the maximum absolute value of the concentration change rate within the current sampling time and the preset number of sampling periods before it with the preset fluctuation threshold.
5. The method for controlling the concentration of raw pulp in the processing of soybean products according to claim 1 or 4, characterized in that, The method also includes an adaptive processing mechanism for new raw materials, specifically: Obtain the protein and moisture content of a new batch of soybean raw materials to form the current raw material feature vector; Retrieve and extract the set of historical raw material feature vectors with the highest cosine similarity to the current raw material feature vector from the historical raw material feature vector database; When the highest cosine similarity is lower than a set threshold, the system enters exploration mode and adds random perturbations to the action commands.
6. The method for controlling the concentration of raw pulp in the processing of soybean products according to claim 5, characterized in that, Entering exploration mode includes: Continuously monitor the deviation between the slurry concentration and the standard target concentration; When the deviation exceeds the set threshold, the control of the reinforcement learning model is paused and switched to PID mode. After collecting a set number of new samples and updating the model in PID mode, it automatically switches back to reinforcement learning control mode.
7. The method for controlling the concentration of raw soy pulp during soybean product processing according to claim 6, characterized in that, The process of issuing the opening adjustment command also includes: The incremental adjustment command is added to the current valve opening to obtain the target absolute opening. Calculate the absolute value of the difference between the target absolute opening and the current valve opening. If the absolute value of the difference is less than the preset physical dead zone threshold, then discard the current opening adjustment command; otherwise, convert the target absolute opening into an analog electrical signal and send it to the valve positioner.
8. The method for controlling the concentration of raw soy pulp during soybean product processing according to claim 5, characterized in that, Add random perturbations to the action commands, including: Based on the initial prediction increment commands for the water supply valve and the initial prediction increment commands for the steam valve output by the reinforcement learning network, small amplitude opening and closing test signals are superimposed on them respectively.
9. The method for controlling the concentration of raw soy pulp during soybean product processing according to claim 1, characterized in that, Apply a resource coordination penalty value, wherein the resource coordination penalty value is a fixed value.
10. A raw pulp concentration control system in the processing of soybean products, comprising a processor and a memory, characterized in that, The memory stores a computer program, and the processor executes the computer program to implement the method for controlling the concentration of raw pulp in the soybean product processing process as described in any one of claims 1-9.