Sewage recycling monitoring method and system based on industrial big data
By constructing a hybrid-driven digital twin system and reinforcement learning agent, the risks and model reliability of the wastewater recycling system are dynamically assessed, solving the problems of system safety and reliability under severe operating conditions, and realizing safe and intelligent collaborative optimization in scenarios such as power plants.
Patent Information
- Application Number
- CN202511344751.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2045-09-19
AI Technical Summary
When faced with drastic fluctuations in operating conditions, existing wastewater recycling systems cannot effectively assess the risk of instantaneous scaling or corrosion using traditional monitoring methods. Artificial intelligence models also suffer from decreased reliability in predicting unseen operating conditions, leading to insufficient system safety and reliability.
A hybrid-driven digital twin system is constructed, which generates a set of candidate control actions through a reinforcement learning agent, calculates the expected cumulative reward using a dynamic saturation index and time-varying confidence, selects the optimal control command, and combines online updates of the reinforcement learning agent to achieve dynamic risk assessment and model reliability assessment.
It enhances the safety assurance capabilities of physical systems under severe operating conditions, ensures the long-term decision reliability and health of artificial intelligence models, and achieves synergistic optimization of short-term safety and long-term intelligence.
Smart Images

Figure CN120831929B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent monitoring and control of industrial wastewater treatment systems, specifically to a wastewater recycling monitoring method and system based on industrial big data. Background Technology
[0002] In industrial production, especially in scenarios requiring deep peak shaving such as power plants, the associated wastewater recovery systems, such as Zero Liquid Discharge (ZLD) systems, face the challenge of drastic fluctuations in operating conditions. Existing monitoring and control methods have the following shortcomings:
[0003] Physical safety risk assessment is lagging: Traditional monitoring methods, such as relying on the Langerile saturation index for water quality assessment, are mainly applicable to stable operating conditions. In dynamic processes with drastic load changes, these static or quasi-static indicators cannot effectively predict and assess the risk of instantaneous scaling or corrosion caused by sudden changes in operating conditions, thus making it difficult to ensure the short-term operational safety of the physical system.
[0004] The reliability of intelligent model decision-making is not guaranteed: When using artificial intelligence, such as reinforcement learning, for control decisions, the limitations of the model itself become a technical challenge. When the system enters a severe operating condition that is rarely seen or has never appeared in the model's training data, the predictive reliability of the artificial intelligence model will decrease significantly. If the system fails to perceive the reduced predictive ability of the model and blindly adopts its decisions, it may lead to erroneous control operations or even system failure. Existing technologies lack a mechanism to evaluate and address the confidence issues of artificial intelligence models' decisions under specific operating conditions in real time, making it difficult to maintain their long-term decision-making capabilities.
[0005] The information disclosed in the background section above is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this invention is to provide a wastewater recycling monitoring method and system based on industrial big data to solve the problems mentioned in the background art.
[0007] The technical solution of the present invention includes: acquiring real-time operating condition data of an industrial system;
[0008] A hybrid-driven digital twin system that maps real-time operating data to construction and acquisition generates a set of candidate control actions through a reinforcement learning agent;
[0009] Based on the hybrid-driven digital twin system, a virtual simulation is performed on each candidate control action in the candidate control action set to generate a simulation state sequence corresponding to each candidate control action.
[0010] Based on each inferred state sequence, calculate the dynamic saturation index characterizing the security risk of the physical system, and calculate the time-varying confidence level characterizing the reliability of the model prediction.
[0011] Based on the dynamic saturation index and time-varying confidence, the expected cumulative reward of each candidate control action is calculated, and the candidate control action that maximizes the expected cumulative reward is selected as the optimal control instruction.
[0012] The system records empirical data generated after executing optimal control commands in the physical system, and updates the reinforcement learning agent online based on this empirical data.
[0013] Preferably, a hybrid-driven digital twin system is constructed, including:
[0014] Construct a mechanistic model based on chemical reaction kinetics and fluid dynamics;
[0015] Construct a data-driven error compensation model to learn and compensate for the error between the mechanistic model and the monitoring data of the physical system in real time;
[0016] By integrating the output of the mechanistic model with the compensation signal of the error compensation model, a high-fidelity digital twin state is generated.
[0017] Preferably, the calculation of the dynamic saturation index includes:
[0018] Obtain the traditional Langerier saturation index;
[0019] Obtain the load change rate, which reflects the degree of fluctuation in operating conditions;
[0020] A dynamic saturation index is generated by combining the Langerier saturation index with the load change rate.
[0021] Preferably, calculating the time-varying confidence level includes:
[0022] Obtain the predicted values of preset key parameters by the data-driven error compensation model;
[0023] Obtain the simulated values of preset key parameters in the simulation;
[0024] Time-varying confidence scores are generated by measuring the deviation between predicted and simulated values.
[0025] Preferably, the expected cumulative reward is calculated, including:
[0026] When the dynamic saturation index is within the preset safety range, a positive physical security reward is generated; when the dynamic saturation index is outside the preset safety range, a negative physical security reward is generated.
[0027] Generate model confidence rewards that are positively correlated with time-varying confidence levels;
[0028] Dynamically calculated adaptive weights are used to balance physical security rewards and model confidence rewards;
[0029] The physical security reward and the model confidence reward are weighted and summed based on adaptive weights to generate the expected cumulative reward.
[0030] Preferably, the adaptive weights are calculated dynamically, including:
[0031] The normalized physical risk score is calculated based on the degree of deviation between the dynamic saturation index and the preset target value.
[0032] Based on the normalized physical risk score and time-varying confidence level, an adaptive weight is calculated using a preset function. The increase in the normalized physical risk score will increase the weight corresponding to the physical security reward.
[0033] Wastewater recycling monitoring methods based on industrial big data record empirical data, including:
[0034] Calculate the actual reward based on the new system state after executing the optimal control command;
[0035] The system state before execution, the optimal control command executed, the actual reward obtained after execution, and the new system state after execution are stored as experience tuples.
[0036] Wastewater recycling monitoring system based on industrial big data includes:
[0037] The data acquisition module is used to acquire real-time operating data of industrial systems;
[0038] The digital twin and action generation module is used to construct and acquire a hybrid-driven digital twin system that maps real-time working condition data, and to generate a set of candidate control actions through a reinforcement learning agent.
[0039] The virtual simulation module is used to perform virtual simulation of each candidate control action in the candidate control action set based on the hybrid-driven digital twin system, and generate the simulation state sequence corresponding to each candidate control action.
[0040] The risk and reliability calculation module is used to calculate the dynamic saturation index that characterizes the security risk of the physical system and the time-varying confidence level that characterizes the reliability predicted by the model, based on each simulated state sequence.
[0041] The decision module is used to calculate the expected cumulative reward of each candidate control action based on the dynamic saturation index and time-varying confidence, and select the candidate control action that maximizes the expected cumulative reward as the optimal control instruction.
[0042] The online update module is used to record the empirical data generated after executing the optimal control instructions in the physical system, and to update the reinforcement learning agent online based on the empirical data.
[0043] This invention provides an improved wastewater recycling monitoring method and system based on industrial big data, which has the following improvements and advantages compared with the prior art:
[0044] 1. This invention abandons the traditional static water quality assessment and innovatively proposes a dynamic saturation index, which makes risk assessment no longer static, but can dynamically and sensitively capture potential scaling or corrosion risks caused by sudden load changes. By predicting the changes of this index in virtual simulation, the system can avoid control actions that will lead to physical risks in advance, which greatly enhances the safety assurance capability of the physical system under drastic fluctuation conditions.
[0045] 2. This invention designs a time-varying confidence index to evaluate the predictive credibility of the artificial intelligence model under the current working conditions in real time, enabling the system to have self-reflection ability and actively avoid decisions that may damage the long-term performance of the model, thus ensuring the long-term health of the artificial intelligence model and the continuous reliability of the decisions.
[0046] 3. An adaptive reward decision-making framework was established. Instead of using fixed weights to balance physical security and model confidence, it dynamically calculates the physical security reward weight and the model confidence reward weight through normalized physical risk scores and time-varying confidence. This dynamic balancing mechanism enables the system to intelligently adjust its decision focus according to the real-time situation of current risks and uncertainties, and ultimately achieve synergistic optimization of short-term security and long-term intelligence.
[0047] 4. After the optimal control command is executed in the physical system, the system will structurally store the state before execution, the executed command, the real reward obtained, and the new state after execution as experience tuples. These experience data derived from real physical world interactions are used to update the reinforcement learning agent online. This ensures that the agent's policy can continuously and rationally evolve in a direction that is more adapted to the real working conditions, thus solving the closed-loop problem of self-optimization of the decision model. Attached Figure Description
[0048] The present invention will be further explained below with reference to the accompanying drawings and embodiments:
[0049] Figure 1 This is a flowchart of the system of the present invention. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.
[0051] Example 1
[0052] Please see Figure 1 This invention provides a wastewater recycling monitoring method and system technical solution based on industrial big data, including: acquiring real-time operating data of industrial systems;
[0053] A hybrid-driven digital twin system that maps real-time operating data to construction and acquisition generates a set of candidate control actions through a reinforcement learning agent;
[0054] Based on the hybrid-driven digital twin system, a virtual simulation is performed on each candidate control action in the candidate control action set to generate a simulation state sequence corresponding to each candidate control action.
[0055] Based on each inferred state sequence, calculate the dynamic saturation index characterizing the security risk of the physical system, and calculate the time-varying confidence level characterizing the reliability of the model prediction.
[0056] Based on the dynamic saturation index and time-varying confidence, the expected cumulative reward of each candidate control action is calculated, and the candidate control action that maximizes the expected cumulative reward is selected as the optimal control instruction.
[0057] The system records empirical data generated after executing optimal control commands in the physical system, and updates the reinforcement learning agent online based on this empirical data.
[0058] In this embodiment, a wastewater recycling monitoring method based on industrial big data is designed to address the technical challenge of coordinating short-term physical system safety with maintaining the long-term decision-making capability of AI models under drastically fluctuating operating conditions such as deep peak shaving in power plants. The method acquires real-time operating data from the industrial ZLD system upon startup, providing real-time data input for the entire decision-making process. It constructs a hybrid-driven digital twin system and drives a reinforcement learning agent to generate a series of candidate control actions. The core advantage of this method lies in that it does not directly apply an action to the physical world, but rather initiates a virtual deduction process of pre-action thinking. Within the digital twin environment, the method performs high-speed processing on each candidate action. The method simulates and anticipates potential future states. By introducing and calculating two key indicators—dynamic saturation index and time-varying confidence level—it can quantitatively evaluate the consequences of each candidate action from both physical safety and model reliability dimensions. Based on this evaluation, the optimal control instruction is selected by calculating the expected cumulative reward, ensuring that the decision is not only safe in the short term but also beneficial to the long-term health of the AI model. After the optimal instruction is executed in the physical system, the system records complete empirical data for online updates to the reinforcement learning agent. This forms a complete closed loop of decision-making, execution, learning, and evolution, enabling the system to continuously enhance its capabilities.
[0059] Example 2
[0060] Building a hybrid-driven digital twin system includes:
[0061] Construct a mechanistic model based on chemical reaction kinetics and fluid dynamics;
[0062] Construct a data-driven error compensation model to learn and compensate for the error between the mechanistic model and the monitoring data of the physical system in real time;
[0063] By integrating the output of the mechanistic model with the compensation signal of the error compensation model, a high-fidelity digital twin state is generated.
[0064] In this embodiment, the process of constructing a hybrid-driven digital twin system provides a solid simulation foundation for achieving high-precision virtual simulation. This system, by constructing a mechanistic model based on chemical reaction kinetics and fluid dynamics, provides a first-principles-compliant underlying description of the ZLD system's operating laws, ensuring that the digital twin system possesses basic physical consistency under any operating condition. However, the mechanistic model alone is insufficient to capture the full complexity of real industrial environments. Therefore, this solution further constructs a data-driven error compensation model, for example, using a long short-term memory network. This model can learn and quantify the deviation between the mechanistic model output and the actual monitoring data of the physical system in real time.
[0065] The error compensation model can employ a structure containing two stacked long short-term memory networks; the input to the model is a vector of real-time monitoring data from the physical system, which may include, for example, the current unit load. Influent flow rate, pH value, temperature, and key ion concentration, such as , The actual measured values, etc.; the model output is the mechanistic model's predicted values for preset key parameters at the next time step, for example... Error prediction of concentration The fusion process is achieved by adding the theoretical output of the mechanistic model to the compensation signal of the error compensation model, thereby generating a high-fidelity digital twin state for the next time step. This process can be represented by the following formula:
[0066] ;
[0067] in, The mechanism model at time The output value, The error compensation model at time 10:00 The output compensation signal, in this way, allows the digital twin system to retain the physical consistency of the mechanism model, while using the data model to correct deviations from the actual working conditions in real time;
[0068] By fusing the theoretical output of the mechanistic model with the real-time compensation signal of the error compensation model, the system can generate a high-fidelity digital twin state that both follows physical laws and closely matches real working conditions. This hybrid-driven architecture greatly improves the accuracy of the virtual simulation environment, providing unprecedented reliability and precision for subsequent reinforcement learning agents to find the optimal control strategy.
[0069] Example 3
[0070] Calculating the dynamic saturation index includes:
[0071] Obtain the traditional Langerier saturation index;
[0072] Obtain the load change rate, which reflects the degree of fluctuation in operating conditions;
[0073] A dynamic saturation index is generated by combining the Langerier saturation index with the load change rate.
[0074] Calculating time-varying confidence includes:
[0075] Obtain the predicted values of preset key parameters by the data-driven error compensation model;
[0076] Obtain the simulated values of preset key parameters in the simulation;
[0077] Time-varying confidence scores are generated by measuring the deviation between predicted and simulated values.
[0078] In this embodiment, the establishment of a dual evaluation index system provides a precise mathematical tool for quantifying the two core dimensions of physical security and model confidence.
[0079] In calculating the dynamic saturation index, this scheme overcomes the limitations of traditional static water quality assessment. While traditional indicators such as the Langerier saturation index are acceptable under stable operating conditions, they cannot effectively assess the instantaneous scaling or corrosion risk of water quality in scenarios with drastic load fluctuations, such as deep peak shaving of generator units. To address this problem, this invention proposes a dynamic saturation index. Its core motivation lies in quantifying the dynamic factor of operating condition fluctuations and incorporating it into the risk assessment model; the index is calculated using the following formula:
[0080] ;
[0081] in, For at any time The dynamic saturation index, which is a dimensionless number; For at any time The traditional Langerile saturation index is calculated in accordance with open standards in the water treatment industry and is also a dimensionless number. This represents the load percentage of the generator set, reflecting the core operating conditions. The load change rate directly characterizes the severity of fluctuations in operating conditions, with the dimension being time. ; This is the load change rate weighting coefficient, a key parameter introduced in this invention. Its value is determined through regression analysis of historical operating data collected before deployment. The purpose is to quantify the impact of drastic load fluctuations on water quality stability, and its dimension is time. This ensures Dimensionless property;
[0082] The calibration process is as follows: Collect historical datasets containing fluctuations in different operating conditions and load changes. The datasets should include generator loads. and Langerier saturation index Define a dimensionless risk indicator, such as a key water quality parameter like turbidity or calcium ion concentration, the probability of it exceeding the normal range, or the relative magnitude of its exceeding the normal range, for example: actual value - upper limit of normal / upper limit of normal. Construct a regression model to... Using the risk index defined above as the dependent variable, we fit the data; the resulting regression coefficient is the calibration value of λ. This method provides an objective and reproducible basis for determining λ.
[0083] During the virtual simulation, the system uses this formula to calculate the possible futures for each candidate action. The sequence; this index can more sensitively capture potential physical risks caused by sudden load changes. When its value exceeds the preset safety threshold, the system can predict the danger. This design makes risk assessment move from static to dynamic, which greatly improves the ability to ensure the physical safety of the ZLD system under extreme working conditions.
[0084] Regarding the calculation of time-varying confidence, this scheme provides a quantitative basis for assessing the health of the AI model itself. When the AI model faces severe conditions that it has never seen or rarely encountered in its training data, the reliability of its predictions will decrease significantly. If the system is unaware of this and blindly trusts the AI's output, it may lead to catastrophic decisions. Time-varying confidence... The design motivation is to enable the system to evaluate the prediction reliability of the AI model in real time under the current operating conditions. To address the problem that different key parameters, such as concentration and pH, have different dimensions and cannot be directly summed, this confidence level is generated by calculating the weighted sum of the relative deviations of each parameter, as shown in the following formula:
[0085] ;
[0086] in, The time-varying confidence level at time t is a dimensionless number. It is a data-driven error compensation model that predicts the value of a key parameter i, such as the Ca²⁺ concentration, at a future moment. This is the high-fidelity simulated value of the key parameter in the virtual simulation, and both have the same dimensions; the relative deviation of each parameter is calculated. This achieves the normalization of deviations of parameters with different dimensions. Since the numerator and denominator have the same dimensions, the ratio is dimensionless; N is the number of key water quality parameters being evaluated. These are dimensionless weighting coefficients for different key parameters, with a total of 1. Their values can be preset by domain experts based on the degree of influence of each parameter on the risk of scaling or corrosion in the system. For example, the Ca²⁺ concentration, which has a more direct impact on scaling, can be given a higher weight. Alternatively, the weight can be objectively determined by using sensitivity analysis during the offline training phase to quantify the impact of each parameter deviation on the final control effect.
[0087] In virtual simulations, the system calculates the difference between the model's predicted values and the twin simulation values by measuring the normalized deviation. ; a low The indicator explicitly warns the system that the current candidate control action will lead the system into an unknown area that the AI model is unfamiliar with and where its predictive ability will be reduced. The introduction of this indicator enables the system to consider not only the physical consequences but also the long-term impact on the AI model's own capabilities when making decisions, thereby proactively avoiding control behaviors that will pollute learning data and impair long-term intelligence.
[0088] Example 4
[0089] Calculate the expected cumulative reward, including:
[0090] When the dynamic saturation index is within the preset safety range, a positive physical security reward is generated; when the dynamic saturation index is outside the preset safety range, a negative physical security reward is generated.
[0091] Generate model confidence rewards that are positively correlated with time-varying confidence levels;
[0092] Dynamically calculated adaptive weights are used to balance physical security rewards and model confidence rewards;
[0093] The physical security reward and the model confidence reward are weighted and summed based on adaptive weights to generate the expected cumulative reward.
[0094] Dynamically calculating adaptive weights, including:
[0095] The normalized physical risk score is calculated based on the degree of deviation between the dynamic saturation index and the preset target value.
[0096] Based on the normalized physical risk score and time-varying confidence level, an adaptive weight is calculated using a preset function. The increase in the normalized physical risk score will increase the weight corresponding to the physical security reward.
[0097] In this embodiment, the calculation of the expected cumulative reward and the dynamic adjustment of adaptive weights together constitute the core of the decision-making mechanism of this invention, reconciling and unifying the contradiction between short-term security and long-term intelligence through a sophisticated mathematical framework. The goal of reinforcement learning is to maximize the cumulative reward; therefore, the design of the reward function directly determines the agent's behavioral tendencies. To enable the agent to simultaneously consider physical security and model health, a reward function that integrates these two objectives must be designed. Furthermore, the priorities of these two objectives change dynamically under different circumstances; to achieve intelligent adjustment of decision-making tendencies, the system introduces adaptive weights to replace fixed weights. The total reward function is composed of a weighted average of physical security rewards and model confidence rewards.
[0098] ;
[0099] in, It is a moment Total reward. Physical security reward. The calculation method is as follows:
[0100] ;
[0101] in, It is a dynamic saturation index; These are the safe range thresholds of the dynamic saturation index. These two values are the core constraints to ensure the safety of the physical system. They are derived from the material tolerance of the key equipment protected by the ZLD system, such as heat exchangers, industry safety regulations, and historical excellent operating data to comprehensively set safe operating boundaries. They are a reflection of the knowledge of domain experts. Adaptive weights used to balance physical security rewards; Physical security reward; Adaptive weights used to balance the confidence reward of the model; Model confidence reward; and These are the hyperparameters for physical security rewards. These two values are determined during the offline training phase of the reinforcement learning model through debugging and hyperparameter optimization methods such as grid search and Bayesian optimization. Much larger This is used to impose a large negative penalty to prevent any behavior from exceeding the safety boundary; model confidence reward The calculation method is as follows:
[0102] ;
[0103] in, It is time-varying confidence level; It is the scaling factor for the model confidence reward, and this factor is related to... , As hyperparameters of reinforcement learning models, they are jointly determined through hyperparameter optimization during the offline training phase of the model. Their main function is to balance the numerical magnitude of physical security rewards and model confidence rewards.
[0104] Core adaptive weights and A normalized physical risk score is calculated by dynamically generating a risk perception mechanism.
[0105] ;
[0106] in, This is the optimal target value for the dynamic saturation index, which is usually set to 0. The physical risk level is mapped dimensionlessly to the [0, 1] interval; the weights are calculated using functions of the Softmax class. The upper and lower limits of the safe range for the dynamic saturation index:
[0107] ;
[0108] in, and It is a dimensionless sensitivity coefficient determined through offline optimization, used to adjust the influence of risk and confidence on weight allocation;
[0109] When setting these two coefficients, the principle of prioritizing physical safety should be followed; parameters The value should be significantly greater than The value, for example, can be set to... and This setup ensures that the normalized physical risk score is accurate. When it increases, its exponential term in the Softmax function It will increase dramatically, thus affecting the weight of physical security rewards. The coefficients rapidly approach 1, forcing the agent to prioritize avoiding physical risks. The specific values of these two coefficients can be evaluated during the offline training phase by systematically testing a series of combinations in a simulation environment. The system's decision-making performance under different risk and uncertainty scenarios can be assessed, and the set of values that achieves the best decision-making balance can be selected as the final configuration.
[0110] The application of this reward and weighting mechanism enables the system's decision-making to exhibit a high degree of intelligence and adaptability; during the deduction process, when a candidate action leads to a predicted... Deviation from the target value, making When it rises, The value of [something] will grow exponentially, thus dominating the total reward function and forcing the agent to abandon the high-risk action; conversely, when the system is in a state with a high safety margin, [something] will grow exponentially. When approaching 0, As their influence increases, AI will be more inclined to choose those that bring higher model confidence. This design addresses the closed-loop problem of decision parameter sources, enabling the system to autonomously and dynamically adjust its decision focus based on the current level of risk and uncertainty while pursuing optimal control, ultimately finding a sustainable optimization path between safety and intelligence.
[0111] Record experiential data, including:
[0112] Calculate the actual reward based on the new system state after executing the optimal control command;
[0113] The system state before execution, the optimal control command executed, the actual reward obtained after execution, and the new system state after execution are stored as experience tuples.
[0114] In this embodiment, the recording method of experience data constructs a structured learning loop for model self-evolution. After the optimal control command is issued to the physical system's actuator and completes the operation, the system immediately collects the new system state and calculates the actual reward obtained from this decision based on the real feedback. The key to this step is that it is not simply data archiving, but rather organizing a complete decision-action-result interaction process into a standardized experience tuple according to a predetermined data structure. ;in, It is the system state before the action is executed. It is the optimal control instruction to be executed. The reward is calculated based on the actual results. This is the new real state that the system enters; by storing these experience tuples containing causal relationships in the experience replay unit, the system provides high-quality structured data for the online updating of reinforcement learning agents; this mechanism ensures that agents can learn from real interactions that occur in the physical world, rather than relying solely on virtual inference, so that their strategies can continuously and causally evolve in a direction that is more adapted to real working conditions, thus completely solving the closed-loop problem of the self-evolution of decision models.
[0115] Example 6
[0116] The wastewater recycling monitoring system based on industrial big data includes: a data acquisition module, used to acquire real-time operating data of the industrial system;
[0117] The digital twin and action generation module is used to construct and acquire a hybrid-driven digital twin system that maps real-time working condition data, and to generate a set of candidate control actions through a reinforcement learning agent.
[0118] The virtual simulation module is used to perform virtual simulation of each candidate control action in the candidate control action set based on the hybrid-driven digital twin system, and generate the simulation state sequence corresponding to each candidate control action.
[0119] The risk and reliability calculation module is used to calculate the dynamic saturation index that characterizes the security risk of the physical system and the time-varying confidence level that characterizes the reliability of the model prediction based on each simulated state sequence.
[0120] The decision module is used to calculate the expected cumulative reward of each candidate control action based on the dynamic saturation index and time-varying confidence, and select the candidate control action that maximizes the expected cumulative reward as the optimal control instruction.
[0121] The online update module is used to record the empirical data generated after executing the optimal control instructions in the physical system, and to update the reinforcement learning agent online based on the empirical data.
[0122] In this embodiment, the reinforcement learning agent adopts the deep deterministic policy gradient algorithm. The reason for choosing the DDPG algorithm is that the control of the sewage recycling system, such as adjusting the frequency of the dosing pump and controlling the opening of the valve, belongs to the continuous action space problem. As a model-free algorithm based on the Actor-Critic framework, DDPG can effectively learn the strategy for making decisions in the continuous action space, making it very suitable for the application scenario of this invention. In this embodiment, the Actor network is responsible for outputting the deterministic control action, and the Critic network is responsible for evaluating the value of the action. The two are iteratively optimized through temporal difference learning.
[0123] This embodiment provides a wastewater recycling monitoring system based on industrial big data. Through a series of tightly coupled modular designs, the aforementioned methods are implemented, forming a complete intelligent control entity. The data acquisition module, serving as the system's real-time data interface, captures operational data in real-time from the industrial site's DCS, PLC, and various sensors. This data is then sent to the digital twin and motion generation module, which integrates high-fidelity simulation environment and control strategy generation functions. On one hand, it constructs a high-fidelity digital twin environment; on the other hand, a built-in reinforcement learning agent proposes multiple possible control schemes. The virtual simulation module, as the computing unit responsible for performing forward simulation to predict the system's dynamic response, simulates the evolution path of each scheme over a future period in parallel and at high speed. The risk and reliability calculation module is responsible for quantifying the simulation results based on preset indicators. The unit employs two core indicators—dynamic saturation index and time-varying confidence level—to rigorously quantify and score each possible future. The decision-making module, as the core computational unit responsible for executing the optimal strategy selection, calculates and selects the control instruction that achieves the best balance between short-term security and long-term intelligence based on the evaluation results and using the aforementioned adaptive reward function. After the instruction is executed, the online update module takes on the responsibility of updating the decision-making model online using real-world interaction data, collecting feedback from the real world, and using this experience to train and improve the reinforcement learning agent in the decision-making module. Through this seamless collaboration and information flow between modules, the entire system forms a complete intelligent closed loop of continuous self-optimization, from perception, simulation, evaluation, decision-making to action and learning, thereby achieving unprecedented robustness and economic benefits in complex industrial environments.
[0124] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A sewage recovery monitoring method based on industrial big data, characterized in that, The method comprises the following steps: acquiring real-time working condition data of an industrial system; constructing a hybrid-driven digital twin system mapped with the acquired real-time working condition data, and generating a candidate control action set through a reinforcement learning agent; based on the hybrid-driven digital twin system, virtually simulating each candidate control action in the candidate control action set to generate a simulation state sequence corresponding to each candidate control action; according to each simulation state sequence, calculating a dynamic saturation index, comprising: acquiring a traditional Langley index; acquiring a load change rate reflecting the fluctuation degree of the working condition; combining the Langley index and the load change rate to generate the dynamic saturation index; the calculation formula of the dynamic saturation index is: ; wherein, is a dynamic saturation index, is a traditional Langley index, is a load change rate, is a load change rate weight coefficient; and calculating a time-varying confidence representing the reliability of the model prediction; Based on the dynamic saturation index and the time-varying confidence, the expected cumulative reward is calculated, including: when the dynamic saturation index is in a preset safety interval, a positive physical safety reward is generated, when the dynamic saturation index is outside the preset safety interval, a negative physical safety reward is generated, and the physical safety reward The calculation method is: ; wherein, is a safety interval threshold for the dynamic saturation index; and is a hyperparameter for the physical safety reward; generating a model confidence reward positively correlated with the time-varying confidence; dynamically calculating an adaptive weight for balancing the physical safety reward and the model confidence reward; weighting and summing the physical safety reward and the model confidence reward based on the adaptive weight to generate an expected cumulative reward; the calculation formula of the expected cumulative reward is: ; wherein, is a desired cumulative reward, is a physical safety reward, is an adaptive weight for the physical safety reward, is a model confidence reward, is an adaptive weight for the model confidence reward; and selecting a candidate control action that maximizes the desired cumulative reward as the optimal control instruction. recording experience data generated after executing the optimal control instruction in the physical system, and updating the reinforcement learning agent online according to the experience data. 2.The industrial big data-based sewage recovery monitoring method of claim 1, wherein, The method for constructing the hybrid-driven digital twin system comprises the following steps: constructing a mechanism model based on chemical reaction kinetics and fluid dynamics; constructing a data-driven error compensation model for learning and compensating errors between the mechanism model and the monitoring data of the physical system in real time; fusing the output of the mechanism model and the compensation signal of the error compensation model to generate a high-fidelity digital twin state. 3.The industrial big data-based sewage recovery monitoring method of claim 1, wherein, calculating the time-varying confidence, comprising: acquiring the predicted value of the data-driven error compensation model for the preset key parameter; acquiring the simulation value of the preset key parameter in the simulation; generating the time-varying confidence by measuring the deviation between the predicted value and the simulation value. 4.The industrial big data-based sewage recovery monitoring method of claim 1, wherein, dynamically calculating the adaptive weight, comprising: based on the deviation of the dynamic saturation index and the preset target value, calculating a normalized physical risk score; based on the normalized physical risk score and the time-varying confidence, calculating the adaptive weight using a preset function, wherein the increase of the normalized physical risk score will increase the weight corresponding to the physical safety reward. 5.The industrial big data-based sewage recovery monitoring method of claim 1, wherein, recording experience data, comprising: calculating the actual reward according to the new system state after executing the optimal control instruction; storing the system state before execution, the optimal control instruction executed, the actual reward obtained after execution, and the new system state after execution as an experience tuple.
6. The wastewater recovery monitoring system based on industrial big data according to any one of claims 1-5, characterized in that, The method comprises the following steps: a data acquisition module for acquiring real-time working condition data of an industrial system; a digital twin and action generation module for constructing a hybrid-driven digital twin system mapped with the acquired real-time working condition data, and generating a candidate control action set through a reinforcement learning agent; a virtual simulation module for virtually simulating each candidate control action in the candidate control action set based on the hybrid-driven digital twin system to generate a simulation state sequence corresponding to each candidate control action; A risk and reliability calculation module is configured to calculate a dynamic saturation index according to each deduction state sequence, including: obtaining a traditional Langley saturation index; obtaining a load change rate reflecting the fluctuation degree of the working condition; combining the Langley saturation index and the load change rate to generate a dynamic saturation index, and calculating a time-varying confidence degree representing the model prediction reliability; A decision module is configured to calculate an expected cumulative reward based on the dynamic saturation index and the time-varying confidence degree, including: generating a positive physical safety reward when the dynamic saturation index is within a preset safety interval, and generating a negative physical safety reward when the dynamic saturation index is outside the preset safety interval; generating a model confidence reward positively correlated with the time-varying confidence degree; dynamically calculating an adaptive weight for balancing the physical safety reward and the model confidence reward; performing weighted summation on the physical safety reward and the model confidence reward based on the adaptive weight to generate the expected cumulative reward; and selecting a candidate control action maximizing the expected cumulative reward as an optimal control instruction; An online updating module is configured to record experience data generated after the optimal control instruction is executed in the physical system, and update the reinforcement learning intelligent agent online according to the experience data.
Citation Information
Patent Citations
Assembly workshop production line construction system based on digital twinning
CN120493592A
Photovoltaic energy storage system power scheduling optimization method based on deep reinforcement learning
CN120582254A