Reverse osmosis energy consumption optimization control method and system based on deep reinforcement learning
Patent Information
- Application Number
- CN202611023939.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-10
- Publication Date
- 2026-08-21
AI Technical Summary
[0005]本申请实施例的目的是提供一种基于深度强化学习的反渗透能耗优化控制方法及系统,解决了现有反渗透控制策略缺乏对工况动态变化的自适应节能调节能力的问题
[0008] Compared to existing technologies, the reverse osmosis energy consumption optimization control method and system provided in this application, by constructing a multi-dimensional state vector that integrates current operating parameters, membrane flux decay coefficient, and feedforward characteristics of water quality fluctuations, enables the control strategy to have a comprehensive perception capability of real-time system operating conditions, membrane element health status, and future feedwater trends. It employs a pre-trained deep reinforcement learning strategy network to infer online the coordinated adjustment of the high-pressure pump and concentrate valve, aiming to reduce unit product water energy consumption while meeting product water quality and quantity constraints. This achieves self-regulation of feedwater quality fluctuations and membrane fouling evolution. It features adaptive and continuous energy saving; at the same time, it uses the reverse osmosis physical mechanism model to calculate the upper limit of the dynamic change of transmembrane pressure difference and the lower limit of the permeate flow rate under the current operating conditions as hard constraint boundaries, and uses the constraint mapping algorithm to project the out-of-bounds initial adjustment amount back into the safe domain with the minimum action correction amount. Thus, while ensuring the structural integrity of the membrane element and the reliability of water supply, it retains the energy-saving optimization intention of the strategy network to the greatest extent, completely solving the technical problem of the separation between safety protection and energy efficiency optimization in traditional control methods, significantly reducing the long-term operating energy consumption of the reverse osmosis system and extending the service life of the membrane element.
Smart Images

Figure CN122613751A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent water treatment control, and in particular to a reverse osmosis energy consumption optimization control method and system based on deep reinforcement learning. Background Technology
[0002] Reverse osmosis, as a core membrane separation technology, is widely used in seawater desalination, advanced industrial wastewater treatment, and ultrapure water production. The energy consumption of a reverse osmosis system is primarily generated by the high-pressure pump, which accounts for a very high proportion of the overall operating cost. In actual production, the raw water quality, temperature, and membrane fouling status exhibit continuous time-varying characteristics, causing the system's optimal operating parameters to dynamically shift. Therefore, how to adjust operating variables such as the high-pressure pump frequency and concentrate regulating valve opening in real time, while meeting the constraints of product water quantity and quality, to ensure the system always operates near its energy-optimal condition, has long been a technical challenge in the reverse osmosis field.
[0003] Currently, the operation and control of reverse osmosis systems mainly rely on manual experience to set fixed operating parameters, or on the use of classic PID controllers to maintain a constant permeate flow rate. Some technical solutions incorporate feedforward compensation into the PID framework to coarsely adjust the high-pressure pump frequency based on changes in feed water quality or temperature, thereby enhancing the response speed to feed water fluctuations. For safety, preset fixed high-pressure protection and low-flow alarm thresholds are typically used. When operating parameters reach these static thresholds, alarms are triggered or the system shuts down to protect the membrane elements from damage. These control strategies constitute the main technical means of current reverse osmosis system operation and management.
[0004] The main shortcomings of existing technologies are as follows: First, control strategies based on fixed parameters or PID control lack the ability to adapt to fluctuations in influent water quality and the evolution of membrane fouling. They cannot proactively adjust the operating point to track the optimal energy consumption range when operating conditions change, resulting in the system operating under conservative conditions for a long time, leading to high energy consumption per unit of produced water. Second, existing solutions lack sufficient state perception dimensions, relying only on limited instrument readings at the current moment for feedback control. They fail to incorporate feedforward information such as membrane flux decay and water quality fluctuation trends into control decisions, resulting in a lack of foresight in optimization strategies and difficulty in making preventative adjustments before membrane fouling intensifies or water quality changes abruptly. Third, the safety assurance mechanism adopts static thresholds and passive shutdown methods, which cannot dynamically generate adaptive safe operating boundaries based on real-time operating conditions. Protection measures and optimization control are disconnected, reducing system availability and failing to maximize energy efficiency while ensuring safety. Summary of the Invention
[0005] The purpose of this application is to provide a reverse osmosis energy consumption optimization control method and system based on deep reinforcement learning, which solves the problem that existing reverse osmosis control strategies lack adaptive energy-saving adjustment capabilities to dynamic changes in operating conditions.
[0006] To address the aforementioned technical problems, the embodiments of this application provide the following technical solutions: The first aspect of this application provides a reverse osmosis energy consumption optimization control method based on deep reinforcement learning, the method comprising: The operating parameters of the reverse osmosis system are acquired in real time. Based on the operating parameters, the membrane flux deviation coefficient, which reflects the flux decay level of the membrane element, is calculated. Based on the historical time series sliding window of the operating parameters, the water quality fluctuation feedforward features for future periods are predicted and extracted online. The operating parameters, the membrane flux deviation coefficient, and the water quality fluctuation feedforward features are constructed into a multi-dimensional state vector. The multidimensional state vector is input into a pre-trained deep reinforcement learning policy network for forward inference, and the high-pressure pump frequency adjustment and concentrate valve opening adjustment of the current control cycle are output as the initial control adjustment. The policy network is obtained through offline training with the reduction of unit water production energy consumption as the incentive target. Substitute the operating parameters into the reverse osmosis physical mechanism model to calculate the hard constraint boundary that ensures the safety of the membrane element under the current operating conditions in real time. The hard constraint boundary includes at least the upper limit of the transmembrane pressure difference and the lower limit of the permeate flow rate. Based on the dynamic response characteristics of the reverse osmosis system, the expected operating state of the next control cycle is predicted by the initial control adjustment amount. If the expected operating state exceeds the hard constraint boundary, the constraint mapping algorithm is used to project the initial control adjustment amount into the hard constraint boundary with the minimum action correction amount to obtain the safety optimization instruction. Otherwise, the initial control adjustment amount is used as the safety optimization instruction. The target control value corresponding to the security optimization instruction is sent to the execution mechanism.
[0007] A second aspect of this application provides a reverse osmosis energy consumption optimization control system based on deep reinforcement learning, the system comprising: The multidimensional state vector generation module is used to acquire the operating parameters of the reverse osmosis system in real time, calculate the membrane flux deviation coefficient reflecting the flux decay level of the membrane element based on the operating parameters, and extract the water quality fluctuation feedforward features for future periods based on the historical time series sliding window of the operating parameters, and construct the operating parameters, the membrane flux deviation coefficient and the water quality fluctuation feedforward features into a multidimensional state vector. The initial control adjustment quantity generation module is used to input the multi-dimensional state vector into a pre-trained deep reinforcement learning policy network for forward inference, and output the high-pressure pump frequency adjustment quantity and concentrate valve opening adjustment quantity of the current control cycle as the initial control adjustment quantity; wherein, the policy network is obtained through offline training with the reduction of unit water production energy consumption as the incentive target. The hard constraint boundary generation module is used to substitute the operating parameters into the reverse osmosis physical mechanism model and calculate the hard constraint boundary that ensures the safety of the membrane element under the current operating conditions in real time. The hard constraint boundary includes at least the upper limit of the transmembrane pressure difference and the lower limit of the permeate flow rate. The safety optimization instruction generation module is used to predict the expected operating state of the next control cycle based on the dynamic response characteristics of the reverse osmosis system and the initial control adjustment amount. If the expected operating state exceeds the hard constraint boundary, the initial control adjustment amount is projected into the hard constraint boundary with the minimum action correction amount using a constraint mapping algorithm to obtain the safety optimization instruction. Otherwise, the initial control adjustment amount is used as the safety optimization instruction. The distribution module is used to distribute the target control value corresponding to the security optimization instruction to the execution mechanism.
[0008] Compared to existing technologies, the reverse osmosis energy consumption optimization control method and system provided in this application, by constructing a multi-dimensional state vector that integrates current operating parameters, membrane flux decay coefficient, and feedforward characteristics of water quality fluctuations, enables the control strategy to have a comprehensive perception capability of real-time system operating conditions, membrane element health status, and future feedwater trends. It employs a pre-trained deep reinforcement learning strategy network to infer online the coordinated adjustment of the high-pressure pump and concentrate valve, aiming to reduce unit product water energy consumption while meeting product water quality and quantity constraints. This achieves self-regulation of feedwater quality fluctuations and membrane fouling evolution. It features adaptive and continuous energy saving; at the same time, it uses the reverse osmosis physical mechanism model to calculate the upper limit of the dynamic change of transmembrane pressure difference and the lower limit of the permeate flow rate under the current operating conditions as hard constraint boundaries, and uses the constraint mapping algorithm to project the out-of-bounds initial adjustment amount back into the safe domain with the minimum action correction amount. Thus, while ensuring the structural integrity of the membrane element and the reliability of water supply, it retains the energy-saving optimization intention of the strategy network to the greatest extent, completely solving the technical problem of the separation between safety protection and energy efficiency optimization in traditional control methods, significantly reducing the long-term operating energy consumption of the reverse osmosis system and extending the service life of the membrane element. Attached Figure Description
[0009] The above and other objects, features, and advantages of exemplary embodiments of this application will become readily understood by reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of this application are illustrated by way of example and not limitation, with the same or corresponding reference numerals denoteing the same or corresponding parts, wherein: Figure 1 A flowchart illustrating a reverse osmosis energy consumption optimization control method based on deep reinforcement learning is shown schematically. Figure 2 A schematic diagram of a reverse osmosis energy consumption optimization control system based on deep reinforcement learning is shown. Detailed Implementation
[0010] Exemplary embodiments of this application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.
[0011] It should be noted that, unless otherwise stated, the technical or scientific terms used in this application shall have the ordinary meaning as understood by one of ordinary skill in the art to which this application pertains.
[0012] The methods described in the embodiments of this application will be explained in detail below.
[0013] This embodiment provides a reverse osmosis energy consumption optimization control method based on deep reinforcement learning. This method can be applied to reverse osmosis membrane separation systems in scenarios such as seawater desalination plants, industrial wastewater deep treatment workshops, and municipal water supply treatment plants. In this embodiment, the reasoning and optimization tasks of the control algorithm can be flexibly deployed in the host computer in the water plant's central control room or in the edge intelligent controller in the local control cabinet, according to the actual operating conditions. Data interaction and control command issuance are completed collaboratively between the host computer and the edge controller, or between the controllers of each process unit, through industrial Ethernet or fieldbus.
[0014] Step S110: Obtain the operating parameters of the reverse osmosis system in real time, calculate the membrane flux deviation coefficient reflecting the flux decline level of the membrane element based on the operating parameters, and extract the feedforward features of water quality fluctuations in the future period based on the historical time series sliding window of the operating parameters. Construct a multi-dimensional state vector from the operating parameters, membrane flux deviation coefficient and feedforward features of water quality fluctuations.
[0015] This step transforms the raw sensor data from the reverse osmosis system into structured input that can be directly used by the deep reinforcement learning policy network. Specifically, this step fuses three types of information: first, the current operating parameters of the system collected in real time by instruments; second, the membrane element flux decay status index obtained by mechanistic calculations of the operating parameters; and third, water quality fluctuation trend information obtained by online prediction of historical time-series data. These three types of information are integrated into a multi-dimensional state vector along the feature dimension, enabling the policy network to simultaneously perceive the system's "current operating condition," "health status," and "future trend," providing a sufficient information foundation for subsequent optimization decisions.
[0016] In a specific implementation, the membrane flux deviation coefficient, which reflects the flux attenuation level of the membrane element, is calculated based on operating parameters, and can be achieved through the following four steps: Obtain the following operating parameters: inlet water temperature, inlet water conductivity, product water conductivity, membrane inlet pressure, concentrate pressure, and product water pressure; The temperature correction coefficient is calculated based on the inlet water temperature, and the effective osmotic pressure difference across the membrane element is estimated based on the inlet water conductivity and the product water conductivity. Based on the average values of the inlet pressure and concentrate pressure, the permeate pressure, and the effective osmotic pressure difference, the net driving pressure under the current operating conditions is calculated. The permeate flow rate is standardized and corrected using a temperature correction factor, and the standardized permeate flux is calculated based on the corrected permeate flow rate and net driving pressure. The standardized permeate flux is compared with the preset membrane element reference flux to obtain the membrane flux deviation coefficient.
[0017] Specifically, six key physical quantities are extracted from the real-time operating parameters: feed water temperature, feed water conductivity, permeate conductivity, inlet pressure, concentrate pressure, and permeate pressure. These six parameters are the basic inputs for subsequent calculations. First, temperature correction is required. The permeate flux of the reverse osmosis membrane is sensitive to temperature; for every 1°C increase in water temperature, the permeate flux increases by approximately 2%-3%. To eliminate the interference of temperature on flux assessment, a temperature correction coefficient is calculated based on the feed water temperature. This coefficient is used to convert the actual permeate flow rate to the equivalent value under standard temperature conditions, making flux data at different water temperatures comparable. Second, the effective osmotic pressure difference across the membrane element is estimated. Osmotic pressure is the inherent resistance that the high-pressure pump must overcome during reverse osmosis, and its magnitude mainly depends on the salt concentration difference across the membrane element. Using the approximately linear relationship between conductivity and salt concentration, the osmotic pressure on the feed water side is estimated from the feed water conductivity, and the osmotic pressure on the permeate side is estimated from the permeate conductivity. Subtracting the two yields the effective osmotic pressure difference. Next, the net drive pressure is calculated. Net driving pressure is the effective force that actually propels water molecules through the reverse osmosis membrane. The average pressure on the feed side of the membrane element is obtained by taking the average of the inlet pressure and the concentrate pressure, subtracting the permeate pressure on the permeate side, and then subtracting the effective osmotic pressure difference estimated in the second step. Next, the standardized permeate flux is calculated. The actual permeate flow rate is multiplied by the temperature correction factor obtained in the first step to complete the temperature standardization correction; then, the corrected permeate flow rate is divided by the effective membrane area of the membrane element and the net driving pressure obtained in the third step to obtain the standardized permeate flux per unit pressure and per unit area. This indicator eliminates the influence of temperature and operating pressure fluctuations and can truly reflect the permeability of the membrane element itself. Finally, the membrane flux deviation coefficient is calculated. The standardized permeate flux obtained in the fourth step is compared with the preset membrane element reference flux (i.e., the nominal flux of a new membrane element under standard test conditions). The ratio of the difference between the two to the reference flux is the membrane flux deviation coefficient. This coefficient typically ranges from 0 to 1; a larger value indicates a more severe decline in membrane flux. When the coefficient is close to 0, it indicates that the membrane element is close to a brand new state; when the coefficient exceeds the preset threshold, it indicates that the membrane fouling or deterioration is quite serious, and a warning for chemical cleaning or replacement needs to be triggered. At the same time, this information can be input into the strategy network to guide it to adopt a more conservative or enhanced operating strategy.
[0018] In a specific implementation, the online prediction and extraction of feedforward features of water quality fluctuations in future periods based on historical time-series sliding windows of operating parameters can be achieved through the following three steps: Extract time-series data sequences containing historical influent water quality parameters and temperature parameters according to a preset sampling frequency; The time series data sequence is input into the pre-trained time series prediction model, the dynamic evolution features in the time dimension are extracted by the time series prediction model, and a high-dimensional feature vector containing future trend information is output through the intermediate feature mapping layer. The high-dimensional feature vector is decoded to generate a feedforward feature sequence that represents the trend of water quality fluctuations within a preset time period in the future, which serves as the feedforward feature for water quality fluctuations.
[0019] Specifically, to endow the control strategy with the ability to predict future changes in influent water quality, this step utilizes a pre-trained time-series prediction model to perform online rolling predictions of the raw water quality fluctuation trend. First, according to a preset sampling frequency (e.g., once per minute), historical influent water quality parameters and temperature parameters from a past time window (e.g., the past 30 minutes) are extracted from a real-time database to form a fixed-length time-series data segment. Influent water quality parameters typically include online measurable indicators such as raw water conductivity, turbidity, or pH value. Second, this time-series data segment is input into a pre-trained time-series prediction model. This model has internal memory or attention mechanisms that can automatically capture the dynamic patterns of water quality parameters evolving over time (e.g., slow upward trends, periodic fluctuations, or sudden jumps). The model abstracts the original time-series data layer by layer through internal nonlinear mapping layers, encoding it into a high-dimensional feature vector containing information about future evolution trends. The decoding and mapping layer at the top of the model decodes the high-dimensional feature vector, mapping it back from the high-dimensional space to the physical quantity space, generating a one-dimensional or multi-dimensional feature sequence representing the water quality fluctuation trend within a preset time period (e.g., the next 5-10 minutes). This sequence is directly used as a feedforward feature for water quality fluctuations; for example, it can be a trend description encoding such as "the conductivity of raw water is expected to increase by 15% within the next 10 minutes." This feature is then input into the policy network along with other state information, enabling decisions to be based not only on the current instantaneous system state but also to respond in advance to future disturbances.
[0020] In an optional implementation, the training of the time-series prediction model is performed offline, separated from the online control process. In simple terms, the training process is as follows: The training data comes from a historical operational database accumulated over long-term operation of the reverse osmosis system (typically 6 months to 2 years). This database comprehensively records operational data under different seasons, time periods, and raw water quality conditions. Each record includes a timestamp and corresponding feed water quality and temperature parameters. Specific feed water quality parameters may include one or more of the online instrument-measurable indicators such as raw water conductivity, turbidity, pH, and oxidation-reduction potential. During data preprocessing, the raw data undergoes the following steps: abnormal jump values caused by instrument malfunctions and missing segments during offline maintenance are removed; missing short-term data are filled using linear interpolation; and all data is smoothed and denoised to eliminate high-frequency measurement noise. The cleaned time-series data is divided into three parts: a training set, a validation set, and a test set.
[0021] The model's input is a fixed-length time-series matrix, with the number of rows equal to the number of time steps in the input window (e.g., 30 steps) and the number of columns equal to the number of categories of selected water quality and temperature parameters. The model's output is a time-series vector of the same length as the number of time steps in the prediction window, where each element represents a predicted value of a water quality parameter at a corresponding future time, such as the predicted value of raw water conductivity or the comprehensive water quality index. The backbone network of the time-series prediction model can utilize a deep learning architecture adept at handling sequential data. An encoder-decoder structure based on Long Short-Term Memory networks or gated recurrent units can be adopted: the encoder receives water quality and temperature data within the historical window step by step, progressively encoding it into a fixed-length context vector containing the temporal evolution pattern; the decoder uses this context vector as its initial state and generates predicted values of water quality parameters for each future time step by step.
[0022] In a specific implementation, the operating parameters, membrane flux deviation coefficient, and water quality fluctuation feedforward characteristics are constructed into a multi-dimensional state vector, which can be achieved through the following four steps: According to the preset control cycle, the sampling frequency synchronization and timing alignment of the operating parameters and membrane flux deviation coefficient are performed. The aligned data is dimensionless to obtain the real-time observation sub-vector and the decay state sub-vector. The water quality fluctuation feedforward features are mapped to a preset feature space to generate a predicted trend subvector that matches the dimension of the real-time observation subvector. The real-time observation sub-vector, decay state sub-vector, and predicted trend sub-vector are concatenated along the feature dimension to construct a multi-dimensional state vector.
[0023] First, sampling frequency synchronization and time alignment are required. Different types of sensors in a reverse osmosis system typically have different sampling frequencies: pressure and flow meters have higher sampling frequencies (e.g., once per second), while water quality meters have lower sampling frequencies (e.g., once every 30 seconds or minutes), and the calculation cycle for the membrane flux deviation coefficient depends on the algorithm execution cycle. Following a preset unified control cycle (e.g., once every 5 seconds), data from different sources are aligned to a time reference: for data with sampling frequencies higher than the control cycle, nearest neighbor values or short-time averages are used for downsampling; for data with sampling frequencies lower than the control cycle, the most recent valid value is used for filling. After this step, all data is fully aligned in the time dimension, ensuring the temporal consistency of subsequent state construction.
[0024] Because the physical dimensions and numerical ranges of different operating parameters vary significantly (for example, influent pressure is typically on the order of several megapascals, while product water conductivity may be on the order of micro-Siemens per centimeter), if directly input into the policy network, parameters with larger values will dominate gradient propagation in the feature space, making it difficult for the policy network to fully learn the coupling relationships between the parameters. Therefore, the aligned operating parameters and membrane flux deviation coefficients are dimensionless. Common processing methods include Z-score standardization (subtracting the historical mean and dividing by the standard deviation) or deviation standardization (scaling to the [0,1] interval). After dimensionless standardization, the part corresponding to the operating parameters constitutes the real-time observation sub-vector, and the part corresponding to the membrane flux deviation coefficient constitutes the decay state sub-vector.
[0025] The feedforward features of water quality fluctuations are output by a time-series prediction model. Their original form may be a sequence of predicted values for multiple future time steps, and their dimensionality does not perfectly match the real-time observation sub-vectors. To ensure dimensionality compatibility between the two types of features during subsequent fusion, this step maps the feedforward features to a predefined feature space that matches the dimensionality of the real-time observation sub-vectors using a predefined feature mapping matrix or a single-layer fully connected network, generating a predicted trend sub-vector. This mapping operation also serves to compress the predicted information and align its semantics.
[0026] The real-time observation sub-vector, decay state sub-vector, and prediction trend sub-vector generated in the first three steps are concatenated end-to-end along the feature dimension. The concatenated vector is the multi-dimensional state vector, whose total dimension is equal to the sum of the dimensions of the three sub-vectors. Structurally, this vector is naturally divided into three semantic blocks: the current operating state block, the membrane decay state block, and the water quality prediction trend block. This facilitates the extraction of features by the nonlinear feature mapping layer within the policy network according to semantic blocks, and also facilitates the interpretability analysis of the policy network's decision-making behavior.
[0027] Step S120: Input the multi-dimensional state vector into the pre-trained deep reinforcement learning policy network for forward inference, and output the high-pressure pump frequency adjustment and concentrate valve opening adjustment of the current control cycle as the initial control adjustment; wherein, the policy network is obtained through offline training with the reduction of unit water production energy consumption as the incentive target.
[0028] This step inputs the multi-dimensional state vector constructed by S110 into the pre-trained deep reinforcement learning policy network. Through a single forward inference, it directly generates the high-pressure pump frequency adjustment and concentrate regulating valve opening adjustment for the current control cycle, serving as the initial control adjustments. In the offline phase, this policy network aims to minimize energy consumption per unit of produced water, learning the control policy through interaction with the environment. During online inference, it eliminates the need for iterative solutions or online learning, enabling it to output optimized actions within milliseconds, meeting the computational efficiency requirements of industrial real-time control.
[0029] In a specific implementation, the multi-dimensional state vector is input into a pre-trained deep reinforcement learning policy network for forward inference, and the output of the high-pressure pump frequency adjustment and concentrate valve opening adjustment for the current control cycle are used as the initial control adjustment quantities, including: By utilizing the shared feature layer of the policy network to extract and fuse features from the multidimensional state vector, comprehensive feature information reflecting the current operating status of the reverse osmosis system is obtained. Based on comprehensive feature information, the initial inference values of the high-pressure pump frequency regulation and the concentrate regulating valve opening regulation are calculated through the dual parallel channels of the policy network. The initial inference values are range-mapped and limited to obtain the high-pressure pump frequency regulation and concentrate regulating valve opening regulation that meet the requirements of the actuator's engineering adjustment step size.
[0030] Specifically, after receiving the multidimensional state vector constructed by S110, the policy network first processes the vector through its internal shared feature layer. The multidimensional state vector simultaneously contains instrument data reflecting the real-time operating conditions of the system (such as pressure, flow rate, and water quality parameters), membrane flux deviation coefficients reflecting the health of membrane elements, and feedforward features reflecting the future evolution trend of influent water quality. These types of information have different physical sources and numerical dimensions. The shared feature layer, through its learned feature transformation capabilities, performs cross-analysis and semantic-level fusion, eliminating redundant information and retaining the most valuable key features for decision-making, ultimately compressing and generating a compact, comprehensive feature information.
[0031] After obtaining the comprehensive feature information, the strategy network enters the decision output stage. Unlike traditional control methods that design separate controllers for the high-pressure pump and concentrate valve, this strategy network adopts a dual-parallel channel structure: one channel is dedicated to calculating how the high-pressure pump frequency should be adjusted based on the comprehensive feature information, while the other channel simultaneously and independently calculates the adjustment direction and magnitude of the concentrate regulating valve opening. The two channels share the underlying comprehensive feature information, but their respective output calculation logics are completely independent. Therefore, the strategy network can simultaneously generate the initial inference values for the high-pressure pump frequency adjustment and the concentrate regulating valve opening adjustment in a single inference operation.
[0032] The initial inference values directly output by the dual channels of the strategy network are typically dimensionless normalized values. Their range is limited to a fixed interval by the network's internal activation function, making them unsuitable for directly driving the actual actuator. Therefore, this stage performs range mapping on the two initial inference values: based on the maximum allowable single adjustment step size of the high-pressure pump frequency converter and the concentrate electric regulating valve, the normalized values are linearly mapped to the corresponding engineering adjustment range. For example, if the single adjustment step size of the high-pressure pump frequency is limited to ±2Hz, the [-1,1] interval of the initial inference value is mapped to [-2,2]Hz. After mapping, the mapped values are further limited to ensure that the final output adjustment amount does not exceed the actuator's physical adjustment capability due to numerical overflow.
[0033] In an optional implementation, the deep reinforcement learning policy network is trained in an offline simulation environment, completely separated from the online control process. The training environment is constructed based on historical operating data of the reverse osmosis system or a high-precision mechanism simulation model, capable of simulating time-varying characteristics such as feedwater quality fluctuations and the gradual evolution of membrane fouling. During training, the policy network takes the current state vector as input and outputs the adjustment amounts of the high-pressure pump frequency and the concentrate regulating valve opening. The environment provides feedback on the state and reward value for the next time step. The reward function uses the reduction in energy consumption per unit of permeate as the core positive incentive term, while applying negative penalties when the permeate flow rate is below standard or the permeate quality exceeds standard, guiding the policy network to minimize energy consumption while satisfying permeate constraints. The training algorithm employs a deep deterministic policy gradient and its variants suitable for continuous action spaces, utilizing empirical replay and a target network to stabilize the training process.
[0034] Step S130: Substitute the operating parameters into the reverse osmosis physical mechanism model and calculate in real time the hard constraint boundary that ensures the safety of the membrane element under the current operating conditions. The hard constraint boundary includes at least the upper limit of the transmembrane pressure difference and the lower limit of the permeate flow rate.
[0035] This step involves independently constructing a safety assurance channel while the strategy network generates initial control and regulation values. Specifically, real-time collected operating parameters are substituted into a preset reverse osmosis physical mechanism model to calculate online the upper limit of the transmembrane pressure difference and the lower limit of the permeate flow rate allowed for safe operation of the membrane element under the current operating conditions, serving as hard constraint boundaries. These boundaries are not fixed thresholds but dynamically change with feed water quality, temperature, and membrane fouling status, providing real-time safety judgment criteria for verifying and correcting the strategy network's output actions in subsequent steps.
[0036] In a specific implementation, the real-time calculation of the hard constraint boundary that ensures the safety of the membrane element under the current operating conditions can be achieved through the following three steps: Using the physical mechanism model of reverse osmosis, combined with the feed water conductivity and feed water temperature in the operating parameters, the real-time osmotic pressure and concentration polarization factor on the surface of the reverse osmosis membrane element are calculated online. Based on the real-time osmotic pressure and the preset mechanical strength threshold of the membrane element, the upper limit of the allowable transmembrane pressure difference under the current operating conditions is determined as the hard constraint boundary of the pressure dimension; Based on the concentration polarization factor and the preset scaling control index, the critical flow rate required to ensure that the membrane surface does not scale is calculated. The critical flow rate is then compared with the preset lower limit of water supply demand, and the larger of the two values is taken as the lower limit of product water flow rate, which serves as the hard constraint boundary of the flow rate dimension.
[0037] Specifically, the feed water conductivity and temperature are extracted from real-time operating parameters. Feed water conductivity can be converted into feed water salt concentration, while feed water temperature is directly used in thermodynamic calculations. The reverse osmosis physical mechanism model is based on the dissolution-diffusion theory and van 't Hoff's osmotic pressure law. It calculates the osmotic pressure on the feed water side of the membrane element based on the feed water salt concentration and temperature. Simultaneously, it estimates the permeate side osmotic pressure by combining the conductivity (i.e., permeate salt concentration) with the permeate side conductivity. The difference between the two is the effective osmotic pressure difference across the membrane. This effective osmotic pressure difference is closely related to the current concentration distribution on the membrane surface. The model further utilizes concentration polarization theory, using parameters such as real-time feed water flow velocity and permeate flux, to calculate the concentration polarization factor on the membrane surface. This factor reflects the degree to which the salt concentration on the membrane surface is higher than the mainstream concentration and is a key intermediate variable for assessing scaling risk and determining minimum flow constraints.
[0038] Transmembrane pressure differential must be greater than osmotic pressure differential to produce water, while it must be lower than the mechanical strength limit to ensure the structural safety of the membrane element. Therefore, the minimum driving pressure differential required to meet the water production demand is added to the osmotic pressure differential under the current operating conditions, and compared with the mechanical strength threshold; the smaller of these two values is taken as the upper limit of transmembrane pressure differential. This upper limit is the hard constraint boundary in the pressure dimension, and the transmembrane pressure differential caused by any operational action must not exceed this value.
[0039] The lower limit of permeate flow rate serves as a hard constraint boundary in the flow rate dimension, and its determination follows the dual principles of "preventing scaling and ensuring safety, and ensuring water supply meets demand." Because the concentration of sparingly soluble salts at the membrane surface is significantly higher than the mainstream influent concentration due to concentration polarization, when this concentration exceeds the solubility product of the sparingly soluble salts, crystals precipitate and deposit on the membrane surface, forming a scaling layer. Therefore, the core of scaling control lies in limiting the salt concentration at the membrane surface to below the critical scaling concentration. The concentration polarization factor is defined as the ratio of the membrane surface salt concentration to the mainstream influent salt concentration, directly quantifying the severity of concentration polarization; a higher factor value indicates a higher risk of scaling. Based on influent water quality analysis or historical operating experience, a scaling control index is preset, such as setting an upper limit for sparingly soluble salt saturation. Combined with the mainstream influent salt concentration, the maximum allowable salt concentration at the membrane surface can be calculated, thus obtaining the maximum allowable concentration polarization factor. Since the concentration polarization factor is directly related to the hydraulic conditions of the membrane surface, the higher the flow velocity in the channel, the lower the factor. Given that the membrane element design parameters are fixed, there is a definite engineering correlation between the membrane surface velocity and the permeate flow rate. Therefore, the maximum allowable concentration polarization factor can be substituted into the concentration polarization model to calculate the minimum membrane surface velocity required to maintain this factor. Then, based on the relationship between the membrane element's cross-sectional area and the system recovery rate, this is finally converted into the critical permeate flow rate value required to prevent membrane fouling. After obtaining this critical flow rate value, it is compared with the preset lower limit of water supply demand, and the larger of the two values is taken as the lower limit of the permeate flow rate.
[0040] Step S140: Based on the dynamic response characteristics of the reverse osmosis system, the expected operating state of the next control cycle is estimated from the initial control adjustment amount; if the expected operating state exceeds the hard constraint boundary, the constraint mapping algorithm is used to project the initial control adjustment amount into the hard constraint boundary with the minimum action correction amount to obtain the safety optimization command; otherwise, the initial control adjustment amount is used as the safety optimization command.
[0041] This step involves performing safety verification and correction on the initial control adjustment quantity generated in S120 and the hard constraint boundary established in S130, and generating the final executable safety optimization instruction.
[0042] In a specific implementation, based on the dynamic response characteristics of the reverse osmosis system, the expected operating state of the next control cycle can be predicted from the initial control adjustment amount, which can be achieved through the following three steps: Obtain the dynamic response gain matrices of high-pressure pump frequency and concentrate regulating valve opening to transmembrane pressure difference and permeate flow rate, respectively; Based on the dynamic response gain matrix, the predicted increments of transmembrane pressure difference and permeate flow rate for the next control cycle are calculated from the initial control adjustment. The expected transmembrane pressure difference is obtained by superimposing the predicted increment of the transmembrane pressure difference with the current transmembrane pressure difference, and the expected product water flow rate is obtained by superimposing the predicted increment of the product water flow rate with the current product water flow rate.
[0043] Specifically, before executing the initial control adjustment, a forward-looking assessment of the consequences of the action is conducted to determine whether the action will cause the system state to touch the hard constraint boundary in the next control cycle.
[0044] In reverse osmosis systems, the high-pressure pump frequency and the concentrate regulating valve opening, two operational variables, have overlapping effects on the two key state variables: transmembrane pressure difference and permeate flow rate. This means that adjusting any actuator will simultaneously cause changes in both state variables. This coupling relationship is quantified using a pre-calibrated offline dynamic response gain matrix. This matrix contains four gain coefficients: the dynamic response gain of the high-pressure pump frequency to the transmembrane pressure difference; the dynamic response gain of the high-pressure pump frequency to the permeate flow rate; the dynamic response gain of the concentrate regulating valve opening to the transmembrane pressure difference; and the dynamic response gain of the concentrate regulating valve opening to the permeate flow rate. The gain matrix is obtained through system identification experiments. For example, under typical operating conditions, small step signals are applied to each actuator, the steady-state changes of the state variables are recorded, and the ratio of these changes is calculated to obtain the corresponding gain coefficient. The gain matrix is stored in the control terminal and can be directly accessed during online operation without repeated identification. Arrange the four gain coefficients into a 2×2 matrix, with rows corresponding to two state variables (transmembrane pressure difference and permeate flow rate) and columns corresponding to two operational variables (frequency regulation and valve opening regulation). Multiply the initial control regulation vector by this matrix to obtain the predicted increments of the two states, i.e., the dynamic response gain matrix.
[0045] After obtaining the initial control adjustment for the current cycle, it is multiplied by the gain matrix to obtain the predicted increments of transmembrane pressure difference and product water flow rate. This calculation essentially involves linearly superimposing the state changes caused by the high-pressure pump frequency adjustment and the concentrate valve opening adjustment, fully reflecting the combined impact of the two actuators' combined actions on the system state. The predicted increment of transmembrane pressure difference obtained in the second step is added to the real-time transmembrane pressure difference measured by the sensor at the current moment to obtain the expected transmembrane pressure difference; the predicted increment of product water flow rate is added to the current real-time product water flow rate to obtain the expected product water flow rate. These two expected values represent the expected operating state of the system in the next control cycle after executing the current initial adjustment, and are directly used for subsequent comparison and judgment with hard constraint boundaries.
[0046] After obtaining the expected operating state, the initial control adjustments output by the strategy network need to be verified for safety, and corrected at the lowest possible cost if necessary. This ensures that the final action commands retain the optimization intent of the strategy network as much as possible while strictly meeting the hard constraints for safe operation of the membrane element. The expected transmembrane pressure difference in the expected operating state is compared with the upper limit of the transmembrane pressure difference in the hard constraint boundary, and the expected permeate flow rate is compared with the lower limit of the permeate flow rate.
[0047] If the expected operating state does not exceed the hard constraint boundaries, i.e., the expected transmembrane pressure difference is not higher than the upper limit and the expected permeate flow rate is not lower than the lower limit, then executing the current initial control adjustment will not cause the system to enter a dangerous operating condition, and the action itself is safe and feasible. In this case, the initial control adjustment is directly output as a safe optimization command without any modification, so as to fully preserve the energy consumption optimization effect of the strategy network.
[0048] If the expected operating state exceeds the hard constraint boundary, then the process of using the constraint mapping algorithm to project the initial control adjustment amount into the hard constraint boundary with the minimum action correction amount is executed to obtain the safety optimization instruction.
[0049] In a specific implementation, the constraint mapping algorithm is used to project the initial control adjustment amount onto the interior of the hard constraint boundary with the minimum action correction amount to obtain the safety optimization instruction. This can be achieved by the following steps: Based on the dynamic response characteristics of the reverse osmosis system, the hard constraint boundary is transformed into a safe adjustment range for the high-pressure pump frequency and the opening of the concentrate regulating valve. Within the safe adjustment range, the parameter combination that minimizes the deviation from the initial control adjustment is determined and used as the safety optimization command.
[0050] Specifically, when the expected operating state exceeds the hard constraint boundary, the constraint mapping algorithm is triggered to safely correct the initial control adjustment. This algorithm does not directly discard the optimization actions output by the policy network, but adjusts them to a safe range with minimal modification cost, thereby achieving the best balance between membrane element protection and operational energy consumption optimization. The algorithm execution consists of the following two stages.
[0051] Phase 1: Transforming hard constraint boundaries into safe adjustment ranges. The upper limit of transmembrane pressure difference and the lower limit of permeate flow rate in the hard constraint boundaries define the safe boundaries of the reverse osmosis system in the state space, rather than direct limits on high-pressure pump frequency and concentrate valve opening. Therefore, it is necessary to utilize the dynamic response characteristics of the reverse osmosis system to transform the constraints of the state space into safe adjustment ranges in the actuator action space. This safe adjustment range is dynamically adjusted in real time according to the current operating conditions: when the feed water quality deteriorates, leading to an increase in osmotic pressure, the safe adjustment range shrinks accordingly; when membrane fouling worsens, leading to a decrease in flux, the safe adjustment range is further limited.
[0052] The second stage involves finding the minimum deviation parameter combination within the safe adjustment range. Within the safe adjustment range determined in the first stage, the algorithm identifies the parameter combination with the smallest deviation from the initial control adjustment value as the safety optimization instruction. When the initial control adjustment value is already within the safe adjustment range, it is directly output as the safety optimization instruction. When the initial control adjustment value exceeds the safe adjustment range, the algorithm automatically finds the position closest to the initial point on the boundary of the range; the adjustment parameter corresponding to this position is the expected adjustment value in the safety optimization instruction.
[0053] First, establish the high-pressure pump frequency regulation. Adjustment amount of concentration regulating valve opening With transmembrane pressure difference and water production flow The local linearization sensitivity matrix between them. Taking the upper limit constraint of transmembrane pressure difference as an example, let the transmembrane pressure difference at the current moment be... The absolute safety upper limit of the mechanistic model solution is The maximum allowable pressure differential increment is: ; Based on the dynamic response characteristics of the system, the hard constraints are transformed into inequality constraints in the actuator action space: ; in, and These are the dynamic partial derivatives of frequency and opening degree with respect to pressure difference under the current operating conditions. By combining the inequalities corresponding to multiple physical boundaries such as the upper limit of transmembrane pressure difference and the lower limit of permeate flow rate, a polygonal feasible region is defined in the action space, which is the safe adjustment range.
[0054] Then, the problem of finding the minimum deviation parameter combination is transformed into a constrained quadratic programming problem. The objective function is to find the safe action that minimizes the deviation from the initial action within the safe adjustment range: ; in, and This is the initial control adjustment output by the policy network. and The safety optimization instructions to be solved are... and This is a weighting coefficient used to balance the adjustment costs of the high-pressure pump and the concentrate valve. The deviation norm can be 1 (corresponding to the absolute value deviation of the L1 norm) or 2 (corresponding to the Euclidean distance deviation of the L2 norm). The constraints are the system of inequalities established in the first stage. The control terminal solves this quadratic programming problem online to obtain a feasible solution that is within the safety boundary and has the minimum Euclidean distance from the initial action. This feasible solution is the expected adjustment value in the safety optimization command.
[0055] Step S150: Send the target control value corresponding to the safety optimization instruction to the execution mechanism.
[0056] This step converts the security optimization instructions generated by S140 after security verification and correction into physical control signals that can be directly responded to and executed by each actuator, and then distributes them. Actuators in a reverse osmosis system may include: High-pressure pump frequency converter: Receives the target frequency value of the high-pressure pump and adjusts the power supply frequency of the high-pressure pump motor to change the pump speed, thereby controlling the inlet water pressure and flow rate into the membrane element. Increasing the frequency increases the inlet water pressure and strengthens the permeate driving force; decreasing the frequency decreases the inlet water pressure and correspondingly reduces energy consumption. The high-pressure pump is the most energy-consuming device in a reverse osmosis system, and its frequency regulation directly determines the system's operating power and permeate capacity.
[0057] The concentrate electric regulating valve receives the target opening value and adjusts the discharge resistance on the concentrate side of the membrane element by changing the valve opening. Increasing the valve opening reduces the back pressure on the concentrate side, slightly lowers the pressure before the membrane, and slightly reduces the permeate flow rate, but the reduced system recovery rate is beneficial for membrane scouring and delaying scaling. Conversely, decreasing the valve opening increases the back pressure on the concentrate side, increases the pressure before the membrane, and increases the permeate flow rate, but the increased recovery rate may exacerbate the risk of membrane fouling. The concentrate valve and the high-pressure pump work together to determine the system's permeate flow rate, recovery rate, and energy consumption.
[0058] The two actuators mentioned above constitute the core operating variables of the reverse osmosis system. After the target control value in the safety optimization command is sent to the corresponding frequency converter and electric valve via industrial Ethernet or fieldbus, the system will operate according to the optimized operating parameters and complete a complete control cycle.
[0059] Preferably, this application embodiment also provides an online adaptive update mechanism for the strategy network to address the problem of slow performance drift of reverse osmosis membrane elements during long-term operation due to factors such as membrane fouling, chemical cleaning, and membrane material aging. This mechanism enables the control strategy obtained through offline training to continuously track the evolution of system characteristics and maintain optimal energy consumption during long-term operation. The specific implementation of this mechanism consists of the following three steps: Within each preset evaluation period, the actual unit water production energy consumption after executing the safety optimization command is statistically analyzed, and the deviation of the average unit water production energy consumption within that evaluation period from the historical best energy consumption is calculated. When the deviation exceeds a preset threshold, the recently accumulated actual operating data is used to construct an incremental training sample set, triggering online fine-tuning of the deep reinforcement learning policy network; By fine-tuning online, the internal weights of the strategy network are updated so that the control strategy adapts to the performance drift of the reverse osmosis membrane element as the performance evolves over time.
[0060] Specifically, the system's operational energy efficiency is assessed periodically using a preset evaluation cycle as the time unit. The duration of the evaluation cycle can be set according to actual operating conditions, typically ranging from several hours to several days. Within each evaluation cycle, actual operating data is continuously recorded after each execution of a safety optimization command, including high-pressure pump power and water production flow rate, and the average unit water production energy consumption for that cycle is calculated accordingly. The average unit water production energy consumption is defined as the ratio of the cumulative power consumption of the high-pressure pump to the cumulative water production within the evaluation cycle, and is the core indicator for measuring the system's operational energy efficiency.
[0061] After obtaining the average energy consumption per unit of permeate for the current evaluation period, it is compared with the historical best energy consumption stored in the terminal. The historical best energy consumption refers to the lowest energy consumption per unit of permeate ever achieved under the same or similar operating conditions (such as similar influent water quality and temperature) during the system's operating history. The degree of deviation is quantified by the relative deviation between the current average energy consumption and the historical best energy consumption, for example, by calculating the difference or ratio between the two. When the degree of deviation does not exceed the preset threshold, it indicates that the strategy network is still well adapted to the current system characteristics, and no update is required; when the degree of deviation exceeds the preset threshold, it indicates that the performance drift of the membrane element has caused a perceptible degradation in the control effect of the original strategy network, and the system energy consumption has systematically deviated from the optimal level, at which point the fine-tuning process is triggered.
[0062] An incremental training sample set is constructed from recently accumulated actual operational data. This sample set consists of online operation records from a period prior to the triggering time, with each sample containing the state vector at that time, the executed adjustment action, and the actual reward value obtained after execution. After the incremental training sample set is constructed, the control terminal performs online fine-tuning of the policy network in a low-priority background thread. Fine-tuning uses the network weights obtained from offline training as initial values, employing the same reinforcement learning algorithm as offline training but applying stronger regularization constraints on the update magnitude. After online fine-tuning, the updated policy network undergoes safety verification: test scenarios of recent typical operating conditions are replayed in a simulation environment to confirm that the updated policy does indeed improve energy consumption indicators while meeting all safety constraints. After successful verification, the control terminal replaces the currently used policy weights with the updated network weights, completing a seamless hot-swap of the policy. Through this closed-loop mechanism, the policy network can continuously track the performance evolution of membrane elements throughout the reverse osmosis system's operational lifespan, maintaining a good fit with the current system characteristics and avoiding continuous energy efficiency degradation due to membrane aging and other reasons.
[0063] The reverse osmosis energy consumption optimization control method provided in this application constructs a multi-dimensional state vector that integrates current operating parameters, membrane flux decay coefficient, and feedforward characteristics of water quality fluctuations. This enables the control strategy to have a comprehensive perception capability of the system's real-time operating conditions, membrane element health status, and future feedwater trends. A pre-trained deep reinforcement learning strategy network is used to infer and output the coordinated adjustment of the high-pressure pump and concentrate valve online. Under the premise of meeting the constraints of product water quality and quantity, the optimization objective is to reduce energy consumption per unit of product water, achieving autonomous adaptation to feedwater quality fluctuations and membrane fouling evolution, and continuous energy saving. Simultaneously, the upper limit of the dynamically changing transmembrane pressure difference and the lower limit of the product water flow rate under the current operating conditions are calculated in real time using a reverse osmosis physical mechanism model as hard constraint boundaries. A constraint mapping algorithm is then used to project the out-of-bounds initial adjustment back into the safe domain with minimal action correction. This maximizes the preservation of the strategy network's energy-saving optimization intent while ensuring the structural integrity of the membrane element and the reliability of the water supply. This completely solves the technical problem of the separation between safety protection and energy efficiency optimization in traditional control methods, significantly reducing the long-term operating energy consumption of the reverse osmosis system and extending the service life of the membrane element.
[0064] Based on the same inventive concept, as an implementation of the above-mentioned reverse osmosis energy consumption optimization control method based on deep reinforcement learning, this application also provides a reverse osmosis energy consumption optimization control system based on deep reinforcement learning. Figure 2 This is a structural diagram of the system in an embodiment of this application. See also... Figure 2 As shown, the system includes: The multidimensional state vector generation module is used to acquire the operating parameters of the reverse osmosis system in real time, calculate the membrane flux deviation coefficient reflecting the flux decline level of the membrane element based on the operating parameters, and extract the feedforward features of water quality fluctuations in the future period based on the historical time series sliding window of the operating parameters. The operating parameters, membrane flux deviation coefficient and feedforward features of water quality fluctuations are constructed into a multidimensional state vector. The initial control adjustment quantity generation module is used to input the multi-dimensional state vector into the pre-trained deep reinforcement learning policy network for forward inference, and output the high-pressure pump frequency adjustment quantity and concentrate valve opening adjustment quantity of the current control cycle as the initial control adjustment quantity; wherein, the policy network is obtained by offline training with the reduction of unit water production energy consumption as the incentive target. The hard constraint boundary generation module is used to substitute operating parameters into the reverse osmosis physical mechanism model and calculate the hard constraint boundary that ensures the safety of the membrane element under the current operating conditions in real time. The hard constraint boundary includes at least the upper limit of the transmembrane pressure difference and the lower limit of the permeate flow rate. The safety optimization instruction generation module is used to predict the expected operating state of the next control cycle based on the dynamic response characteristics of the reverse osmosis system from the initial control adjustment amount. If the expected operating state exceeds the hard constraint boundary, the constraint mapping algorithm is used to project the initial control adjustment amount into the hard constraint boundary with the minimum action correction amount to obtain the safety optimization instruction. Otherwise, the initial control adjustment amount is used as the safety optimization instruction. The distribution module is used to distribute the target control values corresponding to the safety optimization instructions to the execution mechanism.
[0065] The reverse osmosis energy consumption optimization control system provided in this application constructs a multi-dimensional state vector that integrates current operating parameters, membrane flux decay coefficient, and feedforward characteristics of water quality fluctuations. This enables the control strategy to have a comprehensive perception capability of the system's real-time operating conditions, membrane element health status, and future feedwater trends. A pre-trained deep reinforcement learning strategy network is used to infer and output the coordinated adjustment of the high-pressure pump and concentrate valve online. Under the premise of meeting the constraints of product water quality and quantity, the optimization objective is to reduce energy consumption per unit of product water, achieving autonomous adaptation to feedwater quality fluctuations and membrane fouling evolution, and continuous energy saving. Simultaneously, a reverse osmosis physical mechanism model is used to calculate in real time the dynamically changing upper limit of transmembrane pressure difference and the lower limit of product water flow under the current operating conditions as hard constraint boundaries. A constraint mapping algorithm is then used to project the out-of-bounds initial adjustment back into the safe domain with minimal action correction. This maximizes the preservation of the strategy network's energy-saving optimization intent while ensuring the structural integrity of the membrane element and the reliability of the water supply. This completely solves the technical problem of the separation between safety protection and energy efficiency optimization in traditional control methods, significantly reducing the long-term operating energy consumption of the reverse osmosis system and extending the service life of the membrane element.
[0066] This application also provides an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a program for a reverse osmosis energy consumption optimization control method based on deep reinforcement learning. When the processor executes the computer program, it implements all the steps in the aforementioned reverse osmosis energy consumption optimization control method based on deep reinforcement learning, such as S110 to S150 and their sub-steps. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the aforementioned system embodiment.
[0067] For example, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete this application. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the electronic device.
[0068] The electronic device may be a computing device such as a server or cloud server deployed in a deep reinforcement learning-based reverse osmosis energy consumption optimization control system. The electronic device may include, but is not limited to, processors and memory. Those skilled in the art will understand that the above components are merely examples of electronic devices and do not constitute a limitation on the electronic device. It may include more or fewer components than described above, or combine certain components, or different components. For example, the electronic device may also include data acquisition cards and output cards for connecting sensors and actuators, network communication devices, buses, etc.
[0069] The processor can be a central processing unit, or other general-purpose processors, digital signal processors, application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (OPGs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the electronic device, connecting various parts of the entire electronic device via various interfaces and lines.
[0070] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc.; the data storage area may store data created according to system usage, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart memory card, secure digital card, flash memory card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0071] Wherein, if the modules / units integrated in the electronic device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory, a random access memory, an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0072] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A reverse osmosis energy consumption optimization control method based on deep reinforcement learning, characterized in that, The method includes: The operating parameters of the reverse osmosis system are acquired in real time. Based on the operating parameters, the membrane flux deviation coefficient, which reflects the flux decay level of the membrane element, is calculated. Based on the historical time series sliding window of the operating parameters, the water quality fluctuation feedforward features for future periods are predicted and extracted online. The operating parameters, the membrane flux deviation coefficient, and the water quality fluctuation feedforward features are constructed into a multi-dimensional state vector. The multidimensional state vector is input into a pre-trained deep reinforcement learning policy network for forward inference, and the high-pressure pump frequency adjustment and concentrate valve opening adjustment of the current control cycle are output as the initial control adjustment. The policy network is obtained through offline training with the reduction of unit water production energy consumption as the incentive target. Substitute the operating parameters into the reverse osmosis physical mechanism model to calculate the hard constraint boundary that ensures the safety of the membrane element under the current operating conditions in real time. The hard constraint boundary includes at least the upper limit of the transmembrane pressure difference and the lower limit of the permeate flow rate. Based on the dynamic response characteristics of the reverse osmosis system, the expected operating state of the next control cycle is predicted by the initial control adjustment amount. If the expected operating state exceeds the hard constraint boundary, the constraint mapping algorithm is used to project the initial control adjustment amount into the hard constraint boundary with the minimum action correction amount to obtain the safety optimization instruction. Otherwise, the initial control adjustment amount is used as the safety optimization instruction. The target control value corresponding to the security optimization instruction is sent to the execution mechanism.
2. The method according to claim 1, characterized in that, The calculation of the membrane flux deviation coefficient, which reflects the flux attenuation level of the membrane element based on the operating parameters, includes: The following operating parameters are obtained: inlet water temperature, inlet water conductivity, product water conductivity, membrane inlet pressure, concentrate pressure, and product water pressure. The temperature correction coefficient is calculated based on the inlet water temperature, and the effective osmotic pressure difference across the membrane element is estimated based on the inlet water conductivity and the product water conductivity. Based on the average of the inlet pressure and concentrate pressure, the permeate pressure, and the effective osmotic pressure difference, calculate the net driving pressure under the current operating conditions; The product water flow rate is standardized and corrected using the temperature correction coefficient, and the standardized product water flux is calculated based on the corrected product water flow rate and the net driving pressure. The standardized permeate flux is compared with the preset membrane element reference flux to obtain the membrane flux deviation coefficient.
3. The robot obstacle avoidance control method according to claim 1, characterized in that, The online prediction and extraction of water quality fluctuation feedforward features for future periods based on the historical time-series sliding window of the operating parameters includes: Extract time-series data sequences containing historical influent water quality parameters and temperature parameters according to a preset sampling frequency; The time series data sequence is input into a pre-trained time series prediction model, the time series prediction model extracts dynamic evolution features in the time dimension, and outputs a high-dimensional feature vector containing future trend information through an intermediate feature mapping layer. The high-dimensional feature vector is decoded to generate a feedforward feature sequence that represents the water quality fluctuation trend within a preset time period in the future, which serves as the water quality fluctuation feedforward feature.
4. The method according to claim 1, characterized in that, The step of constructing a multidimensional state vector from the operating parameters, the membrane flux deviation coefficient, and the water quality fluctuation feedforward characteristics includes: According to a preset control cycle, the operating parameters and the membrane flux deviation coefficient are subjected to sampling frequency synchronization and timing alignment processing. The aligned data is dimensionless to obtain the real-time observation sub-vector and the decay state sub-vector. The water quality fluctuation feedforward features are mapped to a preset feature space to generate a predicted trend subvector that matches the dimension of the real-time observation subvector. The real-time observation sub-vector, the decay state sub-vector, and the predicted trend sub-vector are concatenated along the feature dimension to construct a multi-dimensional state vector.
5. The method according to claim 1, characterized in that, The step of inputting the multidimensional state vector into a pre-trained deep reinforcement learning policy network for forward inference, and outputting the high-pressure pump frequency adjustment and concentrate valve opening adjustment of the current control cycle as the initial control adjustment quantities, includes: The shared feature layer of the policy network is used to extract and fuse features from the multidimensional state vector to obtain comprehensive feature information reflecting the current operating status of the reverse osmosis system. Based on the comprehensive feature information, the initial inference values of the high-pressure pump frequency regulation and the concentrate regulating valve opening regulation are calculated through the dual parallel channels of the strategy network. The initial inference values are subjected to range mapping and amplitude limiting processing to obtain the high-pressure pump frequency adjustment amount and the concentrate regulating valve opening adjustment amount that meet the requirements of the actuator engineering adjustment step size.
6. The method according to claim 1, characterized in that, The real-time calculation of the hard constraint boundaries that ensure the safety of the membrane element under the current operating conditions includes: Using the aforementioned reverse osmosis physical mechanism model, combined with the feed water conductivity and feed water temperature in the operating parameters, the real-time osmotic pressure and concentration polarization factor on the surface of the reverse osmosis membrane element are calculated online. Based on the real-time osmotic pressure and the preset mechanical strength threshold of the membrane element, the upper limit of the allowable transmembrane pressure difference under the current operating conditions is determined as a hard constraint boundary in the pressure dimension. Based on the concentration polarization factor and the preset scaling control index, the critical flow rate required to ensure that the membrane surface does not scale is calculated, and the critical flow rate is compared with the preset lower limit of water supply demand. The larger of the two values is taken as the lower limit of product water flow rate, which serves as the hard constraint boundary of the flow rate dimension.
7. The method according to claim 1, characterized in that, The method for predicting the expected operating state of the next control cycle based on the dynamic response characteristics of the reverse osmosis system, using the initial control adjustment, includes: Obtain the dynamic response gain matrices of high-pressure pump frequency and concentrate regulating valve opening to transmembrane pressure difference and permeate flow rate, respectively; Based on the dynamic response gain matrix, the predicted increment of transmembrane pressure difference and the predicted increment of permeate flow rate for the next control cycle are calculated from the initial control adjustment. The expected transmembrane pressure difference is obtained by superimposing the predicted increment of the transmembrane pressure difference with the current transmembrane pressure difference, and the expected product water flow rate is obtained by superimposing the predicted increment of the product water flow rate with the current product water flow rate.
8. The method according to claim 1, characterized in that, The method of using a constraint mapping algorithm to project the initial control adjustment amount into the interior of the hard constraint boundary with the minimum action correction amount to obtain a safety optimization instruction includes: Based on the dynamic response characteristics of the reverse osmosis system, the hard constraint boundary is transformed into a safe adjustment range for the high-pressure pump frequency and the opening of the concentrate regulating valve. Within the specified safety adjustment range, the parameter combination that minimizes the deviation from the initial control adjustment is determined and used as the safety optimization instruction.
9. The method according to any one of claims 1-8, characterized in that, The method further includes: Within each preset evaluation period, the actual unit water production energy consumption after executing the safety optimization command is statistically analyzed, and the deviation of the average unit water production energy consumption within that evaluation period from the historical best energy consumption is calculated. When the deviation exceeds a preset threshold, the recently accumulated actual operating data is used to construct an incremental training sample set, triggering online fine-tuning of the deep reinforcement learning policy network; The internal weights of the strategy network are updated through online fine-tuning so that the control strategy adapts to the performance drift of the reverse osmosis membrane element as the performance evolves over time.
10. A reverse osmosis energy consumption optimization control system based on deep reinforcement learning, characterized in that, The system includes: The multidimensional state vector generation module is used to acquire the operating parameters of the reverse osmosis system in real time, calculate the membrane flux deviation coefficient reflecting the flux decay level of the membrane element based on the operating parameters, and extract the water quality fluctuation feedforward features for future periods based on the historical time series sliding window of the operating parameters, and construct the operating parameters, the membrane flux deviation coefficient and the water quality fluctuation feedforward features into a multidimensional state vector. The initial control adjustment quantity generation module is used to input the multi-dimensional state vector into a pre-trained deep reinforcement learning policy network for forward inference, and output the high-pressure pump frequency adjustment quantity and concentrate valve opening adjustment quantity of the current control cycle as the initial control adjustment quantity; wherein, the policy network is obtained through offline training with the reduction of unit water production energy consumption as the incentive target. The hard constraint boundary generation module is used to substitute the operating parameters into the reverse osmosis physical mechanism model and calculate the hard constraint boundary that ensures the safety of the membrane element under the current operating conditions in real time. The hard constraint boundary includes at least the upper limit of the transmembrane pressure difference and the lower limit of the permeate flow rate. The safety optimization instruction generation module is used to predict the expected operating state of the next control cycle based on the dynamic response characteristics of the reverse osmosis system and the initial control adjustment amount. If the expected operating state exceeds the hard constraint boundary, the initial control adjustment amount is projected into the hard constraint boundary with the minimum action correction amount using a constraint mapping algorithm to obtain the safety optimization instruction. Otherwise, the initial control adjustment amount is used as the safety optimization instruction. The distribution module is used to distribute the target control value corresponding to the security optimization instruction to the execution mechanism.