High-precision and low-delay cleaning agent dispensing system

By combining multiple sensors and nonlinear dimensionality reduction technology, along with machine learning and swarm intelligence algorithms, a high-precision, low-latency cleaning agent quantitative delivery system was constructed. This system solved the problems of control lag and accuracy decay in traditional systems under dynamic operating conditions, achieving a high-precision, stable delivery process and ensuring safety.

CN120469197BActive Publication Date: 2026-01-06BEIJING BAICHUAN TECH & TRADE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510822314.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2026-01-06
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Traditional cleaning agent metering delivery systems struggle to accurately characterize the nonlinear interaction between pH fluctuations, changes in concentrate concentration, and actuator response under dynamic operating conditions, leading to control lag and accuracy degradation. Furthermore, they lack physical constraint embedding mechanisms and incremental learning capabilities, resulting in insufficient human-machine collaboration.

Method used

By employing a combination of multiple sensors, nonlinear dimensionality reduction, and machine learning models, and by optimizing control commands through real-time data acquisition, standardization processing, diffusion mapping, and swarm intelligence algorithms, combined with energy function verification and self-learning updates, a high-precision, low-latency closed-loop control architecture is constructed.

Benefits of technology

It significantly improves the adaptability of cleaning agent quantitative delivery under various working conditions, ensures high precision and stability in the delivery process, achieves control precision and safety that are difficult to achieve with traditional methods, and integrates an optimization mechanism that combines human experience with intelligent decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120469197B_ABST
    Figure CN120469197B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of industrial automation control, and discloses a high-precision and low-delay cleaning agent quantitative conveying system, which comprises a clean water tank, an original liquid tank, a soaking tank and a turnover tank, a plurality of tank bodies are connected through pipelines and are provided with pump valve assemblies, a sensor group contains liquid level sensors, PH sensors, temperature sensors and conductivity sensors, and a control module is configured to collect the pH value, temperature, conductivity, liquid level and PH change rate data of the soaking tank in real time, to perform standardization processing on the data, and to generate a low-dimensional feature vector through nonlinear dimension reduction, the application builds a dynamic Q-learning control strategy through state space modeling of fused diffusion mapping features, process time length and original liquid residual amount, significantly improves the working condition adaptability of cleaning agent quantitative conveying, and effectively overcomes the control lag and precision decay problems caused by the modeling deficiency of traditional methods due to nonlinear coupling relationship.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial automation control technology, and in particular to a high-precision and low-latency cleaning agent metering delivery system. Background Technology

[0002] Industrial cleaning agent metering systems are core equipment for achieving precision cleaning in high-end manufacturing. They ensure the stability and consistency of the cleaning process by precisely controlling pump and valve flow rates, pH values, and concentration parameters. Current mainstream technologies are based on closed-loop control architectures using sensor feedback, employing classic control algorithms or rule engines to generate execution instructions. However, as industrial scenarios shift towards flexible production with multiple varieties and small batches, the dynamic coupling of process parameters intensifies, and traditional technologies face the following key bottlenecks:

[0003] Regarding adaptability to dynamic operating conditions, existing control methods rely on linear modeling and single-dimensional parameter feedback, making it difficult to accurately characterize the nonlinear interaction between pH fluctuations, changes in concentrate concentration, and actuator response. For example, when the liquid level in the cleaning tank changes rapidly, the combined effect of the flow sensor feedback delay and the mechanical response lag of the pump and valve causes overshoot oscillations in traditional PID control, leading to problems such as excessive cleaning agent addition or incomplete cleaning. Some improvement schemes attempt to introduce machine learning algorithms, but without establishing a dynamic feature space mapping mechanism, the model's generalization performance significantly decreases under time-varying conditions such as equipment aging and environmental disturbances.

[0004] Regarding system reliability and long-term stability, existing intelligent control solutions generally lack physical constraint embedding mechanisms and model degradation monitoring capabilities. Control commands directly generated by algorithms such as reinforcement learning may violate the safety operating boundaries of equipment, while traditional threshold truncation methods cannot predict the state evolution path after multiple execution steps. In addition, long-term factors such as sensor drift and mechanical wear cause the initial training data distribution to deviate from real-time operating conditions, but existing systems lack incremental learning and version rollback mechanisms, resulting in a continuous decline in control accuracy over time without self-correction.

[0005] In terms of human-machine collaboration and knowledge transfer, the disconnect between automated systems and manual operations has long existed. Most solutions adopt an either-or control switching model, which fails to effectively integrate human experience with the advantages of intelligent decision-making and lacks a structured method for accumulating abnormal handling cases. Optimization parameters and strategies developed by operators during emergency interventions are often stored in unstructured data formats, making it difficult to transform them into reusable control rules for the system, resulting in the loss of knowledge assets and low operational efficiency. Summary of the Invention

[0006] The purpose of this invention is to provide a high-precision and low-latency cleaning agent quantitative delivery system, which solves the problems of control lag and accuracy decay caused by insufficient modeling of nonlinear coupling relationships in traditional methods.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] A high-precision and low-delay cleaning agent dispensing system, including:

[0009] The tanks consist of a clear water tank, a stock solution tank, a soaking tank, and a transfer tank. Multiple tanks are connected by pipelines and equipped with pump and valve assemblies.

[0010] The sensor group includes a level sensor, a pH sensor, a temperature sensor, and a conductivity sensor.

[0011] The control module is configured to:

[0012] The pH value, temperature, conductivity, liquid level and pH change rate of the soaking tank are collected in real time. After the data is standardized, a low-dimensional feature vector is generated by nonlinear dimensionality reduction.

[0013] Based on the low-dimensional feature vector and process state parameters, a machine learning model is used to dynamically generate control commands for pumping raw liquid.

[0014] A swarm intelligence algorithm is used to globally optimize the parameters in the control command, and the stability of the control command is verified by combining the energy function of the dynamic system.

[0015] The verified control command is output to the actuator to drive the pump and valve assembly to complete the quantitative delivery of cleaning agent;

[0016] After each transport cycle, the dimensionality reduction parameters and control model are updated based on the current process data.

[0017] Furthermore, by linking multiple tanks and real-time monitoring by sensor groups, a closed-loop control architecture for cleaning agent delivery is constructed. By combining standardized data preprocessing and nonlinear dimensionality reduction algorithms, low-dimensional feature vectors are extracted to drive machine learning models to generate dynamic control commands. At the same time, swarm intelligence optimization and energy function verification mechanisms are integrated to ensure the global optimality and stability of control commands. Finally, through self-learning updates, continuous iterative optimization of process parameters is achieved, forming a high-precision, adaptive, and disturbance-resistant quantitative delivery system.

[0018] Preferably, the real-time acquired data includes: pH value, temperature, conductivity, liquid level, and pH change rate, and the acquisition frequency of each parameter is once per second. High-frequency data acquisition is achieved through multi-threaded parallel processing technology, synchronized with the control cycle of the pump and valve components, forming a closed-loop feedback mechanism.

[0019] Furthermore, by acquiring high-frequency data once per second and using a closed-loop feedback synchronization mechanism, the real-time performance of sensor data is ensured to meet millisecond-level control requirements, eliminating control delays caused by traditional intermittent sampling. This provides continuous and complete input signals for subsequent standardization and dimensionality reduction algorithms, avoiding fluctuations in the transmission volume caused by missing or delayed data.

[0020] Preferably, the standardization process includes: normalizing each parameter based on its historical mean and standard deviation to eliminate dimensional differences;

[0021] Its normalization formula is:

[0022] ;

[0023] in:

[0024] The original data of the i-th parameter is as follows:

[0025] pH value;

[0026] Temperature value (unit: °C);

[0027] Conductivity value (unit: µS / cm);

[0028] Liquid level (unit: meters);

[0029] pH change rate (unit: pH / second);

[0030] Real-time measured values ​​of the original parameters;

[0031] : No. The arithmetic mean of the historical data for each parameter;

[0032] : The historical data standard deviation of the i-th parameter, with the calculation period being... Consistent.

[0033] Furthermore, to address the issue of dimensional differences in multi-source heterogeneous sensor data, a dynamic sliding window calculation based on historical mean and standard deviation is employed to map parameters with different dimensions to a unified dimensionless space. This eliminates the interference of numerical scale differences on subsequent algorithm weight allocation, provides standardized input for nonlinear dimensionality reduction and machine learning models, and enhances the reliability of cross-parameter correlation analysis.

[0034] Preferably, the nonlinear dimensionality reduction employs a diffusion mapping method, which constructs a high-dimensional spatial topology by calculating the similarity matrix between sensor data.

[0035] The formula for calculating the similarity matrix is:

[0036] ;

[0037] in:

[0038] The first one in a reinforcement learning neural network Input node to the Connection weight parameters for each hidden layer node;

[0039] : No. A standardized sample data vector;

[0040] : Natural exponential function;

[0041] : No. A standardized sample data vector, and Same dimension;

[0042] Euclidean distance operator, representing the spatial distance between two vectors;

[0043] : Adaptive bandwidth parameter, which is the median of the squared Euclidean distances between all samples within the sliding window divided by .

[0044] Furthermore, a high-dimensional spatial topology is constructed based on the diffusion mapping method. The nonlinear similarity between data is quantified by the exponential function. The sensitive range of data association is dynamically adjusted by the adaptive bandwidth parameter to solve the overfitting or underfitting problem of fixed neighborhood radius under complex working conditions, and to provide robust similarity weight input for the generation of Laplacian matrix.

[0045] Preferably, the generation of the low-dimensional feature vector includes: performing eigenvalue decomposition on the Laplacian matrix generated by the diffusion mapping, and selecting the first two largest nontrivial feature vectors as the dimensionality reduction result;

[0046] The Laplace matrix The construction formula is:

[0047] ;

[0048] in:

[0049] : Degree matrix, which is a diagonal matrix, whose diagonal elements Indicates the first The sum of the similarities between a sample and all other samples;

[0050] : is the similarity matrix, with dimensions N×N;

[0051] The inverse square root matrix of the degree matrix, obtained by... It is obtained by taking the reciprocal of the square root of each diagonal element.

[0052] Furthermore, the similarity matrix is ​​normalized using the degree matrix to eliminate the influence of uneven sample density distribution on feature extraction. The dominant manifold direction is extracted through the eigenvalue decomposition of the Laplacian matrix, compressing the high-dimensional sensor data into a low-dimensional feature space, preserving key nonlinear features while reducing the input noise of the machine learning model.

[0053] Preferably, the machine learning model is a deep deterministic policy gradient reinforcement learning model, whose inputs are the dimensionality-reduced feature vector, the remaining amount of raw liquid, and the process running time, and whose output is the increment of raw liquid pumping time.

[0054] The state space of the reinforcement learning is defined as follows:

[0055] ;

[0056] in:

[0057] : The system state vector at any given time;

[0058] : Eigenvalues ​​of low-dimensional manifolds (dimensionless real numbers);

[0059] The remaining volume of the raw material in the tank is measured in real time by a mass flow meter (unit: liters);

[0060] Process running time, accumulated since system startup (unit: seconds).

[0061] Furthermore, the deep deterministic strategy gradient model integrates low-dimensional feature vectors, raw material balance, and process time to simulate a collaborative decision-making mechanism driven by expert experience and data. It dynamically generates pumping time increment commands and triggers a linkage supply protocol when raw material is in short supply, thereby achieving supply interruption early warning and multi-tank collaborative control to ensure the continuity of the transportation process.

[0062] Preferably, the swarm intelligence algorithm is a quantum particle swarm optimization algorithm, wherein each particle encodes a set of control parameters, including proportional, integral, and differential coefficients and a mixing factor, and the particle performance is evaluated by a multi-objective fitness function that includes an error term and its first and second derivatives.

[0063] The fitness function is defined as follows:

[0064] ;

[0065] in:

[0066] : No. Weight fusion strategy for new and old models during incremental learning;

[0067] : No. The set of parameters encoded by each particle is as follows:

[0068] : Proportional control coefficient;

[0069] Integral control coefficient;

[0070] Differential control coefficient;

[0071] Mixing factor (value range 0-1);

[0072] Real-time pH error, calculated as the difference between the current pH value and the target value;

[0073] The first time derivative of pH error;

[0074] : The second time derivative of pH error.

[0075] Furthermore, the quantum particle swarm optimization algorithm globally searches for the optimal solution of control parameters through a multi-objective fitness function, evaluates parameter performance by combining error terms and their higher-order derivatives, dynamically adjusts optimization weights when system oscillations are detected, suppresses overshoot and accelerates convergence, and ensures the stability and adaptability of control commands under complex operating conditions.

[0076] Preferably, the energy function is a linear combination of a quadratic function based on system error and an error integral term, and the control command is determined to be stable when the time derivative of the energy function satisfies a preset convergence condition.

[0077] The core construction formula for the energy function is:

[0078] ;

[0079] in:

[0080] Real-time pH error;

[0081] The integral time variable represents the time span from the start of the process to the current moment.

[0082] The time derivative of the energy function;

[0083] Time infinitesimal element in integral operations;

[0084] : Real-time running time variables of the control system;

[0085] Convergence condition: ,in This is the preset positive convergence rate coefficient.

[0086] Furthermore, the instantaneous error energy is quantified by the quadratic term of the energy function, the historical error energy is accumulated by the integral term, and the dynamic stability of the control command is verified by combining the time derivative criterion. When the system deviates from the convergence condition, the parameter self-learning update is triggered, forming a stability-driven closed-loop optimization mechanism.

[0087] Preferably, the execution of the control command includes: when the second-order rate of change of pH value is detected to exceed a set threshold, automatically switching to a pulse control mode with a fixed frequency and duty cycle.

[0088] Furthermore, when drastic pH fluctuations are detected, the system automatically switches to pulse control mode. By adjusting the fixed frequency and duty cycle, the system can be quickly stabilized, avoiding the risk of divergence of conventional control algorithms under extreme conditions. This provides a buffer time for the reinforcement learning model to recover to a steady state, ensuring the safety of the delivery process.

[0089] Preferably, the self-learning update includes: a process database storing multidimensional sensor data, dimensionality-reduced feature vectors, and control commands, and adaptively adjusting the neighborhood radius parameter of the diffusion map according to the dynamic rate of change of the feature vectors.

[0090] Furthermore, by establishing a dynamically updated process database to store multidimensional sensor data, dimensionality-reduced feature vectors, and control commands, and by adaptively adjusting the neighborhood radius parameter of the diffusion map in conjunction with the dynamic change rate of the low-dimensional feature vectors, the nonlinear dimensionality reduction algorithm and the control model are synergistically optimized. This allows for dynamic adaptation to the data distribution characteristics under different operating conditions, improving the long-term stability and accuracy of the cleaning agent delivery control. At the same time, historical data backtracking supports rapid recovery from abnormal operating conditions and iterative model upgrades, forming a closed-loop self-learning mechanism.

[0091] In summary, the present invention has at least one of the following beneficial technical effects:

[0092] 1. This invention constructs a dynamic Q-learning control strategy by integrating diffusion mapping features, process duration, and remaining concentrate volume into a state-space model, significantly improving the adaptability of quantitative cleaning agent delivery under various operating conditions. Compared to existing control schemes that rely on fixed parameter thresholds or single sensors, this invention effectively overcomes the control lag and accuracy degradation problems caused by insufficient modeling of nonlinear coupling relationships in traditional methods, achieving real-time matching of high-precision flow regulation and process fluctuations.

[0093] 2. This invention is based on a three-layer verification mechanism of pre-action simulation, timing risk prediction, and redundant instruction arbitration. It embeds physical constraints and safety boundary protection into intelligent decision-making. Addressing the risk of instruction infeasibility in existing automated systems, this invention, through the synergy of dynamic confidence thresholds and a conservative strategy library, ensures transmission accuracy while mitigating safety hazards such as equipment overload and pH runaway, establishing a dual guarantee of high-precision control and operational safety.

[0094] 3. This invention, through triggered model updates, weight fusion, and version rollback mechanisms, enables the control system to possess continuous self-optimization capabilities. Compared to the performance degradation problems caused by traditional static models due to equipment aging and environmental disturbances, this solution absorbs new operating condition characteristics through incremental learning and, combined with degradation monitoring, promptly blocks the risk of strategy failure, ensuring that the quantitative delivery system maintains stable control accuracy throughout its entire lifecycle.

[0095] 4. This invention employs weighted fusion instruction generation and case library accumulation technologies to construct a complementary optimization mechanism that combines human experience with AI decision-making. Addressing the shortcomings of existing high-precision conveying systems, such as fragmented human-machine interaction and a black-box decision-making process, this invention utilizes visualized decision tracking and a dual verification process. This retains the efficiency of intelligent algorithms while incorporating the controllability of human intervention, forming a transparent and reusable closed loop of industrial knowledge. Attached Figure Description

[0096] Figure 1 This is a system framework diagram of the present invention;

[0097] Figure 2 This is a flowchart of the method of the present invention. Detailed Implementation

[0098] The following is in conjunction with the appendix Figure 1 -Appendix Figure 2 The present invention will be further described in detail below.

[0099] This invention provides a high-precision and low-delay cleaning agent metering delivery system, comprising:

[0100] The tanks consist of a clear water tank, a stock solution tank, a soaking tank, and a transfer tank. Multiple tanks are connected by pipelines and equipped with pump and valve assemblies.

[0101] The sensor group includes a level sensor, a pH sensor, a temperature sensor, and a conductivity sensor.

[0102] The control module is equipped with a water pump that connects the stock solution tank to the soaking tank. Based on the detected concentration in the soaking tank, it replenishes the soaking tank with a high concentration of cleaning agent stock solution. It also detects the water level in the soaking tank and the pH level, which rises to a certain extent within a specified range. This allows the module to determine the relationship between the pump start time for replenishing the stock solution and the pH change, thereby enabling precise control of the amount of cleaning agent stock solution replenished.

[0103] The control module is configured to include the following steps:

[0104] S1, Real-time Multi-dimensional Data Acquisition

[0105] Inside the soaking tank, level sensors, pH sensors, temperature sensors, and conductivity sensors are deployed, and each sensor is connected to the main control module via a bus. Data on pH value, temperature, conductivity, liquid level, and pH change rate are synchronously collected at a fixed frequency of once per second, ensuring that the data timing is strictly aligned with the pump and valve control cycle to avoid control command delays due to acquisition latency.

[0106] S2, Multi-source data standardization processing

[0107] For each sensor parameter (pH, temperature, conductivity, liquid level, pH change rate), the mean and standard deviation are dynamically calculated based on historical data within a sliding window. The mean is subtracted from the real-time measurement value, and then divided by the standard deviation to eliminate scale differences between parameters of different dimensions. When a sudden change in parameter is detected (such as a sudden jump in pH exceeding the historical fluctuation range), the statistical window is automatically shortened to the most recent 24 hours of data, and the mean and standard deviation are recalculated to avoid normalization distortion.

[0108] S3, Nonlinear Dimensionality Reduction and Feature Extraction

[0109] A diffusion mapping algorithm is used to construct nonlinear similarity relationships in a high-dimensional data space. The specific steps include:

[0110] Similarity weight calculation: Based on the standardized sample vectors, the squared Euclidean distance between each pair of samples is calculated, and the similarity weight between samples is generated by combining the natural exponential function with the adaptive bandwidth parameter.

[0111] Adaptive bandwidth adjustment: The bandwidth parameter is dynamically adjusted based on the median distance between samples within the sliding window. When the rate of change of pH exceeds a set threshold, the bandwidth is reduced to enhance local sensitivity.

[0112] Laplace matrix construction: The degree matrix is ​​generated using the similarity matrix, and the similarity matrix is ​​normalized by the inverse square root of the degree matrix to obtain the probability transition matrix.

[0113] Low-dimensional feature extraction: Perform eigenvalue decomposition on the normalized matrix, select the two largest nontrivial eigenvectors as the dimensionality reduction result, and characterize the dominant nonlinear correlation features in the data.

[0114] S4. Dynamic Control Command Generation

[0115] The reduced low-dimensional feature vector, the remaining volume of the raw material tank, and the process running time are input into a deep deterministic strategy gradient reinforcement learning model. The model outputs pumping time increment control commands by simulating a collaborative mechanism of expert experience and data-driven decision-making. When the remaining raw material volume falls below a preset threshold, a linkage supply protocol between the transfer tank and the raw material tank is triggered to ensure continuous supply.

[0116] S5, Swarm Intelligence Parameter Optimization

[0117] A quantum particle swarm optimization algorithm is used to perform a global search for the control parameters (including proportional, integral, and differential coefficients, and mixing factors). Parameter performance is evaluated using a multi-objective fitness function that integrates the integral terms of real-time error, the first derivative of the error, and the second derivative. When system oscillations are detected, the weights of higher-order derivative terms are dynamically increased to suppress overshoot and accelerate convergence.

[0118] S6. Verification of the stability of the energy function

[0119] An energy function containing a quadratic error term and a historical error integral term is constructed, and its time derivative is calculated to verify the dynamic stability of the control command. If the derivative does not meet the preset convergence condition, a parameter self-learning update mechanism is triggered to re-optimize the control command to ensure system stability.

[0120] S7. Control Command Execution and Abnormal Switching

[0121] In the normal control mode, the optimized pumping time increment is converted into a pump valve drive signal. When the second-order rate of change of pH is detected to exceed the set threshold, the system automatically switches to a fixed-frequency pulse control mode, which quickly stabilizes the system by adjusting the duty cycle until the energy function recovers the convergence condition, at which point it switches back to reinforcement learning control.

[0122] S8, Self-learning parameter update

[0123] A process database is established to store multidimensional sensor data, reduced-dimensional feature vectors, and control command records, supporting historical data backtracking and model iteration. The neighborhood radius parameter of the diffusion map is adaptively adjusted based on the dynamic rate of change of the low-dimensional feature vectors (e.g., sliding window variance).

[0124] When features fluctuate drastically, expand the neighborhood radius to capture global correlations;

[0125] When features tend to plateau, the neighborhood radius is reduced to enhance local resolution.

[0126] The reinforcement learning model is retrained periodically based on new data, and the network weights are updated to adapt to changes in the process.

[0127] Step S1: Real-time multi-dimensional data acquisition

[0128] In this embodiment, the real-time multi-dimensional data acquisition is achieved through the collaborative deployment of multiple types of sensors inside the soaking tank, specifically including a liquid level sensor, a pH sensor, a temperature sensor, and a conductivity sensor. Each sensor adopts a distributed installation structure, with the liquid level sensor vertically fixed to the central axis of the tank's side wall, the pH and conductivity sensors installed at the bottom of the tank using an immersion probe structure, and the temperature sensor installed at a preset temperature measurement point on the inner wall of the tank using an adhesive structure, ensuring that the measured data is consistent with the actual operating conditions of the cleaning agent.

[0129] The sensor array establishes a communication connection with the main control module via a digital bus. Preferably, an industrial-grade fieldbus protocol is used to achieve high-speed data transmission. The main control module has a built-in multi-channel acquisition card that synchronously reads the analog or digital signals output by each sensor at a preset fixed frequency and converts them into process parameter values. The acquisition frequency is strictly aligned with the pump and valve control cycle to ensure the closed-loop synchronization of data acquisition timing and actuator actions, avoiding control command delays caused by timing misalignments.

[0130] For the acquisition of parameters such as pH value, temperature, conductivity, liquid level, and pH change rate, the main control module employs a time-division multiplexing mechanism to poll and read data from each sensor. Preferably, the pH change rate is generated in real time using a differential calculation method, specifically by performing a first-order derivative calculation on the pH values ​​of two adjacent acquisition cycles. The data is expressed as follows:

[0131] ;

[0132] in, This indicates the change in pH value between two consecutive sampling periods;

[0133] This indicates the pH value measured at the current moment.

[0134] p This indicates the pH measurement value at the previous moment;

[0135] This indicates the time interval between two consecutive data collections.

[0136] To ensure data integrity, the main control module has a built-in data verification mechanism. Preferably, the amplitude range and rate of change of each sensor signal are verified: if the parameter value is detected to exceed the preset physical range (e.g., pH value exceeds the range of 0-14), or the rate of change of adjacent cycles exceeds the process allowable threshold (e.g., temperature jump exceeds ±10℃ / s), an abnormal alarm is triggered and invalid data is discarded, and the redundant sensor switching process is started.

[0137] The liquid level sensor preferably employs a non-contact ultrasonic measurement principle. It generates a liquid level height value by emitting a pulse signal, calculating the echo time difference, and combining this with a cleaning agent dielectric constant compensation algorithm. Preferably, the liquid level measurement value is dynamically compensated using the following formula:

[0138] ;

[0139] in, This is the liquid level measurement value after dynamic compensation of dielectric constant. These are values ​​measured directly by ultrasound. To calibrate the dielectric constant, The dielectric constant is estimated in real time based on conductivity and temperature, thereby eliminating the interference of cleaning agent composition fluctuations on liquid level measurement.

[0140] The synchronization of the multi-dimensional data acquisition is achieved through a hardware interrupt triggering mechanism. At the start of each control cycle, the main control module sends a synchronization acquisition command to all sensors. Upon receiving the command, each sensor immediately freezes its current measurement value and uploads it. Preferably, the transmission delay of the synchronization command is compensated for by a clock calibration algorithm to ensure that the time deviation of each parameter acquisition moment is less than 1% of the control cycle.

[0141] The technical solution in this step fully covers the sensor types, acquisition frequencies, and real-time data requirements specified in the code. Through distributed sensor deployment, time-division multiplexing acquisition, dynamic compensation algorithms, and synchronization control, it provides high-precision, low-noise raw data input for subsequent standardization processing and dimensionality reduction algorithms, forming the basic data layer for high-precision control.

[0142] Step S2: Multi-source data standardization processing

[0143] In this embodiment, the multi-source data standardization processing targets parameters such as pH value, temperature, conductivity, liquid level, and pH change rate. Through a dynamic sliding window statistical and normalization mapping mechanism, it eliminates the dimensional differences of multi-source heterogeneous data, providing standardized input for subsequent nonlinear dimensionality reduction and machine learning modeling.

[0144] The core of the standardization process lies in dynamically calculating the historical mean and standard deviation of each parameter. Preferably, the historical mean and standard deviation are updated in real time based on the process data accumulated within the sliding window, and the window length is adaptively adjusted according to the process stability: when the parameter fluctuation is small, a longer period window is used to calculate the statistics to improve robustness; when the parameter undergoes a sudden change or exceeds the historical distribution range, it automatically switches to a shorter period window for recalculation to avoid normalization distortion caused by outdated statistics.

[0145] The normalization mapping process uses the following formula:

[0146] ;

[0147] in:

[0148] The original data of the i-th parameter is as follows:

[0149] pH value;

[0150] Temperature value (unit: °C);

[0151] Conductivity value (unit: µS / cm);

[0152] Liquid level (unit: meters);

[0153] pH change rate (unit: pH / second);

[0154] Real-time measured values ​​of the original parameters;

[0155] : No. The arithmetic mean of the historical data for each parameter;

[0156] : No. The historical data standard deviation of each parameter, the calculation period and Consistent.

[0157] Preferably, for the liquid level height parameter, a nonlinear saturation function is further introduced to limit the range of the normalization result, avoiding excessive impact of extreme liquid level fluctuations on subsequent algorithms. The saturation function is defined as:

[0158] ;

[0159] in, Standardization and saturation processing results of liquid level height measurements;

[0160] This represents the original measured value of the liquid level height;

[0161] and These represent the historical mean and standard deviation of the liquid level parameter, respectively.

[0162] is a sign function used to restrict the normalization result to the interval [-3, 3].

[0163] When the deviation between the real-time value and the historical mean of a certain parameter exceeds three times the standard deviation, it is determined to be a process mutation event. At this time, the statistical reset process is automatically triggered: discarding the historical data within the current sliding window, recalculating the mean and standard deviation based on the latest collected sample, and using a feedforward compensation mechanism to maintain the continuity of the normalized output during the reset period. Preferably, the feedforward compensation is implemented through an exponentially weighted moving average algorithm.

[0164] ;

[0165] in, This represents the historical mean before the reset.

[0166] This represents the real-time measurement value following the current mutation event;

[0167] This is the forgetting factor, used to balance the weights of new and old data. Its value is dynamically adjusted based on the mutation detection results.

[0168] The standardization processing module is deeply coupled with the diffusion mapping algorithm. Preferably, the normalized data needs to be processed by an outlier detection filter to eliminate outlier interference before being input into the dimensionality reduction algorithm. Specifically, Mahalanobis distance is used to calculate the consistency between the sample and the historical distribution.

[0169] ;

[0170] in, It is a real-time data vector containing pH value, temperature, conductivity, liquid level, and pH change rate;

[0171] A vector composed of the mean values ​​of each parameter;

[0172] Let be the covariance matrix of each parameter;

[0173] The transpose operation represents a vector or matrix;

[0174] The result is the Mahalanobis distance calculation, used to quantify the degree of deviation between the current data and the historical distribution.

[0175] If the Mahalanobis distance exceeds a preset threshold, it is identified as an outlier and a data repair mechanism is activated, such as linear interpolation or replacement of adjacent samples, to ensure the input quality of subsequent algorithms.

[0176] For each parameter, normalization is performed based on its historical mean and standard deviation to eliminate dimensional differences;

[0177] Its normalization formula is:

[0178] ;

[0179] in:

[0180] The original data of the i-th parameter is as follows:

[0181] pH value;

[0182] Temperature value (unit: °C);

[0183] Conductivity value (unit: µS / cm);

[0184] Liquid level (unit: meters);

[0185] pH change rate (unit: pH / second);

[0186] Real-time measured values ​​of the original parameters;

[0187] : No. The arithmetic mean of the historical data for each parameter;

[0188] : No. The historical data standard deviation of each parameter, the calculation period and Consistent.

[0189] This step transforms multi-source sensor data into standardized signals with unified mathematical properties through dynamic statistical windows, normalized mapping, mutation detection, and outlier processing, providing highly consistent input for nonlinear dimensionality reduction. At the same time, it enhances robustness to process fluctuations through an adaptive mechanism.

[0190] Step S3: Nonlinear Dimensionality Reduction and Feature Extraction

[0191] In this embodiment, the nonlinear dimensionality reduction and feature extraction are implemented based on the diffusion mapping algorithm. By constructing a high-dimensional spatial topology structure from the standardized multi-source sensor data, low-dimensional feature vectors are extracted to characterize the key nonlinear correlation features in the cleaning agent delivery process, providing a high-information-density input signal for subsequent reinforcement learning control.

[0192] The core of the diffusion mapping algorithm lies in constructing a similarity weight matrix between data samples. Preferably, the similarity weights are calculated using an exponential function combined with an adaptive bandwidth parameter. For any two standardized sample vectors... and Its similarity weight is defined as:

[0193] ;

[0194] in:

[0195] The first one in a reinforcement learning neural network Input node to the Connection weight parameters for each hidden layer node;

[0196] : No. A standardized sample data vector;

[0197] : Natural exponential function;

[0198] : No. A standardized sample data vector, and Same dimension;

[0199] Euclidean distance operator, representing the spatial distance between two vectors;

[0200] : Adaptive bandwidth parameter, which is the median of the squared Euclidean distances between all samples within the sliding window divided by .

[0201] The bandwidth parameter is dynamically adjusted based on the median distance between samples within the sliding window, and the specific calculation method is as follows:

[0202] ;

[0203] in, The kernel function bandwidth adjustment coefficient is used when constructing the feature space of the diffusion map. This represents the number of samples within the current sliding window. The logarithmic scaling factor is the number of samples. This is a median calculation function used to select the median value from a set of data. When the rate of pH change exceeds a preset threshold, it is automatically... The value is reduced to a specific proportion of the original value to enhance the algorithm's ability to capture local mutation features.

[0204] After the similarity matrix is ​​constructed, a probability transition matrix is ​​generated through normalization. Preferably, the normalization process includes degree matrix construction and Laplacian matrix generation. The diagonal elements of the degree matrix are the sum of the similarity weights of each sample. The similarity matrix is ​​normalized by the inverse square root of the degree matrix to obtain a symmetric normalized Laplacian matrix L. The eigenvalue decomposition of the Laplacian matrix is ​​used to extract low-dimensional feature vectors, preferably using the first two largest nontrivial feature vectors as the dimensionality reduction result.

[0205] The Laplace matrix The construction formula is:

[0206] ;

[0207] in:

[0208] : Degree matrix, which is a diagonal matrix, whose diagonal elements Indicates the first The sum of the similarities between a sample and all other samples;

[0209] Similarity matrix, with dimensions N×N;

[0210] The inverse square root matrix of the degree matrix, obtained by... It is obtained by taking the reciprocal of the square root of each diagonal element.

[0211] To enhance the algorithm's adaptability to process fluctuations, the low-dimensional feature vector is dynamically fused with the process running time and the remaining raw material. Preferably, a time decay factor is introduced to perform a weighted average of historical feature vectors to balance the weights of current and historical features. The decay factor is dynamically adjusted according to the process running time to ensure the real-time performance and stability of the feature vector.

[0212] This step compresses high-dimensional sensor data into a low-dimensional space through adaptive bandwidth adjustment, probability transition matrix construction, and dynamic eigenvector fusion. This eliminates redundant information while retaining key nonlinear correlations, providing robust feature inputs for subsequent control command generation.

[0213] Step S4: Dynamic Control Command Generation

[0214] In this embodiment, the reinforcement learning control command generation is based on the dynamic Q-learning algorithm. By integrating low-dimensional feature vectors, process running time, and raw liquid balance parameters, a state-action value function model is constructed, and pump valve start / stop and flow regulation commands are output in real time to achieve closed-loop optimization control of the cleaning agent delivery process.

[0215] The state space of the Q-learning model consists of three parts: the low-dimensional feature vector (dimension 2) extracted by the diffusion mapping algorithm, the cumulative process running time (scalar), and the remaining amount in the raw material storage tank (scalar).

[0216] ;

[0217] in, and These are the eigenvector components after dimensionality reduction;

[0218] for The system state vector at any given time;

[0219] This represents the cumulative running time of the process from its start to the current moment.

[0220] The real-time remaining volume of the raw material storage tank is obtained by converting data from the liquid level sensor.

[0221] The action space is defined as a set of discrete control commands, including main pump start / stop, regulating valve opening, and bypass valve switching operations. Preferably, the action vectors employ a multi-dimensional encoding method, for example:

[0222] Main pump commands (0 / 1 indicates off / on);

[0223] Control valve opening (0%-100% discretized according to preset gradient);

[0224] Bypass valve status (0 / 1 indicates disabled / enabled).

[0225] reward function The design prioritizes process stability and cleaning efficiency as core indicators, specifically including a penalty for pH deviation from the target range, a reward for the rate of concentrate consumption, and an inhibition for frequent control command switching. Preferably, the reward function achieves a multi-objective trade-off through the following formula:

[0226] ;

[0227] in, The weighting coefficients are determined through offline strategy optimization.

[0228] This indicates the amount of original solution consumed per unit time.

[0229] This is the cumulative value of the changes in motion within adjacent control cycles.

[0230] It represents the absolute deviation between the current pH measurement value and the target value, and is used to quantify the pH control accuracy.

[0231] To balance the conflict between exploration and exploitation, an ε-greedy strategy is adopted to dynamically adjust the action selection mechanism. Preferably, the exploration probability ε exhibits a piecewise decay characteristic with the training period: a high exploration rate (e.g., 0.5) is maintained in the initial stage to fully traverse the state space; when the cumulative reward growth rate is lower than the threshold, the exploration rate is gradually reduced to below 0.1, focusing on local optimization of the strategy.

[0232] This step achieves dynamic optimization control of the cleaning agent delivery process through multi-dimensional fusion of state space, composite design of reward function, and adaptive learning rate mechanism. It overcomes the shortcomings of traditional PID control in adaptability to nonlinear and time-varying systems, and constitutes the core decision layer of high-precision intelligent control.

[0233] Step S5: Swarm Intelligence Parameter Optimization

[0234] In this embodiment, the control command verification and dynamic correction module realizes closed-loop verification of reinforcement learning output commands through a multi-level verification mechanism. Combined with real-time sensor data and process constraints, it ensures the safety and feasibility of control commands. At the same time, it corrects commands under abnormal operating conditions online to ensure the stable operation of the cleaning agent delivery system.

[0235] The multi-level verification mechanism includes three levels: pre-action simulation, execution result prediction, and redundant instruction arbitration. Preferably, the pre-action simulation module simulates the changes in process parameters after execution based on the current system state (including pH value, temperature, liquid level, and remaining stock solution) and the control instruction to be executed, using dynamic equations. The dynamic equations are expressed as follows:

[0236] ;

[0237] in, This is a simulated quantity representing the dynamic change of pH value based on the current control parameters. For the pump's flow rate command, This is a command to control the valve opening. This is a pH change rate prediction model trained based on historical data. When the simulation results exceed the preset safety range, a correction procedure is triggered.

[0238] The execution result prediction module assesses the potential risks of instruction execution using a time-series correlation model. Preferably, an LSTM (Long Short-Term Memory)-based prediction model is used, taking the current state vector and control instructions as input, and outputting predicted values ​​for pH fluctuations, raw material consumption rate, and equipment load over the next three control cycles. If any indicator in the predicted values ​​exceeds the process constraint threshold, it is determined to be a high-risk instruction and an arbitration mechanism is initiated.

[0239] The redundant instruction arbitration module generates a set of alternative instructions based on the risk level. Preferably, when a high-risk instruction is detected, the safest instruction with the highest similarity to the current state is retrieved from the historical best instruction library as a candidate; simultaneously, the original instruction is slightly adjusted based on the gradient descent method to generate a corrected instruction.

[0240] To address command distortion caused by sensor anomalies or communication delays, a dynamic confidence threshold is set for anomaly detection. Preferably, the confidence index is calculated using the following formula:

[0241] ;

[0242] in, for Confidence evaluation index for time-state prediction This is the predicted state from the previous cycle. This is the actual observation state. To prevent extremely small constants from being divided by zero, when the confidence level falls below a threshold, the system automatically switches to a conservative control mode, employing a fixed-frequency pump-valve linkage strategy to maintain the basic cleaning process until the system returns to steady state.

[0243] The state space of the reinforcement learning is defined as follows:

[0244] ;

[0245] in:

[0246] in:

[0247] : The system state vector at any given time;

[0248] : Eigenvalues ​​of low-dimensional manifolds (dimensionless real numbers);

[0249] The remaining volume of the raw material in the tank is measured in real time by a mass flow meter (unit: liters);

[0250] Process running time, accumulated since system startup (unit: seconds).

[0251] This step employs a three-layer verification architecture: pre-action simulation to avoid physically infeasible instructions, LSTM prediction model to assess long-term risks, and redundant instruction arbitration to generate safe alternatives. Combined with a dynamic confidence mechanism, this achieves closed-loop verification and adaptive correction of instructions, ensuring the decision-making robustness of the intelligent control system and forming the core guarantee layer for safety control.

[0252] Step S6: Energy Function Stability Verification

[0253] In this embodiment, the adaptive model update and long-term optimization module continuously iterates and strengthens the learning strategy through an online learning mechanism. Combining real-time process data and historical experience database, it dynamically adjusts the control model parameters to solve the impact of time-varying factors such as equipment aging and environmental disturbances on control accuracy, thereby achieving full life cycle optimization of the cleaning agent delivery system.

[0254] The model update mechanism consists of three parts: triggering conditions, incremental learning, and weight fusion. Preferably, the triggering conditions are based on a comprehensive judgment of the average reward value change rate within the sliding window and the state prediction error: when the average reward value decrease rate exceeds a preset threshold for multiple consecutive control cycles, or when the mean square error between the PH prediction model output value and the actual measured value continues to increase, the model update process is automatically initiated.

[0255] The prediction error Calculated using the following formula:

[0256] ;

[0257] in, The number of samples within the sliding window. For predicted values, These are actual measured values. As an index for process stages such as pre-cleaning, pickling, and neutralization.

[0258] The incremental learning process employs a strategy combining experience replay and transfer learning. Preferably, sample data with high similarity to the current working conditions are selected from the historical experience pool and injected into the training set through importance sampling weighting, avoiding the model's loss of generalization ability due to overfitting to local data. The sample similarity calculation uses the feature space projection method, mapping the real-time state vector and historical samples to a low-dimensional space generated by the diffusion mapping algorithm, and measuring similarity using Euclidean distance.

[0259] The energy function is a linear combination of a quadratic function based on the system error and the error integral term, and the control command is considered stable when the time derivative of the energy function satisfies the preset convergence condition.

[0260] The core construction formula for the energy function is:

[0261] ;

[0262] in:

[0263] Real-time pH;

[0264] The integral time variable represents the time span from the start of the process to the current moment.

[0265] The time derivative of the energy function;

[0266] Time infinitesimal element in integral operations;

[0267] : Real-time running time variables of the control system;

[0268] Convergence condition: ,in This is the preset positive convergence rate coefficient.

[0269] To address the issue of data distribution drift, a model degradation monitoring and rollback mechanism is implemented. Preferably, the degradation index is determined by comparing the stability of the model on the retained validation set before and after the update, specifically including the combined changes in instruction switching frequency, pH fluctuation variance, and concentrate consumption rate. When degradation is detected, the model is automatically rolled back to the previous stable version, and an enhanced data acquisition process is triggered to supplement training samples for the current operating conditions.

[0270] This step achieves continuous self-optimization of the control model through dynamic triggering conditions, incremental learning, and weight fusion mechanisms. Combined with degradation monitoring to ensure the long-term stability of the system, it forms an intelligent control closed loop with environmental adaptability, overcoming the performance degradation problem of traditional static models during long-term operation.

[0271] Step S7: Control command execution and exception switching

[0272] In this embodiment, the human-machine collaborative intervention and emergency takeover module realizes dynamic collaboration between the intelligent control system and manual operation through a multimodal interaction interface and an abnormal working condition graded response mechanism. At the same time, it triggers a fully automatic emergency takeover process when an unrecoverable fault is detected, ensuring the safety and fault tolerance of the cleaning agent delivery system.

[0273] The multimodal interaction interface includes a process status visualization panel, risk warning prompts, and a manual command input channel. Preferably, the status visualization panel displays the time-series change curves of core parameters such as pH value, temperature, and liquid level in real time, and overlays a feature space distribution map after dimensionality reduction by diffusion mapping, highlighting abnormal clustering areas to assist manual decision-making. The risk warning prompt generates a graded alarm signal based on the confidence index in step S5 and the model degradation monitoring results in step S6. For example, a low-risk alarm prompts a suggestion for manual review of control commands, while a high-risk alarm forcibly locks automatic control and activates takeover preparation.

[0274] The manual command input channel supports two collaborative modes: priority coverage and weight fusion. Preferably, in priority coverage mode, manually input control commands directly replace the reinforcement learning output for rapid intervention in unexpected situations. In weight fusion mode, manual and automatic commands are dynamically weighted to generate the final control quantity. For example, the human-machine collaborative control command fusion formula is as follows:

[0275] ;

[0276] in, This represents the final control command vector to be executed, with dimensions matching the specific control type (e.g., flow rate in L / min, valve opening in percentage). Represents the dynamic weighting coefficient of manual instructions, a dimensionless scalar (value range 0, 1). This represents the control command vector input manually, which is set by the operator through the interactive panel. This represents the automatic control instruction vector output by the reinforcement learning model, generated by the Q-learning strategy in step S4.

[0277] Emergency takeover triggering conditions are based on a comprehensive judgment of multi-dimensional anomaly detection results. Preferably, fully automatic emergency takeover is activated when the following conditions are met simultaneously:

[0278] The confidence index in step S5 remains below the critical threshold.

[0279] The model degradation monitoring in step S6 shows that the strategy has failed;

[0280] Key parameters (such as pH value and remaining stock solution) deviate from the safe range and automatic correction fails.

[0281] After the takeover process is initiated, the system switches to a predefined conservative control strategy library, employing a rule-based multi-level flow regulation and equipment interlocking mechanism. Preferably, the rules in the conservative strategy library are generated through reverse engineering of historical accident cases and process constraints. For example, when the pH value exceeds the safety upper limit, the main pump is automatically shut down and the dilution valve is opened, while the emergency neutralizing agent dosing device is activated. During the execution of these rules, the parameter recovery status is monitored in real time. If the system does not return to steady state within a preset time, emergency measures are gradually escalated until the system shuts down.

[0282] This step enhances the effectiveness of human supervision through visual interaction, balances the conflict between human and machine decision-making through a weight fusion mechanism, and ensures system security under extreme conditions through tiered emergency takeover. It forms a fault-tolerant system that complements intelligent control and human experience, and solves the black-box risk and rigid response problems in fully automated systems.

[0283] Step S8 Self-learning parameter update

[0284] In this embodiment, the full-process closed-loop monitoring and knowledge accumulation module achieves the traceability of the intelligent control system's status, the iterability of its strategies, and the analyzability of its faults throughout its entire lifecycle through three core functions: multi-source data archiving, strategy version management, and self-diagnostic report generation, thus forming a continuously optimized technical closed loop.

[0285] The multi-source data archiving function collects and stores intermediate process data from steps S4 to S7 in real time, including but not limited to: state-action records of the reinforcement learning model, instruction verification result logs, model update version information, and human-machine collaborative operation history. Preferably, the data storage adopts a hierarchical structure: raw sensor data is stored in a high-speed cache layer in time series form for real-time monitoring; process decision data (such as Q-value tables and intermediate quantities for reward function calculation) is stored in a relational database to support policy backtracking analysis; and the long-term knowledge base stores historical optimal policy sets and typical working condition cases through a distributed storage system.

[0286] The strategy version management mechanism uses semantic tags to control the entire lifecycle tracking of the model. Preferably, each model update (triggered in step S6) generates a version snapshot containing the following metadata: the reason for the update (e.g., reward decline, prediction error exceeding the limit), the feature distribution of the training dataset, the performance metrics of the validation set, and the weight fusion coefficients. The version rollback function allows the system to quickly switch to a historical stable version when performance degradation is detected, while retaining the context data of the failed version for root cause analysis.

[0287] The self-diagnostic report generation module identifies potential system defects by performing correlation analysis on cross-step data anomalies. Preferably, a diagnostic report is automatically generated when the following correlated events are detected:

[0288] The sudden increase in the instruction correction frequency in step S5 and the model update failure event in step S6 occur simultaneously.

[0289] The human-machine collaboration weight remains high and the Q-value table entropy value in step S4 decreases significantly;

[0290] There is a spatiotemporal correlation between the steady-state deviation of process parameters and the abnormal clustering distribution of feature space in step S3.

[0291] The report uses visual charts overlaid with original data snapshots to mark the timestamps and logical connections of abnormal event chains, helping maintenance personnel to quickly locate hardware and software coupling faults.

[0292] The knowledge accumulation mechanism transforms validated control strategies and human intervention experience into a standardized case library. Preferably, the case library entries include: operating condition feature vectors (after dimensionality reduction in step S3), optimal control instruction sets, environmental disturbance types, and boundary constraints. New cases are admitted through a dual verification mechanism: first, their physical feasibility is verified by the redundant arbitration module in step S5, and then their compatibility with the existing knowledge base is verified through the incremental learning process in step S6.

[0293] This step builds decision-making traceability capabilities through end-to-end data archiving, ensures system maintainability through version management mechanisms, and enables experience transfer through self-diagnosis and knowledge accumulation, forming a self-evolutionary closed loop for the intelligent control system and solving the common problems of data silos and knowledge loss in industrial scenarios.

[0294] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A high-precision and low-delay detergent dosing system, characterized by, The application relates to a cleaning agent quantitative delivery system, which comprises the following parts: a clean water tank, a raw solution tank, a soaking tank and a turnover tank, wherein the tank bodies are connected through pipelines and provided with pump valve assemblies; a sensor group containing liquid level sensors, PH sensors, temperature sensors and conductivity sensors; a control module configured to: collect the pH value, temperature, conductivity, liquid level and PH change rate data of the soaking tank in real time, perform standardization processing on the data, and generate a low-dimensional feature vector through nonlinear dimension reduction; generate a raw solution pumping control instruction dynamically through a machine learning model based on the low-dimensional feature vector and process state parameters; globally optimize the parameters in the control instruction by using a swarm intelligence algorithm, and verify the stability of the control instruction in combination with an energy function of a dynamic system; output the verified control instruction to an executing mechanism to drive the pump valve assembly to complete quantitative delivery of the cleaning agent; update the dimension reduction parameters and the control model based on current process data after each delivery cycle ends; the nonlinear dimension reduction adopts a diffusion mapping method to construct a high-dimensional space topology structure by calculating a similarity matrix between sensor data; the calculation formula of the similarity matrix is as follows: wherein: jth normalized sample data vector; exp: a natural exponential function; the kth normalized sample data vector, with the same dimensionality; ||.||: a Euclidean distance calculation operator, representing the spatial distance between two vectors; ∈: an adaptive bandwidth parameter, which is the median of the Euclidean distance squares between all samples in a sliding window divided by ln(N+1), wherein ln(N+1) is a logarithmic scaling factor of the sample quantity, and N is the total number of samples in the current sliding window; the generation of the low-dimensional feature vector comprises: performing eigenvalue decomposition on a Laplace matrix generated by the diffusion mapping, and selecting the first two maximum non-trivial eigenvectors as the dimension reduction result; the construction formula of the Laplace matrix L is as follows: L = D -1 / 2 WD -1 / 2 ; wherein: D: degree matrix, which is a diagonal matrix with diagonal elements D jj represents the sum of the similarities of the jth sample with all other samples; W: a similarity matrix with a dimension of N*N; D -1 / 2 : inverse square root matrix of the degree matrix D jj , obtained by taking the inverse square root of each diagonal element of D the swarm intelligence algorithm is a quantum particle swarm optimization algorithm, wherein each particle encodes a group of control parameters, including proportional integral differential coefficients and mixing factors, and the particle performance is evaluated through a multi-objective fitness function containing error items and their first and second derivatives; the multi-objective fitness function is defined as follows: wherein: x m : the parameter set encoded by the mth particle, specifically: K p : proportional control coefficient; K i : integral control coefficient; K d : differential control coefficient; alpha: a mixing factor, with a value range of 0-1; e: a real-time pH error, calculated as the difference between the current pH value and the target value; first order time derivative of the pH error; The second time derivative of pH error; the energy function is a linear combination of a quadratic function based on system error and an error integral item, and the control instruction is determined to be stable when the time derivative of the energy function meets a preset convergence condition; the core construction formula of the energy function is as follows: wherein: e: a real-time pH error; tau: an integral time variable, representing the time span from the beginning of the process to the current time; dtau: a time infinitesimal in the integral operation; t: a real-time running time variable of the control system; Convergence condition: where η is a preset positive convergence rate coefficient, Time derivative of the energy function.

2. The high-precision and low-delay rinse agent dosing system according to claim 1, characterized in that, the collected data includes the pH value, temperature, conductivity, liquid level height and PH change rate, and the collection frequency of each parameter is one per second, and high-frequency data collection is realized through a multi-thread parallel processing technology, which is synchronized with the control cycle of the pump valve assembly to form a closed-loop feedback mechanism.

3. The high-precision and low-delay rinse agent dosing system according to claim 1, characterized in that, the standardization processing includes: performing normalization calculation on each parameter based on the historical mean and standard deviation of each parameter to eliminate the dimensional difference; the normalization formula is as follows: wherein: raw data of the ith parameter, specifically: i = 1: pH value; i = 2: temperature value (unit: ℃); i = 3: conductivity value (unit: μS / cm); i = 4: liquid level height (unit: meter); i = 5: pH value change rate (unit: pH / s); X i : real-time measured value of the original parameter; μ i : arithmetic mean of the historical data of the i-th parameter; σ i : Standard deviation of historical data of the i-th parameter, calculated over the same period as μ i .

4. The high-precision and low-delay rinse agent dosing system according to claim 1, characterized in that, The machine learning model is a deep deterministic policy gradient reinforcement learning model, and the input of the machine learning model is the reduced dimension feature vector, the stock solution residual amount and the process running time, and the output of the machine learning model is the stock solution pumping time increment; The state space of the reinforcement learning is defined as: s t = [v1, v2, Q remain , t proc ]; Wherein: s t : system state vector at time t; v1, v2: low-dimensional manifold eigenvalue (dimensionless real number); Q remain : remaining amount of stock solution tank, measured in real time by mass flow meter (unit: liter); t proc : Process up time, accumulated from system start (in seconds).

5. The high-precision and low-delay rinse agent dosing system according to claim 1, characterized in that, The control instruction execution comprises: when the second-order change rate of the pH value exceeds a set threshold, automatically switching to a pulse control mode with a fixed frequency and duty cycle.

6. The high-precision and low-delay rinse agent dosing system according to claim 1, characterized in that, The process of updating the dimension reduction parameter and the control model based on the current process data comprises: storing a process database of multi-dimensional sensor data, a dimension reduction feature vector and a control instruction, and adaptively adjusting a neighborhood radius parameter of the diffusion mapping according to a dynamic change rate of the feature vector.

Citation Information

Patent Citations

  • Intelligent water-saving pipeline cleaning optimization system based on artificial intelligence

    CN118153783A

  • Parameter optimization method and device of coupling type transformer

    CN119644765A