Intelligent automatic dosing method based on complex sewage treatment

CN122520148APending Publication Date: 2026-08-07INNER MONGOLIA DONGYUAN ENVIRONMENTAL PROTECTION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INNER MONGOLIA DONGYUAN ENVIRONMENTAL PROTECTION TECH CO LTD
Filing Date
2026-07-13
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]现有技术需要解决的主要技术问题在于:在复杂污水的表观检测指标相近而污染组分及反应需求不同,且进水自然波动、前序残余药效与当前加药响应在时间上相互叠加的条件下,控制系统无法从常规运行数据中辨识当前水体的真实药耗状态

Benefits of technology

1.以水质状态向量的变化方向、持续长度、跨指标联动关系和加药响应滞后划分水质同质区间,使后续辨识建立在反应关系连续的时间范围内,减少不同水质阶段的数据相互混入。针对隐性药耗状态不确定度较高的区间,将原计划药量重排为控制时域内累计药量相等的基准投加序列和微扰投加序列,通过前段正向微扰与后段等量补偿形成区别于自然进水波动的受控剂量变化,同时不增加控制时域的总计划药量。微扰产生的实际响应经流量传播时差校正后,与由相似历史水质区间构建的反事实响应轨迹进行比较,响应斜率差、拐点时差、累计响应面积差和恢复时长差共同构成带有药剂类型及水质区间标识的因果响应指纹。所述因果响应指纹把当前加药动作对应的响应变化从进水组成变化和前序残余药效形成的共同变化中分离,并为电荷中和需求、氧化还原需求、酸碱缓冲需求和难降解干扰强度的修正提供可追溯的观测依据。隐性药耗状态由此不再只依赖表观检测值的静态映射,而是依据受控微扰后的实际响应进行更新;在酸碱度、浊度、电导率或目标污染物指标相近但污染组分不同的情况下,控制过程仍可区分对应药剂需求的变化方向,并据此求解多药剂需求向量及分段释放次序。上述数据划分等总量时域微扰、反事实比较和状态修正连续衔接,使自然水质波动不易被误记为药剂剂量响应,减少模型在水质组分切换后沿原错误方向连续补药的情况,使投加决策与当前反应需求保持一致。当当前因果响应指纹与预测响应指纹不一致时,控制过程不沿既有状态直接放大药量,而是转入候选主导状态验证,使状态判断具有再次校验路径,避免一次错误估计固化为连续控制偏差。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122520148A_ABST
    Figure CN122520148A_ABST
Patent Text Reader

Abstract

The present application relates to sewage treatment automatic control and wastewater treatment technical field, specifically to an intelligent automatic dosing method based on complex sewage treatment. The method collects the inflow, water quality data, historical dosing data and effluent response data in the control cycle, divides the water quality homogeneity interval, and constructs the implicit drug consumption state including charge neutralization demand, oxidation-reduction demand, acid-base buffer demand and refractory interference intensity; when the state is uncertain, the planned dosing amount is rearranged into a benchmark dosing sequence and a perturbation dosing sequence with equal cumulative drug amount, the actual response trajectory is collected and the counterfactual response trajectory is generated, the causal response fingerprint is formed according to the time sequence difference of the two trajectories, the implicit drug consumption state is corrected, and the multi-drug constrained dosing amount and the segmented release order are determined. The method can distinguish complex water quality with similar apparent indicators but different drug consumption needs, reduce repeated drug supplement and dosing oscillation, and reduce residual drug, additional salt load and chemical sludge increment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of automatic control and wastewater treatment technology, specifically to an intelligent automatic dosing method for complex wastewater treatment. Background Technology

[0002] Complex wastewater treatment processes typically determine reagent dosage based on influent flow rate, pH, oxidation-reduction potential, conductivity, turbidity, and target pollutant concentration. Existing automatic dosing systems generally collect this process data continuously at both the influent and effluent ends. The base dosage is obtained by multiplying the influent flow rate by a preset unit water volume reagent consumption coefficient, and then feedback correction is performed based on the deviation between the effluent detection value and the control target. For continuous treatment stages such as pH adjustment, oxidation-reduction, and coagulation, a common control method involves setting reagent dosing curves, allowable dosing ranges, sampling periods, and response delays in the controller, and using proportional-integral-derivative control, fuzzy control, or piecewise rule control to calculate pumping commands at each moment. When the detection value exceeds the set range, the system increases the reagent dosage according to the magnitude of the deviation; when the detection value returns to the set range, the system reduces the reagent dosage or maintains the current dosage. To avoid frequent start-ups and shutdowns caused by instantaneous fluctuations, some solutions also perform moving averages on the detection data and set dead zones, hysteresis zones, or maximum adjustment amounts per cycle. The control methods of this type are mainly based on apparent water quality indicators and empirical chemical consumption coefficients. They assume that the same detection indicators correspond to similar chemical reagent requirements and treat subsequent changes in effluent as a direct feedback result of the combined effect of current dosing and influent changes. This approach is suitable for wastewater with relatively stable composition and well-defined reaction patterns. However, when pollutant types, colloidal charges, buffering capacity, and reducing components continuously change, the empirical mapping relationship is prone to shift.

[0003] To adapt to complex water quality conditions, some existing technologies introduce data-driven predictive models on top of conventional feedback control. These models typically align influent flow rate, water quality test values, chemical dosage, and effluent indicators from several historical control cycles according to preset residence times to form training samples. Regression models, neural networks, or time-series predictive models are then used to establish a mapping relationship between water quality status and chemical dosage. During online operation, the predictive model receives current water quality data and previous dosing records, outputs the recommended chemical dosage for the next control cycle, and the controller then limits the dosage based on the chemical dosing upper limit, effluent control boundary, and equipment execution capability. Other schemes incorporate state observers, treating unmeasurable pollution loads or reaction intensities as implicit states. The state estimates are corrected based on the residual between the predicted and actual responses, and various chemical dosing combinations are obtained through rolling optimization. To handle reaction lag, fixed delays, historical windows, or the moment of maximum correlation are typically used to determine the correspondence between dosing actions and effluent responses. To update the model, newly added data during operation is re-added to the sample set, and model parameters are periodically adjusted. These technologies still primarily rely on statistical correlations formed from naturally occurring operational data. Changes in influent water quality, residual effects of preceding chemicals, and effects of current dosing often co-occur on the detection curve within a similar timeframe. Training samples are unlikely to provide controlled information that can individually identify the true effect of a particular dosing action. Consequently, the model may treat accompanying changes as dose responses and continue to use the original mapping relationship even after the water quality composition changes.

[0004] The main technical problem that existing technologies need to solve is that, under conditions where the apparent detection indicators of complex wastewater are similar but the pollutant components and reaction requirements differ, and where natural fluctuations in influent, residual effects of preceding reagents, and the current dosing response are superimposed over time, the control system cannot identify the true reagent consumption status of the water body from routine operating data. When the surface charge, reducing substance content, acid-base buffering substance content, and recalcitrant interfering components of colloidal particles change, the pH, turbidity, conductivity, or target pollutant concentration may still be within a similar range, but the charge neutralization, redox equivalent, and acid-base adjustment required to achieve the same treatment target are not consistent. If the reagent dosage is calculated solely based on current detection values ​​or historical correlation models, it is difficult to determine whether changes in subsequent detection curves are caused by changes in influent composition, delayed arrival of preceding reagents, or reactions of the current reagent, thus lacking verifiable evidence for the direction of correction of implicit states. Even increasing the sampling frequency, expanding the training sample, or shortening the control cycle only increases the number of observations of the same mixed causal process and cannot provide mutually distinguishable response information for different reagent consumption states. Therefore, when the water quality is the same but the demand is different or when the water quality components are switched, the system may still continuously adjust along the wrong dosage direction, which will prevent the real reagent demand from being identified in time and cause the dosing decision to be inconsistent with the actual response demand. Summary of the Invention

[0005] The purpose of this invention is to provide an intelligent automatic dosing method for complex wastewater treatment, which can solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: The intelligent automatic dosing method for complex wastewater treatment includes: collecting influent flow rate, water quality characterization data, historical dosing data and effluent response data during the control cycle; dividing the water quality characterization data into homogeneous water quality intervals; and constructing a hidden dosing state for each homogeneous water quality interval, including charge neutralization requirements, redox requirements, acid-base buffering requirements and the intensity of recalcitrant interference. The state uncertainty is determined based on the estimated dispersion of each latent state component. When the state uncertainty meets the perturbation triggering condition, the planned dosage in the control time domain is rearranged into a baseline dosage sequence and a perturbation dosage sequence with equal cumulative dosage but different timing. The perturbation is applied according to the perturbation application sequence and the actual response trajectory is collected. A counterfactual response trajectory is generated based on the baseline application sequence. A causal response fingerprint is generated based on the temporal difference between the two response trajectories. The latent drug consumption state is corrected based on the causal response fingerprint. The multi-drug constraint dosage and segmented release sequence are determined based on the corrected latent drug consumption state, and the dosage is executed.

[0007] Preferably, the water quality characterization data is divided into homogeneous water quality intervals, including: combining influent flow rate, pH, oxidation-reduction potential, conductivity, turbidity and target pollutant indicators into a water quality state vector according to the sampling time; For adjacent water quality state vectors, calculate the consistency of change direction, length of continuous change, and cross-index linkage relationship respectively. The sampling time when the index change direction reverses and the linkage index combination switches between the front and back windows is determined as the candidate boundary time. Based on the response lag between historical dosing data and effluent response data, the candidate boundary time is corrected by shifting it forward or backward. After correction, adjacent intervals with continuous effluent response slopes, connected response lags, and no further switching of cross-index linkage relationships in the water quality state vector are merged to obtain water quality homogeneous intervals corresponding to the dosing response, and an independent state evolution record is configured for each water quality homogeneous interval.

[0008] Preferably, the planned dosage is rearranged into a baseline dosing sequence and a perturbation dosing sequence, including: constructing a state covariance based on the estimated dispersion of each latent state component, and determining the state component to be identified by combining the recent dose response residual and the response sensitivity of each drug. Within the current control time domain, a candidate perturbation position corresponding to the state component to be identified is selected. A portion of the drug amount located at the candidate perturbation position in the baseline dosing sequence is moved forward to form a positive perturbation, and an equal amount of drug amount is deducted at subsequent positions to form a compensation perturbation. The sequence combinations that cannot be executed are screened out according to the instantaneous dosage limit, the interval between adjacent dosages, the effluent control boundary, and the residual efficacy of the preceding drug. Select the perturbation delivery sequence with the maximum response discrimination from the remaining sequence combinations, and keep the cumulative drug dosage of the perturbation delivery sequence and the reference delivery sequence equal in the current control time domain.

[0009] Preferably, the generation of counterfactual response trajectory and causal response fingerprint includes: selecting reference intervals from historical water quality homogeneous intervals divided according to the same steps that meet the matching conditions with the current water quality homogeneous interval in terms of inflow trend, initial value of hidden drug consumption state, previous residual drug effect and flow propagation characteristics. Remove response segments within the reference interval that contain non-corresponding reagent adjustments or sudden changes in influent, and align the remaining response segments in time according to the flow propagation time difference; The aligned response fragments are used to generate the counterfactual response trajectory corresponding to the baseline delivery sequence; The difference in response slope, inflection point time, cumulative response area, and recovery time between the actual response trajectory and the counterfactual response trajectory are extracted respectively. The differences are then combined according to the matching confidence of the reference interval to form a causal response fingerprint with reagent type and water quality interval identifiers.

[0010] Preferably, the construction of the implicit drug consumption state includes: for each water quality homogeneous interval, forming an observation residual vector by composing the response increments of pH, redox potential, turbidity, and target pollutant indicators relative to the start of the interval; Extract the historical response directions corresponding to charge neutralization demand, redox demand, acid-base buffering demand and recalcitrant interference intensity from the state evolution record. Subtract the projection components of each historical response direction onto the previous response direction and normalize them to form mutually distinguishable state-sensitive basis vectors. The observation residual vector is projected onto each state-sensitive basis vector to obtain the observation increment of each latent state component; By fusing the observed increment with the state estimate of the previous homogeneous water quality interval based on the continuity of the water quality state vector, the previous drug efficacy attenuation, and the interval response reliability, the implicit drug consumption state and state covariance of the current homogeneous water quality interval are formed.

[0011] Preferably, the selection of the perturbation delivery sequence includes: for each candidate perturbation position, establishing a pair of dose segments consisting of a positive perturbation and a compensation perturbation, and determining the time interval between the two dose segments based on the historical response delay corresponding to the state component to be identified; A state response model is established using the historical dose response trajectories of each latent state component. Each pair of dose segments is combined with the two ends of the uncertainty interval of the state component to be identified to obtain candidate response trajectories. Calculate the separability of candidate response trajectories in terms of response direction, inflection point order and duration, and use the effluent control boundary, previous residual efficacy and cumulative dosage conservation as sequence feasibility constraints. When the compensating perturbation does not meet the feasibility constraints at its original position, the compensating perturbation is moved along the response propagation direction within the current control time domain until a perturbation dosing sequence with the highest cumulative drug dosage and the same as the baseline dosing sequence is formed.

[0012] Preferably, the filtering and time-series alignment of the reference interval includes: combining the influent flow rate and water quality characterization data of multiple consecutive sampling times before the perturbation occurs into a water quality state vector, and forming a prior historical segment together with the reagent dosing trajectory and the effluent change trajectory. Search for candidate reference intervals in historical water quality homogeneous intervals that are similar to previous historical segments and have not experienced corresponding perturbations; Based on the changes in influent flow rate and the sequential response relationship of various water quality indicators, the propagation time axis of the current homogeneous water quality interval and the candidate reference interval are estimated respectively, and the candidate reference interval is mapped to the current propagation time axis. Based on the similarity of previous historical fragments, the similarity of the initial value of implicit drug consumption, and the continuous length of no non-corresponding drug adjustment, a combination weight is assigned to the mapped candidate reference interval. The weighted median response trajectory is used as the counterfactual response trajectory, and the response dispersion at each sampling time is recorded as the matching confidence.

[0013] Preferably, the step of correcting the latent drug consumption status based on the causal response fingerprint includes: writing the historical causal response fingerprint and the corresponding latent status component into the fingerprint status corresponding record according to the drug type identifier and the water quality interval identifier, and extracting the fingerprint response template with the same identifier as the current causal response fingerprint; Align the current causal response fingerprint with each fingerprint response template to obtain the state correction candidate quantity corresponding to each hidden state component; The correction weights are determined based on the matching credibility of the counterfactual response trajectory, the continuity of the actual response, and the degree of conflict of each state correction candidate quantity; The implicit drug consumption state is updated based on the modified weight and the dispersion of the state modification candidate quantity, and the state covariance is generated. The updated implicit state components are combined with the drug causal response coefficients in the fingerprint response template to form a constraint relationship. The multi-drug demand vector is solved, and the multi-drug constrained dosage is generated by combining the drug reaction constraints and the remaining release positions.

[0014] Preferably, updating the latent drug consumption state and generating the state covariance further includes: weighting and combining each latent state component before the update with the corresponding fingerprint response template to form a predicted response fingerprint, and comparing the response direction, inflection point time difference and cumulative response area difference between the predicted response fingerprint and the current causal response fingerprint. When consistency does not meet the state acceptance condition, each latent state component is set as a candidate dominant state. Initial posterior weights are generated based on the matching degree between the current causal response fingerprint and the fingerprint response template of each candidate dominant state, and the corresponding verification drug amount is divided from the multi-drug constraint dosage. The validation doses were released in descending order of initial posterior weights, and local response fragments were collected after each release to form validation fingerprints. The posterior weights are updated based on the matching relationship between the verification fingerprint and each fingerprint response template. The candidate dominant state with the highest posterior weight is selected to recalculate the latent drug consumption state and the multi-drug constraint dosage.

[0015] Preferably, the segmented release order includes: dividing the multi-drug constrained dosage into a basic release amount, a validation dosage, and a reserved release amount according to the posterior weight of the candidate dominant state, and configuring them to the remaining release positions in the current control time domain in the order of basic release amount first, validation dosage in the middle, and reserved release amount last. After each release, extract the local response fragment corresponding to the corresponding drug response delay, and update the causal response fingerprint, implicit drug consumption state, and remaining drug demand vector. When the updated candidate dominant state changes, the unreleased reserved release amount is redistributed among the new multi-agent demand vector, while keeping the sum of the released and unreleased amounts of each agent from not exceeding the corresponding agent constraint dosage. At the end of the current control time domain, the final latent drug consumption state, verification fingerprint, and actual release sequence are written into the record corresponding to the fingerprint state.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Water quality homogeneous intervals are divided based on the direction of change, duration, cross-index linkage, and dosing response lag of the water quality state vector. This ensures that subsequent identification is based on a continuous time range of response relationships, reducing data contamination between different water quality stages. For intervals with high uncertainty in implicit chemical consumption, the original planned dosage is rearranged into a baseline dosing sequence and a perturbation dosing sequence with equal cumulative dosage within the control time domain. A controlled dose change, distinct from natural influent fluctuations, is formed through a positive perturbation in the initial stage and equal compensation in the subsequent stage, without increasing the total planned dosage in the control time domain. The actual response generated by the perturbation, after being corrected for flow propagation time difference, is compared with the counterfactual response trajectory constructed from similar historical water quality intervals. The difference in response slope, inflection point time difference, cumulative response area difference, and recovery time difference together constitute a causal response fingerprint with chemical type and water quality interval identifiers. The causal response fingerprint separates the response change corresponding to the current dosing action from the combined changes in influent composition and the formation of residual effects from previous dosing, and provides traceable observational evidence for the correction of charge neutralization demand, redox demand, acid-base buffering demand, and the intensity of persistent degradation interference. The implicit dosing state thus no longer relies solely on the static mapping of apparent detection values, but is updated based on the actual response after controlled perturbation. Even when pH, turbidity, conductivity, or target pollutant indicators are similar but the pollutant components differ, the control process can still distinguish the direction of change in corresponding reagent demand and solve for the multi-reagent demand vector and the segmented release sequence accordingly. The continuous connection between the above-mentioned data partitioning, total amount time-domain perturbation, counterfactual comparison, and state correction makes it less likely that natural water quality fluctuations will be misrepresented as reagent dosage responses, reducing the possibility of continuous dosing in the original erroneous direction after water quality component switching, and ensuring that dosing decisions are consistent with the current response requirements. When the current causal response fingerprint is inconsistent with the predicted response fingerprint, the control process does not directly increase the dosage along the existing state, but instead switches to candidate dominant state verification, so that the state judgment has a path for re-verification, avoiding the solidification of a single incorrect estimate into a continuous control deviation.

[0017] 2. At the secondary control level, the positions of candidate perturbations are determined based on the estimated dispersion of the state components to be identified, historical response delays, and reagent response sensitivity. Positive perturbations and compensating perturbations are paired into dose segments. Each candidate sequence is simultaneously constrained by the instantaneous dosing limit, adjacent dosing intervals, effluent control boundaries, residual prior efficacy, and cumulative dosage conservation. Candidate response trajectories generated from different state uncertainty intervals are compared for separability in response direction, inflection point order, and duration. Therefore, the selected perturbation sequences can provide relatively clear state differentiation information without exceeding the predetermined dosing boundaries. When the compensating perturbation does not meet the execution conditions, its release position is adjusted along the response propagation direction, maintaining consistency between the cumulative dosage in the control time domain and the baseline sequence, thus avoiding additional total dosage deviations during the identification process. The counterfactual response trajectory is constructed from historical intervals with similar preceding historical segments, similar initial values ​​of implicit drug consumption states, and no non-corresponding drug adjustments. After propagation time axis mapping, a weighted median response is synthesized, and the matching confidence is represented by the response dispersion at each sampling time, which can reduce the impact of random fluctuations in a single historical segment on the causal response fingerprint. The observation residual vector is decomposed into mutually distinguishable state-sensitive basis vectors, so that charge neutralization demand, redox demand, acid-base buffering demand, and the intensity of persistent degradation interference obtain their respective observation increments. Then, the state covariance is updated by combining water quality continuity, drug efficacy decay, and interval response confidence, so that the subsequent dosage is adjusted according to the completeness of state evidence. When the predicted response fingerprint is inconsistent with the current causal response fingerprint, the control process splits the multi-drug constraint dosage into basic release, verification dosage, and reserved release. The candidate dominant state and posterior weight are updated through local response segments, and the dosage that has not yet been released is redistributed. Therefore, when there is insufficient state evidence, the entire calculated dose will not be released at once, and the original remaining sequence will not be executed after the state judgment changes. This can limit the spread of erroneous state estimation to subsequent control cycles, reduce repeated drug replenishment, dose overshoot, and dosing oscillations caused by continuous reverse correction. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating the overall process of an intelligent automatic dosing method for complex wastewater treatment. Figure 2 A flowchart for classifying homogeneous water quality zones and establishing hidden chemical consumption status is constructed. Figure 3 Flowchart for generating time-domain perturbations and causal response fingerprints with equal total amount; Figure 4 This is a flowchart of state correction and segmented release based on causal response fingerprints. Detailed Implementation

[0019] In one embodiment, reference Figure 1A smart automatic dosing method for complex wastewater treatment includes: collecting influent flow rate, water quality characterization data, historical dosing data, and effluent response data during the control cycle; dividing the water quality characterization data into homogeneous water quality intervals; and constructing implicit dosing states for each homogeneous water quality interval, including charge neutralization requirements, redox requirements, acid-base buffering requirements, and the intensity of persistent degradation interference; determining the state uncertainty based on the estimated dispersion of each implicit state component; when the state uncertainty meets the perturbation triggering condition, rearranging the planned dosing amount in the control time domain into a baseline dosing sequence and a perturbation dosing sequence with equal cumulative dosing amounts but different time sequences; executing dosing according to the perturbation dosing sequence and collecting the actual response trajectory; generating a counterfactual response trajectory based on the baseline dosing sequence; generating a causal response fingerprint based on the time sequence difference between the two response trajectories; correcting the implicit dosing state based on the causal response fingerprint; determining the multi-agent constrained dosing amount and segmented release order based on the corrected implicit dosing state; and executing the dosing.

[0020] In this embodiment, the control cycle is the data range for forming a complete state update and dosing decision, and the control time domain is the time range that expands backward from the current decision time and allows for the arrangement of several agent releases. The two are established according to the response delay of wastewater in the current treatment process. The influent flow rate is used to calculate the pollution load entering the treatment process per unit time and correct the response propagation position. The water quality characterization data includes at least pH, oxidation-reduction potential, conductivity, turbidity and target pollutant indicators. Historical dosing data records the agent type, release time, release amount and residual amount of agent that has not yet completed the reaction. The effluent response data records the continuous change trajectory of each characterization indicator after the agent action. During data processing, the sampling time axis is first unified, and the missing time is filled in by the local trend formed by the adjacent effective data. Isolated mutation points are retained or removed according to the cross-indicator linkage relationship at the same time. The propagation time difference caused by the change of influent flow rate is corrected so that the data entering the same state estimation process correspond to the same batch or adjacent mixed batches of wastewater.

[0021] Specifically, the homogeneous water quality interval is not divided according to whether a single detection value crosses a fixed limit, but is determined by the direction of change, duration of change, linkage sequence, and delayed response after dosing of multiple indicators. The starting point of the interval corresponds to the moment when the water quality composition or reaction requirements undergo a identifiable switch. Within the interval, the detection values ​​are allowed to drift slowly, but the linkage relationship of the main indicators and the direction of dose response must remain continuous. The charge neutralization requirement in the latent drug consumption state is used to characterize the consumption tendency of colloidal and charged pollutant components for coagulants. The redox requirement is used to characterize the equivalent requirement of reducing or oxidizing components for the corresponding agents. The acid-base buffer requirement is used to characterize the water body's ability to resist changes in acidity and alkalinity. The intensity of persistent interference is used to characterize the degree to which the target pollutant indicator is insufficient to respond to conventional agents and to mask other state components. Each state component is obtained by inversion from the observable response and is not directly equivalent to a certain detection indicator.

[0022] Table 1 below shows the complex wastewater treatment data fields and state mapping relationships used in this embodiment. The preprocessing results in the table serve as the common input for water quality homogeneity interval division, implicit chemical consumption state estimation, and counterfactual trajectory construction. Fields are saved according to a unified sampling time and are accompanied by data source identifiers, so that the original value, correction value, and state value used for decision-making of the same indicator can be traced back to each other.

[0023] Table 1. Data Fields and Status Mapping Relationships for Complex Wastewater Treatment ; The initial value of the latent chemical consumption state is jointly determined by the preceding historical segments and similar historical segments of the current water quality homogeneous interval. The preceding historical segments cover multiple consecutive sampling times before the perturbation occurs. First, the water quality characterization data are converted into changes relative to the starting point of the interval, and then converted into unit water volume changes according to the influent flow rate. Subsequently, the changes are projected onto the sensitive directions corresponding to each state component based on the response direction after the historical chemical release. The initial state covariance is jointly formed by the state dispersion of similar historical intervals, the completeness of the current data, and the uncertainty of the previous residual chemical effect. The state uncertainty can be the weighted value of the state covariance, the maximum eigenvalue, or the combination of the normalized variances of each state component. The perturbation triggering conditions simultaneously require that the state uncertainty reaches the set identification requirements, the current effluent prediction does not exceed the safety control boundary, there is an arrangeable compensation release position in the control time domain, and the previous chemical effect is not in the inseparable strong response stage.

[0024] After the perturbation is triggered, the baseline dosing sequence maintains the planned release amounts and release times of each agent given by the original control strategy. The perturbation dosing sequence selects agents with a high degree of coupling with the state component to be identified from the baseline dosing sequence, moves part of the planned agent amount forward to form a positive perturbation, and deducts the same amount of agent at subsequent executable positions to form a compensating perturbation. The sum of the positive perturbation and the compensating perturbation is 0. Therefore, the cumulative agent amount in the control time domain of the perturbation dosing sequence and the baseline dosing sequence is equal. After the system issues the dosing command according to the perturbation dosing sequence, it continuously collects the effluent response. The actual response trajectory retains the original time sequence and records the influent changes synchronously. The counterfactual response trajectory represents the expected response when the baseline dosing sequence is still executed under the current water quality conditions. The two trajectories are compared after being aligned with the propagation time axis.

[0025] The causal response fingerprint is composed of multiple temporal differences, including the response slope difference reflecting the deviation of the rate of change of the index after perturbation from the baseline situation, the inflection point time difference reflecting the advance or delay of the reaction initiation or turning point, the cumulative response area difference reflecting the difference in the overall response quantity within an observation window, and the recovery time difference reflecting the difference in the time required for the index to return to the stable range. Each difference is also accompanied by a reagent type identifier, a water quality range identifier, a matching confidence level, and a response continuity identifier. When correcting the state, the correction direction and correction magnitude of each state component are determined based on the correspondence record between the historical fingerprint and the latent state change. The corrected state is then combined with the reagent causal response coefficient to solve the multi-reagent demand vector. The solution is constrained by the reaction relationship between reagents, the remaining release location, the single allowable adjustment range, and the cumulative reagent dosage boundary.

[0026] The multi-agent constrained dosage is not released all at a single moment, but is divided into basic release, validation release, and reserved release based on the state credibility. The basic release corresponds to the common demand with a high degree of credibility in the current state estimate, the validation release corresponds to the candidate state that needs to be further distinguished through local response, and the reserved release is not released before obtaining subsequent response. After each release, the local response segment corresponding to the corresponding agent response delay is extracted and the causal response fingerprint is updated. When the new local fingerprint is consistent with the current dominant state, the release continues according to the remaining plan. When the new local fingerprint points to other dominant states, the unreleased dosage is redistributed. This embodiment obtains identifiable response through controlled and cumulative dosage conservation time-series changes, and uses counterfactual comparison to separate the agent effect from the natural fluctuation of influent and the residual effect of previous agents, so that wastewater with similar apparent detection indicators but different actual response needs can form different state correction results.

[0027] In a preferred embodiment, reference Figure 2The process of dividing water quality characterization data into homogeneous water quality intervals includes: combining influent flow rate, pH, oxidation-reduction potential, conductivity, turbidity, and target pollutant indicators into a water quality state vector according to the sampling time; calculating the consistency of change direction, duration of change, and cross-indicator linkage relationship for adjacent water quality state vectors; determining the sampling time when the change direction of indicators reverses and the linkage indicator combination switches between the preceding and following windows as the candidate boundary time; correcting the candidate boundary time by shifting it forward or backward based on the response lag between historical dosing data and effluent response data; merging the intervals in the corrected adjacent intervals where the effluent response slope is continuous, the response lag is continuous, and the cross-indicator linkage relationship of the water quality state vector has not switched again to obtain a homogeneous water quality interval corresponding to the dosing response, and configuring an independent state evolution record for each homogeneous water quality interval.

[0028] In this embodiment, each component in the water quality state vector is first dimensionless using the median and interquartile range within the current rolling window, so that detection data of different dimensions can jointly participate in the calculation of direction and linkage. The consistency of change direction is determined according to the consistency ratio of the difference signs of each index in adjacent windows. The continuous change length records the number of samples that appear consecutively in the same direction. The cross-index linkage is represented by the temporal correlation order between index differences. The comprehensive discrimination value of the candidate boundary moment is calculated using the following formula.

[0029] ; in, This represents the interval switching discriminant at sampling time t. This indicates the consistency of the direction of change within the preceding and following windows, and its value ranges from 0 to 1. This represents the value after normalizing the continuously changing length using the window length. This indicates the degree of switching between cross-indicator linkage combinations, according to the formula. Calculation, where This represents the number of metrics within the current window whose interaction relationships have changed compared to the previous window. The total number of indicators involved in determining the linkage relationship; This indicates the degree of lag discontinuity obtained based on historical dosing data and effluent response data, according to the formula... Calculation, where To estimate the actual response lag time based on current data, This is the theoretical response lag time based on historical data statistics. This represents the maximum permissible hysteresis deviation. , , and For non-negative weights that sum to 1, for example, when turbidity and redox potential are simultaneously reversed and the order of conductivity changes, we can let , , , And setting the four weights to 0.30, 0.20, 0.30, and 0.20 respectively, we obtain... =0.25×0.30+0.70×0.20+0.85×0.30+0.60×0.20=0.59, when When the value exceeds a preset threshold of 0.55, that moment is listed as a candidate dividing point; if only a single indicator shows isolated fluctuations while other indicators remain unchanged, then... and If the values ​​are all less than 0.3, the candidate boundary points corresponding to the specified times are eliminated, and no new homogeneous water quality intervals are formed.

[0030] After the candidate boundary moment is formed, the response lag distribution is estimated by using the time of drug release and the time when each effluent index begins to change continuously in the historical dosing records. The candidate boundary moment is then moved forward or backward along the propagation time axis to the most likely water quality composition switching position. Subsequently, the effluent response slope, response lag continuity and cross-index linkage combination on both sides of the boundary point are checked. If the slopes on both sides can be connected by the same continuous curve and the linkage combination has not switched again, the boundary is canceled and the intervals are merged. If there is a continuous directional difference on both sides or the response lag is recalculated, the boundary is retained. Each final interval saves the initial state value, state covariance, drug response template, causal response fingerprint and interval termination reason, thereby avoiding data with different reaction requirements from entering the same state evolution record and ensuring that subsequent perturbation and counterfactual comparisons have a consistent water quality basis.

[0031] In a preferred embodiment, reference Figure 3 The process of rearranging the planned dosage into a baseline dosing sequence and a perturbation dosing sequence includes: constructing a state covariance based on the estimated dispersion of each latent state component, and determining the state component to be identified by combining the recent dose response residual and the response sensitivity of each agent; selecting candidate perturbation positions corresponding to the state component to be identified within the current control time domain, shifting a portion of the dosage in the baseline dosing sequence located at the candidate perturbation position forward to form a positive perturbation, and deducting an equal amount of dosage at subsequent positions to form a compensating perturbation; filtering out unexecutable sequence combinations according to the instantaneous dosing limit, adjacent dosing interval, effluent control boundary, and previous residual efficacy; selecting the perturbation dosing sequence with the maximum response discrimination of the state component to be identified from the remaining sequence combinations, and maintaining the cumulative dosage of the perturbation dosing sequence and the baseline dosing sequence equal within the current control time domain.

[0032] The estimated dispersion of each latent state component is represented by the diagonal elements of the state covariance. The recent dose response residual is the difference between the actual response trajectory and the state response model output at the same propagation position. The response sensitivity of each agent is calculated based on the change in state component corresponding to the unit dose change in the historical homogeneous interval. When a certain state component has a large estimated dispersion, a continuously unidirectional dose response residual, and a sensitivity that can be distinguished by the current agent, it is identified as the state component to be identified. Candidate perturbation positions are generated from the reference release positions that have not yet been executed in the current control time domain. Each position forms several positive perturbation quantities and subsequent equal compensation positions. Combinations that would cause the cumulative dosage at any time to exceed the control boundary, would overlap with the response of the previous residual drug effect, or would not be compensated before the end of the control time domain are eliminated.

[0033] The remaining sequence combinations are used to calculate the response differentiation score according to the following formula, and the sequence with the highest score is selected as the perturbation delivery sequence.

[0034] ; in, This represents the response discrimination score for the j-th candidate sequence. The directional separability of the candidate responses formed by the state components to be identified at both ends of the uncertainty interval is expressed by the formula: Calculation; where and These are the candidate response trajectories obtained when the upper and lower boundary values ​​of the state component to be identified are taken, respectively, and t0 and t1 are the start and end times of the observation window, respectively. To observe the total duration of the window, This represents the maximum range of response change that this state component may cause. The degree of difference in the inflection point order of candidate responses is represented by the formula. Calculation, where The number of candidate response trajectories whose inflection points appear in different orders. This represents the total number of inflection points between the two trajectories. The degree of difference in the duration of candidate responses is indicated by the formula. Calculate, where, where and These represent the time required for the two candidate trajectories to recover to stability from the start of the response; This represents the combined constraint cost resulting from proximity to the effluent control boundary, overlap of prior residual effects, and insufficient compensation location, expressed by the formula... Calculation, where For the proportion of drug dosage exceeding the dosing limit, The proportion of time that overlaps with the preceding drug effect in the time domain. For the proportion of time when there is a lack of available compensation locations, , , represents the corresponding non-negative weighting coefficient. , , and Non-negative weights, such as those of a candidate sequence. , , and When the values ​​are 0.80, 0.65, 0.70, and 0.20 respectively, and the four weights are 0.35, 0.25, 0.25, and 0.15 respectively, the candidate sequence is retained because it has a clear directional difference and a low constraint cost. Another candidate sequence, even if its directional difference is similar, is... When the value is 0.75, the score is reduced and not used, thus allowing the perturbation location to serve state identification instead of being converted into additional drug input.

[0035] In a preferred embodiment, generating the counterfactual response trajectory and causal response fingerprint includes: selecting reference intervals from historical water quality homogeneous intervals divided according to the same steps that meet the matching conditions with the current water quality homogeneous interval in terms of influent trend, initial value of latent drug consumption state, previous residual drug efficacy, and flow propagation characteristics; removing response segments within the reference intervals that do not correspond to drug adjustments or influent abrupt changes, and aligning the remaining response segments temporally according to the flow propagation time difference; generating a counterfactual response trajectory corresponding to the baseline dosing sequence using the aligned response segments; extracting the response slope difference, inflection point time difference, cumulative response area difference, and recovery time difference between the actual response trajectory and the counterfactual response trajectory, and combining each difference according to the matching confidence of the reference interval into a causal response fingerprint with a drug type identifier and a water quality interval identifier.

[0036] The selection of reference intervals adopts a two-level constraint. The first level requires that the influent trend direction, the order of initial state components, and the flow propagation pattern be consistent. The second level calculates the distance between previous historical segments and retains intervals with smaller distances. Non-corresponding agent adjustment refers to the release of other agents that can change the same water quality index within the target agent response observation window. Influent mutation refers to the water quality state vector crossing the current reference interval boundary. Segments that meet any of these conditions do not participate in counterfactual construction. The remaining segments are stretched or compressed according to their respective propagation time axes to align the agent release point, the expected response start point, and the recovery segment start point in sequence. Subsequently, a weighted median value is used to form a counterfactual response value at each aligned sampling position, and the matching confidence is formed by the degree of dispersion between reference segments.

[0037] The causal response fingerprint is represented by the following formula.

[0038] ; in, This represents the causal response fingerprint of the nth homogeneous water quality interval. Indicates the confidence level of the reference interval match. The actual response trajectory slope difference is divided by the average slope difference benchmark value of the same type of historical intervals (the median of the absolute values). The inflection point time difference between the two trajectories is divided by the average inflection point time difference baseline value. The cumulative response area difference divided by the average cumulative area difference benchmark value. The recovery time difference is divided by the average recovery time difference baseline value; each baseline value is taken from the median absolute value of the corresponding component of the historical high-confidence fingerprint under the same water quality interval type, so that all four components are dimensionless values, for example... When the value is 0.90 and the four dimensionless differences exhibit positive slope difference, early inflection point, positive cumulative area difference, and shortened recovery time, respectively, the fingerprint is recorded as a response consistent with the positive perturbation. If we set it to 0.35, each fingerprint component will shrink synchronously and its weight in state correction will be reduced, so that the latent drug-consuming state will not be dominated by the counterfactual trajectory with low matching quality.

[0039] In another preferred embodiment, when reference fragments are insufficient, the counterfactual trajectory is not directly replaced by a single historical interval. Instead, it starts from the previous historical fragments before the perturbation of the current homogeneous water quality interval, and forms a candidate reference trajectory by recursively according to the reference addition sequence and the historical state response template. Then, a small number of historical fragments that can pass the first-level constraint are used to correct the recursive deviation. The correction amount is only released within the historical discrete range of the corresponding index. The actual response trajectory and the candidate reference trajectory still form a causal response fingerprint in the same way. This embodiment uses reference interval cleaning, propagation time axis alignment and matching confidence weighting to make the counterfactual response reflect the comparable processing process when no perturbation is implemented, and to suppress the mixing of fingerprints by influent mutations or other agent adjustments.

[0040] Furthermore, the construction of the implicit drug consumption state includes: for each homogeneous water quality interval, forming an observation residual vector from the response increments of pH, redox potential, turbidity, and target pollutant indicators relative to the interval's starting point; extracting historical response directions corresponding to charge neutralization demand, redox demand, acid-base buffering demand, and recalcitrant interference intensity from the state evolution record, sequentially subtracting the projection components of each historical response direction onto the previous response direction and normalizing them to form mutually distinguishable state-sensitive basis vectors; projecting the observation residual vectors onto each state-sensitive basis vector to obtain the observation increments of each implicit state component; and fusing the observation increments with the state estimates of the previous homogeneous water quality interval according to the continuity of the water quality state vector, the amount of prior drug efficacy decay, and the interval response reliability to form the implicit drug consumption state and state covariance of the current homogeneous water quality interval.

[0041] The state-sensitive basis vectors are generated in a fixed order. First, the historical directions with a relatively concentrated response to charge neutralization requirements are selected as the first basis vector. Then, the projection of the redox demand historical directions onto the first basis vector is subtracted. Subsequently, the projection subtraction process is repeated for acid-base buffer requirements and the intensity of non-degradable interference. After normalization, the sensitive directions that are orthogonal or approximately orthogonal to each other are obtained. The projection values ​​of the observation residual vectors on each sensitive direction constitute the observation increment. The state fusion is performed using the following formula.

[0042] , ; in, This represents the implicit chemical consumption state vector for the nth homogeneous water quality interval. This represents the estimated state value of the previous interval. Indicates the current observation increment. This represents the observation mapping matrix composed of state-sensitive basis vectors. This represents the fusion matrix determined based on the continuity of the water quality state vector, the pre-order efficacy attenuation, and the interval response confidence. and Let I represent the state covariance before and after the update, respectively. For example, when the current interval is continuous with the previous interval and the response confidence is high, we can... The corresponding diagonal element is set to 0.75 to allow for a larger observation increment in the state update. If the preceding drug effect decay is uncertain and the response fragment is incomplete, the corresponding diagonal element is set to 0.30 to keep the state change restricted and retain a larger value. In the subsequent control time domain, perturbation identification for the state components is triggered.

[0043] In a preferred embodiment, the state-sensitive basis vector is stored separately according to the water quality homogeneous interval type. The acid-base response direction formed by high buffer water quality is not directly used for low buffer water quality. When a new high-confidence causal response fingerprint is accumulated in the same type interval, the sensitive basis vector is updated with the observation direction corresponding to the fingerprint, and projection subtraction and normalization are performed again. The state covariance records the residual correlation between each state component. This embodiment decomposes the common changes of multiple apparent indicators into mutually distinguishable drug consumption observation increments, reducing the situation where the same acid-base or turbidity change is repeatedly included in different latent state components.

[0044] In another preferred embodiment, the selection of the perturbation dosing sequence includes: for each candidate perturbation position, establishing paired dose segments consisting of a positive perturbation and a compensating perturbation, and determining the time interval between the two dose segments based on the historical response delay corresponding to the state component to be identified; establishing a state response model using the historical dose response trajectory of each latent state component, and combining each paired dose segment with the two ends of the uncertainty interval of the state component to be identified to obtain candidate response trajectories; calculating the separability between candidate response trajectories in terms of response direction, inflection point order, and duration, and using the effluent control boundary, the previous residual efficacy, and the cumulative dosage conservation as sequence feasibility constraints; when the compensating perturbation does not meet the feasibility constraints at the original position, moving the compensating perturbation along the response propagation direction within the current control time domain until a perturbation dosing sequence with the same cumulative dosage as the benchmark dosing sequence and the highest separability is formed.

[0045] The positive perturbation of paired dose segments is drawn from the baseline release amount. The compensation perturbation is equal to its absolute value and opposite in direction. The time interval at least covers the response start delay of the state component to be identified and is shorter than the expected remaining length of the current water quality homogeneous interval. The state response model uses the dose change, response direction, inflection point position and duration in the historical dose response trajectory as condition index. Trajectories that cross the water quality interval boundary are not used. Candidate response trajectories are obtained by substituting the upper and lower ends of the uncertain interval of the state component to be identified. When two candidate trajectories are always in the same direction and the inflection points are close within the observation window, the corresponding dose segment is not selected. When two candidate trajectories form an identifiable direction difference or inflection point order difference, the feasibility constraint is checked again. When the compensation position moves, the positive perturbation position remains unchanged and the subsequent release positions are checked one by one until the cumulative dosage conservation and the effluent control boundary are satisfied at the same time. This embodiment enables the perturbation sequence to produce a distinguishable response to the most uncertain dosage component at present, and limits the change of the identification action on the total dosage in the control time domain by equal compensation.

[0046] Furthermore, the selection and time-series alignment of the reference intervals include: combining the influent flow rate and water quality characterization data from multiple consecutive sampling times before the perturbation occurs into a water quality state vector, and forming a prior historical segment together with the reagent dosing trajectory and the effluent change trajectory; retrieving candidate reference intervals from historical water quality homogeneous intervals that are similar to the prior historical segment and have not experienced corresponding perturbations; estimating the propagation time axis of the current water quality homogeneous interval and the candidate reference intervals according to the sequential response relationship of the influent flow rate change and each water quality indicator, and mapping the candidate reference intervals to the current propagation time axis; assigning combined weights to the mapped candidate reference intervals according to the similarity of the prior historical segment, the similarity of the initial value of the implicit reagent consumption state, and the continuous length without non-corresponding reagent adjustments, using the weighted median response trajectory as the counterfactual response trajectory, and recording the response dispersion at each sampling time as the matching confidence.

[0047] The preceding historical segments adopt a uniform length and use the perturbation planning time as the end. The influent flow rate, water quality state vector, benchmark reagent dosing trajectory, and effluent change trajectory are normalized according to their respective dimensions. The segment similarity is determined by the consistency of trend direction, the order of local inflection points, and the normalized distance. The candidate reference interval must have a complete pre-perturbation observation window and counterfactual response observation window and must not have undergone time-domain rearrangement of the same type as the current perturbation. The propagation time axis is estimated based on the cumulative influent flow rate and the order in which each water quality index begins to respond. During mapping, the relative order from the reagent release point to the response start point remains unchanged. At each mapping sampling position, the combined weights are accumulated from small to large, and the response value when the cumulative weight exceeds half is taken as the weighted median response value. The response dispersion is determined by the weighted absolute deviation of each reference value relative to the weighted median response value. When the number of reference intervals is small or the dispersion is large, the matching confidence is reduced. In this embodiment, the counterfactual trajectory is supported by multiple similar running segments, and a single abnormal historical segment is avoided from changing the current efficacy judgment.

[0048] In a preferred embodiment, reference Figure 4 The method of correcting latent drug consumption status based on causal response fingerprints includes: writing historical causal response fingerprints and corresponding latent state components into the fingerprint state corresponding record according to the drug type identifier and water quality interval identifier; extracting fingerprint response templates with the same identifier as the current causal response fingerprint; aligning the current causal response fingerprint with each fingerprint response template to obtain the state correction candidate quantity corresponding to each latent state component; determining the correction weight based on the matching credibility of the counterfactual response trajectory, the continuity of the actual response, and the degree of conflict of each state correction candidate quantity; updating the latent drug consumption status and generating the state covariance according to the correction weight and the dispersion of the state correction candidate quantity; forming a constraint relationship between the updated latent state components and the drug causal response coefficients in the fingerprint response template; solving the multi-drug demand vector; and generating the multi-drug constrained dosage by combining the inter-drug reaction constraints and the remaining release position.

[0049] The fingerprint state correspondence record uses the agent type identifier and water quality interval identifier as the primary index, and the perturbation direction, state change direction and fingerprint confidence as secondary indexes. The fingerprint response template is formed by high confidence fingerprints under the same index according to the component median and discrete range. When the components are aligned, the response direction sign is unified and the slope difference, cumulative area difference and recovery time difference are converted according to the length of each observation window. When the current fingerprint falls into the discrete range of a certain template, the state correction candidate quantity is generated according to the state change quantity recorded in the template. When the correction directions given by multiple templates are opposite, the degree of conflict is increased and the common correction weight is reduced. The updated state covariance absorbs the discreteness of the candidate quantity, so that the state components with template conflicts maintain a high degree of uncertainty in subsequent control cycles. The multi-agent demand vector is solved by the following formula.

[0050] ; in, This represents the multi-reagent demand vector for the nth homogeneous water quality interval. This represents the feasible region comprised of inter-agent reaction constraints, cumulative dosing boundaries, and remaining release locations. This represents the drug causal response coefficient matrix (dimensions are state quantity / drug quantity) in the fingerprint response template. We will continue to use the current implicit drug consumption state vector (the dimension is state quantity). and All are diagonal weight matrices (dimensionless) composed of the confidence levels of each state component. This represents a non-negative coefficient indicating an unfounded jump in drug dosage between adjacent control cycles. This represents the actual demand vector used in the previous interval, after... After mapping, both terms are quadratic forms of the state vector, with consistent dimensions (state vector). 2 ); This represents the matrix transpose, for example, when the confidence level of charge neutralization requirement is high while the confidence level of recalcitrant interference strength is low, it can be made... The diagonal value corresponding to the former is set to 0.85 and the diagonal value corresponding to the latter is set to 0.30. The solution results prioritize the drug consumption component supported by the causal response fingerprint. For drug amounts with insufficient evidence, the space for subsequent verification is reserved. In this embodiment, the state correction, dosage solution and release constraint adopt the same set of causal response relationships and do not form independent control conclusions.

[0051] Furthermore, updating the latent drug consumption state and generating the state covariance also includes: weighting and combining each latent state component before the update with its corresponding fingerprint response template to form a predicted response fingerprint; comparing the response direction, inflection point time difference, and cumulative response area difference between the predicted response fingerprint and the current causal response fingerprint; when the consistency does not meet the state acceptance condition, setting each latent state component as a candidate dominant state; generating initial posterior weights based on the matching degree between the current causal response fingerprint and the fingerprint response templates of each candidate dominant state; and dividing the corresponding verification drug amount from the multi-drug constraint dosage; releasing the verification drug amount in descending order of the initial posterior weights; collecting local response fragments after each release and forming verification fingerprints; updating the posterior weights based on the matching relationship between the verification fingerprints and each fingerprint response template; and recalculating the latent drug consumption state and the multi-drug constraint dosage by selecting the candidate dominant state with the highest posterior weight.

[0052] The predicted response fingerprint is formed by multiplying the pre-update state components by the corresponding drug causal response coefficient and then accumulating the fingerprint components. The consistency comparison first checks whether the response directions are opposite, and then checks whether the inflection point time difference and the cumulative response area difference fall within the template's allowable discrete range. If any high-confidence component has an opposite direction or multiple components exceed the allowable range at the same time, it is determined that the state acceptance condition is not met. The candidate dominant states are assumed to be the main source of the current deviation, namely charge neutralization requirement, redox requirement, acid-base buffer requirement, or the intensity of non-degradable interference. The initial posterior weight is obtained by normalizing the inverse of the distance between the current fingerprint and the corresponding template. The verification drug quantity is divided from the required amount of the corresponding drug that has not yet been released and does not change the executed drug quantity. After the local response fragment forms the verification fingerprint, the corresponding posterior weight is increased or decreased according to the template matching degree. When the posterior weights of two candidate states are close, the remaining verification drug quantity is retained instead of being released directly. This embodiment replaces the continuous addition of drugs along a single erroneous state with small-scale backtrackable verification, so that the dominant state can be reselected when the state model is inconsistent with the actual drug effect.

[0053] In a preferred embodiment, the segmented release sequence includes: dividing the multi-agent constraint dosage into a basic release amount, a verification dosage, and a reserved release amount based on the posterior weight of the candidate dominant state, and configuring them to the remaining release positions in the current control time domain in the order of basic release amount first, verification dosage in the middle, and reserved release amount last; after each release, extracting the local response fragment corresponding to the corresponding drug response delay, and updating the causal response fingerprint, implicit drug consumption state, and remaining drug demand vector; when the updated candidate dominant state changes, redistributing the unreleased reserved release amount among the new multi-agent demand vector, and keeping the sum of the released and unreleased amounts of each drug not exceeding the corresponding drug constraint dosage; at the end of the current control time domain, writing the final implicit drug consumption state, verification fingerprint, and actual release sequence into the record corresponding to the fingerprint state.

[0054] The basic release amount is the intersection of the dosages supported by all high-posterior candidate states. The verification dosage is the dosage that has a distinguishing effect between different candidate states and can form a local response within the remaining observation window. The reserved release amount is the balance after deducting the first two from the multi-dosage constraint dosage. The release position is arranged according to the response delay of each dosage to avoid releasing the next dosage that would change the same index before the previous verification dosage has formed an observable response. After each local response update, the remaining dosage demand vector is recalculated. When the dominant state does not change, the release order is adjusted only within the original reserved release amount. When the dominant state changes, the unexecuted dosages that are inconsistent with the new state are canceled and the available balance is allocated to the new demand vector. The released amount is retained as an irreversible constraint. The final state, verification fingerprint, and actual release sequence saved at the end of the control time domain are used to update the state sensitive basis vector, fingerprint response template, and dosage causal response coefficient of the corresponding water quality interval. This embodiment forms a closed control process from state identification, verification dosing, local response verification to the redistribution of the remaining dosage, so that different actual dosage consumption requirements under the same apparent water quality conditions can be completed according to their respective verification results.

Claims

1. An intelligent automatic dosing method for complex wastewater treatment, characterized in that, include: Collect influent flow rate, water quality characterization data, historical dosing data and effluent response data during the control cycle, divide the water quality characterization data into homogeneous water quality intervals, and construct a hidden chemical consumption state for each homogeneous water quality interval, including charge neutralization requirements, redox requirements, acid-base buffering requirements and the intensity of recalcitrant interference. The state uncertainty is determined based on the estimated dispersion of each latent state component. When the state uncertainty meets the perturbation triggering condition, the planned dosage in the control time domain is rearranged into a baseline dosage sequence and a perturbation dosage sequence with equal cumulative dosage but different timing. The perturbation is applied according to the perturbation application sequence and the actual response trajectory is collected. A counterfactual response trajectory is generated based on the baseline application sequence. A causal response fingerprint is generated based on the temporal difference between the two response trajectories. The latent drug consumption state is corrected based on the causal response fingerprint. The multi-drug constraint dosage and segmented release sequence are determined based on the corrected latent drug consumption state, and the dosage is executed.

2. The intelligent automatic dosing method for complex wastewater treatment according to claim 1, characterized in that, The water quality characterization data is divided into homogeneous water quality intervals, including: combining influent flow rate, pH, redox potential, conductivity, turbidity and target pollutant indicators into a water quality state vector according to the sampling time; For adjacent water quality state vectors, calculate the consistency of change direction, length of continuous change, and cross-index linkage relationship respectively. The sampling time when the index change direction reverses and the linkage index combination switches between the front and back windows is determined as the candidate boundary time. Based on the response lag between historical dosing data and effluent response data, the candidate boundary time is corrected by shifting it forward or backward. After correction, adjacent intervals with continuous effluent response slopes, connected response lags, and no further switching of cross-index linkage relationships in the water quality state vector are merged to obtain water quality homogeneous intervals corresponding to the dosing response, and an independent state evolution record is configured for each water quality homogeneous interval.

3. The intelligent automatic dosing method for complex wastewater treatment according to claim 1, characterized in that, The planned dosage is rearranged into a baseline dosing sequence and a perturbation dosing sequence, including: constructing the state covariance based on the estimated dispersion of each latent state component, and determining the state component to be identified by combining the recent dose response residual and the response sensitivity of each drug. Within the current control time domain, a candidate perturbation position corresponding to the state component to be identified is selected. A portion of the drug amount located at the candidate perturbation position in the baseline dosing sequence is moved forward to form a positive perturbation, and an equal amount of drug amount is deducted at subsequent positions to form a compensation perturbation. The sequence combinations that cannot be executed are screened out according to the instantaneous dosage limit, the interval between adjacent dosages, the effluent control boundary, and the residual efficacy of the preceding drug. Select the perturbation delivery sequence with the maximum response discrimination from the remaining sequence combinations, and keep the cumulative drug dosage of the perturbation delivery sequence and the reference delivery sequence equal in the current control time domain.

4. The intelligent automatic dosing method for complex wastewater treatment according to claim 1, characterized in that, The generation of counterfactual response trajectories and causal response fingerprints includes: selecting reference intervals from historical water quality homogeneous intervals divided according to the same steps that meet the matching conditions with the current water quality homogeneous intervals in terms of inflow trend, initial value of hidden drug consumption state, previous residual drug efficacy and flow propagation characteristics. Remove response segments within the reference interval that contain non-corresponding reagent adjustments or sudden changes in influent, and align the remaining response segments in time according to the flow propagation time difference; The aligned response fragments are used to generate the counterfactual response trajectory corresponding to the baseline delivery sequence; The difference in response slope, inflection point time, cumulative response area, and recovery time between the actual response trajectory and the counterfactual response trajectory are extracted respectively. The differences are then combined according to the matching confidence of the reference interval to form a causal response fingerprint with reagent type and water quality interval identifiers.

5. The intelligent automatic dosing method for complex wastewater treatment according to claim 2, characterized in that, The construction of the implicit drug consumption state includes: for each water quality homogeneous interval, the response increments of pH, redox potential, turbidity and target pollutant index relative to the starting point of the interval are used to form an observation residual vector; Extract the historical response directions corresponding to charge neutralization demand, redox demand, acid-base buffering demand and recalcitrant interference intensity from the state evolution record. Subtract the projection components of each historical response direction onto the previous response direction and normalize them to form mutually distinguishable state-sensitive basis vectors. The observation residual vector is projected onto each state-sensitive basis vector to obtain the observation increment of each latent state component; By fusing the observed increment with the state estimate of the previous homogeneous water quality interval based on the continuity of the water quality state vector, the previous drug efficacy attenuation, and the interval response reliability, the implicit drug consumption state and state covariance of the current homogeneous water quality interval are formed.

6. The intelligent automatic dosing method for complex wastewater treatment according to claim 3, characterized in that, The selection of the perturbation delivery sequence includes: for each candidate perturbation position, establishing a pair of dose segments consisting of a positive perturbation and a compensation perturbation, and determining the time interval between the two dose segments based on the historical response delay corresponding to the state component to be identified; A state response model is established using the historical dose response trajectories of each latent state component. Each pair of dose segments is combined with the two ends of the uncertainty interval of the state component to be identified to obtain candidate response trajectories. Calculate the separability of candidate response trajectories in terms of response direction, inflection point order and duration, and use the effluent control boundary, previous residual efficacy and cumulative dosage conservation as sequence feasibility constraints. When the compensating perturbation does not meet the feasibility constraints at its original position, the compensating perturbation is moved along the response propagation direction within the current control time domain until a perturbation dosing sequence with the highest cumulative drug dosage and the same as the baseline dosing sequence is formed.

7. The intelligent automatic dosing method for complex wastewater treatment according to claim 4, characterized in that, The selection and time-series alignment of the reference interval includes: combining the influent flow rate and water quality characterization data of multiple consecutive sampling times before the occurrence of the perturbation into a water quality state vector, and forming a prior historical segment together with the reagent dosing trajectory and the effluent change trajectory. Search for candidate reference intervals in historical water quality homogeneous intervals that are similar to previous historical segments and have not experienced corresponding perturbations; Based on the changes in influent flow rate and the sequential response relationship of various water quality indicators, the propagation time axis of the current homogeneous water quality interval and the candidate reference interval are estimated respectively, and the candidate reference interval is mapped to the current propagation time axis. Based on the similarity of previous historical fragments, the similarity of the initial value of implicit drug consumption, and the continuous length of no non-corresponding drug adjustment, a combination weight is assigned to the mapped candidate reference interval. The weighted median response trajectory is used as the counterfactual response trajectory, and the response dispersion at each sampling time is recorded as the matching confidence.

8. The intelligent automatic dosing method for complex wastewater treatment according to claim 7, characterized in that, The method of correcting latent drug consumption status based on causal response fingerprints includes: writing historical causal response fingerprints and corresponding latent status components into the fingerprint status corresponding record according to drug type identifier and water quality interval identifier, and extracting fingerprint response templates with the same identifier as the current causal response fingerprint. Align the current causal response fingerprint with each fingerprint response template to obtain the state correction candidate quantity corresponding to each hidden state component; The correction weights are determined based on the matching credibility of the counterfactual response trajectory, the continuity of the actual response, and the degree of conflict of each state correction candidate quantity; The implicit drug consumption state is updated based on the modified weight and the dispersion of the state modification candidate quantity, and the state covariance is generated. The updated implicit state components are combined with the drug causal response coefficients in the fingerprint response template to form a constraint relationship. The multi-drug demand vector is solved, and the multi-drug constrained dosage is generated by combining the drug reaction constraints and the remaining release positions.

9. The intelligent automatic dosing method for complex wastewater treatment according to claim 8, characterized in that, The process of updating the latent drug consumption status and generating state covariance also includes: weighting and combining each latent state component before the update with the corresponding fingerprint response template to form a predicted response fingerprint, and comparing the response direction, inflection point time difference and cumulative response area difference between the predicted response fingerprint and the current causal response fingerprint. When consistency does not meet the state acceptance condition, each latent state component is set as a candidate dominant state. Initial posterior weights are generated based on the matching degree between the current causal response fingerprint and the fingerprint response template of each candidate dominant state, and the corresponding verification drug amount is divided from the multi-drug constraint dosage. The validation doses were released in descending order of initial posterior weights, and local response fragments were collected after each release to form validation fingerprints. The posterior weights are updated based on the matching relationship between the verification fingerprint and each fingerprint response template. The candidate dominant state with the highest posterior weight is selected to recalculate the latent drug consumption state and the multi-drug constraint dosage.

10. The intelligent automatic dosing method for complex wastewater treatment according to claim 9, characterized in that, The segmented release order includes: based on the posterior weight of the candidate dominant state, dividing the multi-agent constrained dosage into basic release amount, validation dosage and reserved release amount, and configuring them to the remaining release positions in the current control time domain in the order of basic release amount first, validation dosage in the middle and reserved release amount last. After each release, extract the local response fragment corresponding to the corresponding drug response delay, and update the causal response fingerprint, implicit drug consumption state, and remaining drug demand vector. When the updated candidate dominant state changes, the unreleased reserved release amount is redistributed among the new multi-agent demand vector, while keeping the sum of the released and unreleased amounts of each agent from not exceeding the corresponding agent constraint dosage. At the end of the current control time domain, the final latent drug consumption state, verification fingerprint, and actual release sequence are written into the record corresponding to the fingerprint state.