Multi-heater cooperative and hierarchical fault-tolerant control system and method
By employing a multi-heater collaborative and hierarchical fault-tolerant control system, the problem of rapid approach and stable constant temperature in instantaneous water heaters when facing fluctuations in inlet water temperature, ambient temperature, and flow rate is solved. This achieves rapid response and steady-state control, reduces energy consumption and equipment differences, and improves system safety and controllability, making it suitable for low-cost platforms.
Patent Information
- Application Number
- CN202511905051.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-17
- Publication Date
- 2026-02-06
AI Technical Summary
Existing instant hot water devices struggle to quickly approach and maintain a stable temperature when faced with fluctuations in inlet water temperature, ambient temperature, and flow rate. They also suffer from issues such as uneven distribution of multiple heaters, feedback lag and overshoot due to thermal inertia, insufficient sensor reliability, and an inadequate balance between energy efficiency and lifespan.
A multi-heater collaborative and hierarchical fault-tolerant control system is adopted. Through the collaborative work of sensor units, actuator units and main control units, it realizes zoned adaptive control, short-term trend prediction and feedforward compensation, multi-objective power allocation, gradual constraint and multi-layer anomaly detection. It constructs a three-layer detection and four-level fault-tolerant mechanism of sensors, actuators and system, combined with user preference learning and energy-saving preheating.
Under typical prototype operating conditions, the system rapidly approaches the target temperature range, maintains small fluctuations during the isothermal phase, improves response speed and stability, reduces energy consumption and equipment differences, enhances system safety and controllability, and is compatible with low-cost MCUs.
Smart Images

Figure CN121474724A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of instant hot water, in particular to a multi-heater collaborative and hierarchical fault-tolerant control system and method. BACKGROUND
[0002] In the field of smart home and light commercial bathroom, instant hot water devices (such as intelligent toilet instant module, instant faucet, etc.) are increasingly popular due to the absence of a heat storage tank, fast response, and small footprint. The core goal is to quickly approach the target temperature, maintain a stable constant temperature, and ensure energy safety. However, existing solutions have the following problems.
[0003] 1. Strong time-varying working conditions: In the domestic scenario, the water inlet temperature, ambient temperature, and flow rate fluctuate significantly (seasonal / night / day changes in water inlet temperature, flow rate steps, etc.), making it difficult for fixed parameter PID to balance "fast" and "stable".
[0004] 2. Insufficient multi-heater distribution and smoothness: Commonly used are 2-3 different power heaters; fixed gear / simple ON / OFF ignores efficiency and life balance, and the sudden change in duty cycle can cause power grid impact, EMC, and water temperature fluctuations.
[0005] 3. Thermal inertia leading to feedback lag and overshoot: The instant heating element and pipeline have short-term thermal inertia, and pure feedback is prone to deviating and then returning under rapid disturbance, resulting in poor user experience.
[0006] 4. Weak measurement reliability and fault tolerance: Sensors are prone to jitter or drift under low flow, installation deviation, or aging, and if there is no layered detection and hierarchical fault tolerance, the risk can be amplified.
[0007] 5. Insufficient energy efficiency and life balance: Lack of load balancing and preheating strategies leads to excessive aging of individual heaters and standby energy waste. High-order MPC can improve it, but the computational power and calibration cost are higher, which is not conducive to low-cost platforms.
[0008] Therefore, there is an urgent need for a control framework that can be implemented on a general low-cost MCU: to improve "fast-stable" compatibility with partitioning adaptation, to reduce overshoot risk with short-term trend prediction and feedforward, to balance efficiency / life / stability with multi-objective power distribution and gradual constraint, and to configure three-layer detection and hierarchical fault tolerance with optional preference learning / energy-saving preheating, forming a controllable, rollbackable, and calibrated solution. SUMMARY
[0009] The purpose of the present application is to provide a multi-heater collaborative and hierarchical fault-tolerant control system and method that does not rely on high-cost sensors and complex calculations to achieve the following goals: 1. Rapid approach and stable constant temperature: quickly approach the target temperature range under typical sample working conditions, and maintain small fluctuations in the constant temperature stage.
[0010] 2. Multi-target power allocation: efficiency priority and life balance are realized, and a smooth constraint is imposed on single-cycle power / duty cycle variation.
[0011] 3. Hierarchical fault tolerance and rollback: a three-layer sensor / actuator / system detection and four-level fault tolerance (NORMAL / DEGRADED / ESTIMATED / SAFE_STOP) are constructed.
[0012] 4. Intelligent self-adaptation and energy saving: user preference learning and energy saving preheating. Related thresholds, tolerances and timing parameters can be set by calibration / table lookup, not limited to a single fixed value.
[0013] The technical solution of the present application is: A multi-heater collaborative and hierarchical fault-tolerant control system, comprising: a sensor unit for collecting outlet water temperature signals and flow signals; an actuator unit including at least two paths of heaters capable of independent power modulation; a main control unit in communication connection with the sensor unit and the actuator unit, and configured to perform the following operations: collect and verify the signals of the sensor unit, and perform data validity determination; based on the data validity determination result, a partition adaptive control strategy is used to calculate the total power demand, wherein corresponding control parameters are selected according to different intervals of temperature error; short-term trend prediction is performed on temperature data in the near period, and according to the prediction validity determination result, feedforward compensation or rollback to conservative strategy is selectively applied; the total power demand is distributed to each path of the heater, and the distribution is based on the efficiency priority principle and the life balance principle; a gradual change constraint is imposed on the single-cycle power or duty cycle variation amplitude, and a hysteresis logic is set to avoid frequent switching; the distribution result is converted into a PWM signal to drive the actuator unit; abnormal detection is performed on the sensor layer, the actuator layer and the system layer; based on the abnormal detection result, the system state is switched to the corresponding fault tolerance level, and the fault tolerance level is associated with different power upper limits and / or runtime limits.
[0014] Preferably, the partition adaptive control strategy comprises: when the temperature error is greater than a first threshold, a first set of control parameters is used to achieve rapid approach to the target temperature; when the temperature error is between the first threshold and the second threshold, a second set of control parameters is used to balance the response speed and stability; when the temperature error is less than or equal to the second threshold value, a third set of control parameters is adopted to fine-tune and suppress overshoot; wherein the first threshold value and the second threshold value are coupled and corrected according to flow conditions.
[0015] Preferably, the indicators for the prediction effectiveness determination by the main control unit include at least one of the following: whether the prediction residual exceeds a preset residual threshold value; variance or consistency of the prediction residual within a preset time window; consistency of the prediction trend and the measured trend.
[0016] Preferably, the life balance principle is based on the cumulative operating hours and / or health assessment results of each heater, when the cumulative operating hours difference between different heaters exceeds a first trigger threshold value, triggering load transfer to balance the life, and using a double threshold hysteresis logic containing a trigger threshold value and a recovery threshold value to suppress oscillation.
[0017] Preferably, the fault tolerance level includes: normal level, corresponding to complete control strategy and no power or time limit; degraded level, switching to backup sensors and limiting system maximum power in response to main sensor abnormalities; estimation level, entering a model-based estimation mode and simultaneously limiting system maximum power and single run duration in response to abnormalities in both main and backup sensors; safety shutdown level, immediately stopping heating and alarming in response to estimation timeout, actuator failure or temperature out of control.
[0018] Preferably, the main control unit is further configured to, at the degraded level or the estimation level, if the abnormal condition is removed, the system state is rolled back to a higher level of fault tolerance level.
[0019] Preferably, the main control unit is further configured to manage the operating state of the system, the operating state including: standby state, monitoring user proximity signals or preheating trigger signals; preheating state, allowing the water temperature to approach the target temperature interval in an energy-saving manner; ready state, maintaining the water temperature near the target interval and waiting for the water outlet signal; running state, performing the partition adaptive control, trend prediction, power distribution and fault tolerance processing; stop state, performing emptying or slow stop processing after water outlet is stopped; fault state, entering a hierarchical fault tolerance process and executing a safety protection strategy.
[0020] A control method for a multi-heater cooperative and hierarchical fault tolerance control system, comprising the following steps: S1, collect and verify sensor signals, and perform data validity determination; S2, based on the determination result, calculate the total power demand using a partition adaptive control strategy; S3, perform short-term trend prediction on temperature data, and selectively apply feedforward compensation according to the prediction validity determination result; S4, based on the efficiency priority and life balance principle, distribute the total power demand to each heater; S5, during power distribution, apply gradual change constraint and hysteresis logic to single cycle power change; S6, convert the distribution result into a PWM signal to drive the heater; S7, real-time multi-layer anomaly detection; S8, according to the anomaly detection result, switch the system fault tolerance level, and execute the power and time limit strategy corresponding to the level.
[0021] Preferably, the short-term trend prediction uses the grey prediction GM(1, 1) model or equivalent method, and the prediction time domain is 2 to 5 control periods.
[0022] Preferably, the life balance principle is realized by monitoring the cumulative working hours of each heater, and when the working hour difference exceeds the set threshold, the power distribution strategy is dynamically adjusted to promote the cumulative working hours to be balanced. Advantages
[0023] The multi-heater cooperative and hierarchical fault-tolerant control system and method of the application has the following advantages through comparative observation under typical prototype working conditions: 1. Response and steady state: The speed is improved by about 15% to 25% compared with the fixed PID in the warming-up stage, and the constant temperature fluctuation amplitude is reduced by about 40% to 60%.
[0024] 2. Energy consumption and life: Feedforward reduces overshoot energy consumption, and life balance reduces device differences; the overall energy saving is about 5% to 10%, and the life balance improvement is about 30%.
[0025] 3. Safety and controllability: three-layer detection + four-level fault tolerance and linkage with execution boundary, limited power / time service or orderly shutdown in failure scenarios.
[0026] 4. Low-cost implementation: suitable for general-purpose MCUs (such as Cortex-M0 / M3, etc.), parameters can be calibrated and configured by table lookup.
[0027] 5. Engineering flexibility: threshold, window and coupling relationship support linear / segmented / table lookup expression, suitable for different power segments and scenarios. BRIEF DESCRIPTION OF DRAWINGS
[0028] The application will be further described below with reference to the accompanying drawings and embodiments. Figure 1 The schematic diagram of the core control main flow of the control system of the application; Figure 2 The state machine diagram of the control system of the application. Figure 3 The hierarchical fault-tolerant strategy diagram of the control system of the application. DETAILED DESCRIPTION
[0029] The specific embodiments of the application will be described in detail below with reference to the accompanying drawings. It should be pointed out that the following embodiments are only exemplary and do not constitute a limitation on the protection scope of the application.
[0030] In a preferred embodiment of the application, the multi-heater cooperative and hierarchical fault-tolerant control system of the application mainly includes the following units: Sensor unit: water outlet temperature sensors (main sensor and backup sensor) and flow sensors are used. The temperature sensor is preferably PT1000 or NTC thermistor, and the flow sensor is preferably Hall effect flowmeter or turbine flowmeter. The main and backup temperature sensors are installed in close proximity but independently wired to improve reliability.
[0031] Actuator unit: contains three independent heaters with power configurations of 800W, 1200W and 1500W. Each heater supports duty cycle modulation, and power control is performed through solid-state relays or IGBT modules.
[0032] Master control unit: a general MCU based on ARM Cortex-M3 core (such as STM32F103 series) is used, which is responsible for signal acquisition, control algorithm execution, state management and fault-tolerant processing. The working frequency of the MCU is set to 72MHz, and a 12-bit ADC is provided for sensor signal acquisition.
[0033] Data storage unit: a FLASH memory (such as W25Q16) with SPI interface is used, and a wear leveling algorithm is used to manage persistent storage of parameter configuration, user preferences, running statistics and fault logs.
[0034] Safety protection link: independent over-temperature protection switch and fuse are set at the hardware level; power limitation, time limitation and safety shutdown strategy are implemented at the software level for double protection.
[0035] The specific configuration of the system is as follows.
[0036] I. Control framework of the system.
[0037] See Figure 1The core control main flowchart is shown, and the system adopts interlayer coupling and priority coordination. Specifically, a "partition-prediction-optimization" three-level cooperative control is adopted, and the layers have specific order and priority: a. Data validity determination: collect and verify temperature / flow / primary and backup consistency, and decide to enter the control loop or return to the estimation / filter path.
[0038] b. Partition adaptive control: select strategy according to error interval; at least one parameter is coupled with flow and other working conditions for correction.
[0039] c. Prediction validity determination: determine whether the prediction is reliable based on residual error, variance / consistency window, and trend consistency.
[0040] d. Feedforward / conservative strategy switching: apply feedforward if prediction is valid; otherwise, return to conservative (weaken / stop feedforward).
[0041] e. Multi-objective power distribution: consider efficiency priority and life balance (cumulative working hours / health degree, etc.) at the same time.
[0042] f. Gradual constraint and hysteresis: limit single-cycle duty cycle / power change amplitude, and set hysteresis to avoid frequent switching and temperature jump / power grid impact.
[0043] g. PWM execution: convert the distribution result into PWM output and record statistics / working hours.
[0044] h. Abnormality detection: three-layer detection of sensor layer, actuator layer, and system layer to identify out-of-range and failure.
[0045] i. Hierarchical fault tolerance: NORMAL / DEGRADED / ESTIMATED / SAFE_STOP; fault tolerance level can interrupt or limit power distribution and return to the collection / determination loop for closed-loop control.
[0046] Core features (arrangement relationship): data validity→partition adaptive→prediction determination→feedforward / conservative→power distribution→gradual / hysteresis→PWM execution→abnormality detection→hierarchical fault tolerance (fault tolerance can interrupt or limit power distribution and return to the collection / determination loop). The arrangement of interlayer trigger conditions, return paths, and mutual constraints constitutes the key cooperative features of the invention.
[0047] In specific implementation: 1.1. Error partition: large error interval (e.g. error> threshold 1): larger proportion coefficient and moderate integration, fast approach to target. Medium error interval (e.g. threshold 2< error≤ threshold 1): balance speed and stability, avoid overshoot. Small error interval (e.g. error≤ threshold 2): strengthen fine tuning and anti-jitter, reduce overshoot probability. Threshold can be determined by sample machine calibration / table lookup, not fixed value.
[0048] 1.2 Short-term trend prediction and feedforward compensation: Perform univariate short-term prediction on recent temperature data (such as grey prediction GM(1,1) or equivalent methods), and generate appropriate feedforward compensation within the limited prediction time domain (Example 2 to 5 control cycles) to offset thermal inertia deviation in advance.
[0049] 1.3. Prediction Validity Determination: Boolean decision is made based on one or more indicators such as prediction residual threshold, residual variance / consistency window, and consistency with actual trends. When any indicator exceeds the limit or lacks stability, a conservative strategy is adopted (stopping or reducing feedforward), adhering to the principle of "safety first, performance second". The criterion can be calibrated / looked up in a table and is not limited to a single fixed value.
[0050] 1.4 Multi-objective power allocation and gradual constraint: Distributing the total power demand to multiple heaters, comprehensively including: Efficiency First: Prioritize high-efficiency combinations; Lifetime balancing: Load transfer is triggered based on cumulative working hours / health status, and oscillation is suppressed by dual threshold hysteresis of "trigger threshold / recovery threshold".
[0051] An upper limit is set for single-cycle duty cycle / power variation (example Δ ≤ 20%~40% of rated value, which can be calibrated according to power supply capability / EMC), and hysteresis logic is superimposed to reduce the impact of abrupt changes and frequent switching. Lifetime equalization and gradual / hysteresis are effective simultaneously in the same allocation layer, forming a composite constraint.
[0052] Figure 1 In the control process, the priority of anomaly detection covers the power layer, and fault tolerance can be triggered at any stage to limit or interrupt power output.
[0053] II. State Machine Management.
[0054] like Figure 2 As shown, the system state machine includes: IDLE (Standby): Triggered by detecting user proximity / warm-up; PREHEAT (optional): Energy-saving method close to the target range; READY: Remain near the target area, waiting for water to be released; RUNNING: Executes "partitioning-prediction-optimization" collaborative control; STOP (Stop): Drainage / slow-stop treatment after water flow stops; FAULT (Fault): Enters graded fault tolerance and takes protective measures.
[0055] If any critical anomaly is detected in any state (such as sensor failure, overcurrent, or temperature control instability), the system can immediately switch to FAULT; once the fault is resolved, the system can return to IDLE.
[0056] III. Anomaly Detection and Graded Fault Tolerance.
[0057] like Figure 3 As shown, the system's anomaly detection and graded fault tolerance, including fault tolerance level → power / time boundary mapping, specifically includes: Three-layer detection: Sensor layer (range / rate of change / master / backup consistency); Actuator layer (current / response / duty cycle anomaly); System level (temperature control instability trend / energy consistency anomaly).
[0058] Level 4 fault tolerance: L1 NORMAL: Complete control strategy with no power / time limitations.
[0059] L2 DEGRADED: Main sensor malfunction → switch to standby; mapped to power limit (example 60% to 80% of rated).
[0060] L3 ESTIMATED: Both primary and backup are abnormal → estimation mode; mapped to power limit (example 40%~60% of rated) + time limit (example 3~10 minutes).
[0061] L4 SAFE_STOP: Estimation timeout / Actuator serious fault / Temperature control failure → Immediate shutdown alarm.
[0062] The fault tolerance level not only switches the use of sensors, but also links to execution boundaries such as maximum power and operating time limits, forming a closed-loop mapping of "fault tolerance level → power / time boundary". Each level can roll back according to recovery conditions (e.g., L3 → L2 → L1).
[0063] IV. Examples
[0064] Example 1: Rapid approach to and stable constant temperature.
[0065] Conditions: Inlet water temperature 15°C, set temperature 38°C, flow rate 1.5 L / min, ambient temperature 20°C (example).
[0066] Process: Large error → partition adaptive fast convergence; after the prediction passes the validity judgment, feedforward is applied; small error → fine-tuning; the allocation layer is smoothed out by gradual / hysteresis.
[0067] Results: The temperature approached the target range (approximately 37–39°C) within the exemplary time (approximately 20 s), with fluctuations of approximately ±0.5°C during the isothermal phase (example).
[0068] Example 2: Hierarchical fault tolerance for sensor anomalies.
[0069] Main sensor abnormality→L2: judge, backup and limit power (e.g. 70%)→restore to L1; if both main and backup fail→L3: estimate (e.g. 50%+5 minutes)→timeout or out of control→L4: safe shutdown.
[0070] Results: service uninterrupted or orderly exit, avoid scalding / equipment damage.
[0071] Example 3: life balance allocation effect.
[0072] Three-way heater (example: 800W / 1200W / 1500W), cumulative working hour difference→trigger life balance and hysteresis strategy→converge cumulative working hour difference after exemplary period (e.g. <100 h).
[0073] Results: observed life balance improvement and maintenance cost reduction.
[0074] Under multiple groups of working conditions (example: water inlet temperature 10-25°C, set 35-42°C, flow rate 0.8-2.5 L / min), it is observed that: The average temperature rise is about 18-25 s; The constant temperature fluctuation is about ±0.3-0.8°C; The total energy saving is about 7%; Life balance: cumulative working hour difference <80 h after 1000 h; Fault tolerance: main sensor failure automatically downgraded; when both sensors fail, the estimation mode is limited to (5 minutes) before safe shutdown.
[0075] The embodiments are used to illustrate the feasibility and effectiveness, and do not constitute a limitation on the scope of protection.
[0076] The above embodiments are only for illustrating the technical concept and characteristics of the present application, the purpose is to enable those skilled in the art to understand the content of the present application and implement it, and cannot limit the protection scope of the present application. Any modification made according to the spirit and essence of the main technical solution of the present application should be covered within the protection scope of the present application.
Claims
1. A multi-heater coordinated and staged fault-tolerant control system, characterized by, The application relates to a water heater, comprising: a sensor unit for collecting water outlet temperature signals and flow signals; an actuator unit comprising at least two heating paths capable of independent power modulation; a main control unit in communication with the sensor unit and the actuator unit, and configured to perform the following operations: collect and verify the signals of the sensor unit, and make a data validity determination; based on the data validity determination result, calculate the total power demand using a partition adaptive control strategy, wherein the corresponding control parameters are selected according to the different intervals of the temperature error; make a short-term trend prediction of the temperature data in the recent period, and selectively apply feedforward compensation or fall back to a conservative strategy according to the prediction validity determination result; distribute the total power demand to each heating path, which is based on the efficiency priority principle and the life balance principle; apply a gradual change constraint to the power or duty cycle change amplitude of a single cycle, and set a hysteresis logic to avoid frequent switching; convert the distribution result into a PWM signal to drive the actuator unit; perform abnormality detection of the sensor layer, the actuator layer and the system layer; based on the abnormality detection result, switch the system state to the corresponding fault tolerance level, which is associated with different power upper limits and / or running time limits.
2. The multi-heater synergistic and staged fault-tolerant control system of claim 1, wherein, The partition adaptive control strategy comprises: when the temperature error is greater than a first threshold value, a first set of control parameters is used to achieve rapid approach to the target temperature; when the temperature error is between the first threshold value and a second threshold value, a second set of control parameters is used to balance the response speed and stability; when the temperature error is less than or equal to the second threshold value, a third set of control parameters is used for fine tuning and suppression of overshoot; wherein the first threshold value and the second threshold value are coupled and corrected according to the flow condition.
3. The multi-heater synergistic and staged fault-tolerant control system of claim 1, wherein, The indexes for the main control unit to make a prediction validity determination include at least one of the following: whether the prediction residual exceeds a preset residual threshold value; the variance or consistency of the prediction residual within a preset time window; the consistency of the prediction trend and the measured trend.
4. The multi-heater synergistic and staged fault-tolerant control system of claim 1, wherein, The life balance principle is based on the cumulative running hours and / or health degree evaluation results of each heating path. When the cumulative working hours difference between different heaters exceeds a first trigger threshold, load transfer is triggered to balance the life, and a double threshold hysteresis logic containing a trigger threshold and a recovery threshold is used to suppress oscillation.
5. The multi-heater synergistic and hierarchical fault-tolerant control system of claim 1, wherein, The fault tolerance levels include: a normal level corresponding to the complete control strategy without power or time limit; a degraded level, in response to a main sensor abnormality, switching to a backup sensor and limiting the maximum power of the system; an estimation level, in response to both the main sensor and the backup sensor being abnormal, entering a model-based estimation mode, and simultaneously limiting the maximum power of the system and the single running time; a safe shutdown level, in response to estimation timeout, actuator serious failure or temperature out of control, immediately stopping heating and alarming.
6. The multi-heater synergistic and hierarchical fault-tolerant control system of claim 5, wherein, The main control unit is further configured to, at the degraded level or the estimation level, if the abnormal condition is removed, make the system state fall back to a higher fault tolerance level.
7. The multi-heater synergistic and hierarchical fault-tolerant control system of claim 1, wherein, The main control unit is further configured to manage the running state of the system, wherein the running state comprises: a standby state for monitoring user proximity signals or preheating trigger signals; a preheating state for making the water temperature approach the target temperature interval in an energy-saving manner; Ready state, maintaining water temperature near target interval, waiting for water-out signal; Running state, performing partition adaptive control, trend prediction, power distribution and fault-tolerant processing; Fault state, entering hierarchical fault-tolerant process and executing safety protection strategy.
8. A control method for the multi-heater synergic and hierarchical fault-tolerant control system of any one of claims 1-7, characterized in that, Comprising the following steps: S1, collecting and verifying sensor signals, performing data validity determination; S2, based on the determination result, using partition adaptive control strategy to calculate total power demand; S3, performing short-term trend prediction on temperature data, and selectively applying feedforward compensation according to the prediction validity determination result; S4, based on the efficiency priority and life balance principle, distributing total power demand to each heater; S5, during power distribution, applying gradual change constraint and hysteresis logic to single-cycle power change; S6, converting the distribution result into PWM signal to drive the heater; S7, performing real-time multi-layer anomaly detection; S8, according to the anomaly detection result, switching the system fault-tolerant level, and executing the power and time limit strategy corresponding to the level.
9. The control method according to claim 8, characterized by, The short-term trend prediction uses the grey prediction GM(1,1) model or equivalent method, and the prediction time domain is 2 to 5 control periods.
10. The control method according to claim 8, characterized by The life balance principle is realized by monitoring the cumulative working hours of each heater, and when the working hour difference exceeds the set threshold, the power distribution strategy is dynamically adjusted to promote the cumulative working hours to be balanced.