Intelligent defrosting control method and system for low-temperature variable-frequency air source heat pump and computer program product

By employing an intelligent defrosting control method that integrates multimodal frost information fusion and deep temporal prediction, combined with distributed reinforcement learning, the defrosting control problem of low-temperature variable frequency air source heat pumps under low temperature and high humidity conditions was solved. This method achieves the goals of timely defrosting, continuous heating, minimum energy consumption, and safe and reliable control, thereby improving the system's operating efficiency and reliability.

CN122015358APending Publication Date: 2026-05-12XINLEI COMPRESSOR CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XINLEI COMPRESSOR CO LTD
Filing Date
2026-03-27
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing low-temperature variable frequency air source heat pumps suffer from high energy consumption, severe heating attenuation, and untimely and inflexible control during defrosting under low temperature and high humidity conditions, making it difficult to achieve the four objectives of timely defrosting, continuous heating, energy economy, and control reliability.

Method used

An intelligent defrosting control method is adopted, which integrates multimodal frost information fusion, deep temporal prediction, and distributed reinforcement learning. The optimal control vector is predicted by a convolutional-LSTM deep network, and the compressor frequency, fan air volume, and expansion valve opening are optimized by combining safety redundancy judgment and online reinforcement learning. This achieves timely defrosting, continuous heating, minimum energy consumption, and safe and reliable control.

Benefits of technology

It significantly reduces defrosting energy consumption by 8%–15%, shortens defrosting downtime, improves heating continuity, enhances low-temperature heating capacity and seasonal energy efficiency, and strengthens system operation safety and cross-climate adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122015358A_ABST
    Figure CN122015358A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of air source heat pumps, in particular to an intelligent defrosting control method and system for a low-temperature variable-frequency air source heat pump and a computer program product. According to the method, characteristics of coil temperature, frost layer thickness, environment humidity and the like collected by a multi-modal sensor are fused, optimal control parameters of a next control period are predicted through a convolution-LSTM deep network, and security redundancy judgment is executed in combination with confidence and a hardware boundary; and triggering a rollback strategy based on the temperature slope and the frost thickness when the confidence coefficient is insufficient. Meanwhile, a distributed reinforcement learning framework is adopted to optimize a prediction model and a rollback control table online, and dynamic balance of defrosting energy consumption and integrity is achieved. Experiments prove that the defrosting energy consumption and the shutdown time can be remarkably reduced, and the stability and the heating performance of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of air source heat pump technology, and in particular to an intelligent defrosting control method, system and computer program product for a low-temperature variable frequency air source heat pump. Background Technology

[0002] Air source heat pumps have been widely used in residential and commercial HVAC systems due to their advantages such as dual heating / cooling modes, energy saving, and flexible installation. However, when the outdoor ambient temperature is below 5°C and the humidity is high, frost easily forms on the surface of the outdoor finned heat exchanger (evaporator). The frost layer significantly reduces the heat transfer coefficient and airflow, leading to reduced heating capacity, compressor exhaust overheating, and a sharp drop in system efficiency, which can even trigger low-pressure protection shutdown in severe cases. Therefore, an efficient and reliable defrosting control strategy is one of the core bottlenecks for the commercialization of low-temperature variable frequency air source heat pumps.

[0003] Early reverse-cycle defrosting typically used a fixed-time method: the compressor would run for 30–90 minutes, then forcibly switch the four-way valve to enter defrosting mode, lasting 4–10 minutes before forcibly exiting. This method requires no additional sensors but has two drawbacks: 1. Under-defrosting: if the defrost layer hasn't completely melted before exiting, residual frost will rapidly regenerate; 2. Over-defrosting: even when the external environment improves, defrosting continues for a fixed duration, leading to prolonged heating interruptions and wasted power, resulting in a ≥10% decrease in overall high performance factor (HSPF). Improvements have been made by using coil temperature difference as a trigger / termination criterion. For example, US Patent 4373349 proposes an adaptive defrosting control system for heat pump systems. However, because it only monitors a single point temperature, this method is greatly affected by wind speed, frost distribution, and sensor drift; furthermore, using a fixed-frequency compressor, it cannot optimize frequency and energy consumption in real time under variable-frequency operating conditions.

[0004] In the 1980s, microcontroller-based self-learning timing strategies emerged, such as the US4573326A and US4751825A. The controller records the time t of the previous frost-defrost-re-frost cycle and correlates it with the target frost thickness time t. ref The proportional coefficient is calculated to update the timing threshold for the next cycle. This scheme can be adjusted according to the season and unit aging, but it is still based solely on time characteristics and lacks utilization of dynamic information such as ambient wet-bulb temperature and coil temperature rise slope; at the same time, it uses a single sensor, making it difficult to detect non-uniform frost.

[0005] The paper "Deep-learning-based prediction on performance change of ASHP" uses CNN-LSTM to predict the COP of heat pumps, with an RMS error of approximately 2.4%. Other LSTM studies have incorporated meteorological fields and operating conditions as inputs to predict airflow. These works demonstrate the effectiveness of deep temporal networks in multi-source nonlinear regression, but the research focus remains limited to energy consumption or capacity prediction, without coupling with the defrost control loop. In recent years, deep reinforcement learning (DRL) has begun to show promise in building energy scheduling: DQN and DDPG are used for air conditioning setpoint optimization, and PPO is used for the economical operation of heating systems. However, no open-source / patent literature has yet found applications of distributed PPO in low-temperature variable frequency heat pump-defrosting-real-time safety constraint scenarios.

[0006] Based on the above research, existing methods either suffer from significant energy loss, are overly conservative, rely solely on empirical thresholds making them difficult to adapt to different seasons, or lack a unified framework for collaborative control of multiple actuators and online self-learning. In scenarios where ultra-low temperatures (below -25°C) and high humidity frost coexist, a smart defrosting solution that integrates deep prediction, reinforcement learning, and safety redundancy is needed to achieve four objectives: timely defrosting, continuous heating, energy efficiency, and control reliability. Summary of the Invention

[0007] To address the aforementioned technical problems, the technical objective of this invention is to provide an intelligent defrosting control method for low-temperature variable frequency air source heat pumps. By integrating multimodal frosting information, deep time-series prediction, and distributed reinforcement learning, the compressor frequency, fan air volume, and expansion valve opening are adaptively and collaboratively optimized throughout the defrosting process, achieving the goals of timely defrosting, continuous heating, minimum energy consumption, and safe and reliable control.

[0008] To achieve the above objectives, the present invention adopts the following technical solution:

[0009] A method for intelligent defrosting control of a low-temperature variable frequency air source heat pump, the method comprising:

[0010] S1 Data Fusion Acquisition:

[0011] 1.1 Obtain the temperature T of the defrost coil of the finned heat exchanger def ;

[0012] 1.2 Obtain the frost thickness characteristic L in real time using an ice thickness detection sensor or visual imaging device. ice ;

[0013] 1.3 Collect the compressor's current frequency F, outdoor fan air volume Q, electronic expansion valve opening θ, and ambient wet-bulb temperature T. wb Relative humidity (RH);

[0014] S2 Nonlinear Prediction Decision:

[0015] Using a pre-trained convolutional-LSTM (C-LSTM) deep network model, the multimodal feature sequence {T} from step S1 is input. def ,L ice ,F,Q,θ,T wb ,RH}, output the optimal control vector {F} for the next control cycle. pred Q pred ,θ pred} and give the confidence level P. conf ;

[0016] S3 security redundancy assessment:

[0017] When P conf ≥P thr And when the output vector satisfies the hardware safety boundary, predictive control is executed, P thr Set the confidence threshold; otherwise, trigger a fallback strategy.

[0018] a) If T def ≥T defT -T set And dT def / dt≥K slope T defT T is the defrost exit temperature threshold. set To reduce the frequency temperature difference offset in advance, K slope If the temperature slope threshold is met, the compressor frequency will be reduced to the first-stage frequency reduction target frequency F according to the preset curve. set1 Simultaneously adjust the target air volume Q set With target opening θ set ;

[0019] b) If T def ≥T defT or L ice ≤L thr L thr To achieve the minimum permissible residual frost thickness threshold, the compressor frequency is reduced to the secondary throttling target frequency F. set2 And switch the four-way valve back to heating;

[0020] S4 Online Reinforcement Learning Update:

[0021] Using defrosting energy consumption-time integral and remaining frost thickness as reward functions, a distributed reinforcement learning agent is used to fine-tune the parameters of C-LSTM and the backoff control table online, with an update interval of no more than N defrosting cycles;

[0022] S5 End Judgment:

[0023] When the compressor, fan, and expansion valve are all running stably, Δtend And T def L ice When the exit threshold is met, the current defrosting control cycle ends.

[0024] Preferably, the convolutional-LSTM network includes:

[0025] a) Two cascaded two-dimensional convolutional layers, each with a kernel size of 3×3, the first layer having ≥32 output channels and the second layer having ≥64 output channels, and batch normalization and ReLU activation are used between the two layers;

[0026] b) Set the time step to T after the convolutional layer. seq A bidirectional LSTM layer with 128 hidden units and a dropout rate of 0.3 on the output side;

[0027] c) Convolutional features and one-dimensional features such as temperature and humidity are concatenated along the channel dimension before LSTM.

[0028] Preferably, the convolutional-LSTM network further incorporates a multi-head self-attention fusion layer at the LSTM output, with h=4 heads, to assign different weights to the frost thickness sequence and the temperature sequence, thereby improving the response speed to sudden frost growth.

[0029] Preferably, the reinforcement learning unit adopts a distributed PPO architecture, including:

[0030] a) At least two Actor nodes are deployed on multiple heat pump edge control boards to collect interaction trajectories in parallel and execute the current strategy;

[0031] b) A Learner node, deployed on the gateway server, is responsible for calculating the global policy gradient and updating parameters;

[0032] c) A parameter server for asynchronously synchronizing policy network weights between Actor nodes and Learner nodes, with an update frequency of not less than f. sync =1Hz;

[0033] Furthermore, the distributed PPO experience replay employs a priority experience replay cache, with priority determined by time difference error. Weighted, the cache size is ≥10,000 entries, and the temperature distribution and frost thickness distribution during random sampling must meet the principle of being the same as the actual working conditions;

[0034] Further optimization, the priority empirical replay sampling probabilities are as follows:

[0035] .

[0036] Preferably, the reinforcement learning reward function R satisfies Where E is the energy consumption for this defrosting cycle, and L res N represents the remaining frost thickness upon exiting. switch The number of four-way valve switching times, α, β, and γ are adaptively adjusted by the Learner node based on the energy consumption-performance Pareto front of the most recent M defrost cycles to achieve a dynamic balance between energy consumption and defrost integrity;

[0037] Furthermore, the Pareto adaptive weight update satisfies:

[0038] .

[0039] Preferably, Fset1 and Fset2 in the rollback strategy are obtained by calculating the energy consumption curve of the previous defrosting cycle using an exponentially weighted average, so as to ensure the adaptability of the rollback control.

[0040] Preferably, the visual imaging device is a ToF depth camera or a millimeter-wave radar with a resolution ≥0.1mm, and is equipped with a thermal compensation algorithm to reduce condensation-fog interference.

[0041] Furthermore, the present invention also provides an intelligent defrosting control system for a low-temperature variable frequency air source heat pump, comprising:

[0042] A multimodal sensor array is used to acquire data such as fin temperature, frost thickness, and ambient temperature and humidity.

[0043] The actuators include a variable frequency compressor, an outdoor fan, and an electronic expansion valve;

[0044] The depth prediction unit, with a built-in convolutional LSTM network, is used to predict the optimal control vector based on sensor data.

[0045] Reinforcement learning units are used to update the parameters of the depth prediction unit in real time based on energy consumption and remaining frost thickness;

[0046] A safety redundancy control unit is used to execute a fallback strategy when the prediction confidence is insufficient or exceeds the limit;

[0047] The central controller is used to implement all the steps of the method described above and to coordinate the aforementioned units.

[0048] Furthermore, the present invention also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the method.

[0049] Furthermore, the present invention also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the method.

[0050] This invention, by employing the aforementioned technical solution, integrates multimodal frosting information, deep temporal prediction, and distributed reinforcement learning to adaptively and collaboratively optimize compressor frequency, fan airflow, and expansion valve opening throughout the defrosting process. This achieves the goals of timely defrosting, continuous heating, minimum energy consumption, and safe and reliable control. It possesses the following technical advantages:

[0051] 1. Significantly reduce defrosting energy consumption: By using a convolutional LSTM deep network to predict the optimal compressor frequency F, fan air volume Q, and expansion valve opening θ for the next control cycle, the energy consumption per unit of heat during defrosting is reduced. Compared to the traditional ΔT-timed dual-threshold method, energy consumption is reduced by 8%–15%. Distributed PPO reinforcement learning fine-tunes model parameters online based on the real-time energy consumption–frost thickness reward / penalty function, which can further reduce energy consumption by 2%–4% within 3–5 defrosting cycles.

[0052] 2. Shorten defrosting downtime and improve heating continuity: Utilize frost thickness L ice With temperature slope dT def The dual criterion of / dt accurately identifies the transient state immediately after frost removal; combined with a two-stage frequency reduction curve, the switching timing of the four-way valve is advanced by 30–60 seconds. Actual testing shows that the duration of a single defrost cycle is reduced from 6–8 minutes to 4–6 minutes, and the cumulative heating interruption time throughout the year is reduced by approximately 12 hours (based on a 2HP unit and a 150-day heating season in cold regions).

[0053] 3. Improve low-temperature heating capacity and seasonal energy efficiency (HSPF): Pre-emptive frequency reduction prevents compressor overheating and high exhaust temperatures during the defrosting process; simultaneously reducing fan airflow and valve opening mitigates system pressure fluctuations and maintains higher evaporation temperatures. Under −25℃ / 85%RH conditions, the heating capacity reduction is reduced from 35% to 22%; the seasonal weighted HSPF increases by 6%–9%.

[0054] 4. Enhanced cross-climate adaptability: The fusion of multimodal perception (temperature, humidity, frost thickness) and deep learning features enables the model to be transferred to different models and climate zones. Parameter self-tuning can be completed in just 1–2 days of online learning, avoiding the need for multiple manual calibrations required by traditional thresholding schemes.

[0055] 5. Improve system operational security and robustness: Import prediction confidence level P conf With a dual verification mechanism for hardware security boundaries, the controller automatically reverts to a safe frequency reduction table when front-end sensors fail, network latency occurs, or model drift occurs, preventing compressor liquid slugging, overcurrent, or high-pressure shutdown. This ensures an annual defrosting failure rate (the percentage of residual frost requiring secondary defrosting) of less than 1%, significantly better than the industry standard of 3%–5%.

[0056] In summary, this invention has achieved significant improvements in multiple dimensions, including energy consumption, defrosting time, low-temperature heating capacity, cross-regional adaptability, and system safety, providing key technical support for the efficient and stable operation of low-temperature variable frequency air source heat pumps in extremely cold and humid regions. Attached Figure Description

[0057] Figure 1 This is a schematic diagram of the overall hardware framework.

[0058] Figure 2 This is a flowchart of the method.

[0059] Figure 3 This is a flowchart of the data synchronization and preprocessing process.

[0060] Figure 4 This is a diagram of the convolutional-LSTM-attention deep network structure.

[0061] Figure 5 State machine diagram for safety redundancy decision-making.

[0062] Figure 6 This is a diagram illustrating a distributed PPO architecture and data flow. Detailed implementation details are provided.

[0063] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.

[0064] I. System Hardware Overall Architecture

[0065] 1.1 Unit Structure

[0066] Compressor: 180V~380V DC inverter scroll compressor, rated capacity 7.5kW; drive supports stepless speed regulation between 15Hz and 120Hz.

[0067] Finned heat exchanger: Three rows of copper tubes and aluminum fins, with a hydrophilic coating and low thermal conductivity. .

[0068] Outdoor fan: EC brushless DC fan, rated air volume Rotation speed .

[0069] Electronic expansion valve: Stepper motor type, 480 pulses at full opening.

[0070] Four-way directional valve: DN20, switching time <0.8s.

[0071] Main control board: Rockchip RK3568 + NPU0.8TOPS; supplemented by STM32G474 MCU real-time layer.

[0072] Edge computing module (Edge-AIBox): NVIDIA Jetson Orin Nano 8GB, dedicated to inference / learning, with a maximum power consumption of 15W.

[0073] Communication: The master controller communicates with Edge-AI via TSN-Ethernet; slave nodes (temperature / pressure / radar) communicate via CAN-FD2Mbps star topology.

[0074] 1.2 Sensor Configuration

[0075]

[0076] II. Data Synchronization and Preprocessing

[0077] 1. Bus scheduling and timestamp calibration

[0078] 1.1 All sensor nodes (temperature RTD, ToF camera, humidity probe, current transformer, etc.) are connected to a 2Mbps CAN-FD bus. The main control board MCU acts as the bus arbitration node, sending a SYNC frame every 100ms to instruct each node to report in the next time slot.

[0079] 1.2 The main controller integrates an IEEE-1588 Precision Time Protocol (PTP) hardware timestamp unit. After power-on, the main controller first broadcasts PTP_SYNC, PTP_FOLLOW_UP, and PTP_DELAY_REQ / RESP messages to the edge AI module (Jetson). After two complete handshake cycles, the system clock of Jetson is adjusted to have an error of ≤0.5ms with the MCU.

[0080] 1.3 Subsequently, the MCU allocates a 32-bit periodic time stamp t via a CAN-FD remote frame. ref Provide this to each sensor node. When reporting data, the node will send t... ref Write back the first 4 bytes to the DLC so that the master can align multi-source data in the buffer along a uniform timeline.

[0081] 2. Second-frame data acquisition

[0082] 2.1 The MCU triggers a data acquisition cycle every 1 second. The system time is latched immediately at the start of the cycle. .

[0083] 2.2 Temperature at 18 points on RTD The ADC sampling and reporting are completed within t0+50–120ms; the compressor, fan, and current transformer complete sampling and reporting within t0+50–120ms; the humidity probe completes sampling and reporting within t0+50–120ms. Completed internally Sample and report.

[0084] 2.3 The ToF depth camera is triggered by an external exposure at 1Hz using a Jetson sensor, which then interrupts and returns to the original 64-pixel distance matrix D. raw Jetson immediately writes a 3×3 median filter kernel onto the NPU circuitry for denoising, resulting in a smoothing matrix D.

[0085] 3. Depth map denoising and normalization

[0086] 3.1 Regarding D raw The mathematical expression for using 3×3 median filtering is:

[0087] ;

[0088] 3.2 The filtered distance matrix is ​​uniformly scaled to the 0–1 range based on pixel intensity:

[0089] .

[0090] 4. Frost thickness calculation

[0091] 4.1 The system records a depth reference frame D0 30 seconds after the initial defrosting is completed. Thereafter, the instantaneous frost thickness L is calculated every second. ice :

[0092] ;

[0093] Where: D 0,i,j - The baseline distance between the i and j-th pixels; D i,j - Current distance in pixels.

[0094] 5. Temperature drift compensation

[0095] If the internal temperature T of the camera is detected cam Compared with the current outdoor wet-bulb temperature T wb Difference Then perform linear temperature drift correction:

[0096] .

[0097] 6. Merge to form the original state frame

[0098] The MCU concatenates the following thirteen types of fields into the raw state frame S(t0) for the current second: 18-point coil temperature vector; temperature slope dT def / dt (Obtained from 3s difference); Frost thickness L after camera compensation ice corr Compressor frequency F, fan speed Q, expansion valve opening θ; wet-bulb temperature T wbThe relative humidity (RH) is measured; the instantaneous current of the compressor / fan and the energy consumption per second integral are measured. This state frame is then pushed to the edge AI module Ring-Buffer to construct a sliding window tensor of length 16 and input it into the convolutional-LSTM prediction network.

[0099] Through the above six steps, the system achieves millisecond-level time synchronization, second-level sampling, and multimodal fusion, providing stable, accurate, and reproducible input data for subsequent nonlinear prediction and reinforcement learning control.

[0100] III. Feature Sequence Construction

[0101] 1. Maintain the sixteen-frame sliding window edge AI module to save the most recent 16 seconds of the original state frames: Whenever a new frame appears Write to the end of the queue, the oldest extra frame Automatically removed from the front of the queue.

[0102] 2. Scalar feature zero-mean normalization

[0103] Let the original value of a certain scalar characteristic be... The global mean has been obtained during the offline phase. with standard deviation The normalization formula is:

[0104] ;

[0105] Offline statistics Calculate and solidify each scalar value, such as coil temperature, wet bulb temperature, relative humidity, and current, into the firmware.

[0106] 3. Depth map normalization

[0107] For the depth matrix at the current time step: First, find the minimum and maximum pixel values: Then follow the formula:

[0108] ;

[0109] Mapped to Interval.

[0110] 4. Channel Dimension Stitching

[0111] Arrange all normalized scalars at the same moment into a column vector: Then put Copy to expand size Tensor: Finally, the feature is stitched together with the depth map on the channel axis to form a fused feature: Its dimensions are .

[0112] 5. Constructing temporal tensors

[0113] For each frame within the sliding window ( (Indicating from old to new) Execute step 4 to obtain... Stacking them along the newly introduced time dimension yields the input tensor: Its overall shape is This perfectly matches the three-dimensional structure of the convolutional-LSTM network: time × channel × space.

[0114] 6. Online statistical adaptive

[0115] At the end of each day, the system re-estimates the mean based on all samples from that day. with standard deviation Using an exponential decay coefficient (Recommended value: 0.05) Update global statistics: This allows for the slow tracking of data distribution drift caused by seasonality and unit aging without compromising model stability.

[0116] The main parameters are defined as follows: Sixteen-frame sliding window; Single frame original state; Normalization result for any scalar; Total number of scalar channels; Normalized depth map; Scalar 3D copy tensor; Single-frame features after image and scalar fusion; The final input tensor of the convolutional LSTM is the tensor.

[0117] Through the above six consecutive steps, the system completes data alignment, noise suppression, scale unification, and spatiotemporal stitching within a second-level cycle, providing high-quality, format-consistent temporal features for subsequent nonlinear prediction and reinforcement learning processes.

[0118] IV. Deep Prediction Network Design

[0119] 1. Convolutional Feature Extraction

[0120] Let the single-frame fusion tensor be: The two-stage 3x3 convolution operation is performed sequentially:

[0121] ;

[0122] Convolution kernel , Batch Normalization Layer (BN) k Normalization and learning scaling shift, activation function ReLU is used. The output tensor shape is fixed. .

[0123] 2. Convolution output and scalar concatenation

[0124] Normalize scalar vectors at the same time step Repeated expansion The channel axis is spliced ​​to obtain .

[0125] 3. Time series expansion

[0126] A sliding window of length sixteen Rearranged into a sequence along the time dimension Each frame First, a one-dimensional vector is obtained by global average pooling. This step compresses the spatial dimensions, retaining only the channel features.

[0127] 4. Bidirectional LSTM coding

[0128] will sequence Input bidirectional LSTM: The number of hidden units is 128, and the output is constructed by concatenating forward and reverse hidden states. A dropout rate is then applied to the output of all time steps. .

[0129] 5. Four-head self-attention fusion

[0130] Stack the full-time outputs of the LSTM into a matrix. Let the number of heads be... Key and value dimensions Attention calculation formula: ;in The fused feature vector is obtained. .

[0131] 6. Fully Connected Mapping and Output

[0132] ;

[0133] vector This refers to the target compressor frequency, target fan airflow, and target expansion valve opening for the next control cycle. (Scalar) To predict confidence levels.

[0134] Parameter definition: Convolution kernel weights; Convolution bias; Scalar channel count; BiLSTM(*,128) is a bidirectional LSTM with 128 hidden cells; The m-th head attention transformation matrix; Fully connected layer parameters (256→64); Output mapping parameters (64→3); Confidence score output weights and biases.

[0135] Through the above six consecutive steps, the network completes spatial convolution feature extraction, time series modeling, attention-weighted fusion, and generates continuous action vectors and credibility, providing core decision-making basis for subsequent safety redundancy judgment and collaborative control of the execution mechanism.

[0136] V. Offline Training and Quantization

[0137] 1. Dataset Construction

[0138] 1.1 Actual Machine Data

[0139] The system was continuously operated for 210 days in an environmental chamber at -5℃ to -30℃ and 60%RH to 95%RH, and a total of [data / records / records] were recorded. ; original state.

[0140] 1.2 Synthesis of extreme freezing rain

[0141] By overlaying the frozen water film thickness curve onto a multiphysics CFD-DEM model, a 400-hour depth map and temperature and humidity trajectory were generated; the label distribution was kept consistent with the actual action distribution, thus expanding the sample diversity.

[0142] 1.3 Division

[0143] The sequence is randomly divided into 80% training, 10% validation, and 10% testing; the sequence splitting step is 16 seconds to ensure that the time of each subset does not overlap.

[0144] 2. Loss Function

[0145] Let the continuous control label be The model output is Confidence level label (Data is accurate and reliable). Loss definition:

[0146] .

[0147] 3. Training details

[0148] Batch size 64, optimizer AdamW, initial learning rate The learning rate is calculated by Cosine Annealing over 120 epochs. Descending to If the MAE on the validation set does not decrease over ten consecutive epochs, the training will terminate early. At the end of training, the MAE on the test set is 0.87Hz (compressor frequency), 38m³h⁻¹ (air volume), and 4.2 pulses (valve opening).

[0149] 4. Quantitative Awareness Training (QAT)

[0150] 4.1 Instrumentation: Insert pseudo-quantization nodes before and after convolution, fully connected layers, and activation layers. Use symmetric per-channel quantization for weights and symmetric per-tensor quantization for activation layers.

[0151] 4.2 Calibration: Randomly sample 256 batches and statistically analyze the 99.9th percentile of the activation distribution, mapping it to the INT8 dynamic range. .

[0152] 4.3 Fine-tuning: Inherit floating-point weights from the quantized network, and then adjust the learning rate. Fine-tune for 10 epochs; keep the MAE (Maximum Evidence) error on the validation set to <1%.

[0153] 4.4 Export: Through serialization with the TensorRTINT8 engine, the inference latency was reduced from 48ms FP16 to 24ms INT8, the video memory usage was reduced from 62MB to 18MB, and the precision decreased by <0.6%.

[0154] By employing the above process, the model maintains prediction accuracy while meeting the real-time and low-power requirements of edge devices, providing a reliable core for the online inference-control process.

[0155] VI. Online Inference and Security Redundancy

[0156] 1. Inference Scheduling

[0157] The MCU sends a PREDICT trigger signal to the edge AI module at 1000ms intervals. The edge AI then reads the sliding window tensor. Complete INT8 inference within 25ms and output: continuous action vector. Confidence scalar The reasoning result is accompanied by a timestamp. Return to MCU via TSN-Ethernet.

[0158] 2. Confidence level and boundary checks

[0159] The MCU maintains the following thresholds:

[0160] Confidence threshold;

[0161] Compressor frequency hard limit;

[0162] Fan speed limit;

[0163] Valve opening limits.

[0164] MCU performs the following judgment:

[0165] ;

[0166] If all four conditions are met, then... Write directly to the target register and enter execution path A; otherwise, enter the rollback logic execution path B.

[0167] 3. Rollback Logic Judgment

[0168] During operation, three status variables are measured in real time: coil temperature. Temperature rise slope Average frost thickness .

[0169] Branch rollback rules:

[0170] Condition 1 (Level 1 frequency reduction)

[0171] ;

[0172] implement Keep the defrost mode active.

[0173] Condition 2 (Level 2 frequency reduction and exit)

[0174] ;

[0175] implement .

[0176] Scheduling order: First check condition 1; if it is not met, then check condition 2; if neither is met, then temporarily maintain the current frequency and re-evaluate in the next second.

[0177] 4. Target command issuance

[0178] When the MCU enters a 10ms control interrupt, the hardware is updated in the following way:

[0179] Compressor → SVPWM duty cycle mapping target frequency .

[0180] Fan → 0–10V or PWM duty cycle mapped speed .

[0181] Electrically controlled valve → driven to target opening degree by pulse counting .

[0182] Four-way valve → relay switches between heating / defrosting circuits.

[0183] 5. Safety counter

[0184] If the confidence level is below the threshold or the target exceeds the limit for three consecutive cycles (3s), the MCU automatically downgrades to a fixed safety table. The fault code is recorded in the log and awaits reset by maintenance personnel.

[0185] Through the above six steps, the system completes a closed-loop decision-making process of inference, verification, execution, or rollback within a one-second cycle. It fully utilizes the high-precision prediction of deep networks while ensuring that it can quickly degrade to the safety curve in any scenario with insufficient confidence or out-of-bounds actions, avoiding risks such as compressor overcurrent, liquid slugging, or insufficient defrosting.

[0186] VII. Distributed Reinforcement Learning Updates

[0187] 1. Actor Node Trajectory Acquisition

[0188] Each heat pump prototype loads the current policy network during operation. And it acts as an independent Actor. Within the MCU–AI internal loop, the Actor records moments. of (Temporal tensor input) (Perform the action) (Instant rewards) and (Probability of the old strategy).

[0189] A local trajectory of length 32 was continuously collected. Packets can then be pushed to the Learner node via gRPC without waiting for the entire batch to accumulate, enabling simultaneous sampling and uploading.

[0190] 2. Reward function calculation: For each defrosting cycle, the total reward is generated online in real time.

[0191] ;

[0192] in: Energy consumption for this round of defrosting; Thick frost remained upon exiting; Number of times the four-way valve is switched during the process. Default weight The Pareto rules will then automatically fine-tune the settings.

[0193] 3. Learner node experience storage

[0194] After receiving the trajectory, the Learner server splits it into single-step triples and writes them to the priority replay cache. The cache capacity is set to 10,000 entries; for the first... Time difference error in sample calculation:

[0195] ;

[0196] Will As a basis for priority.

[0197] The probability of a sample being selected:

[0198] .

[0199] 4. Proximity Optimization (PPO)

[0200] Learner initiates a policy update every 2048 accumulated samples.

[0201] 1) Dominance estimation uses generalized dominance estimation (GAE-λ):

[0202] ;

[0203] 2) Objective function

[0204] ;

[0205] 3-gradient update

[0206] Using the Adam optimizer, learning rate The value network and policy network are iterated for five rounds simultaneously. If the KL divergence of a single batch exceeds 0.02, the iteration is stopped early to prevent policy runaway.

[0207] 5. Weighted Asynchronous Broadcast

[0208] After the Learner completes the batch update, it will assign the new weights. Write the parameters to the server cache and record the version number. Each Actor compares the version number using a 1Hz polling method: if the server version is higher, download the incremental weight package and hot-swap the local network; if the download fails three times in a row, the Actor automatically downgrades to the fixed security policy table and stores an alarm in the log.

[0209] 6. Pareto weight self-adjustment

[0210] At midnight each day, Learner compiles data from all defrosting cycles of the previous day. , , Calculate the Pareto front using three-dimensional samples. If energy consumption improves while residual frost worsens by more than 5%, then... Increase by 0.05 and decrease by the same amount. If both improve, then... The adjustment was lowered by 0.02 to encourage more aggressive defrosting; the weighting remains at one for all three factors after the adjustment.

[0211] 7. Communication and Fault Tolerance Mechanisms

[0212] The gRPC stream channel sets the compressor serial number as metadata to help the Learner distinguish the source; if the Actor's heartbeat times out by 3 seconds, the Learner will pause pushing weights to it and mark it as offline; if the Learner crashes, the Actor can continue to run for up to 24 hours. When the accumulated trajectory exceeds the cache limit, the oldest data will be overwritten to avoid memory leaks.

[0213] The above seven steps form a complete closed loop: multi-machine parallel sampling → priority storage → batch PPO update → asynchronous distribution. This mechanism combines high sample efficiency with real-time deployability, and can continuously improve defrosting control strategies under rare operating conditions such as extreme low temperatures, high humidity, and sudden load changes.

[0214] 8. Adaptive Backtracking Curve

[0215] Two-stage down-frequency targets in the backoff table (Level 1) and (Level 2) It is essential to ensure thorough defrosting while avoiding power waste. Therefore, this invention reduces the energy consumption of the previous defrosting cycle. Compared with reference energy consumption For comparison, the target frequency is adjusted in real time using an exponentially weighted average method; at the same time, an energy consumption coefficient is enabled. The host computer can remotely adjust the power supply based on real-time electricity prices or carbon emission factors, enabling linkage with the building energy management system (EMS) or virtual power plant (VPP).

[0216] 1. Exponentially weighted update formula

[0217] 1) Energy consumption deviation coefficient

[0218] ;

[0219] 2) Primary frequency reduction target

[0220] ;

[0221] 3) Secondary frequency reduction target

[0222] ;

[0223] 2. Parameter Explanation

[0224]

[0225] like but The frequency automatically increases to accelerate defrosting; if but The frequency is automatically reduced to save energy.

[0226] 3. Overview of Execution Process

[0227] Calculate immediately after each defrosting cycle And store it in EEPROM. Refresh in real time using the above formula. and Write the data to the MCU rollback table; the next defrost cycle will then take effect. If there are three consecutive cycles of energy consumption deviation... If the value is less than 5%, it will automatically decrease. When the value reaches 0.1, it enters the steady-state oscillation suppression mode. Through this adaptive mechanism, the rollback curve can dynamically converge to the optimal balance point of energy consumption and performance while ensuring the integrity of defrosting, based on unit aging, climate change, and energy signals.

[0228] IX. Determination of Defrosting Exit

[0229] 1. Stability determination

[0230] 1.1 Error band

[0231] make To measure the compressor frequency, The current target frequency; , as well as , Similarly, define relative error:

[0232] ;

[0233] 1.2 Stability Threshold

[0234] System Settings (i.e., 3%). When all three relative errors simultaneously satisfy: Once this is achieved, the implementing agency is considered to be stable.

[0235] 1.3 Timer

[0236] Start the timer when the system reaches a stable state. If any error exceeds the threshold, then immediately... Reset to zero.

[0237] 2. Dual thresholds for thermal properties and frost thickness

[0238] set up Current temperature of the coil Defrosting exit temperature threshold Average frost thickness The residual frost thickness threshold and defrost release conditions are as follows:

[0239] ;

[0240] Typical values .

[0241] 3. Exit delay

[0242] System settings exit delay Only if steps 1 and 2 are simultaneously satisfied: Only then is defrosting considered truly complete. If any condition fails during this period, the timer resets to zero, and the monitoring process restarts.

[0243] 4. Action execution

[0244] 4.1 Valve switching

[0245] When the timer reaches 30 seconds, the four-way valve's solenoid coil is immediately triggered, switching the refrigerant flow from defrost mode back to heating mode. The valve's actuation time must be less than 0.8 seconds; the MCU must maintain the compressor frequency locked for 3 seconds after the switch. Once the heating is stable, it can be switched to normal control.

[0246] 4.2 Window Reset

[0247] Clear the depth reference frame buffer to prepare for updating the new reference 30 seconds after the next defrost cycle ends; set all 16 frames of data in the sliding window to zero to prevent residual frost thickness from triggering erroneously; reset the stabilization timer and frost thickness accumulation integral.

[0248] 5. Safe rollback

[0249] If the exit conditions are not met within 10 minutes of entering the defrost cycle, the controller will forcibly execute a second-level frequency reduction and shut off the valve to exit, recording the alarm code DEFROST_OVERTIME; the background analysis model will then improve the next cycle. and To prevent the frost layer from becoming too heavy and exceeding the time limit again.

[0250] 6. Key parameters are adjustable

[0251] The factory setting is 3%, which can be adjusted from 1 to 5% via the maintenance interface. Select 0–5℃ based on the refrigerant and model; 0.10–0.20 mm; 20–40 seconds; timeout protection limit 600 seconds.

[0252] By employing the above six points, the controller can ensure that the frost layer has been completely removed while preventing frequent jumps in the compressor, fan, and expansion valve due to command oscillations, thus ensuring a smooth, reliable defrosting exit with minimal energy consumption.

[0253] Test case

[0254] This study verifies whether the multimodal-deep prediction + distributed reinforcement learning + safety redundancy defrosting control method of this invention can reduce defrosting energy consumption, shorten defrosting downtime, and improve low-temperature heating capacity under extremely cold and high humidity conditions compared with the traditional ΔT-timed dual threshold algorithm (Baseline). Repeatable experimental data are provided.

[0255] I. Experimental Platform and Measurement Methods

[0256] 1. Prototype: Two identical 2HP low-temperature variable frequency air source heat pumps.

[0257] Unit #1 loads Baseline control logic;

[0258] Unit #2 is loaded with the algorithm of this invention (AI-Defrost).

[0259] Environmental chamber: -35℃~10℃, humidity adjustable from 30%RH to 95%RH, wind speed 2.5ms⁻¹.

[0260] 2. Operating Condition Settings:

[0261] Operating Condition A: -15℃ / 70%RH

[0262] Operating Condition B: -25℃ / 85%RH

[0263] Operating condition C: -30℃ / 50%RH.

[0264] 3. Measurement Indicators

[0265] Defrosting energy consumption E: Integral energy meter (accuracy class 0.2);

[0266] Defrosting duration t_def: The time it takes for the four-way valve to switch to exit.

[0267] Residual frost thickness L_res: measured by ToF camera, accuracy 0.1mm;

[0268] Heating capacity decay rate R_cap: Comparison of heating capacity stable for 10 minutes before and after defrosting;

[0269] Long-term failure rate: The number of times residual frost regenerates and requires secondary defrosting during continuous 360 hours of operation / total number of times.

[0270] Sampling frequency: 1Hz for all power levels, temperature, and frost thickness; 10Hz for control commands and operating status logs.

[0271] II. Experimental Procedure

[0272] Pretreatment: Both units were run at rated heat in the environmental chamber until all fins were clean.

[0273] Operating condition switching: Adjust the temperature and humidity to the target operating condition and wait for it to stabilize for 30 minutes.

[0274] Cyclic defrosting: The monitoring unit automatically enters and exits defrosting, recording a complete cycle of data.

[0275] Repeat: Record 10 defrosting cycles for each operating condition; defrost again to zero before changing operating conditions.

[0276] Long-term test: Continuous operation at -25℃ / 85%RH for 360 hours, failure rate was statistically analyzed.

[0277] III. Experimental Results

[0278]

[0279] Note: The average of 10 cycles is taken for each item; the energy consumption error is guaranteed to be <0.5% by power meter calibration, and the time error is <0.1s.

[0280] Long-term 360-hour results: Baseline required 28 defrost cycles to rebuild residual frost, with a failure rate of 3.9%. AI-Defrost required only 5 cycles, with a failure rate of 0.7%.

[0281] IV. Results Analysis

[0282] 1. Energy consumption reduction of 8%–17%: Convolutional-LSTM prediction enables compressors, fans and valves to reduce frequency and shut down as soon as the frost layer melts; at the same time, distributed PPO further fine-tunes for an additional 2%–4% energy saving.

[0283] 2. Defrosting downtime is reduced by 25%–35%: the system triggers shutdown 30–60 seconds earlier based on the dual criteria of frost thickness and temperature rise slope, significantly reducing heating interruptions.

[0284] 3. Improved low-temperature heating capacity: The heating attenuation rate in operating condition B decreased from 35% to 22%, and the weighted average HSPF increased by 7% throughout the season.

[0285] 4. Significantly improved reliability: The 0.7% residual frost regeneration rate is far lower than the industry standard of 3-5%, proving that the confidence level-backoff mechanism effectively avoids misjudgment.

[0286] The foregoing description of embodiments of the present invention, through which those skilled in the art are able to implement or use the present invention, will be readily apparent to those skilled in the art. Various modifications to these embodiments will be readily apparent to those skilled in the art. The general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novelty disclosed herein.

[0287] This application may be described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0288] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0289] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0290] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0291] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0292] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

Claims

1. A method for intelligent defrosting control of a low-temperature variable frequency air source heat pump, characterized in that, The method includes: S1 Data Fusion Acquisition: 1.1 Obtaining the temperature T of the defrosting coil of the finned heat exchanger def ; 2. Obtain the frost thickness characteristics L in real time using an ice thickness detection sensor or visual imaging device. ice 1.3 Collect the compressor's current frequency F, outdoor fan air volume Q, electronic expansion valve opening θ, and ambient wet-bulb temperature T. wb Relative humidity (RH); S2 Nonlinear Prediction Decision: Using a pre-trained convolutional-LSTM (C-LSTM) deep network model, the multimodal feature sequence {T} from step S1 is input. def ,L ice ,F,Q,θ,T wb ,RH}, output the optimal control vector {F} for the next control cycle. pred Q pred ,θ pred } and give the confidence level P. conf ; S3 Safety Redundancy Judgment: When P conf ≥P thr And when the output vector satisfies the hardware safety boundary, predictive control is executed, P thr The confidence threshold is used; otherwise, a fallback strategy is triggered: a) If T def ≥T defT -T set And dT def / dt≥K slope T defT T is the defrost exit temperature threshold. set To reduce the frequency temperature difference offset in advance, K slope If the temperature slope threshold is met, the compressor frequency will be reduced to the first-stage frequency reduction target frequency F according to the preset curve. set1 Simultaneously adjust the target air volume Q set With target opening θ set b) If T def ≥T defT or L ice ≤L thr L thr To achieve the minimum permissible residual frost thickness threshold, the compressor frequency is reduced to the secondary throttling target frequency F. set2 And switch the four-way valve back to heating; S4 Online Reinforcement Learning Update: Using defrosting energy consumption-time integral and remaining frost thickness as reward functions, a distributed reinforcement learning agent is used to fine-tune the parameters and backoff control table of C-LSTM online, with an update interval of no more than N defrosting cycles; S5 End Judgment: When the compressor, fan, and expansion valve are all running stably for Δt end And T def L ice When the exit threshold is met, the current defrosting control cycle ends.

2. The method according to claim 1, characterized in that, The convolutional LSTM network comprises: a) two cascaded two-dimensional convolutional layers, each with a kernel size of 3×3, the first layer having ≥32 output channels and the second layer having ≥64 output channels, with batch normalization and ReLU activation used between the two layers; b) a time step of T is set after the convolutional layers. seq The bidirectional LSTM layer has 128 hidden units and a dropout rate of 0.3 is added to the output side; c) Convolutional features and one-dimensional features such as temperature and humidity are concatenated by the channel dimension before LSTM.

3. The method according to claim 2, characterized in that, The convolutional-LSTM network further incorporates a multi-head self-attention fusion layer at the LSTM output, with h=4 heads, to assign different weights to the frost thickness sequence and the temperature sequence, thereby improving the response speed to sudden frost growth.

4. The method according to claim 1, characterized in that, The reinforcement learning unit adopts a distributed PPO architecture, including: a) at least two Actor nodes, deployed on multiple heat pump edge control boards, for parallel acquisition of interaction trajectories and execution of the current policy; b) one Learner node, deployed on the gateway server, responsible for calculating the global policy gradient and updating parameters; c) one parameter server, used for asynchronously synchronizing policy network weights between Actor nodes and Learner nodes, with an update frequency of not less than f. sync =1Hz; Preferably, the experience replay of the distributed PPO adopts a priority experience replay cache, the priority is weighted according to the time difference error |δ|, the cache size is ≥10000 records, and the temperature distribution and frost thickness distribution during random sampling must meet the principle of being the same as the actual working conditions. Further optimization, the priority empirical replay sampling probabilities are as follows: 。 5. The method according to claim 4, characterized in that, The reinforcement learning reward function R satisfies R=−α·E−β·L res −γ·N switch , Where E represents the energy consumption of this defrosting cycle, and L... res N represents the remaining frost thickness upon exiting. switch The number of four-way valve switching times, α, β, and γ are adaptively adjusted by the Learner node based on the energy consumption-performance Pareto front of the most recent M defrost cycles to achieve a dynamic balance between energy consumption and defrost integrity; Preferably, the Pareto adaptive weight update satisfies: 。 6. The method according to claim 1, characterized in that, F in the rollback strategy set1 With F set2 The energy consumption curve of the previous defrosting cycle is obtained by exponential weighted averaging to ensure the adaptability of the rollback control.

7. The method according to claim 1, characterized in that, The visual imaging device is a ToF depth camera or millimeter-wave radar with a resolution ≥0.1mm and is equipped with a thermal compensation algorithm to reduce condensation-fog interference.

8. An intelligent defrosting control system for a low-temperature variable frequency air source heat pump, characterized in that, include: A multimodal sensor array is used to acquire data such as fin temperature, frost thickness, and ambient temperature and humidity. The actuators include a variable frequency compressor, an outdoor fan, and an electronic expansion valve; The depth prediction unit, with a built-in convolutional LSTM network, is used to predict the optimal control vector based on sensor data. Reinforcement learning units are used to update the parameters of the depth prediction unit in real time based on energy consumption and remaining frost thickness; A safety redundancy control unit is used to execute a fallback strategy when the prediction confidence is insufficient or exceeds the limit; A central controller is used to implement all the steps of the method according to any one of claims 1-8 and to coordinate the aforementioned units.

9. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method described in any one of claims 1-5.

10. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method described in any one of claims 1-5.