Control system and method for free operation performance test laboratory of variable-frequency air conditioner
By employing a dual-layer closed-loop control architecture consisting of a dynamic load simulation module, a model predictive control module, and a deep reinforcement learning agent, the problems of low control precision, slow response speed, and insufficient automation in the performance testing laboratory for variable frequency air conditioners were solved. This enabled a high-precision, fast-response, and automated testing process, improving testing efficiency and result accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-07
AI Technical Summary
Existing variable frequency air conditioner performance testing laboratories suffer from problems such as low control precision, slow response speed, insufficient automation, and low testing efficiency in free operation mode, which cannot meet the requirements of high precision, fast response, and realistic simulation.
A two-layer closed-loop control architecture consisting of a dynamic load simulation module, a model predictive control module, and a deep reinforcement learning agent is adopted. The long short-term memory network predicts future temperature and humidity changes, and the deep reinforcement learning agent makes dynamic decisions on thermal and humidity load disturbances to achieve high-precision, fast-response, and automated control.
It significantly improves control accuracy and response speed, with temperature control accuracy improved to ±0.2℃, humidity control accuracy improved to ±1%RH, response time shortened to within 8 minutes, dynamic performance optimized, automation level improved, testing efficiency improved, and the system has self-optimization capabilities.
Smart Images

Figure CN121806497A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of air conditioner testing technology, and more specifically to a control system and method for a laboratory for testing the free-running performance of variable frequency air conditioners. Background Technology
[0002] With the continuous development of inverter air conditioning technology, its energy-saving and comfort advantages are becoming increasingly prominent. However, accurately evaluating the performance of inverter air conditioners in actual use environments has become a significant challenge for the industry. Performance testing of inverter air conditioners is a crucial step in evaluating their energy efficiency, comfort, and reliability. Free-running performance testing, as the testing method that best reflects actual operating conditions, places higher demands on the control precision and response speed of the testing environment.
[0003] Currently, there is considerable research on control strategies for variable frequency air conditioners. For example, CN113339941B discloses a control method for variable frequency air conditioners. This method obtains a load calculation function of indoor load as indoor temperature and humidity change, calculates the indoor load corresponding to each temperature and humidity combination in a temperature and humidity parameter library, and determines the set temperature and humidity combination based on a pre-established relationship model between indoor load and air conditioner operating status, achieving low energy consumption while meeting indoor comfort requirements. However, this method mainly focuses on the control optimization of the air conditioner itself and does not address the precise control of the test environment.
[0004] Regarding air conditioning system load forecasting, CN119642336B proposes a central air conditioning unit energy-saving optimization system based on load forecasting. It introduces a VMD-Transformer-GRU hybrid deep learning model, improves cooling load forecasting accuracy through multi-source data fusion, and employs multi-agent deep reinforcement learning (MADRL) to achieve dynamic coordination and joint optimization among devices. While this technology offers innovations in load forecasting and system collaboration, it does not specifically address the high-precision temperature and humidity control requirements of a test laboratory environment.
[0005] For intelligent control of air conditioning systems, CN120702068A discloses a machine learning-based control strategy optimization method for a novel air conditioning system. This method establishes a state space, action space, and reward function suitable for the air conditioning system and optimizes the control strategy using deep reinforcement learning. CN120160246A proposes an intelligent energy-saving control method and system for variable frequency air conditioning. By introducing reinforcement learning algorithms, it leverages the adaptive capabilities of the agent in dynamic environments and continuously optimizes the control strategy based on real-time feedback. While these methods apply advanced machine learning techniques, they primarily focus on the control optimization of the air conditioning system itself, rather than the precise simulation and control of the test environment.
[0006] Regarding load anomaly monitoring, CN119934656B proposes a method and system for monitoring and intelligently controlling abnormal heat and humidity loads of room air conditioning systems. It utilizes an improved lightweight Informer model trained in the cloud, combined with transfer learning to achieve rapid model transfer, enabling real-time identification of abnormal operating conditions such as open windows, large gatherings of people, and high heat sources. While this technology innovates in load anomaly identification, it does not address how to actively generate simulated loads to test air conditioning performance.
[0007] However, existing variable frequency air conditioner performance testing laboratories have the following technical problems in free operation mode.
[0008] First, traditional testing laboratories generally use PID control systems, which are passive response systems and cannot predict future temperature and humidity changes. This results in large overshoots and long settling times during testing, typically exceeding 30 minutes to reach a stable state. This lag severely impacts testing efficiency and result accuracy.
[0009] Secondly, the free-running test of variable frequency air conditioners needs to simulate dynamic heat and humidity load changes in a real environment, while traditional test systems struggle to handle multi-variable coupling and nonlinear dynamic loads such as temperature and humidity. Especially under sudden load changes, the control system's response is often untimely, failing to maintain the stability of the test environment and resulting in significant deviations in the test results.
[0010] Third, traditional laboratories rely heavily on manual intervention for load simulation and operating condition switching, resulting in low levels of automation, long testing cycles, and high costs. Operators need to manually adjust the load equipment based on experience, which is not only inefficient but also makes it difficult to ensure the consistency and repeatability of the tests.
[0011] Fourth, existing technologies lack intelligent testing methods for the free-running characteristics of variable frequency air conditioners, failing to meet the demands for high precision, fast response, and realistic simulation. In particular, existing testing systems lack adaptability and flexibility in simulating the impact of user behavior and environmental changes on air conditioner performance.
[0012] In summary, there is an urgent need for a control system and method for a variable frequency air conditioner free-running performance testing laboratory that can accurately control the testing environment, quickly respond to load changes, and automatically simulate real working conditions, in order to improve testing efficiency and result accuracy. Summary of the Invention
[0013] To address the problems of low control precision, slow response speed, insufficient automation, and low testing efficiency in traditional variable frequency air conditioner performance testing laboratories under free operation mode, and to achieve significantly improved control precision, greatly enhanced response speed, optimized dynamic performance, increased automation, improved testing efficiency, and enhanced system self-optimization capabilities, this invention provides a control system and method for a variable frequency air conditioner free operation performance testing laboratory.
[0014] According to one aspect of the present invention, a control system for a test laboratory for the free-running performance of a variable frequency air conditioner is provided, comprising: a dynamic load simulation module for generating thermal and humidity load disturbances simulating the free-running operation of the variable frequency air conditioner; a model predictive control module deployed at an edge node, which predicts the trajectory of laboratory temperature and humidity changes within a predetermined time domain based on a long short-term memory network, and generates control commands through rolling optimization to compensate for the thermal and humidity load disturbances; and a deep reinforcement learning agent for dynamically deciding the intensity and timing of the thermal and humidity load disturbances based on the operating state of the air conditioner under test, so that the load disturbances simulate real free-running conditions; wherein the deep reinforcement learning agent and the model predictive control module constitute a two-layer closed-loop control architecture: the deep reinforcement learning agent acts as an outer-loop disturbance generator, and the model predictive control module acts as an inner-loop compensation controller, and the two work together to achieve stable control of the laboratory environment operating parameters under dynamic load disturbances.
[0015] Optionally, the dynamic load simulation module includes: a heat load simulation unit for regulating the laboratory ambient temperature using a resistance heater group; a humidity load simulation unit for regulating the laboratory ambient humidity using an ultrasonic humidifier and a condensation dehumidification device; and an air circulation unit for ensuring uniform mixing of air inside the laboratory using a variable frequency fan.
[0016] Optionally, the heat load simulation unit implements PWM regulation through a solid-state relay or power regulator; the steam diffusion rate of the wet load simulation unit is controlled by a variable frequency fan in a closed loop.
[0017] Optionally, the model prediction control module includes: an LSTM prediction unit, used to input multi-dimensional time-series data including historical laboratory temperature, historical laboratory humidity, historical inverter air conditioner power consumption and load equipment status, and output the laboratory temperature and humidity change trajectory within a predetermined future time domain; and an MPC optimization unit, used to solve for the optimal control sequence based on the laboratory temperature and humidity change trajectory within the predetermined future time domain, so that the laboratory temperature and humidity follow the set trajectory and the control sequence satisfies the constraints of the laboratory environmental control actuator.
[0018] Optionally, the LSTM prediction unit includes: a forgetting gate structure for deciding the degree of retention of historical information through a sigmoid function; an input gate structure for controlling the proportion of new candidate cell states included; a cell state update mechanism for updating long-term memory through a weighted combination of the forgetting gate structure and the input gate structure; and an output gate structure for generating hidden states through a tanh activation function to output predicted values, wherein the predicted values are the trajectory of laboratory temperature and humidity changes within a predetermined time domain in the future, and the predicted values are directly used as the input of the MPC optimization unit.
[0019] Optionally, the deep reinforcement learning agent includes: a state perception unit for real-time acquisition of the operating frequency, real-time power, and set operating parameters of the air conditioner under test; and an action decision unit for outputting continuous control commands for changes in heat load power and changes in humidity load.
[0020] Optionally, the deep reinforcement learning agent is trained using a deep deterministic policy gradient algorithm, and the reward function of the deep reinforcement learning agent includes: a temperature and humidity tracking error term, a system energy consumption term, and a penalty term for the frequency change rate of the tested air conditioner.
[0021] Optionally, it also includes a cloud-based collaborative platform for receiving environmental operating parameter data from multiple laboratories, training and optimizing the long short-term memory network through federated learning, generating incremental model update packages, encrypting them, pushing them to edge nodes, and triggering hot updates after verifying model stability at the edge nodes.
[0022] Optionally, it also includes a multi-layered sensor network deployed in the laboratory, including temperature sensors, humidity sensors, pressure sensors, wind speed sensors, and data acquisition systems.
[0023] According to a second aspect of the present invention, a control method for a test laboratory for the free-running performance of a variable frequency air conditioner is provided, comprising: generating a thermal and humidity load disturbance simulating the free-running operation of the variable frequency air conditioner through a dynamic load simulation module; predicting the trajectory of laboratory temperature and humidity changes within a predetermined time domain based on a long short-term memory network through a model predictive control module deployed at edge nodes, and generating control commands through rolling optimization to compensate for the thermal and humidity load disturbance; and dynamically deciding the intensity and timing of the thermal and humidity load disturbance based on the operating state of the air conditioner under test through a deep reinforcement learning agent, so that the load disturbance simulates the real free-running conditions; wherein the deep reinforcement learning agent and the model predictive control module constitute a two-layer closed-loop control architecture: the deep reinforcement learning agent acts as an outer loop disturbance generator, and the model predictive control module acts as an inner loop compensation controller, and the two work together to achieve stable control of the laboratory environment operating parameters under dynamic load disturbance.
[0024] The control system and method for the free-running performance testing laboratory of variable frequency air conditioners provided by this invention have the following beneficial effects.
[0025] 1. Significantly improved control accuracy: Temperature control accuracy is improved from the traditional ±1.0℃ to ±0.2℃, and humidity control accuracy is improved from ±3%RH to ±1%RH. Compared with traditional PID control systems, the LSTM-MPC model predictive control module of this invention can predict the future behavior of the system and optimize the control sequence, thereby achieving higher precision temperature and humidity control.
[0026] 2. Significantly improved response speed: Temperature adjustment time has been reduced from over 30 minutes to less than 8 minutes, and humidity adjustment time has been reduced from over 30 minutes to less than 10 minutes. This is mainly due to the LSTM long short-term memory network's ability to anticipate system changes and generate optimal control sequences through the MPC optimization unit, thus significantly shortening the system response time.
[0027] 3. Dynamic performance optimization: The temperature step response overshoot is reduced from 1.0℃ to <0.2℃, and the humidity step response overshoot is reduced from 3%RH to below 1.5%RH. This invention effectively suppresses system overshoot and improves dynamic performance through a two-layer closed-loop control architecture consisting of a deep reinforcement learning agent and a model predictive control module based on LSTM-MPC.
[0028] 4. Increased Automation: Achieves automatic decoupling of multiple variables without manual intervention; intelligent load simulation adaptively handles dynamic testing requirements. This invention integrates the decoupling function into the MPC optimization objective and uses a deep deterministic policy gradient algorithm to train a deep reinforcement learning agent to simulate dynamic thermal and humidity loads, thus realizing a highly automated testing process.
[0029] 5. Improved testing efficiency: The rise time (temperature) is shortened from 10-15 minutes to less than 5 minutes, and the stabilization time is significantly reduced, improving testing efficiency and reducing energy consumption and time costs. The collaborative work of the dynamic load simulation module, model prediction control module, and deep reinforcement learning agent in this invention significantly improves testing efficiency.
[0030] 6. System Self-Optimization Capability: Through an edge-cloud collaborative architecture, deep data mining and system self-optimization are achieved, enabling continuous evolution of the control system's capabilities. The edge nodes of this invention are responsible for real-time data processing and control, while the cloud platform is responsible for complex model training and optimization. Their collaborative work enables continuous evolution of model capabilities, giving the system self-optimization capabilities. Attached Figure Description
[0031] The accompanying drawings, which are provided to further illustrate this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application.
[0032] Figure 1 This is a system framework diagram of the present invention.
[0033] Figure 2 This is the overall flowchart of the present invention.
[0034] Figure 3 This is a schematic diagram of an electronic device. Detailed Implementation
[0035] The embodiments of this application will now be described in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application. Furthermore, the following embodiments and features can be combined with each other unless otherwise specified. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0036] Example 1
[0037] In one embodiment of the present invention, please refer to Figure 1 This paper presents a control system for a performance testing laboratory of variable frequency air conditioners under dynamic load conditions. The system utilizes multiple functional modules working collaboratively to perform performance testing on variable frequency air conditioners. The system includes a dynamic load simulation module, a model predictive control module, and a deep reinforcement learning agent.
[0038] The dynamic load simulation module consists of a heat load simulation unit, a humidity load simulation unit, and an air circulation unit. The heat load simulation unit uses a resistance heater group to achieve adjustable power from 0 to 20 kW, with a power adjustment accuracy of ±25 W. The resistance heaters are divided into multiple independently controlled groups, with PWM regulation achieved through solid-state relays or power regulators. The system response time is less than 0.5 seconds, enabling rapid adjustment of the heat load level in the laboratory. The humidity load simulation unit uses a combination design of an ultrasonic humidifier and a condensation dehumidifier. The humidification rate can be adjusted from 0 to 5 kg / h, and the dehumidification rate from 0 to 3 kg / h. The vapor diffusion rate of the humidity load simulation unit is controlled by a variable frequency fan in a closed loop, ensuring the uniformity and controllability of humidity changes. The air circulation unit uses a variable frequency fan to achieve uniform mixing of airflow velocity from 0.1 to 20 m / s within the laboratory, ensuring consistent temperature and humidity distribution. It should be noted that the dynamic load simulation module is uniformly managed by a modular PLC controller, equipped with analog input / output modules, and the data acquisition cycle can be set from 1 to 10 seconds, meeting the data acquisition requirements for high-precision control.
[0039] The control system also includes a multi-layered sensor network deployed in the laboratory, comprising: temperature sensors, humidity sensors, pressure sensors, wind speed sensors, and a data acquisition system. The temperature sensor uses a PT1000 platinum resistance temperature sensor with a measurement accuracy of ±0.05℃. The humidity sensor uses a dew point meter for humidity measurement with an accuracy of ±0.5%RH, including but not limited to one of the following: thin-film capacitive, resistive, and cold mirror type dew point meters. The pressure sensor uses a micro differential pressure transmitter with a measurement accuracy of ±0.25%FS. The wind speed sensor uses a high-precision hot-wire anemometer with a measurement range of 0.1–20 m / s and a measurement accuracy of ±0.05 m / s. The data acquisition system has a sampling frequency of no less than 1 Hz, and measures parameters including: air temperature (accuracy ±0.05℃), relative humidity (accuracy ±0.5%), air pressure (accuracy ±1 Pa), wind speed (accuracy ±0.05 m / s), and electrical parameters (power, current, voltage). In addition, to ensure that the indoor-outdoor heat exchange coefficient is less than 0.4 W / (m²·K) and to minimize the interference of the external environment on the testing process, the laboratory enclosure structure must be made of high-strength insulation materials (such as polyurethane composite panels).
[0040] In a preferred embodiment, the laboratory involved in this application adopts a distributed data acquisition system with an architecture divided into three layers, specifically: (1) Equipment layer: various sensors such as PT1000 temperature sensor, dew point sensor, micro differential pressure transmitter, hot wire anemometer, and power acquisition module (for power, current, voltage), as well as actuators such as resistive heater group, solid-state relay, power regulator, ultrasonic humidifier, condensation dehumidification device, and variable frequency fan. (2) Control layer: PLC controller and edge computing gateway. (3) Management layer: upper-level industrial control computer running machine learning algorithms.
[0041] In a preferred embodiment, the system uses the OPC UA standard for communication to ensure data interoperability between different devices. The industrial control computer is equipped with a database storage system (SQL Server) to record all sensor data and actuator status. The sampling period can be set to 1 second (fastest) or customized each time as needed, providing sufficient historical data support for machine learning models.
[0042] The model prediction and control module is deployed on edge nodes and includes an LSTM prediction unit and an MPC optimization unit. The LSTM prediction unit takes into account multi-dimensional time-series data of historical temperature, humidity, air conditioning power consumption, and load equipment status, processes it using a deep learning algorithm, and outputs the trajectory of laboratory temperature and humidity changes within a predetermined future time domain. This prediction unit contains a complete long short-term memory network structure, specifically including a forgetting gate structure, an input gate structure, a cell state update mechanism, and an output gate structure. The forgetting gate structure determines the degree of retention of historical information through a sigmoid function; the input gate structure controls the proportion of new candidate cell states included; the cell state update mechanism updates long-term memory through a weighted combination of the forgetting gate and the input gate; and the output gate structure generates hidden states using a tanh activation function to output the predicted value, which is the trajectory of laboratory temperature and humidity changes within the predetermined future time domain and directly serves as the input to the MPC optimization unit. Based on the laboratory temperature and humidity change trajectory within the predetermined future time domain, the MPC optimization unit solves for the optimal control sequence to ensure that the laboratory temperature and humidity follow the set trajectory, and that the control sequence satisfies the constraints of the laboratory environmental control actuator. The MPC optimization unit employs multiple constraints, including but not limited to a heater power limit of no more than 20kW, a humidifier steam diffusion rate of no more than 5kg / h, and a dehumidifier condensation power of no more than 3kW, to ensure that the system operates within safe limits.
[0043] In a preferred embodiment, this application employs an LSTM-MPC-based model predictive control module as the core strategy, replacing traditional PID control, to achieve high-precision, fast-response dynamic control of laboratory temperature and humidity. In this invention, the internal model of this control strategy is a long short-term memory network, and its workflow includes the following steps.
[0044] (1) Prediction: In each control cycle t, the system collects current and historical status information (such as temperature, humidity, air conditioning power consumption, load equipment status, etc.) as input. The built-in LSTM prediction unit predicts the future based on this information. Each step size ( To predict the trajectory of laboratory temperature and humidity changes within the time domain. .
[0045] (2) Optimization: The MPC optimization unit solves for a future Each step size ( To control the time domain, ) optimal control sequence This ensures that the predicted output closely approximates the desired setpoint trajectory (reference trajectory) while satisfying the constraints of the actuator (such as power limits). The mathematical description of this optimization problem is given by Equations 1 and 2, where Equation 1 is: Formula 2 satisfies: ;in, For the prediction time domain (number of future prediction steps); To control the time domain (number of future prediction steps); This is the system's predicted output (e.g., temperature and humidity) at time t+i. For reference trajectory (set value); To control the increment (such as power change); Q is the output weight matrix (emphasizing tracking accuracy); R is the control matrix (emphasizing control smoothness).
[0046] (3) Rolling implementation: The first control variable in the optimized control sequence Apply the controllable object (e.g., adjust the heater power). Repeat the above steps in the next sampling period, and re-predict and optimize based on the new measurements to achieve "rolling optimization".
[0047] In a preferred embodiment, the LSTM prediction unit uses an LSTM network, which has excellent time-series data processing capabilities. Its structure and formula are as follows: (1) Forget gate structure, which determines which information to discard from the previous cell state. The mathematical description of the forget gate structure is shown in Formula 3, which is: ,in, Output the forget gate structure (values range from 0 to 1, where 1 means to retain all values and 0 means to retain none). This is the hidden state from the previous moment; For current input features (such as historical temperature, humidity, load equipment status, etc.); This is the weight matrix; For bias term. (2) Input gate structure, which determines which new information will be stored in the cell state. The specific mathematical description is shown in Formula 4, which is: ,in, The input gate structure output controls the degree to which new information is retained. , Both are weight matrices; , All are bias terms; Candidate cell states are defined in the range [-1, 1]. (3) Cell state updates, combining forgetting gate and input gate structures, to update long-term memory. New cell states The old state is preserved by the forget gate. New candidate states included in input gate control This is a joint decision, and the specific mathematical description is given in Formula 5, which is: ,in, This reflects the current cellular state and carries long-term dependence. This represents the cell state at the previous moment; (4) Output gate structure: Based on the current loading, a hidden state is generated. The specific mathematical description is given in Formula 6, which is: ,in, It is an output gate structure that controls the degree to which the current state is exposed; This is the weight matrix; For bias terms; The current hidden state is used as the prediction input. (5) The final prediction layer will... The load value is output through the fully connected layer, as detailed in Formula 7. Formula 7 is as follows: ,in, The predicted value at time t; This is the weight matrix; This is the bias term. The LSTM prediction unit takes into account multi-dimensional time-series data of historical laboratory temperature, historical laboratory humidity, historical inverter air conditioner power consumption, and load equipment status, with a sampling interval of 1 second, covering multivariate time-series data during dynamic testing, and outputs predicted temperature and humidity values for the next 5–10 minutes. This predicted value is directly used as the input to the MPC optimization unit.
[0048] In a preferred embodiment, the training data for the LSTM network comes from a historical laboratory test database, specifically including: ① Data composition: collected from devices such as a PT1000 temperature sensor (accuracy ±0.05℃), dew point sensor (accuracy ±0.5%RH), and power acquisition module, covering multi-dimensional time-series data such as temperature, humidity, air conditioning power consumption, and load equipment status. ② Data period: using test data accumulated over the past 12 months, covering standard operating conditions (such as GB / T7725-2022) and dynamic test operating conditions, with a total data volume of approximately 100,000 sequences. ③ Preprocessing process: the data is first imputed for missing values and filtered for outliers (based on the 3σ criterion), and then normalized to the [-1, 1] interval using Z-score to eliminate the influence of dimensions. ④ Sequence partitioning: the input sequence length is 1800 (corresponding to 30 minutes of historical data, with a sampling interval of 1 second), and the output sequence length is 300-600 (corresponding to 5-10 minutes of predicted values).
[0049] In a preferred embodiment, after the LSTM network is trained, temporal cross-validation is used to evaluate its performance. The core metrics include: calculating the root mean square error (RMSE) according to Formula 8, which is: , where n is the number of samples in the validation set; The true value at time t; Let be the predicted value at time t. RMSE measures the absolute deviation between the predicted and actual values, reflecting the accuracy of the model's prediction. Temperature prediction RMSE ≤ 0.05℃, humidity prediction RMSE ≤ 0.8%RH, ensuring prediction accuracy is better than the control targets (±0.2℃, ±1%RH). The Mean Absolute Percentage Error (MAPE) is calculated according to Formula 9, which is: MAPE is used to assess the relative magnitude of prediction error, avoiding the influence of absolute dimensions, and is suitable for comparisons of multiple variables (such as temperature and humidity). Temperature MAPE < 3%, humidity MAPE < 5%, reflects the relative error level. The coefficient of determination is calculated according to formula 10. Formula 10 is: ,in, The average of the true values is used as the baseline model. Quantifiable improvements in a model's predictive power relative to a simple benchmark (such as the mean) measure the proportion of data variability explained by the model. A value >0.95 indicates a high degree of agreement between the predicted trajectory and the actual value. Only after the LSTM network passes verification is it deployed to the MPC controller for use.
[0050] A deep reinforcement learning agent dynamically determines the intensity and timing of heat and humidity load disturbances based on the operating status of the air conditioner under test, enabling the load disturbances to simulate real-world free-running conditions. This deep reinforcement learning agent comprises a state perception unit and an action decision unit. The state perception unit collects real-time data on the operating frequency, real-time power, and set operating parameters of the air conditioner under test, providing a data foundation for intelligent decision-making. The action decision unit outputs continuous control commands for changes in heat load power and humidity load, achieving precise control of load disturbances. The deep reinforcement learning agent is trained using a deep deterministic policy gradient algorithm, and its reward function includes: a temperature and humidity tracking error term, a system energy consumption term, and a penalty term for the frequency change rate of the air conditioner under test. This reward function comprehensively considers temperature and humidity control accuracy and energy consumption, guiding the deep reinforcement learning agent to learn the optimal disturbance strategy.
[0051] In a preferred embodiment, this application employs a multivariable decoupling and intelligent load integration strategy, specifically including the following: (1) Integrated dynamic decoupling: Due to the strong coupling relationship between temperature and humidity, this invention directly integrates the decoupling function into the optimization objective of the MPC. The MPC controller, through its multiple inputs and multiple outputs (MIMO) characteristics, considers both temperature and control variables during optimization. (2) Virtual load intelligent decision-making based on deep reinforcement learning: In order to simulate dynamic heat and humidity load, a deep deterministic policy gradient (DDPG) algorithm is used to train an agent (i.e., a deep reinforcement learning agent). The deep reinforcement learning agent works in conjunction with the LSTM-MPC controller: ① State space: the current temperature and humidity of the laboratory, the operating frequency of the air conditioner under test, power, set operating conditions, etc. ② Action space: the input of heat load and humidity load. ③ The reward function is: Where R is the reward value, with the goal of maximizing the reward; and T is the actual temperature in the laboratory. H represents the set temperature value; H represents the actual humidity in the laboratory. To set the humidity value; These are weighting coefficients used to balance error and energy consumption. The deep reinforcement learning agent decides the load input based on the state; this action is considered an external disturbance to the system. The LSTM-MPC controller can quickly sense this disturbance and compensate by predictively optimizing the control actuators (heaters, humidifiers, etc.) to maintain stable operating conditions. Load decision-making and temperature and humidity control constitute a higher-level intelligent closed loop.
[0052] A deep reinforcement learning agent and a model predictive control module constitute a two-layer closed-loop control architecture. The deep reinforcement learning agent acts as the outer-loop disturbance generator, while the model predictive control module acts as the inner-loop compensation controller. The outer-loop disturbance generation cycle of this two-layer closed-loop control architecture is 1–5 minutes, and the inner-loop compensation control cycle is 1–10 seconds. Synchronization between the two is achieved via the OPC UAover MQTT protocol. This two-layer closed-loop design enables the system to maintain high-precision and stable control of laboratory temperature and humidity while simulating real-world load disturbances.
[0053] The system also includes a cloud-based collaboration platform to enable edge-cloud collaboration, facilitating continuous model optimization and updates. Edge nodes deploy lightweight long short-term memory networks with no more than 64 hidden layer units and a sampling frequency of 1Hz to meet real-time control requirements. The cloud-based collaboration platform (i.e., the cloud server) aggregates environmental operating parameter data from multiple laboratories through federated learning, generates incremental model update packages, and pushes them to the edge nodes after encryption. Edge nodes employ a parallel verification mechanism, automatically rolling back to the old version when the new model's prediction error exceeds 5%, ensuring stable and reliable system operation.
[0054] In a preferred embodiment, edge nodes are deployed locally in the enthalpy difference laboratory, using industrial control computers or embedded AI computing boxes (such as the NVIDIA Jetson series) with sufficient computing power. These nodes are responsible for real-time data acquisition, preprocessing, lightweight model inference, and real-time closed-loop control. They utilize lightweight models for real-time load prediction and make control decisions accordingly, ensuring low-latency response during testing. A lightweight LSTM network, pre-trained and pruned in the cloud, is deployed to the edge nodes. This model has a simpler structure (e.g., fewer hidden layer units), sacrificing a small amount of accuracy for extremely fast inference speed.
[0055] The cloud platform is deployed in a remote data center as a high-performance GPU server cluster, responsible for massive data storage, complex model training, global model optimization, and long-term data analysis. Edge nodes encrypt data before uploading it to the cloud. Edge-cloud communication is implemented based on the OPC UA over MQTT protocol, with specific configurations including: 1) Transport layer security: using a TLS 1.3 encrypted channel and two-way certificate authentication (edge nodes and the cloud server exchange X.509 certificates) to prevent man-in-the-middle attacks. 2) Data encapsulation: sensor data and control commands are encapsulated in JSON format, including timestamps, device IDs, and data signatures (SHA-256 algorithm) to ensure integrity. 3) Real-time guarantee: the MQTT topic subscription / publish mode supports millisecond-level latency to meet control loop requirements (sampling frequency 1Hz).
[0056] The cloud leverages aggregated data from multiple devices over extended periods to train more complex and accurate models. These cloud-optimized lightweight models are then securely and efficiently deployed to the edge, enabling continuous evolution of model capabilities. The cloud-to-edge model update process follows these steps: 1) Triggering conditions: Periodic updates (every 24 hours) or event triggers (when the edge control error exceeds a threshold for 10 consecutive minutes). 2) Cloud optimization: The cloud GPU cluster uses federated learning to aggregate data from multiple laboratories, training a lightweight LSTM model (hidden layer units pruned from 128 to 64), generating incremental update packages (.onnx format). 3) Secure transmission: The update package is encrypted with AES-256 and pushed to the edge gateway via HTTPS, with MD5 hash values verified before and after transmission. 4) Edge deployment: Hot model updates at edge nodes: First, the old model is backed up, then the new model is loaded and run in parallel for verification (comparing 5-minute prediction results). If the error is less than 5%, the control loop is switched, with the entire process taking less than 30 seconds. 5) Rollback mechanism: If the new model causes control instability (e.g., overshoot > 0.3℃), it will automatically roll back to the old version and issue an alarm.
[0057] The control accuracy and performance analysis of the system includes: 1) The laboratory temperature and humidity control accuracy can reach the following indicators, as shown in Table 1. 2) Time delay optimization strategy: In order to reduce the time delay from issuing the command to the adjustment, the system adopts the following strategies: (1) Predictive control: Based on multi-step prediction of LSTM network, the control quantity is calculated in advance to compensate for the system inertial delay; (2) Feedforward compensation: The position of the actuator is adjusted in advance according to the load change trend; (3) Adaptive sampling period: High frequency sampling (1Hz) is used in the transient process of the system, and low frequency sampling (0.1Hz) is used in the steady state process; (4) Actuator optimization: High-response electric regulating valve and variable frequency fan are used to shorten the mechanical response time. Through these measures, the system response time can be shortened from 10-15min of traditional control to 3-5min, and the adjustment time (from the change of the set value to the new steady state) is reduced from more than 30min to 10-15min. 3) Dynamic tracking performance analysis: The dynamic tracking performance of the laboratory control system to the change of set value is evaluated by rise time, overshoot and settling time. Tests show that after introducing machine learning algorithms: (1) Temperature following: When the setpoint changes by 2℃, the rise time is ≤5min, the overshoot is <0.2℃, and the settling time is ≤8min; (2) Humidity following: When the setpoint changes by 10%RH, the rise time is ≤6min, the overshoot is <1.5%RH, and the settling time is ≤10min. This excellent following performance ensures that the operating conditions can be quickly adjusted during the test, meeting the requirements for dynamic load changes in non-frequency locking tests.
[0058] Table 1. Achievable Precision Targets for Laboratory Temperature and Humidity Control
[0059]
[0060] To visually demonstrate the advantages of this invention, this application also provides comparative data on the effects in two typical application scenarios. The comparison is between a traditional PID control laboratory and the LSTM-MPC-based control system of this invention.
[0061] Scenario 1: Step Change Test of Setpoint
[0062] Scenario Description: In a simulated inverter air conditioner test, the laboratory temperature and humidity setpoints are suddenly changed (e.g., temperature jumps from 25°C to 27°C, humidity jumps from 50%RH to 60%RH) to evaluate the system's dynamic tracking performance. This is a common scenario in free-running tests, where the air conditioner frequency may switch accordingly. See Table 2 for detailed performance comparison data.
[0063] Table 2 Comparison data of scene 1
[0064]
[0065] In step temperature changes, traditional PID systems suffer from delayed response, leading to temperature overshoot of up to 1°C and distorting air conditioner performance assessments. This invention, however, utilizes LSTM prediction and early compensation to reduce overshoot to <0.3°C and shorten the settling time to within 8 minutes, ensuring the accuracy of test data. For example, tests show that this invention reaches a new steady state within 5 minutes when the temperature setpoint changes by 2°C, while traditional systems require more than 30 minutes, significantly improving testing efficiency.
[0066] Scenario 2: Dynamic Load Simulation Test
[0067] Scenario Description: This scenario simulates the random changes in heat and humidity load faced by an air conditioner in a real-world environment (such as load fluctuations caused by indoor human activity). Traditional laboratory methods rely on fixed load programs, while this invention uses a DDPG agent to adaptively decide on load input. Detailed performance comparison data is shown in Table 3.
[0068] Table 3 Comparison data of the effects in Scene 2
[0069]
[0070] When simulating dynamic frequency switching of air conditioners, traditional systems require manual parameter reconfiguration, leading to test interruptions. This invention uses a DDPG agent to adjust the load (such as heater power) in real time and LSTM-MPC for rapid compensation, maintaining stable operating conditions. For example, when the air conditioner frequency switches from low to high, the load suddenly increases. This invention stabilizes temperature and humidity within 3-5 minutes, while traditional systems may overshoot and require manual intervention, extending test time. Furthermore, based on an edge-cloud architecture, historical data is used for model training, enabling the system to adapt to testing new types of air conditioners and improving repeatability.
[0071] Through the comparison of the effects in the above typical application scenarios, it can be seen that the present invention is significantly superior to traditional technologies in terms of control accuracy, response speed, automation level and testing efficiency.
[0072] In summary, this control system, through deep learning and reinforcement learning techniques combined with precise load simulation equipment, can accurately simulate the free-running conditions of inverter air conditioners in real-world environments. This provides a reliable experimental environment for performance testing of inverter air conditioners, significantly improving the accuracy and repeatability of test results. Furthermore, it employs an edge-cloud collaborative architecture to achieve in-depth data mining and system self-optimization.
[0073] Example 2
[0074] This embodiment, based on Embodiment 1 above, provides a control method for a variable frequency air conditioner free-running performance testing laboratory. Please refer to [link to previous document]. Figure 2 This method, applied to the control system of the variable frequency air conditioner free-running performance testing laboratory in Example 1, uses intelligent algorithms to simulate load fluctuations under real-world usage conditions, thereby achieving accurate performance evaluation of the variable frequency air conditioner. The specific implementation steps are as follows.
[0075] S1. Generate the thermal and humidity load disturbance of the inverter air conditioner during free operation through the dynamic load simulation module.
[0076] S2. The model prediction control module deployed at the edge node predicts the trajectory of laboratory temperature and humidity changes in the future within a predetermined time domain based on the long short-term memory network, and generates control commands through rolling optimization to compensate for the heat and humidity load disturbance.
[0077] S3. Through deep reinforcement learning, the intelligent agent dynamically decides the intensity and timing of heat and humidity load disturbances based on the operating status of the air conditioner under test, so that the load disturbances simulate real free operation conditions.
[0078] The deep reinforcement learning agent and the model prediction control module form a two-layer closed-loop control architecture: the deep reinforcement learning agent acts as an outer-loop disturbance generator, and the model prediction control module acts as an inner-loop compensation controller. The two work together to achieve stable control of the laboratory environment operating parameters under dynamic load disturbances.
[0079] Example 3
[0080] Based on Embodiment 2 described above, this embodiment also provides an electronic device, please refer to the appendix. Figure 3 , Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0081] like Figure 3As shown, an electronic device may include a processing unit (such as a central processing unit, graphics processing unit, etc.) that can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) or a program loaded from a storage device into random access memory (RAM). The RAM also stores various programs and data required for the operation of the electronic device. The processing unit, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0082] Typically, the following devices can be connected to an I / O interface: input devices such as touchscreens, touchpads, keyboards, mice, and cameras; output devices such as liquid crystal displays (LCDs) and speakers; storage devices such as magnetic tapes and hard drives; and communication devices. Communication devices allow electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.
[0083] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processing device, it performs the functions defined above in the methods of some embodiments of this disclosure.
[0084] Example 4
[0085] Based on Embodiment 2 above, this embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above method.
[0086] It should be noted that, in some embodiments of this disclosure, the computer-readable medium described above may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), or any suitable combination thereof.
[0087] In some embodiments, the client and server may communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and may interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0088] The aforementioned computer-readable medium may be included in the aforementioned device or may exist independently without being assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: generate a simulated thermal and humidity load disturbance during the free operation of a variable frequency air conditioner via a dynamic load simulation module; predict the trajectory of laboratory temperature and humidity changes within a predetermined time domain based on a long short-term memory network via a model prediction control module deployed at edge nodes, and generate control commands through rolling optimization to compensate for the thermal and humidity load disturbance; and dynamically decide the intensity and timing of the thermal and humidity load disturbance based on the operating state of the air conditioner under test via a deep reinforcement learning agent, so that the load disturbance simulates real free-running conditions.
[0089] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0090] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0091] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a data acquisition unit, a rule determination unit, a weight calculation unit, and an anomaly determination unit. The names of these units do not necessarily limit the specific unit; for example, a dynamic load simulation module may also be described as "generating a simulated thermal and humidity load disturbance during the free operation of a variable frequency air conditioner."
[0092] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), etc.
[0093] Obviously, those skilled in the art will understand that the various steps of the present invention described above can be performed in a manner different from that described above, and the simulation methods and experimental equipment include, but are not limited to, the above description. The steps of the present invention described above can be performed in a different order in certain circumstances, and the steps shown or described above can be performed separately. Therefore, the present invention is not limited to any particular combination of hardware and software.
[0094] The above description, in conjunction with specific embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such deductions or substitutions should be considered within the scope of protection of the present invention.
Claims
1. A control system for a variable frequency air conditioner free-running performance testing laboratory, characterized in that, include: The dynamic load simulation module is used to generate thermal and humidity load disturbances that simulate the free operation of a variable frequency air conditioner. The model prediction control module is deployed on the edge node. It predicts the trajectory of laboratory temperature and humidity changes in a predetermined time domain based on a long short-term memory network, and generates control commands through rolling optimization to compensate for the heat and humidity load disturbance. A deep reinforcement learning agent is used to dynamically decide the intensity and timing of thermal and humidity load disturbances based on the operating status of the air conditioner under test, so that the load disturbances simulate real free operation conditions. The deep reinforcement learning agent and the model prediction control module form a two-layer closed-loop control architecture: the deep reinforcement learning agent acts as an outer-loop disturbance generator, and the model prediction control module acts as an inner-loop compensation controller. The two work together to achieve stable control of the laboratory environment operating parameters under dynamic load disturbances.
2. The control system of the variable frequency air conditioner free-running performance testing laboratory as described in claim 1, characterized in that, The dynamic load simulation module includes: The heat load simulation unit is used to regulate the laboratory ambient temperature using a resistance heater assembly. The humidity load simulation unit is used to regulate the humidity of the laboratory environment using an ultrasonic humidifier and a condensation dehumidification device. An air circulation unit is used to ensure uniform air mixing inside the laboratory by employing a variable frequency fan.
3. The control system of the variable frequency air conditioner free-running performance testing laboratory as described in claim 2, characterized in that, The heat load simulation unit achieves PWM regulation through a solid-state relay or power regulator; the steam diffusion rate of the wet load simulation unit is controlled by a variable frequency fan in a closed loop.
4. The control system of the variable frequency air conditioner free-running performance testing laboratory as described in claim 1, characterized in that, The model prediction control module includes: The LSTM prediction unit is used to take into input multi-dimensional time-series data including historical laboratory temperature, historical laboratory humidity, historical inverter air conditioner power consumption and load equipment status, and output the trajectory of laboratory temperature and humidity changes within a predetermined future time domain. The MPC optimization unit is used to solve for the optimal control sequence based on the laboratory temperature and humidity change trajectory within the predetermined future time domain, so that the laboratory temperature and humidity follow the set trajectory and the control sequence satisfies the constraints of the laboratory environmental control actuator.
5. The control system of the variable frequency air conditioner free-running performance testing laboratory as described in claim 4, characterized in that, The LSTM prediction unit includes: The forget gate structure is used to determine the degree of retention of historical information through the sigmoid function; An input gate structure is used to control the proportion of new candidate cell states included; Cell state update mechanisms are used to update long-term memory through a weighted combination of forget gate structures and input gate structures; An output gate structure is used to generate hidden states through a tanh activation function to output predicted values. The predicted values are the trajectory of laboratory temperature and humidity changes within a predetermined time domain in the future, and the predicted values are directly used as inputs to the MPC optimization unit.
6. The control system of the variable frequency air conditioner free-running performance testing laboratory as described in claim 1, characterized in that, The deep reinforcement learning agent includes: The status sensing unit is used to collect the operating frequency, real-time power, and set operating parameters of the air conditioner under test in real time. The action decision unit is used to output continuous control commands for changes in heat load power and changes in wet load.
7. The control system of the variable frequency air conditioner free-running performance testing laboratory as described in claim 1, characterized in that, The deep reinforcement learning agent is trained using a deep deterministic policy gradient algorithm. The reward function of the deep reinforcement learning agent includes: a temperature and humidity tracking error term, a system energy consumption term, and a penalty term for the frequency change rate of the tested air conditioner.
8. The control system of the variable frequency air conditioner free-running performance testing laboratory as described in claim 1, characterized in that, It also includes a cloud-based collaborative platform, which receives environmental operating parameter data from multiple laboratories, trains and optimizes the long short-term memory network through federated learning, generates incremental model update packages, pushes them to edge nodes after encryption, and triggers hot updates after verifying model stability at the edge nodes.
9. The control system of the variable frequency air conditioner free-running performance testing laboratory as described in claim 1, characterized in that, It also includes a multi-layered sensor network deployed in the laboratory, including temperature sensors, humidity sensors, pressure sensors, wind speed sensors, and data acquisition systems.
10. A control method for a test laboratory of the free-running performance of a variable frequency air conditioner, characterized in that, include: The dynamic load simulation module generates heat and humidity load disturbances that simulate the free operation of a variable frequency air conditioner. The model prediction control module deployed at the edge node predicts the trajectory of laboratory temperature and humidity changes within a predetermined time domain based on the long short-term memory network, and generates control commands through rolling optimization to compensate for the heat and humidity load disturbance. The deep reinforcement learning agent dynamically decides the intensity and timing of heat and humidity load disturbances based on the operating status of the air conditioner under test, so that the load disturbances simulate real free operation conditions. The deep reinforcement learning agent and the model prediction control module form a two-layer closed-loop control architecture: the deep reinforcement learning agent acts as an outer-loop disturbance generator, and the model prediction control module acts as an inner-loop compensation controller. The two work together to achieve stable control of the laboratory environment operating parameters under dynamic load disturbances.
Citation Information
Patent Citations
A method for controlling a variable frequency air conditioner
CN113339941B
Intelligent energy-saving control method and system for air conditioner frequency conversion
CN120160246A