Transformer station terminal box temperature and humidity adaptive control system
By combining physical information neural networks and multi-objective reinforcement learning decision-makers with fuzzy controllers, the control of heaters and fans is dynamically optimized, solving the problems of accuracy and energy consumption in temperature and humidity control of substation terminal boxes, and improving the operational reliability and lifespan of the equipment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUBEI DONGCHENG TECHNOLOGY CO LTD
- Filing Date
- 2026-04-08
- Publication Date
- 2026-05-15
AI Technical Summary
Existing methods for controlling temperature and humidity in substation terminal boxes cannot cope with transient risks caused by drastic changes in the external environment, cannot accurately cover sensor blind spots, and have high energy consumption, thus failing to meet the needs for refined and forward-looking control.
By employing a physical information neural network model combined with a multi-objective reinforcement learning decision maker and a fuzzy controller, real-time data is collected to simulate the three-dimensional temperature and humidity field and identify risks. The control commands for heaters and fans are dynamically optimized to achieve a balance between condensation suppression and overheating prevention.
It achieves precise temperature and humidity control within the terminal box, reduces energy consumption, avoids control oscillations, and significantly improves operational reliability and service life.
Smart Images

Figure CN122043964A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of xx, specifically relating to an adaptive temperature and humidity control system for a substation terminal box. Background Technology
[0002] As a critical piece of equipment in the power system, the internal temperature and humidity environment of substation terminal boxes directly affects the operational reliability and service life of electrical components. Existing methods for controlling the temperature and humidity of terminal boxes mostly adopt a passive control mode based on fixed thresholds, only activating heaters or fans to remedy the situation when the average temperature and humidity inside the box exceed the standard, which has significant drawbacks.
[0003] Such methods struggle to address transient risks caused by drastic changes in the external environment (such as sudden rain, intense sunlight, and large diurnal temperature variations), and cannot resolve the conflicting objectives of condensation suppression and overheat control in high-humidity and high-temperature scenarios. Furthermore, sensor response hysteresis and the thermal inertia of the enclosure lead to control oscillations or failures. Traditional data-driven models lack generalization ability under extreme conditions, easily generating unreasonable control commands. In addition, discrete actuator control modes consume high energy and struggle to accurately cover localized risk areas within sensor blind spots, failing to meet the refined and forward-looking control requirements of terminal boxes. Summary of the Invention
[0004] To address this issue, the present invention provides a temperature and humidity adaptive control system for substation terminal boxes to solve the aforementioned technical problems.
[0005] According to one aspect of the present invention, a method for adaptive temperature and humidity control of a substation terminal box is provided, comprising the following steps:
[0006] S1: Real-time acquisition of monitoring data from multiple locations inside the terminal box, external environmental data, and the current status data of the actuator, wherein the actuator includes at least a heater and a fan;
[0007] S2: Input the monitoring data, external environment data and actuator status data into a pre-trained physical information neural network model. The physical information neural network model is trained by embedding the physical laws of thermodynamics and mass conservation as constraints into its loss function to output the real-time distribution and future prediction sequence of the temperature field, humidity field and dew point temperature field in the three-dimensional space inside the terminal box.
[0008] S3: Based on the real-time and predicted physical field distribution output by the physical information neural network model, dynamically identify and locate the core area of condensation risk and the hot spot area of overheating risk inside the box, and calculate the corresponding condensation risk index and overheating risk index.
[0009] S4: The condensation risk index, overheating risk index, external environment change trend and historical energy consumption data are used as state inputs and input into the multi-objective reinforcement learning decision-maker. The decision-maker outputs a joint optimization control command for the heater and fan by maximizing the reward function that comprehensively considers condensation suppression, overheating prevention and energy consumption minimization.
[0010] S5: The joint optimization control command is modulated using a fuzzy controller to generate a smooth pulse width modulation signal to control the power of the heater and to generate a speed control signal to control the speed of the fan.
[0011] Preferably, in step S2, the physical information neural network model adopts an encoder-decoder structure, wherein the encoder consists of multiple fully connected layers and multiple convolutional layers to encode sparse sensor monitoring data, actuator status and time series information into a low-dimensional feature vector, and the decoder consists of multiple deconvolutional layers and multiple fully connected layers to combine the low-dimensional feature vector with the three-dimensional spatial coordinates inside the box to output the temperature value and humidity value of the spatial coordinate point.
[0012] Preferably, the loss function for training the physical information neural network model includes data fitting loss and physical equation residual loss, wherein the physical equation residual loss is determined by randomly sampling a large number of collision points in the computational domain and calculating the degree to which their predicted values violate the embedded energy conservation equation and mass conservation equation.
[0013] Preferably, in step S3, the dynamic identification of the core area of condensation risk specifically involves finding, in the dew point temperature field, the location area with the smallest temperature difference between its temperature and the local dew point temperature and the difference being lower than the first safety threshold among all surface locations inside the box; and the dynamic identification of the hot spot area of overheating risk specifically involves finding, in the temperature field, the location area with the highest temperature and the temperature being higher than the second safety threshold.
[0014] Preferably, the action space of the multi-objective reinforcement learning decision-maker is a two-dimensional continuous space, directly defining the target power of the heater and the target speed of the fan.
[0015] Preferably, the reward function specifically includes a first reward term negatively correlated with the condensation risk index, a second reward term negatively correlated with the overheating risk index, a third reward term negatively correlated with the actuator energy consumption, and a contradictory adjustment penalty term used to punish the overuse of a single actuator when both the condensation risk and the overheating risk are higher than a specific threshold.
[0016] Preferably, the conflict adjustment penalty is configured such that when both the condensation risk and the overheating risk are at a high level, if the action command output by the decision-maker is to significantly increase the heater power and shut down the fan, or to significantly increase the fan speed and shut down the heater, then a negative reward is applied to the action.
[0017] Preferably, the fuzzy controller in step S5 converts the joint optimization control command and its changing trend into fuzzy linguistic variables, applies preset fuzzy rules for inference, and defuzzifies the inference result into the final actuator control signal.
[0018] Preferably, the condensation risk index is a time series function of the minimum difference between the temperature and the dew point temperature in the core condensation risk area, and the overheating risk index is a time series function of the difference between the temperature and the safety threshold in the overheating risk hotspot area.
[0019] In another aspect, this application also provides a substation terminal box temperature and humidity adaptive control system, comprising:
[0020] The acquisition module is used to acquire monitoring data from multiple locations inside the terminal box, external environmental data, and the current status data of the actuator in real time. The actuator includes at least a heater and a fan.
[0021] The temperature and humidity distribution prediction module is used to input the monitoring data, external environment data and actuator status data into a pre-trained physical information neural network model. The physical information neural network model is trained by embedding the physical laws of thermodynamics and mass conservation as constraints into its loss function to output the real-time distribution and future prediction sequence of the temperature field, humidity field and dew point temperature field in the three-dimensional space inside the terminal box.
[0022] The risk area positioning module is used to dynamically identify and locate the core condensation risk area and the overheating risk hotspot area inside the box based on the real-time and predicted physical field distribution output by the physical information neural network model, and to calculate the corresponding condensation risk index and overheating risk index.
[0023] The control command generation module is used to input the condensation risk index, overheating risk index, external environment change trend and historical energy consumption data as state inputs into the multi-objective reinforcement learning decision-maker. The decision-maker outputs a joint optimized control command for the heater and fan by maximizing the reward function that comprehensively considers condensation suppression, overheating prevention and energy consumption minimization.
[0024] The control command modulation module is used to modulate the joint optimized control command using a fuzzy controller, generate a smooth pulse width modulation signal to control the power of the heater, and generate a speed control signal to control the speed of the fan.
[0025] This invention uses a physical information neural network to accurately simulate and predict the three-dimensional temperature, humidity, and dew point field inside the terminal box. It combines multi-objective reinforcement learning to dynamically balance the risks of condensation, overheating, and energy consumption. Then, through fuzzy control, it achieves smooth adjustment of the actuator. This effectively solves the problems of passivity, insufficient generalization under extreme conditions, and decision-making defects in contradictory scenarios of traditional fixed threshold control. It can proactively suppress transient risks, avoid control oscillations caused by sensor hysteresis, significantly reduce energy consumption, and accurately cover local risks in sensor blind spots, greatly improving the operational reliability and service life of the terminal box. Attached Figure Description
[0026] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0027] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0028] Figure 1 A flowchart of a substation terminal box temperature and humidity adaptive control method provided in an embodiment of the present invention.
[0029] Figure 2 This is a schematic diagram of the physical information neural network model architecture provided in an embodiment of the present invention.
[0030] Figure 3 This is a schematic diagram of the fuzzy reasoning and defuzzification process provided in an embodiment of the present invention.
[0031] Figure 4 This is a schematic diagram of a substation terminal box temperature and humidity adaptive control system provided in an embodiment of the present invention. Detailed Implementation
[0032] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0033] It should be noted that all user information (including but not limited to user device information, user personal information, object information corresponding to device usage data, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, device usage data, etc.) involved in all embodiments of this disclosure are information and data authorized by the user or fully authorized by all parties.
[0034] like Figure 1 As shown in the figure, an embodiment of the present invention discloses a substation terminal box temperature and humidity adaptive control method 100, which includes the following method steps:
[0035] S1: Real-time acquisition of monitoring data from multiple locations inside the terminal box, external environmental data, and the current status data of the actuator, wherein the actuator includes at least a heater and a fan;
[0036] S2: Input the monitoring data, external environment data and actuator status data into a pre-trained physical information neural network model. The physical information neural network model is trained by embedding the physical laws of thermodynamics and mass conservation as constraints into its loss function to output the real-time distribution and future prediction sequence of the temperature field, humidity field and dew point temperature field in the three-dimensional space inside the terminal box.
[0037] S3: Based on the real-time and predicted physical field distribution output by the physical information neural network model, dynamically identify and locate the core area of condensation risk and the hot spot area of overheating risk inside the box, and calculate the corresponding condensation risk index and overheating risk index.
[0038] S4: The condensation risk index, overheating risk index, external environment change trend and historical energy consumption data are used as state inputs and input into the multi-objective reinforcement learning decision-maker. The decision-maker outputs a joint optimization control command for the heater and fan by maximizing the reward function that comprehensively considers condensation suppression, overheating prevention and energy consumption minimization.
[0039] S5: The joint optimization control command is modulated using a fuzzy controller to generate a smooth pulse width modulation signal to control the power of the heater and to generate a speed control signal to control the speed of the fan.
[0040] In some embodiments, for step S1, multi-dimensional data is collected by sensing devices deployed in and around the terminal box.
[0041] Specifically, regarding the collection of monitoring data inside the terminal box, the monitoring data inside the terminal box is mainly based on temperature and humidity data. To ensure that the temperature and humidity differences in different areas inside the box can be accurately reflected, this embodiment adopts a non-uniform arrangement of sensors inside the box. For example, at least 4 temperature and humidity sensors are deployed. The arrangement of the sensors is determined by preliminary computational fluid dynamics (CFD) analysis to cover key physical field change areas inside the box. It is understood that the number of sensors can be adjusted according to the actual application scenario, and this invention does not impose any limitations.
[0042] For example, the first sensor is deployed in the air circulation area near the cabinet door, an area susceptible to external environmental influences, and can quickly capture temperature and humidity changes caused by air exchange between the inside and outside of the cabinet. The second sensor is deployed on the shaded inner wall of the cabinet, a location where temperatures are typically lower and condensation is common. The third sensor is deployed at the far end of the heater to monitor the coverage of the heater's heating effect and the temperature gradient distribution. The fourth sensor is deployed in the densely packed terminal block area, where concentrated heat dissipation from components easily leads to localized overheating hotspots, and is used to monitor temperature changes. Each temperature and humidity sensor is set to a sampling frequency of once per minute, with sampling accuracy meeting the requirements of a temperature error not exceeding ±0.5℃ and a humidity error not exceeding ±3%RH. The collected data is transmitted in real time to the local data processing unit to form time-series data of temperature and humidity at multiple locations within the cabinet.
[0043] Regarding external environmental data, the collection scope covers key meteorological parameters affecting temperature and humidity changes inside the terminal box, specifically including external ambient temperature, external ambient humidity, wind speed, solar radiation intensity, and precipitation probability. In some embodiments, integrated meteorological sensors can be deployed at suitable locations on the top or around the terminal box to directly collect external temperature, humidity, and wind speed data. The temperature collection range is set to -40℃ to 85℃, the humidity collection range is set to 0%RH to 100%RH, and the wind speed collection range is set to 0 to 30 m / s. Solar radiation intensity data can be obtained by deploying a solar radiation intensity sensor, with a collection range of 0 to 2000 W / m². Precipitation probability data is obtained by connecting to a local weather station or an internet meteorological API interface, with the acquisition frequency consistent with the sampling frequency of the sensors inside the box, i.e., once per minute, to ensure the time synchronization between external environmental data and internal monitoring data.
[0044] In addition, to enable subsequent forward control, short-term weather forecast data for the next 15-60 minutes is obtained through this API interface, including predicted temperature, humidity, wind speed and solar radiation intensity at each time point in the future. This forecast data will serve as the input for the physical information neural network model to predict the future distribution of physical fields.
[0045] For the current status data of the actuator, for example, the actuator includes at least a heater and a fan, and the operating status parameters of both are collected respectively. Specifically, for the heater, the current operating current or actual output power of the heater is collected by deploying a current sensor or power sensor in its power supply circuit. If the heater supports multi-level control, the information of its current level is also collected, with a collection frequency of 1 time / minute.
[0046] For fans, the current actual speed of the fan is collected by deploying a speed sensor in the fan drive circuit or by reading the output signal of the fan controller. If the fan supports stepless speed regulation, its current speed regulation voltage or PWM (pulse width modulation) duty cycle information is also collected, with the collection frequency being 1 time / minute.
[0047] The collected actuator status data is fed back to the data processing unit in real time. On the one hand, it is used by the physical information neural network model to perceive the current active control state inside the box, so as to improve the accuracy of the model's simulation of the physical field inside the box. On the other hand, it is used by the multi-objective reinforcement learning decision-maker to judge the current operating trend of the actuator, providing a reference for the generation of subsequent optimized control commands.
[0048] It is understandable that all the data collected in step S1 needs to undergo preprocessing operations, including data cleaning (removing outliers and filling missing values), data standardization (converting data of different scales to the same scale), and data synchronization (ensuring that the timestamps of data from different sources are consistent). The preprocessed data is stored in a local database for subsequent steps, which will not be elaborated here.
[0049] In some embodiments, for step S2, according to an embodiment of the present invention, see [link to relevant documentation]. Figure 2 The physical information neural network model adopts an encoder-decoder deep neural network structure, which can effectively handle the mapping relationship between sparse sensor data and high-dimensional physical field output.
[0050] Specifically, the encoder input consists of sparse sensor monitoring data (temperature and humidity data at four locations inside the box and temperature and humidity data outside the box) collected and preprocessed in step S1, actuator status data (heater power / level, fan speed), and time series information (time stamp of data acquisition). For example, the encoder encodes the above high-dimensional, multi-source input data into a low-dimensional feature vector through a combination structure of three fully connected layers and two convolutional layers. The feature vector is set to 64 dimensions, which is essentially an abstract representation of the current comprehensive state inside and outside the terminal box and can cover the key information affecting the changes in the physical field inside the box.
[0051] The decoder is input to the low-dimensional feature vector output by the encoder and the three-dimensional spatial coordinates (x, y, z) inside the terminal box. The three-dimensional spatial coordinates are obtained by dividing the internal space of the terminal box into a grid. For example, the internal space of the terminal box is divided into several spatial grid points with a precision of 1cm×1cm×1cm. Each grid point corresponds to a unique (x, y, z) coordinate. The decoder uses a combination structure of 4 deconvolution layers and 2 fully connected layers to fuse the low-dimensional feature vector and the three-dimensional spatial coordinates, and finally outputs the temperature value T(x, y, z) and humidity value H(x, y, z) of each spatial grid point, thereby realizing the reconstruction of the three-dimensional temperature field and humidity field inside the box.
[0052] In one embodiment, preferably, the training process of the physical information neural network model specifically includes the following steps:
[0053] First, a training dataset is constructed. This dataset contains historical monitoring data of the terminal box under different seasons (spring, summer, autumn, winter), different weather conditions (sunny, rainy, cloudy, hazy), and different operating states (heater start-stop, fan start-stop). For example, the amount of data should be no less than 10,000 sets. Each set of data includes various input data collected in step S1 and the corresponding measured temperature and humidity data of key locations inside the box (obtained by additional temporary sensors).
[0054] Preferably, the loss function for training the physical information neural network model includes data fitting loss L_data and physical equation residual loss L_physics. The expression for the total loss function L_total is: L_total = α×L_data + β×L_physics, where α and β are weighting coefficients. In this embodiment, α is set to 0.4 and β is set to 0.6 to highlight the influence of physical law constraints on the model's prediction accuracy. L_data is the mean square error between the temperature and humidity values at the sensor location predicted by the model and the actual sensor readings. Its calculation formula is: L_data = (1 / N)×∑(i=1 to N)[(T_pred,i -T_meas,i)² + (H_pred,i - H_meas,i)²] In the formula, N is the number of sensors (N=4 in this embodiment), T_pred,i is the temperature value of the i-th sensor location predicted by the model, T_meas,i is the actual temperature value of the i-th sensor, H_pred,i is the humidity value of the i-th sensor location predicted by the model, and H_meas,i is the actual humidity value of the i-th sensor.
[0055] L_physics is the residual loss of the physics equations, which is determined by randomly sampling a large number of collision points within the computational domain inside the terminal box and calculating the extent to which the predicted values of these collision points violate the embedded energy conservation equations and mass conservation equations.
[0056] Specifically, firstly, M collision points are randomly sampled in the three-dimensional space inside the box (M=1000 in this embodiment), and the coordinates of each collision point are (x_j,y_j,z_j) (j=1,2,...,M). For the energy conservation equation, considering the influence of heat conduction and internal heat source (heater), the governing equation is: ρc(∂T / ∂t) = k∇²T + Q_heater Where, ρ is the air density (taken as 1.29kg / m³), c is the specific heat capacity of air (taken as 1005J / (kg·K)), k is the air thermal conductivity coefficient (taken as 0.024W / (m·K)), ∇² is the Laplace operator, Q_heater is the heat flux density generated by the heater at the collision point (calculated based on the heater power and position), and ∂T / ∂t is the partial derivative of temperature with respect to time. Substitute the temperature field T(x_j,y_j,z_j,t) predicted by the model into the above equation and calculate the difference between the left and right sides of the equation. This difference is the residual Res_energy,j of the energy conservation equation at the j-th collision point.
[0057] For the mass conservation equation, considering the diffusion and condensation processes of water vapor, the governing equation is: ∂(ρ_v)∂t =D∇²ρ_v - S_cond where ρ_v is the water vapor density (calculated from the relative humidity and temperature predicted by the model), D is the water vapor diffusion coefficient (taken as 2.5×10^-5 m² / s), and S_cond is the condensation source sink term (S_cond is positive when the collision point temperature is lower than the dew point temperature, representing a decrease in water vapor due to condensation; otherwise, it is 0). Similarly, substituting the water vapor density corresponding to the humidity field predicted by the model into the above equation, and calculating the difference between the left and right sides of the equation, we obtain the residual Res_mass,j of the mass conservation equation at the j-th collision point.
[0058] The formula for calculating the residual loss L_physics of the physical equation is: L_physics = (1 / M)×∑(j=1 toM)[Res_energy,j² + Res_mass,j²];
[0059] During training, the Adam optimizer is used. For example, the learning rate is initially set to 1×10^-4, and after every 100 epochs, the learning rate is reduced to 0.8 times the original value. The total number of training epochs is set to 500, until the total loss function L_total converges to a stable value (such as less than 1×10^-3). At this point, offline training is completed, and a preliminary physical information neural network model is obtained.
[0060] In one embodiment, after the physical information neural network model receives the real-time data and short-term weather forecast data input in step S1, it first calculates the real-time distribution of the temperature field T(x,y,z) and humidity field H(x,y,z) in the three-dimensional space inside the terminal box based on the real-time data; then, according to the real-time distribution of the temperature and humidity fields, it calculates the dew point temperature T_dp(x,y,z) of each spatial grid point (the dew point temperature is calculated using the Magnus-Tetens equation, i.e., T_dp = (b×ln(RH / 100) + a×T) / (a - ln(RH / 100) - (b×T), where a=17.625, b=243.04℃, RH is relative humidity, and T is temperature), thus obtaining the real-time distribution of the dew point temperature field; finally, based on future short-term weather forecast data, the model predicts the distribution of temperature field, humidity field, and dew point temperature field corresponding to each time step in the next 15-30 minutes through rolling time domain prediction with a time step of 5 minutes, forming a future prediction sequence, which will serve as the basis for risk identification in step S3.
[0061] In some embodiments, for step S3, preferably, dynamically identifying the core area of condensation risk, specifically, in the dew point temperature field, finding the location area in which the temperature difference between the surface location inside the box and the local dew point temperature is the smallest and the difference is lower than a first safety threshold.
[0062] Specifically, the scope of "internal surface location" is first defined, which includes the inner walls of the terminal box (front, rear, left, right, upper, and lower walls), the surface of the terminal blocks, the surface of the heater, and the surface of the fan, as well as other solid surfaces that are in direct contact with the air. By extracting the temperature value T_surface(x,y,z) corresponding to the above surface locations in the three-dimensional temperature field output in step S2, and the dew point temperature value T_dp_surface(x,y,z) corresponding to these locations, the difference between the temperature and the dew point temperature of each surface location, ΔT_cond = T_surface(x,y,z) - T_dp_surface(x,y,z), is calculated.
[0063] Then, a first safety threshold ΔT_cond_th is set. In this embodiment, ΔT_cond_th is set to 2℃. When ΔT_cond < ΔT_cond_th, it indicates that there is a risk of condensation at that location. From all surface locations where ΔT_cond < ΔT_cond_th, the location with the smallest ΔT_cond is selected. A spherical area with a radius of 5cm centered on this location is selected as the core area of condensation risk. For example, if the ΔT_cond of a point on the inner wall of the terminal box is 0.5℃ (less than 2℃) and is the smallest ΔT_cond value among all surface locations, then the area within 5cm centered on this point is the core area of condensation risk.
[0064] Similarly, in a preferred embodiment, overheat risk hotspot areas are dynamically identified, specifically by finding the location area with the highest temperature in the temperature field that is higher than the second safety threshold.
[0065] First, a second safety threshold T_overheat_th is set, which is determined based on the temperature tolerance of the components inside the terminal box; in this embodiment, it is set to 40℃. Then, the temperature values T(x,y,z) of all spatial grid points in the three-dimensional temperature field output in step S2 are extracted, and grid points where T(x,y,z) > T_overheat_th are selected. From these grid points, the point with the highest temperature value, T_max(x,y,z), is identified. A spherical region with a radius of 5cm, centered on this point, is selected as the overheat risk hotspot area. For example, if the temperature at a point in a densely populated area of the terminal blocks is 45℃ (higher than 40℃) and is the highest temperature inside the box, then the area within 5cm of this point is the overheat risk hotspot area.
[0066] It should be noted that in some embodiments, if there are multiple points with the same and highest temperature values or multiple points with the same and lowest ΔT_cond values within the box, the 5cm spherical regions corresponding to these points are merged to form a continuous risk area to ensure the completeness of risk identification.
[0067] Preferably, in one embodiment, the condensation risk index is a time series function of the minimum difference between the temperature and the dew point temperature in the core condensation risk area; the overheating risk index is a time series function of the difference between the temperature and the safety threshold in the overheating risk hotspot area.
[0068] Specifically, the condensation risk index Risk_cond is calculated as follows: Risk_cond(t) = -k1×∑(τ=t-Δt to t)ΔT_cond_min(τ) where t is the current time, Δt is the time window length (set to 10 minutes in this embodiment), ΔT_cond_min(τ) is the minimum difference between the temperature and dew point temperature in the core condensation risk area at time τ, and k1 is the weighting coefficient (set to 0.1). This formula shows that the smaller ΔT_cond_min(τ) is (the higher the condensation risk), and the longer it lasts within the time window, the smaller the value of the condensation risk index Risk_cond(t) (a negative indicator; the smaller the value, the higher the risk).
[0069] The overheat risk index Risk_overheat is calculated as follows: Risk_overheat(t) = k2×∑(τ=t-Δt to t)(T_max(τ) - T_overheat_th) where T_max(τ) is the highest temperature of the overheat risk hotspot at time τ, T_overheat_th is the second safety threshold (40℃), k2 is the weighting coefficient (set to 0.05), and Δt is 10 minutes. This formula shows that the larger the difference between T_max(τ) and T_overheat_th (the higher the overheat risk), and the longer the difference persists within the time window, the larger the value of the overheat risk index Risk_overheat(t) (a positive indicator; the larger the value, the higher the risk).
[0070] The condensation risk index and overheating risk index calculated in the above manner can comprehensively reflect the intensity and duration of the risk, providing a quantitative basis for the decision-making of the multi-objective reinforcement learning decision-maker in step S4.
[0071] In some embodiments, for step S4, a multi-objective reinforcement learning decision-maker is used to achieve a dynamic trade-off between condensation suppression, overheating prevention, and energy consumption minimization, and output an optimized joint control command.
[0072] Specifically, for the state space definition of a multi-objective reinforcement learning decision-maker, according to an embodiment of the present invention, the state space S of the multi-objective reinforcement learning decision-maker consists of four key state variables, namely S = [Risk_cond, Risk_overheat, Delta_Energy, Weather_Trend]. For example, the specific definitions and values of each state variable are as follows:
[0073] Risk_cond: This is the condensation risk index calculated in step S3. Its value range is determined according to the actual calculation results. In this embodiment, it is [-5, 2]. The smaller the value, the higher the condensation risk.
[0074] Risk_overheat: This is the overheating risk index calculated in step S3. Its value ranges from [0, 10]. The larger the value, the higher the overheating risk.
[0075] Delta_Energy: The recent cumulative energy consumption trend, calculated as follows: Delta_Energy = E_current - E_prev, where E_current is the cumulative energy consumption of the actuator (heater + fan) in the past hour, and E_prev is the cumulative energy consumption in the previous hour. The energy consumption is calculated as follows: E_heater = ∑(τ=1 to 60)P_heater(τ)×1 / 60 (P_heater(τ) is the power of the heater in minute τ, in kW), E_fan = ∑(τ=1 to 60)(k_fan×R_fan(τ)³)×1 / 60 (k_fan is the fan power consumption coefficient, set to 1×10^-6 kW·min / (r³) in this embodiment, and R_fan(τ) is the fan speed in minute τ, in r / min), and the total energy consumption E = E_heater + E_fan. The value of Delta_Energy ranges from [-0.5, 0.5]. A positive value indicates that the current energy consumption has increased compared to before, and a negative value indicates that the current energy consumption has decreased compared to before.
[0076] Weather_Trend: The trend of external environmental changes, obtained through analysis of short-term weather forecast data, is specifically divided into three levels: "improving," "stable," and "deteriorating," represented by 1, 0, and -1, respectively. For example, if the forecast shows that the outside temperature will drop by 5°C and the humidity will decrease by 10%RH within the next 30 minutes, then Weather_Trend is determined to be "improving" (value 1); if the forecast shows that the changes in outside temperature and humidity within the next 30 minutes are both less than 1°C and 2%RH, then Weather_Trend is determined to be "stable" (value 0); if the forecast shows that the outside temperature will rise by 3°C and the humidity will rise by 15%RH within the next 30 minutes, then Weather_Trend is determined to be "deteriorating" (value -1).
[0077] It is understandable that the four state variables mentioned above need to be standardized before being input into the decision-maker, and their values need to be mapped to the interval [-1, 1] to ensure that the influence weight of each state variable on the decision is balanced. This will not be elaborated further here.
[0078] In one embodiment, regarding the definition of the action space of the multi-objective reinforcement learning decision-maker, preferably, the action space of the multi-objective reinforcement learning decision-maker is a two-dimensional continuous space, directly defining the target power of the heater and the target speed of the fan.
[0079] Specifically, the action space A = [P_heater_target, R_fan_target], where P_heater_target is the target power of the heater, and its value range is determined according to the rated power of the heater. In this embodiment, the rated power of the heater is 1kW, so the value range of P_heater_target is [0, 1]kW (0 represents the heater being off, and 1kW represents the heater operating at full power); R_fan_target is the target speed of the fan, and its value range is determined according to the rated speed of the fan. In this embodiment, the rated speed of the fan is 2000r / min, so the value range of R_fan_target is [0, 2000]r / min (0 represents the fan being off, and 2000r / min represents the fan operating at full speed).
[0080] The advantage of using a continuous motion space is that it enables precise adjustment of heater power and fan speed, avoiding the problem of insufficient control precision caused by traditional discrete motion spaces (such as heaters only "on / off" and fans only "high / medium / low" speeds), and is more adaptable to the dynamic changes in condensation and overheating risks.
[0081] In one embodiment, preferably, the reward function specifically includes: a first reward term negatively correlated with the condensation risk index, a second reward term negatively correlated with the overheating risk index, a third reward term negatively correlated with actuator energy consumption, and a contradictory adjustment penalty term used to penalize the overuse of a single actuator when both the condensation risk and overheating risk are above a specific threshold. The overall expression of the reward function is: R = R_cond + R_overheat + R_energy + R_penalty;
[0082] Specifically, the first reward item, R_cond, is negatively correlated with the condensation risk index, Risk_cond. That is, the smaller Risk_cond is (the higher the condensation risk), the smaller the value of R_cond; conversely, the larger Risk_cond is (the lower the condensation risk), the larger the value of R_cond. Its calculation formula is: R_cond = k_cond × Risk_cond, where k_cond is the reward coefficient, set to 2 in this embodiment to ensure that the value range of R_cond matches that of other reward items.
[0083] The second reward item, R_overheat, is negatively correlated with the overheating risk index, Risk_overheat. That is, the larger the Risk_overheat (the higher the overheating risk), the smaller the value of R_overheat; conversely, the smaller the Risk_overheat (the lower the overheating risk), the larger the value of R_overheat. Its calculation formula is: R_overheat = k_overheat × (R_overheat_max - Risk_overheat) where k_overheat is the reward coefficient (set to 1), and R_overheat_max is the maximum possible value of the overheating risk index (10 in this embodiment). Therefore, the value range of R_overheat is [0, 10].
[0084] The third reward item, R_energy, is negatively correlated with the actuator's energy consumption. That is, the higher the energy consumption, the smaller the value of R_energy; the lower the energy consumption, the larger the value of R_energy. Its calculation formula is: R_energy = k_energy × (E_max - E_current) where k_energy is the reward coefficient (set to 0.5), E_max is the actuator's maximum hourly energy consumption (in this embodiment, 1kW×1h + (1×10^-6 kW·min / (r³)×(2000r / min)³)×60min = 1 + 4.8 = 5.8kWh), and E_current is the cumulative energy consumption for the current hour. Therefore, the value range of R_energy is [0, 2.9].
[0085] Contradictory adjustment penalty R_penalty: Preferably, it is configured as follows: when both condensation risk and overheating risk are high, if the action command output by the decision-maker is to significantly increase the heater power and shut down the fan, or to significantly increase the fan speed and shut down the heater, then a negative reward is applied to the action.
[0086] Specifically, firstly, a high threshold for condensation risk, Risk_cond_high (set to -2 in this embodiment), and a high threshold for overheat risk, Risk_overheat_high (set to 5 in this embodiment), are set. When Risk_cond ≤ Risk_cond_high and Risk_overheat ≥ Risk_overheat_high, it is determined that both condensation risk and overheat risk are at a high level.
[0087] If the following conditions are met, R_penalty will be triggered if the action command meets one of the following two conditions: Condition 1: P_heater_target ≥ 0.8kW (significantly increase heater power) and R_fan_target = 0r / min (turn off fan), in which case R_penalty = -k_penalty1 (k_penalty1 is set to 5); Condition 2: R_fan_target ≥ 1600r / min (significantly increase fan speed) and P_heater_target = 0kW (turn off heater), in which case R_penalty = -k_penalty2 (k_penalty2 is set to 5); If the action command does not meet either of the above two conditions, then R_penalty = 0;
[0088] By designing the above reward function, we can guide the multi-objective reinforcement learning decision-maker to make optimal decisions under different risk scenarios: when the risk is low, we prioritize minimizing energy consumption; when a single risk is high, we prioritize suppressing that risk; when both risks are high, we avoid "brutal" control and learn a fine strategy that balances both.
[0089] In one embodiment, the multi-objective reinforcement learning decision-maker is optionally trained using the Deep Deterministic Policy Gradient (DDPG) algorithm, which is suitable for reinforcement learning tasks in a continuous action space.
[0090] The training process specifically includes, firstly, constructing an experience replay pool, storing the (S, A, R, S') experience samples generated during the system operation in the experience replay pool, where S is the current state, A is the current action, R is the reward obtained by the current action, and S' is the next state entered after executing action A. The capacity of the experience replay pool is set to 100,000.
[0091] The network structure is as follows: the decision-maker consists of a policy network (Actor) and a value network (Critic). The policy network takes the state space S as input and outputs the action space A. Its structure consists of three fully connected layers (with 128, 64, and 32 neurons respectively), and uses ReLU as the activation function. The value network takes the concatenated vector of state S and action A as input and outputs the action value Q(S,A). Its structure also consists of three fully connected layers (with 128, 64, and 32 neurons respectively), and uses ReLU as the activation function.
[0092] During training, a dual-network structure (i.e., a main network and a target network) is adopted. The main network is responsible for generating actions and evaluating action values, while the target network is responsible for providing a stable target Q-value. During training, 32 experience samples are randomly sampled from the experience replay pool, and the target Q-value Q_target = R + γ×Q_target(S', A_target) is calculated (where γ is a discount factor set to 0.95, and A_target is the action output by the target policy network). Then, the parameters of the main value network are updated by minimizing the loss function L_critic = ∑(Q(S,A) - Q_target)² of the value network.
[0093] Simultaneously, the parameters of the main policy network are updated using policy gradient ascent to maximize the action value Q(S,A) output by the value network. Every 100 training steps, the parameters of the main network are soft-updated to the target network (update coefficient τ is set to 0.01). The total number of training steps is set to 100,000 until the actions output by the policy network can stably maximize the cumulative reward.
[0094] In one embodiment, during the decision-making phase, in each control cycle (set to 5 minutes in this embodiment), the multi-objective reinforcement learning decision-maker receives the state variable S input in step S3, outputs the target power P_heater_target of the heater and the target speed R_fan_target of the fan through the trained policy network, forms a joint optimization control command, and transmits it to the fuzzy controller in step S5.
[0095] In some embodiments, for step S5, see Figure 3 By using a fuzzy controller to smoothly modulate the joint optimization control commands, mechanical wear and current surges caused by frequent start-stop of the actuator are avoided. The specific fuzzy inference and defuzzification process is as follows.
[0096] Specifically, in S301, the fuzzy controller receives the input of the joint optimization control command output by the multi-objective reinforcement learning decision-maker, namely the heater target power P_heater_target and the fan target speed R_fan_target, as well as the changing trends of these two input quantities ΔP_heater = P_heater_target - P_heater_current (P_heater_current is the current heater power) and ΔR_fan = R_fan_target - R_fan_current (R_fan_current is the current fan speed).
[0097] Therefore, the fuzzy controller is a dual-input dual-output structure. The input variables are [P_heater_target, ΔP_heater] and [R_fan_target, ΔR_fan], and the output variables are the heater's PWM duty cycle D_heater (value range 0~100%) and the fan's speed control voltage U_fan (value range 0~12V, corresponding to fan speed 0~2000r / min).
[0098] Specifically, S302, the fuzzification of input variables includes, firstly, defining fuzzy linguistic variables and their corresponding membership functions for each input variable. Taking the target power of the heater P_heater_target as an example, its value range is [0, 1] kW. Five fuzzy linguistic variables are defined: "zero" (Z), "small" (S), "medium" (M), "large" (L), and "maximum" (VL). The membership functions of each fuzzy linguistic variable adopt triangular membership functions, and the specific universe of discourse is divided as follows:
[0099] "Zero" (Z): 0~0.2kW, with the center point being 0kW;
[0100] "Small" (S): 0.1~0.3kW, with a center point of 0.2kW;
[0101] "Middle" (M): 0.2~0.5kW, with a center point of 0.35kW;
[0102] "Large" (L): 0.4~0.8kW, with a center point of 0.6kW;
[0103] Maximum power (VL): 0.7~1.0kW, with a center point of 0.85kW.
[0104] Similarly, for ΔP_heater (within the range [-0.5, 0.5]kW), five fuzzy linguistic variables are defined: "Negative Large" (NB), "Negative Small" (NS), "Zero" (Z), "Positive Small" (PS), and "Positive Large" (PB). The universe of discourse is divided as follows:
[0105] "Negative-high" (NB): -0.5~-0.3kW, with a center point of -0.4kW;
[0106] "Negative Small" (NS): -0.4~-0.1kW, with a center point of -0.25kW;
[0107] "Zero" (Z): -0.2~0.2kW, with the center point at 0kW;
[0108] "Small Power" (PS): 0.1~0.4kW, with a center point of 0.25kW;
[0109] "Zhengda" (PB): 0.3~0.5kW, with a center point of 0.4kW.
[0110] The fuzzification processing of the fan target speed R_fan_target (value range [0, 2000] r / min) and ΔR_fan (value range [-1000, 1000] r / min) is similar to that of the heater, both defining 5 fuzzy linguistic variables, which will not be elaborated here.
[0111] During the fuzzification process, for each input variable's actual value, the membership degree value belonging to each fuzzy linguistic variable is calculated using the membership degree function. For example, if the actual value of P_heater_target is 0.3kW, then its membership degree belonging to "small" (S) is 0.5, its membership degree belonging to "medium" (M) is 0.5, and its membership degree belonging to other fuzzy linguistic variables is 0.
[0112] In one embodiment, S303, a fuzzy rule base is constructed. The fuzzy controller applies preset fuzzy rules for inference. The fuzzy rule base is constructed based on engineering experience and control requirements, and adopts the "IF-THEN" rule format. Taking the PWM duty cycle control of the heater as an example, some fuzzy rules are as follows:
[0113] IF P_heater_target is Z and ΔP_heater is Z, THEN D_heater is Z;
[0114] IF P_heater_target is S and ΔP_heater is PS, THEN D_heater is S;
[0115] IF P_heater_target is M and ΔP_heater is Z, THEN D_heater is M;
[0116] IF P_heater_target is L and ΔP_heater is PB, THEN D_heater is L;
[0117] IF P_heater_target is VL and ΔP_heater is Z, THEN D_heater is VL;
[0118] IF P_heater_target is M and ΔP_heater is NB, THEN D_heater is S;
[0119] IF P_heater_target is L and ΔP_heater is NS, THEN D_heater is M;
[0120] Similarly, for fan speed control voltage regulation, some fuzzy rules may include:
[0121] IF R_fan_target is Z and ΔR_fan is Z, THEN U_fan is Z;
[0122] IF R_fan_target is S and ΔR_fan is PS, THEN U_fan is S;
[0123] IF R_fan_target is M and ΔR_fan is Z, THEN U_fan is M;
[0124] IF R_fan_target is L and ΔR_fan is PB, THEN U_fan is L;
[0125] IF R_fan_target is M and ΔR_fan is NB, THEN U_fan is S;
[0126] IF R_fan_target is L and ΔR_fan is NS, THEN U_fan is M;
[0127] In addition, to address scenarios where condensation and overheating risks coexist, priority adjustment rules need to be added to the fuzzy rules. For example: IF P_heater_target is M and R_fan_target is M and Risk_cond isHigh, THEN D_heater is M+ and U_fan is M- (where "M+" represents an output slightly higher than "medium" and "M-" represents an output slightly lower than "medium"), to ensure that while balancing risks, higher priority control needs are met first.
[0128] In this embodiment, the fuzzy rule base contains a total of 5×5×2=50 rules (5 fuzzy linguistic variables for each input variable and two output variables), covering all possible combinations of input states.
[0129] In one embodiment, S304, fuzzy reasoning is performed based on the Mamdani inference method. Optionally, the fuzzy reasoning uses the Mamdani inference method, and the specific steps are as follows:
[0130] For each fuzzy rule, the membership values of the input variables are ANDed (and the minimum value is taken) to obtain the trigger strength of the rule;
[0131] Based on the trigger strength of the rule, the fuzzy linguistic variables of the output variables are "pruned" to obtain the output fuzzy set corresponding to the rule;
[0132] Perform an OR operation on the output fuzzy sets corresponding to all rules (taking the maximum value) to obtain the final output fuzzy set.
[0133] S305, based on the centroid method, defuzzifies the output fuzzy set into a control signal. The defuzzification uses the centroid method to convert the output fuzzy set into a precise control signal value. Taking the PWM duty cycle D_heater of the heater as an example, its calculation formula is: D_heater = ∑(i=1 to n)(μ_i × D_i) / ∑(i=1 to n)μ_i. In the formula, n is the number of fuzzy linguistic variables contained in the output fuzzy set, μ_i is the membership value of the i-th fuzzy linguistic variable, and D_i is the center value corresponding to the i-th fuzzy linguistic variable (e.g., the center value of "zero" (Z) is 0%, the center value of "small" (S) is 20%, the center value of "medium" (M) is 50%, the center value of "large" (L) is 80%, and the center value of "maximum" (VL) is 100%).
[0134] Through the above fuzzy reasoning and defuzzification process, the PWM control signal of the heater and the speed control signal of the fan are finally obtained. These signals are transmitted to the drive circuit of the actuator to control the heater to output power according to the PWM duty cycle and the fan to adjust the speed according to the speed regulation voltage, so as to achieve smooth, accurate and adaptive look-ahead control of the temperature and humidity of the terminal box.
[0135] In summary, by combining the precise modeling capability of physical information neural networks, the dynamic decision-making capability of multi-objective reinforcement learning, and the smooth execution capability of fuzzy control through the coordinated operation of steps S1 to S5, this invention effectively solves the defects of traditional terminal box temperature and humidity control methods in transient processes, contradictory scenarios, and sensor hysteresis, and achieves proactive suppression of condensation and overheating risks and minimization of energy consumption.
[0136] Figure 4 An adaptive temperature and humidity control system 400 for a substation terminal box is shown. The system embodiment is similar to... Figure 1 Corresponding to the illustrated method embodiments, the specific methods include:
[0137] The acquisition module 401 is used to acquire monitoring data from multiple locations inside the terminal box, external environmental data, and the current status data of the actuator in real time. The actuator includes at least a heater and a fan.
[0138] The temperature and humidity distribution prediction module 402 is used to input the monitoring data, external environment data and actuator status data into a pre-trained physical information neural network model. The physical information neural network model is trained by embedding the physical laws of thermodynamics and mass conservation as constraints into its loss function to output the real-time distribution and future prediction sequence of the temperature field, humidity field and dew point temperature field in the three-dimensional space inside the terminal box.
[0139] The risk area positioning module 403 is used to dynamically identify and locate the core area of condensation risk and the hot spot area of overheating risk in the box based on the real-time and predicted physical field distribution output by the physical information neural network model, and to calculate the corresponding condensation risk index and overheating risk index.
[0140] The control command generation module 404 is used to input the condensation risk index, overheating risk index, external environment change trend and historical energy consumption data as state inputs into the multi-objective reinforcement learning decision-maker. The decision-maker outputs a joint optimized control command for the heater and fan by maximizing the reward function that comprehensively considers condensation suppression, overheating prevention and energy consumption minimization.
[0141] The control command modulation module 405 is used to modulate the joint optimized control command using a fuzzy controller, generate a smooth pulse width modulation signal to control the power of the heater, and generate a speed regulation signal to control the speed of the fan.
[0142] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for adaptive temperature and humidity control of a substation terminal box, characterized in that, Includes the following steps: S1: Real-time acquisition of monitoring data from multiple locations inside the terminal box, external environmental data, and the current status data of the actuator, wherein the actuator includes at least a heater and a fan; S2: Input the monitoring data, external environment data and actuator status data into a pre-trained physical information neural network model. The physical information neural network model is trained by embedding the physical laws of thermodynamics and mass conservation as constraints into its loss function to output the real-time distribution and future prediction sequence of the temperature field, humidity field and dew point temperature field in the three-dimensional space inside the terminal box. S3: Based on the real-time and predicted physical field distribution output by the physical information neural network model, dynamically identify and locate the core area of condensation risk and the hot spot area of overheating risk inside the box, and calculate the corresponding condensation risk index and overheating risk index. S4: The condensation risk index, overheating risk index, external environment change trend and historical energy consumption data are used as state inputs and input into the multi-objective reinforcement learning decision-maker. The decision-maker outputs a joint optimization control command for the heater and fan by maximizing the reward function that comprehensively considers condensation suppression, overheating prevention and energy consumption minimization. S5: The joint optimization control command is modulated using a fuzzy controller to generate a smooth pulse width modulation signal to control the power of the heater and to generate a speed control signal to control the speed of the fan.
2. The substation terminal box temperature and humidity adaptive control method according to claim 1, characterized in that, It also includes, In step S2, the physical information neural network model adopts an encoder-decoder structure. The encoder consists of multiple fully connected layers and multiple convolutional layers to encode sparse sensor monitoring data, actuator status, and time series information into a low-dimensional feature vector. The decoder consists of multiple deconvolutional layers and multiple fully connected layers to combine the low-dimensional feature vector with the three-dimensional spatial coordinates inside the box and output the temperature and humidity values of the spatial coordinate points.
3. The substation terminal box temperature and humidity adaptive control method according to claim 2, characterized in that, The loss function for training the physical information neural network model includes data fitting loss and physical equation residual loss. The physical equation residual loss is determined by randomly sampling a large number of collision points in the computational domain and calculating the degree to which their predicted values violate the embedded energy conservation equation and mass conservation equation.
4. The substation terminal box temperature and humidity adaptive control method according to claim 1, characterized in that, include: In step S3, the core area of condensation risk is dynamically identified. Specifically, in the dew point temperature field, the location area with the smallest temperature difference between the surface location inside the box and the local dew point temperature, and where the difference is lower than the first safety threshold, is found. Dynamically identify overheating risk hotspots, specifically by finding the location region with the highest temperature in the temperature field that is above the second safety threshold.
5. The substation terminal box temperature and humidity adaptive control method according to claim 1, characterized in that, The action space of the multi-objective reinforcement learning decision-maker is a two-dimensional continuous space, which directly defines the target power of the heater and the target speed of the fan.
6. The substation terminal box temperature and humidity adaptive control method according to claim 1, characterized in that, The reward function specifically includes a first reward term negatively correlated with the condensation risk index, a second reward term negatively correlated with the overheating risk index, a third reward term negatively correlated with the actuator energy consumption, and a contradictory adjustment penalty term used to punish the overuse of a single actuator when both the condensation risk and the overheating risk are above a specific threshold.
7. The substation terminal box temperature and humidity adaptive control method according to claim 6, characterized in that, The conflict adjustment penalty is configured such that when both condensation risk and overheating risk are high, if the action command output by the decision-maker is to significantly increase the heater power and shut down the fan, or to significantly increase the fan speed and shut down the heater, then a negative reward is applied to that action.
8. The substation terminal box temperature and humidity adaptive control method according to claim 1, characterized in that, The fuzzy controller described in step S5 transforms the joint optimization control command and its changing trend into fuzzy linguistic variables, applies preset fuzzy rules for inference, and defuzzifies the inference result into the final actuator control signal.
9. The substation terminal box temperature and humidity adaptive control method according to claim 1, characterized in that, The condensation risk index is a time series function of the minimum difference between the temperature and the dew point temperature in the core area of condensation risk, and the overheating risk index is a time series function of the difference between the temperature and the safety threshold in the hot spot area of overheating risk.
10. A temperature and humidity adaptive control system for a substation terminal box, characterized in that, include: The acquisition module is used to acquire monitoring data from multiple locations inside the terminal box, external environmental data, and the current status data of the actuator in real time. The actuator includes at least a heater and a fan. The temperature and humidity distribution prediction module is used to input the monitoring data, external environment data and actuator status data into a pre-trained physical information neural network model. The physical information neural network model is trained by embedding the physical laws of thermodynamics and mass conservation as constraints into its loss function to output the real-time distribution and future prediction sequence of the temperature field, humidity field and dew point temperature field in the three-dimensional space inside the terminal box. The risk area positioning module is used to dynamically identify and locate the core condensation risk area and the overheating risk hotspot area inside the box based on the real-time and predicted physical field distribution output by the physical information neural network model, and to calculate the corresponding condensation risk index and overheating risk index. The control command generation module is used to input the condensation risk index, overheating risk index, external environment change trend and historical energy consumption data as state inputs into the multi-objective reinforcement learning decision-maker. The decision-maker outputs a joint optimized control command for the heater and fan by maximizing the reward function that comprehensively considers condensation suppression, overheating prevention and energy consumption minimization. The control command modulation module is used to modulate the joint optimized control command using a fuzzy controller, generate a smooth pulse width modulation signal to control the power of the heater, and generate a speed control signal to control the speed of the fan.