Central air conditioner intelligent control method and system based on user behavior prediction
By combining the double-delay deep deterministic policy gradient algorithm model with the behavior-density zoning and emotion load correction model, closed-loop intelligent control of the central air-conditioning system is achieved, which solves the problems of regional environmental imbalance and lack of adaptability in traditional control methods and improves energy efficiency and thermal comfort.
Patent Information
- Application Number
- CN202511074599.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-09-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional central air-conditioning control methods lack deep integration of the dynamic characteristics of human behavior and are unable to adjust air-conditioning output parameters in real time, resulting in regional environmental imbalance and energy waste. In addition, the control algorithm lacks adaptability and is prone to overshoot and response lag.
An intelligent control method based on user behavior prediction is adopted. Through the dual-delay deep deterministic policy gradient algorithm model combined with the behavior-density zoning and emotional load correction model, central air-conditioning operating parameters and personnel behavior data are collected in real time, and the control strategy is dynamically adjusted to form a closed-loop intelligent control mechanism.
The central air-conditioning system can meet the thermal comfort of users while improving energy efficiency, reducing overcooling or overheating, enhancing the adaptability to the strong coupling and nonlinear characteristics of the system, and optimizing the control strategy.
Smart Images

Figure CN120684794A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of central air-conditioning intelligent control, and in particular to a central air-conditioning intelligent control method and system based on user behavior prediction. Background Art
[0002] With the increasing level of intelligent buildings, central air conditioning systems, as a core component of building energy consumption, have attracted much attention for their control accuracy and energy efficiency. Dynamic changes in user behavior have a significant impact on indoor thermal environment requirements. Traditional central air conditioning control methods often rely on preset parameters or simple feedback adjustments, which are difficult to adapt to fluctuations in occupancy density in different spatial areas and individual differences in temperature perception. In this context, intelligent control technology that integrates user behavior prediction has become a research hotspot. By integrating central air conditioning operating parameters, occupant behavior data, and environmental perception information, a dynamically responsive control logic is constructed to achieve the coordinated optimization of system operating efficiency and user comfort.
[0003] Existing technologies have two significant shortcomings: First, the control strategy lacks deep integration of the dynamic characteristics of personnel behavior, and mostly adopts fixed partitioning or static load calculation methods. It is impossible to adjust the air-conditioning output parameters in real time according to the flow trajectory of personnel and changes in density distribution, resulting in overcooling or overheating in some areas, causing energy waste; second, the control algorithm is not adaptable enough. Traditional PID control or single intelligent algorithm has difficulty in handling the strong coupling and nonlinear characteristics of the central air-conditioning system. Overshoot or response lag is prone to occur during the dynamic adjustment of parameters. In addition, a precise correlation mechanism between behavioral load and system operating parameters has not been established, making it difficult to achieve optimal energy efficiency while meeting users' thermal comfort needs. Summary of the Invention
[0004] In order to overcome the shortcomings and deficiencies of the prior art, the present invention provides a central air-conditioning intelligent control method and system based on user behavior prediction.
[0005] The technical solution adopted by the present invention is a central air-conditioning intelligent control method based on user behavior prediction, comprising the following steps:
[0006] Step S1: collecting real-time parameters during the operation of the central air conditioner, including return air temperature, supply air temperature, supply air volume, cooling capacity, heating capacity, fan speed, compressor frequency, humidifier working status, dehumidifier working status, feedback values of temperature sensors in each area, user behavior, and output signals of flow monitoring devices;
[0007] Step S2: Based on user behavior, the output signal of the flow monitoring device, and the feedback value of the temperature sensor in each area, the input variable set of the behavior-density zoning and emotional load correction model is constructed to divide the human density levels in different spatial areas and determine the initial emotional load coefficient corresponding to each area;
[0008] Step S3: Input the real-time operating parameters of the central air conditioning, the occupant density level, and the initial emotional load coefficient into the experience replay pool of the double-delayed deep deterministic policy gradient algorithm model. A preliminary control action sequence is generated through the actor network of the model. The preliminary control action sequence includes the compressor frequency adjustment amplitude, the fan speed change, and the air supply volume adjustment value.
[0009] Step S4: Using the behavior-density zoning and emotional load correction model, the preliminary control action sequence is corrected. According to the changes in the density level of people in each area and the real-time temperature fluctuations, the emotional load coefficient is dynamically updated to obtain the corrected control action sequence.
[0010] Step S5: Apply the modified control action sequence to the central air-conditioning actuator, and collect the real-time operating parameters of the central air-conditioning and user behavior feedback data after execution, and store them in the experience replay pool;
[0011] Step S6: Evaluate the data in the experience replay pool through the Critic network of the double-delay deep deterministic policy gradient algorithm model, update the parameters of the Actor network and the Critic network, and repeat steps S3 to S5 to achieve closed-loop intelligent control of the central air conditioner.
[0012] Furthermore, when constructing the behavior-density zoning and emotional load correction model in step S2, the personnel density level classification adopts the following formula:
[0013]
[0014] Among them, D i is the population density level of the ith area, α and β are weight coefficients, N i is the real-time number of people in the i-th area, A i is the area of the i-th region, ΔT i is the temperature change of the i-th region within the time interval Δt;
[0015] The initial emotional load coefficient is determined using the following formula:
[0016] E i0 =γ×(T i -T set )+δ×D i
[0017] Among them, E i0 is the initial emotional load coefficient of the i-th region, γ and δ are correction coefficients, T i is the real-time temperature of the ith region, T set Set the temperature for the area;
[0018] In step S4, when dynamically updating the emotional load coefficient, the cooling capacity Q of the central air conditioner is combined with the c And heating capacity Q h Make adjustments when Q c >0, When Q h >0, Among them E i is the updated emotional load coefficient, ε and ζ are adjustment coefficients, Q cmax is the maximum cooling capacity of the central air conditioner, Q hmax This is the maximum heating capacity of the central air conditioner.
[0019] Furthermore, in step S3, during the experience replay pool data processing of the double-delayed deep deterministic policy gradient algorithm model, the central air conditioning supply air volume L and return air temperature T are introduced when the Actor network generates the preliminary control action sequence. r As a constraint, the following formula is used to optimize the action output:
[0020] a t =μ(s t ∣θ μ )×[1+η×(LL set ) / L set ]×[1+λ×(T r -T rset ) / T rset ]
[0021] Among them, a t is the optimized control action, μ(s t ∣θ μ ) is the basic action output by the Actor network, θ μ is the Actor network parameter, η and λ are adjustment coefficients, L is the real-time air supply volume, L set To set the air supply volume, T r is the real-time return air temperature, T rset To set the return air temperature;
[0022] Step S3 also includes prioritizing the data in the experience replay pool, assigning corresponding priority weights to different data samples based on the temperature adjustment effect after the control action is executed and the importance of the user behavior feedback data, and preferentially extracting high-priority data samples for network parameter updating.
[0023] Furthermore, in step S4, when the preliminary control action sequence is corrected using the behavior-density partitioning and emotional load correction model, the following formula is used to correct the compressor frequency adjustment amplitude:
[0024] Δf i ′=Δf i×(1+κ×ΔE i )×(1+ω×D i / D max )
[0025] Where Δf i ′ is the frequency adjustment amplitude of the compressor corresponding to the i-th region after correction, Δf i is the initial compressor frequency adjustment amplitude, σ and ω are correction parameters, ΔE i is the change in the emotional load coefficient of the i-th region, D max is the maximum personnel density level;
[0026] Step S4 also includes a secondary correction of the air supply volume adjustment value based on the deviation between the real-time temperature of each area and the set temperature. When the deviation value is positive, the air supply volume adjustment value is increased proportionally. When the deviation value is negative, the air supply volume adjustment value is reduced proportionally. The correction proportional coefficient is positively correlated with the absolute value of the deviation value.
[0027] Furthermore, the real-time operating parameters of the central air conditioner collected after execution in step S5 include compressor operating power, evaporator inlet and outlet temperatures, condenser inlet and outlet temperatures, humidifier water consumption, and temperature uniformity index of each area;
[0028] The user behavior feedback data in step S5 includes changes in the time people stay in each area and changes in their location and movement trajectory; different parameters and data are associated and stored according to timestamps to form a complete data chain including control actions, execution results, and feedback data. When stored in the experience replay pool, a hierarchical storage structure is adopted, and data storage units are divided by area. Data is arranged in chronological order within each unit.
[0029] Furthermore, in step S6, when the critic network evaluates the data in the experience replay pool, the following evaluation formula is used:
[0030] Q(s t , a t ∣θ Q )=Q0(s t , a t ∣θ Q )+ξ×(Q c -Q cmin ) / (Q cmax -Q cmin )+v×(Q h -Q hmin ) / (Q hmax -Q hmin )
[0031] Among them, Q(s t , a t ∣θ Q) is the revised evaluation value, Q0(s t , a t ∣θ Q ) is the initial evaluation value, θ Q is the Critic network parameter, ξ and v are evaluation coefficients, Q cmin is the minimum cooling capacity, Q cmax is the maximum cooling capacity, Q hmin is the minimum heating capacity, Q hmax is the maximum heating capacity;
[0032] When updating the parameters of the Actor network and the Critic network in step S6, a dual-delay update mechanism is adopted, and different update frequencies are set. The update frequency of the Actor network is lower than that of the Critic network. A soft update strategy of the target network parameters is introduced during the update process, and the degree of approximation of the current network parameters to the target network parameters is controlled by setting the update rate parameters.
[0033] Furthermore, step S3 includes the following sub-steps:
[0034] S3.1: Convert the real-time central air conditioning operating parameters, including return air temperature, supply air temperature, cooling capacity, heating capacity, fan speed, and compressor frequency, into standardized feature vectors. These vectors, along with the occupant density level and initial emotional load coefficient, form a multidimensional input matrix, which is then fed into the experience replay pool of the double-delayed deep deterministic policy gradient algorithm model, ensuring that the input data dimensions match the model's required input layer dimensions.
[0035] S3.2: The experience replay pool divides the input multi-dimensional input matrix into multiple data segments according to the preset time window length. Each data segment contains all parameter information within the time window.
[0036] S3.3: The actor network of the double-delayed deep deterministic policy gradient algorithm model reads data fragments from the experience replay pool and, through calculations in a multi-layer neural network, generates a preliminary control action sequence, including the compressor frequency adjustment amplitude, fan speed change, and air supply volume adjustment value. The activation function of each network layer uses the ReLU function.
[0037] S3.4: Limit the range of the generated preliminary control action sequence, limit the compressor frequency adjustment range to ±20% of the rated frequency, limit the fan speed change to ±15% of the rated speed, and limit the air supply volume adjustment value to ±30% of the rated air supply volume.
[0038] Furthermore, step S4 includes the following sub-steps:
[0039] S4.1: The behavior-density zoning and emotional load correction model receives the preliminary control action sequence and the real-time occupant density level and temperature data of each area, and calculates the rate of change of the occupant density level of each area. The rate of change is the ratio of the difference between the occupant density level at the current moment and the previous moment to the time interval;
[0040] S4.2: Dynamically adjust the update rate of the emotional load coefficient based on the calculated rate of change of the occupant density level. The greater the rate of change, the higher the update rate, and vice versa. Calculate the real-time emotional load coefficient for each area based on this update rate.
[0041] S4.3: Substitute the real-time emotional load coefficient into the correction formula and correct each parameter in the preliminary control action sequence one by one. First, correct the compressor frequency adjustment amplitude, then correct the fan speed change, and finally correct the air supply volume adjustment value.
[0042] S4.4: After the correction is completed, the corrected control action sequence is checked for consistency to ensure that the change in the control action at two adjacent moments is within the preset smoothing threshold range. If it exceeds the range, smoothing is performed.
[0043] Furthermore, step S5 includes the following sub-steps:
[0044] S5.1: Convert the modified control action sequence into electrical signals recognizable by the central air-conditioning actuators, and send them to the compressor, fan, and air outlet adjustment device, respectively, to control these actuators to operate according to the control action sequence;
[0045] S5.2: During the operation of the actuator, sensors installed on various components of the central air conditioner collect real-time operating parameters such as compressor operating frequency, actual fan speed, actual air supply volume, real-time temperature of each zone, return air temperature, and supply air temperature. The collection interval is 10 seconds.
[0046] S5.3: Collect user behavior and mobility monitoring devices to collect feedback data on movement trajectories, dwell time, and gathering status in various areas. Associate this data with the central air conditioning operating parameters at the same time.
[0047] S5.4: Organize the associated marked operating parameters and behavioral feedback data into data records in chronological order, each record including a timestamp, all operating parameter values, and all behavioral feedback data, and store them in a designated storage area of the experience replay pool.
[0048] The central air-conditioning intelligent control system based on user behavior prediction includes:
[0049] The parameter acquisition and preprocessing integrated unit has its input connected to the central air conditioner's temperature sensor, pressure sensor, flow sensor, user behavior, flow monitoring device, compressor operation status monitor, and fan operation status monitor, and its output connected to the behavior-density partition parameter generation unit to collect and integrate various real-time parameters.
[0050] A behavior-density partition parameter generation unit, the input end of which is connected to the output end of the parameter acquisition and preprocessing integration unit, and the output end of which is connected to the control strategy generation and correction unit, and is used to divide the personnel density level according to the input parameters and determine the initial emotional load coefficient;
[0051] A double-delayed deep deterministic policy gradient algorithm operation unit, the input end of which is connected to the output end of the parameter acquisition and preprocessing integration unit and the output end of the behavior-density partition parameter generation unit, and the output end is connected to the control strategy generation and correction unit to generate a preliminary control action sequence;
[0052] A control strategy generation and correction unit, the input end of which is respectively connected to the output end of the dual-delay deep deterministic policy gradient algorithm operation unit and the output end of the behavior-density partition parameter generation unit, and the output end is connected to the actuator drive unit, for correcting the preliminary control action sequence and generating the final control strategy;
[0053] An actuator drive unit, whose input is connected to the output of the control strategy generation and correction unit, and whose output is connected to the compressor, fan, air outlet adjustment device, humidifier, and dehumidifier of the central air conditioner, and is used to convert the control strategy into a drive signal;
[0054] The data storage and network parameter update unit has its input connected to the output of the parameter acquisition and preprocessing integrated unit and the output of the actuator drive unit, and its output connected to the dual-delay deep deterministic policy gradient algorithm operation unit, and is used to store operating data and update algorithm model parameters.
[0055] Beneficial Effects: The present invention proposes a central air conditioning intelligent control method and system based on user behavior prediction. By integrating a dual-delay deep deterministic policy gradient algorithm model with a behavior-density zoning and emotional load correction model, a closed-loop intelligent control mechanism is formed, which has many significant beneficial effects. By collecting central air conditioning operating parameters and personnel behavior data in real time, using the behavior-density zoning model to dynamically divide regions and adjust the emotional load coefficient, the control strategy can accurately adapt to changes in personnel density and temperature requirements, solving the problem of regional environmental imbalance caused by ignoring the dynamic characteristics of behavior in traditional control. At the same time, relying on the parameter iterative update mechanism of the dual-delay deep deterministic policy gradient algorithm, the system enhances its adaptability to the strong coupling and nonlinear characteristics of the system, reducing overshoot and lag in control actions. This method uses behavior-density zoning and emotional load correction models to associate personnel flow, density level and temperature fluctuations, and dynamically corrects control parameters such as compressor frequency and fan speed to avoid energy waste caused by overcooling or overheating. To address the problem of insufficient algorithm adaptability, the method uses the experience replay and network parameter update of the double-delay deep deterministic policy gradient algorithm to achieve continuous optimization of the control strategy, thereby improving system energy efficiency while ensuring thermal comfort in each area, and forming a precise linkage between behavior prediction and control execution. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 is a flow chart of the method steps of the present invention;
[0057] Figure 2 It is a diagram of the system unit composition of the present invention. DETAILED DESCRIPTION
[0058] It should be noted that, unless there is a conflict, the embodiments in this application and the features described in the embodiments can be combined with each other. The application is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0059] like Figure 1 As shown, the central air-conditioning intelligent control method based on user behavior prediction includes the following steps:
[0060] Step S1: collecting real-time parameters during the operation of the central air conditioner, including return air temperature, supply air temperature, supply air volume, cooling capacity, heating capacity, fan speed, compressor frequency, humidifier working status, dehumidifier working status, feedback values of temperature sensors in each area, user behavior, and output signals of flow monitoring devices;
[0061] Specifically, step S1 is the data foundation of the entire intelligent control method. Its core lies in the comprehensive and accurate collection of various real-time parameters related to the control of the central air-conditioning system during operation. These parameters cover multiple dimensions. The return air temperature and supply air temperature directly reflect the output effect of the air-conditioning system and its interaction with the environment. The supply air volume, cooling capacity, and heating capacity reflect the energy output capacity of the system. The fan speed and compressor frequency are related to the energy consumption level of the system. The working status of the humidifier and dehumidifier affects the indoor humidity environment. The feedback values of the temperature sensors in each area can reflect the actual environmental conditions of different spaces. The output signals of the personnel flow monitoring device provide a basis for subsequent behavioral analysis. The collection of these parameters provides raw data support for the calculation of subsequent models and the generation of control strategies, which is the prerequisite for achieving precise control.
[0062] During the specific implementation process, parameters are collected through sensors installed in key parts of the central air conditioner. The return air temperature is collected by a temperature sensor installed in the return air outlet, with a collection range of 16-30℃ and an accuracy of ±0.5℃; the supply air temperature is obtained by a temperature sensor in the supply air outlet, also maintaining an accuracy of ±0.5℃. The supply air volume is measured by an air volume sensor with a measurement range of 200-2000m 3 / h, with an accuracy of ±5%. The cooling and heating capacities are collected through energy metering devices, with ranges of 5-100kW and 3-80kW respectively. The fan speed is monitored by a speed sensor with a range of 800-1500r / min, and the compressor frequency is collected by a frequency sensor with a range of 30-100Hz. The working status of the humidifier and dehumidifier is determined by the current sensor of its control circuit. If there is current output, it is in working state, and if there is no current output, it is in stopped state. The temperature sensors in each area are arranged at a density of one per 50 square meters, and the feedback value range is 18-32℃. The personnel flow monitoring device uses an infrared sensor or camera to output a signal reflecting the number and movement of people, with a sampling interval of 10 seconds.
[0063] Step S2: Based on user behavior, the output signal of the flow monitoring device, and the feedback value of the temperature sensor in each area, the input variable set of the behavior-density zoning and emotional load correction model is constructed to divide the human density levels in different spatial areas and determine the initial emotional load coefficient corresponding to each area;
[0064] Specifically, step S2 combines human behavior with spatial regional characteristics, providing a zoning basis and load correction foundation for subsequent control strategy generation. Constructing an input variable set based on the output signals of the human flow monitoring device and the feedback values from the temperature sensors in each area links human activity to ambient temperature. Categorizing human density levels clarifies the degree of human concentration in different areas, while determining the initial emotional load coefficient takes into account the differences in people's subjective temperature perception in different environments. This step lays the foundation for zoning and load assessment for precise control based on user behavior, making subsequent control actions more targeted.
[0065] In specific implementation, the output signals from the crowd flow monitoring device are first processed to extract the real-time occupancy information for each area. Feedback from the temperature sensors in each area is also aggregated, and these two data sets are integrated into the input variable set for the behavior-density zoning and emotional load correction model. Density levels are categorized based on the ratio of occupancy to area: 0-0.1 people / m2 is considered low density, 0.1-0.3 people / m2 is medium density, and 0.3 people / m2 or higher is high density. The initial emotional load coefficient is determined based on the deviation between the real-time temperature and the set temperature in each area. When the real-time temperature is above the set temperature, the deviation is positive, and the initial emotional load coefficient increases with the deviation. When the real-time temperature is below the set temperature, the deviation is negative, and the initial emotional load coefficient increases with the absolute value of the deviation. The coefficient ranges from 0.8 to 1.5, with 1.0 as the baseline value. After zoning, the density level and initial emotional load coefficient for each area are stored and serve as key parameters for the next model input.
[0066] Step S3: Input the real-time operating parameters of the central air conditioning, the occupant density level, and the initial emotional load coefficient into the experience replay pool of the double-delayed deep deterministic policy gradient algorithm model. A preliminary control action sequence is generated through the actor network of the model. The preliminary control action sequence includes the compressor frequency adjustment amplitude, the fan speed change, and the air supply volume adjustment value.
[0067] Specifically, step S3 is the key link between data acquisition and control strategy generation. Its core lies in using a dual-delay deep deterministic policy gradient algorithm model to process input data and generate preliminary control actions. Inputting the real-time operating parameters of the central air conditioner, the occupant density level, and the initial emotional load coefficient into the experience replay pool enables data accumulation and reuse, providing material for learning and optimization of the algorithm model. The preliminary control action sequence generated by the Actor network is directly related to the core operating components of the central air conditioner. The reasonable generation of the compressor frequency adjustment amplitude, fan speed change, and air supply volume adjustment value is the first step in achieving system regulation. This step determines the general direction of the control strategy and has a significant impact on subsequent corrections and execution.
[0068] During the implementation process, the real-time parameters of the central air conditioner collected in step S1, including the return air temperature, supply air temperature, etc., and the personnel density level and initial emotional load coefficient obtained in step S2 are first sorted according to the preset data format and input into the experience replay pool of the double-delay deep deterministic policy gradient algorithm model. The experience replay pool caches the input data, and the cache capacity is set to 10,000 items. The data is stored in chronological order, and the earliest data is overwritten when new data enters. After receiving the data in the experience replay pool, the Actor network generates a preliminary control action sequence through the calculation of a multi-layer neural network. Among them, the calculation of the compressor frequency adjustment amplitude is based on the difference between the current frequency and the target frequency, and the adjustment range is ±10Hz; the fan speed change is determined according to the change in the required air volume, and the range is ±200r / min; the supply air volume adjustment value is calculated based on the regional load change, and the range is ±300m 3 / h. The generated preliminary control action sequence will be temporarily stored and wait for the next step to be corrected.
[0069] Step S4: Using the behavior-density zoning and emotional load correction model, the preliminary control action sequence is corrected. According to the changes in the density level of people in each area and the real-time temperature fluctuations, the emotional load coefficient is dynamically updated to obtain the corrected control action sequence.
[0070] Specifically, the core function of step S4 is to optimize the initial control action sequence to better align it with actual occupant behavior and environmental changes. The application of the behavior-density zoning and emotional load correction model dynamically adjusts the emotional load coefficient based on changes in occupant density levels and real-time temperature fluctuations. This ensures that control actions not only take into account the system's initial state but also adapt to dynamic changes. The resulting control action sequence is more accurate and adaptable than the initial control action sequence, reducing control deviations caused by occupant movement and temperature fluctuations, thereby improving the control effectiveness of the central air conditioning system.
[0071] In implementation, the behavior-density zoning and emotional load correction model receives real-time data on changes in occupant density and temperature fluctuations in each zone. When the occupant density changes from low to medium, the emotional load coefficient is adjusted upward by 0.1-0.2; when it changes from medium to high, it is adjusted upward by 0.2-0.3; conversely, when the density decreases, the emotional load coefficient is adjusted downward accordingly. Furthermore, when the real-time temperature fluctuation in a zone exceeds ±0.5°C, the emotional load coefficient is adjusted based on the direction of the fluctuation: upward if the temperature exceeds the limit, and downward if it exceeds the limit, with an adjustment range of 0.05-0.15. The updated emotional load coefficient is substituted into the correction logic to modify the initial control action sequence. The compressor frequency adjustment range is adjusted proportionally to the coefficient change, with a correction ratio of 50%-80% of the coefficient change; the fan speed change is corrected by a correction ratio of 40%-60% of the coefficient change; and the air volume adjustment value is corrected by a correction ratio of 60%-90% of the coefficient change. Once the correction is complete, a revised control action sequence is generated and ready for input into the next step.
[0072] Step S5: Apply the modified control action sequence to the central air-conditioning actuator, and collect the real-time operating parameters of the central air-conditioning and user behavior feedback data after execution, and store them in the experience replay pool;
[0073] Specifically, step S5 is a crucial step in achieving a closed-loop control system. Its purpose is to implement the revised control actions and collect feedback data from execution, providing a basis for model optimization. Applying the revised control action sequence to the central air conditioning actuators, allowing the system to operate according to the optimized strategy, is a key step in achieving the actual results of the control method. Simultaneously, collecting real-time operating parameters and personnel behavior feedback data from execution and storing them in an experience replay pool enables data recycling, enabling the algorithm model to continuously learn from actual operation and continuously optimize the control strategy, thereby improving the system's control accuracy and adaptability.
[0074] During implementation, the revised control action sequence is first converted into control signals recognizable by the actuators and sent to the central air conditioner's compressor, fan, and air flow control device. The compressor changes its operating frequency based on the frequency adjustment signal, with a response time of 3-5 seconds. After receiving the speed change signal, the fan completes its speed adjustment within 2-4 seconds, and the air flow control device adjusts the air flow within 3-6 seconds. During implementation, real-time operating parameters of the central air conditioner are collected every 5 seconds, including the adjusted return air temperature, supply air temperature, actual cooling capacity, actual heating capacity, actual fan speed, and actual compressor frequency. Simultaneously, personnel behavior feedback data is collected every 10 seconds via a personnel flow monitoring device, including the duration of stay and movement direction of personnel in each area. These parameters and data are chronologically linked to form a complete data record containing control actions, execution results, and feedback. This data is then stored in a designated area of the experience replay pool, categorized by region, with a separate data sub-database established for each region.
[0075] Step S6: Evaluate the data in the experience replay pool through the Critic network of the double-delay deep deterministic policy gradient algorithm model, update the parameters of the Actor network and the Critic network, and repeat steps S3 to S5 to achieve closed-loop intelligent control of the central air conditioner.
[0076] Specifically, step S6 is the core step to ensure continuous optimization of the algorithm model and closed-loop operation of the control method. The critic network of the dual-delay deep deterministic policy gradient algorithm model evaluates data in the experience replay pool, determining the effectiveness and rationality of control actions and providing a basis for updating network parameters. Updating the parameters of the actor and critic networks enables the model to continuously adapt to new operating data and environmental changes, improving the model's predictive accuracy and control capabilities. By repeating steps S3 to S5, closed-loop intelligent control of the central air conditioner is achieved, ensuring that the system maintains good operating status and control effects over the long term.
[0077] In specific implementation, the Critic network extracts the most recent 2000 data records from the experience replay pool and evaluates the control actions, post-execution operating parameters, and feedback data in each record. Evaluation metrics include temperature control accuracy, energy consumption change rate, and operator behavior stability. An evaluation value is generated based on the evaluation results. Based on the evaluation values, the parameters of the Actor and Critic networks are updated every 500 data records collected. During the update, the parameter adjustment direction is determined based on the Critic network's evaluation results. Then, the weight parameters related to control action generation in the Actor network and the weight parameters related to evaluation calculation in the Critic network are adjusted according to a preset learning rate, set between 0.001 and 0.01. After the parameter update is complete, step S3 is executed again, using the updated network to generate a new control action sequence. Step S4 is then followed by correction, and step S5 is executed and data collected. This cycle repeats, forming a continuous closed-loop control process, with each cycle lasting 30 seconds.
[0078] Preferably, when constructing the behavior-density zoning and emotional load correction model in step S2, the personnel density level division adopts the following formula:
[0079]
[0080] Among them, D i is the population density level of the ith area, α and β are weight coefficients, N i is the real-time number of people in the i-th area, A i is the area of the i-th region, ΔT i is the temperature change of the i-th region within the time interval Δt;
[0081] The initial emotional load coefficient is determined using the following formula:
[0082] E i0 =γ×(T i -T set )+δ×D i
[0083] Among them, E i0 is the initial emotional load coefficient of the i-th region, γ and δ are correction coefficients, T i is the real-time temperature of the ith region, T set Set the temperature for the area;
[0084] In step S4, when dynamically updating the emotional load coefficient, the cooling capacity Q of the central air conditioner is combined with the c And heating capacity Q h Make adjustments when Q c >0, When Qh >0, Among them E i is the updated emotional load coefficient, ε and ζ are adjustment coefficients, Q cmax is the maximum cooling capacity of the central air conditioner, Q hmax This is the maximum heating capacity of the central air conditioner.
[0085] Specifically, the application of the behavior-density zoning and emotional load correction model is explored. Its significance lies in the ability to more precisely quantify occupant density levels and determine emotional load coefficients, providing a reliable basis for the generation of subsequent control actions. During implementation, the weighting coefficients for occupant density grading are set based on different building types and usage scenarios. Generally, the ratio of occupant number to area is weighted slightly higher than the temperature change rate, with values ranging from 0.6-0.8 and 0.2-0.4, respectively. The real-time occupant count is obtained through infrared sensors or camera counts, while the area of the area is a preset fixed value. The time interval is typically set to 5 minutes, and the temperature change is calculated by the temperature difference between two previous and subsequent time points. When determining the initial emotional load coefficient, the correction coefficient is also adjusted based on actual needs. The deviation between the real-time temperature and the set temperature is weighted less, ranging from 0.3-0.5, while the occupant density level is weighted more, ranging from 0.5-0.7. The set temperature is the user's pre-set comfort level, ranging from 22-26°C. When the emotional load coefficient is dynamically updated, the adjustment coefficient is set according to the cooling and heating modes respectively. The value is 0.1-0.3 for cooling and 0.15-0.35 for heating. The maximum cooling capacity and heating capacity are the rated parameters of the central air conditioner and are determined according to the equipment model.
[0086] Preferably, in the experience replay pool data processing process of the double-delayed deep deterministic policy gradient algorithm model in step S3, when the Actor network generates the preliminary control action sequence, the air supply volume L and return air temperature T of the central air conditioner are introduced. r As a constraint, the following formula is used to optimize the action output:
[0087] a t =μ(s t ∣θ μ )×[1+η×(LL set ) / L set ]×[1+λ×(T r -T rset ) / T rset ]
[0088] Among them, a t is the optimized control action, μ(s t ∣θ μ ) is the basic action output by the Actor network, θ μis the Actor network parameter, η and λ are adjustment coefficients, L is the real-time air supply volume, L set To set the air supply volume, T r is the real-time return air temperature, T rset To set the return air temperature;
[0089] Step S3 also includes prioritizing the data in the experience replay pool, assigning corresponding priority weights to different data samples based on the temperature adjustment effect after the control action is executed and the importance of the user behavior feedback data, and preferentially extracting high-priority data samples for network parameter updating.
[0090] Specifically, the preliminary control action sequence generated by the double-delay deep deterministic policy gradient algorithm model is optimized, and the rationality of the control action and the efficiency of model learning are improved by introducing constraints and data priority sorting. During implementation, the adjustment coefficient is dynamically adjusted according to the operating status of the central air conditioner. When the supply air volume deviates greatly from the set value, the adjustment coefficient related to the supply air volume is 0.2-0.4. When the return air temperature deviates greatly, the adjustment coefficient related to the return air temperature is 0.25-0.45. The supply air volume and return air temperature are set according to the regional function and the number of personnel. For example, the supply air volume in the office area is set to 500-800m 3 / h, with a return air temperature of 24°C. Basic actions are calculated by the Actor network based on input data, and network parameters are optimized through multiple training iterations. When prioritizing data in the experience replay pool, weights are determined based on the degree of deviation in temperature control effects and the significance of employee behavior feedback. Data samples with smaller temperature control effect deviations and more significant employee behavior feedback receive higher priority weights, ranging from 0.6 to 1.0. Other data samples have weights ranging from 0.1 to 0.5. High-priority data samples are preferentially extracted for network parameter updates, with each extraction accounting for 20% to 30% of the total sample size.
[0091] Preferably, when the behavior-density partitioning and emotional load correction model is used to correct the preliminary control action sequence in step S4, the following formula is used to correct the compressor frequency adjustment amplitude:
[0092] Δf i ′=Δf i ×(1+κ×ΔE i )×(1+ω×D i / D max )
[0093] Where Δf i ′ is the frequency adjustment amplitude of the compressor corresponding to the i-th region after correction, Δf i is the initial compressor frequency adjustment amplitude, κ and ω are correction parameters, ΔE iis the change in the emotional load coefficient of the i-th region, D max is the maximum personnel density level;
[0094] Step S4 also includes a secondary correction of the air supply volume adjustment value based on the deviation between the real-time temperature of each area and the set temperature. When the deviation value is positive, the air supply volume adjustment value is increased proportionally. When the deviation value is negative, the air supply volume adjustment value is reduced proportionally. The correction proportional coefficient is positively correlated with the absolute value of the deviation value.
[0095] Specifically, the initial control action sequence is precisely modified using a behavior-density zoning and emotional load correction model to ensure that the control actions align with the actual conditions in the area. During implementation, correction parameters are set based on the area's importance and the range of occupant density. The correction parameter for the emotional load coefficient variation ranges from 0.1 to 0.3, and the correction parameter for the occupant density level ranges from 0.15 to 0.35. The maximum occupant density level is calculated based on the area's maximum occupancy and area. For example, the maximum occupant density level for conference rooms is 1.0, and for corridors is 0.3. When performing a secondary correction on the air flow adjustment value, the deviation is the difference between the actual temperature and the setpoint temperature. When the deviation is positive, the correction coefficient increases from 0.1 to 0.5 as the absolute value of the deviation increases. When the deviation is negative, the correction coefficient increases from 0.1 to 0.4 as the absolute value of the deviation increases. The correction sequence is strictly based on compressor frequency, fan speed, and air flow, ensuring coordinated parameter adjustments and avoiding system fluctuations caused by improper parameter adjustment sequence.
[0096] Preferably, the real-time operating parameters of the central air conditioner collected after execution in step S5 include compressor operating power, evaporator inlet and outlet temperatures, condenser inlet and outlet temperatures, humidifier water consumption, and temperature uniformity index of each area;
[0097] The user behavior feedback data in step S5 includes changes in the time people stay in each area and changes in their location and movement trajectory; different parameters and data are associated and stored according to timestamps to form a complete data chain including control actions, execution results, and feedback data. When stored in the experience replay pool, a hierarchical storage structure is adopted, and data storage units are divided by area. Data is arranged in chronological order within each unit.
[0098] Specifically, the data collection and storage methods after implementation are standardized to provide high-quality data support for continuous model optimization. During implementation, compressor operating power was collected in the range of 5-50kW, monitored in real time via power sensors. Evaporator inlet and outlet temperatures ranged from 5-15°C, and condenser inlet and outlet temperatures ranged from 30-45°C, both collected via temperature sensors. Humidifier water consumption was determined based on humidifier type, with electrode-type humidifiers consuming 0.5-2 L / h. Temperature uniformity was calculated by calculating the temperature difference between multiple temperature sensors within a region, with a requirement of no more than ±1°C. Changes in dwell time in personnel behavioral feedback data were recorded using a timing device, while changes in movement trajectory were captured using cameras and trajectory analysis algorithms. Data storage was precisely linked by timestamp, with timestamp accuracy down to the second level. A hierarchical storage structure was employed, with each region's data storage unit independent and data within the unit arranged chronologically. The storage format was binary for fast read and write access. The storage period was set to 7-30 days, depending on requirements; data exceeding the period was archived.
[0099] Preferably, in step S6, when the critic network evaluates the data in the experience replay pool, the following evaluation formula is used:
[0100] Q(s t , a t ∣θ Q )=Q0(s t , a t ∣θ Q )+ξ×(Q c -Q cmin ) / (Q cmax -Q cmin )+ν×(Q h -Q hmin ) / (Q hmax -Q hmin )
[0101] Among them, Q(s t , a t ∣θ Q ) is the revised evaluation value, Q0(s t , a t ∣θ Q ) is the initial evaluation value, θ Q is the Critic network parameter, ξ and v are evaluation coefficients, Q cmin is the minimum cooling capacity, Q cmax is the maximum cooling capacity, Q hmin is the minimum heating capacity, Q hmax is the maximum heating capacity;
[0102] When updating the parameters of the Actor network and the Critic network in step S6, a dual-delay update mechanism is adopted, and different update frequencies are set. The update frequency of the Actor network is lower than that of the Critic network. A soft update strategy of the target network parameters is introduced during the update process, and the degree of approximation of the current network parameters to the target network parameters is controlled by setting the update rate parameters.
[0103] Specifically, a critic network accurately evaluates data and updates network parameters appropriately, ensuring model learning effectiveness and control performance stability. During implementation, evaluation coefficients are set based on the actual output range of cooling and heating capacity. The cooling coefficient ranges from 0.2 to 0.35, while the heating coefficient ranges from 0.25 to 0.4. The minimum and maximum cooling and heating capacities are the rated parameters of the central air conditioner. For example, for a certain central air conditioner model, the minimum cooling capacity is 5kW and the maximum is 50kW, and the minimum heating capacity is 3kW and the maximum is 40kW. Initial evaluation values are calculated by the critic network using a basic algorithm, and the revised evaluation values serve as an important basis for network parameter updates. When updating network parameters, the actor network uses a double-delay update mechanism, with an update frequency of once every 10 critic network updates. The update rate parameter ranges from 0.001 to 0.01, controlling the speed at which the current network parameters approach the target network parameters. This ensures smooth parameter updates and avoids model oscillation caused by overly rapid updates. During each update, the parameter error is calculated first, and then the network weights are adjusted according to the error and update rate. After the parameter update is completed, the model performance is tested to ensure that it meets the control requirements.
[0104] Preferably, step S3 includes the following sub-steps:
[0105] S3.1: Convert the real-time central air conditioning operating parameters, including return air temperature, supply air temperature, cooling capacity, heating capacity, fan speed, and compressor frequency, into standardized feature vectors. These vectors, along with the occupant density level and initial emotional load coefficient, form a multidimensional input matrix, which is then fed into the experience replay pool of the double-delayed deep deterministic policy gradient algorithm model, ensuring that the input data dimensions match the model's required input layer dimensions.
[0106] S3.2: The experience replay pool divides the input multi-dimensional input matrix into multiple data segments according to the preset time window length. Each data segment contains all parameter information within the time window.
[0107] S3.3: The actor network of the double-delayed deep deterministic policy gradient algorithm model reads data fragments from the experience replay pool and, through calculations in a multi-layer neural network, generates a preliminary control action sequence, including the compressor frequency adjustment amplitude, fan speed change, and air supply volume adjustment value. The activation function of each network layer uses the ReLU function.
[0108] S3.4: Limit the range of the generated preliminary control action sequence, limit the compressor frequency adjustment range to ±20% of the rated frequency, limit the fan speed change to ±15% of the rated speed, and limit the air supply volume adjustment value to ±30% of the rated air supply volume.
[0109] Preferably, step S4 includes the following sub-steps:
[0110] S4.1: The behavior-density zoning and emotional load correction model receives the preliminary control action sequence and the real-time occupant density level and temperature data of each area, and calculates the rate of change of the occupant density level of each area. The rate of change is the ratio of the difference between the occupant density level at the current moment and the previous moment to the time interval;
[0111] S4.2: Dynamically adjust the update rate of the emotional load coefficient based on the calculated rate of change of the occupant density level. The greater the rate of change, the higher the update rate, and vice versa. Calculate the real-time emotional load coefficient for each area based on this update rate.
[0112] S4.3: Substitute the real-time emotional load coefficient into the correction formula and correct each parameter in the preliminary control action sequence one by one. First, correct the compressor frequency adjustment amplitude, then correct the fan speed change, and finally correct the air supply volume adjustment value.
[0113] S4.4: After the correction is completed, the corrected control action sequence is checked for consistency to ensure that the change in the control action at two adjacent moments is within the preset smoothing threshold range. If it exceeds the range, smoothing is performed.
[0114] Preferably, step S5 includes the following sub-steps:
[0115] S5.1: Convert the modified control action sequence into electrical signals recognizable by the central air-conditioning actuators, and send them to the compressor, fan, and air outlet adjustment device, respectively, to control these actuators to operate according to the control action sequence;
[0116] S5.2: During the operation of the actuator, sensors installed on various components of the central air conditioner collect real-time operating parameters such as compressor operating frequency, actual fan speed, actual air supply volume, real-time temperature of each zone, return air temperature, and supply air temperature. The collection interval is 10 seconds.
[0117] S5.3: Collect user behavior and mobility monitoring devices to collect feedback data on movement trajectories, dwell time, and gathering status in various areas. Associate this data with the central air conditioning operating parameters at the same time.
[0118] S5.4: Organize the associated marked operating parameters and behavioral feedback data into data records in chronological order, each record including a timestamp, all operating parameter values, and all behavioral feedback data, and store them in a designated storage area of the experience replay pool.
[0119] like Figure 2 As shown in FIG, the central air-conditioning intelligent control system based on user behavior prediction includes:
[0120] The parameter acquisition and preprocessing integrated unit has its input connected to the central air conditioner's temperature sensor, pressure sensor, flow sensor, user behavior, flow monitoring device, compressor operation status monitor, and fan operation status monitor, and its output connected to the behavior-density partition parameter generation unit to collect and integrate various real-time parameters.
[0121] A behavior-density partition parameter generation unit, the input end of which is connected to the output end of the parameter acquisition and preprocessing integration unit, and the output end of which is connected to the control strategy generation and correction unit, and is used to divide the personnel density level according to the input parameters and determine the initial emotional load coefficient;
[0122] A double-delayed deep deterministic policy gradient algorithm operation unit, the input end of which is connected to the output end of the parameter acquisition and preprocessing integration unit and the output end of the behavior-density partition parameter generation unit, and the output end is connected to the control strategy generation and correction unit to generate a preliminary control action sequence;
[0123] A control strategy generation and correction unit, the input end of which is respectively connected to the output end of the dual-delay deep deterministic policy gradient algorithm operation unit and the output end of the behavior-density partition parameter generation unit, and the output end is connected to the actuator drive unit, for correcting the preliminary control action sequence and generating the final control strategy;
[0124] An actuator drive unit, whose input is connected to the output of the control strategy generation and correction unit, and whose output is connected to the compressor, fan, air outlet adjustment device, humidifier, and dehumidifier of the central air conditioner, and is used to convert the control strategy into a drive signal;
[0125] The data storage and network parameter update unit has its input connected to the output of the parameter acquisition and preprocessing integrated unit and the output of the actuator drive unit, and its output connected to the dual-delay deep deterministic policy gradient algorithm operation unit, and is used to store operating data and update algorithm model parameters.
[0126] A central air conditioning intelligent control method and system based on user behavior prediction accurately adapts to the dynamic characteristics of human behavior. Using a behavior-density zoning and emotional load correction model, it divides density levels based on occupant flow monitoring signals and regional temperature feedback, and dynamically updates the emotional load coefficient. This mechanism breaks the limitations of traditional fixed zoning, allowing control actions to adjust in real time as occupant distribution changes, avoiding regional heat and cold imbalances caused by occupant movement and overcoming the shortcomings of existing technologies that lack dynamic integration of occupant behavior.
[0127] In terms of algorithmic performance, the application of a dual-delayed deep deterministic policy gradient algorithm model improves control accuracy and adaptability. Its experience replay pool stores and evaluates historical data, achieving closed-loop control through parameter updates in the actor and critic networks. This model can handle the strong coupling and nonlinear characteristics of central air conditioning systems, reducing overshoot and lag during parameter adjustment. This addresses the lack of adaptability of traditional algorithms and ensures that control actions are more closely aligned with the system's actual operating conditions.
[0128] The various units in the system work together to form a highly efficient control system. The parameter acquisition unit integrates multiple types of real-time data, the behavior-density partitioning unit provides the basis for partitioning, the algorithm calculation unit generates control strategies, the correction unit optimizes action sequences, the execution unit drives device operation, and the storage and update unit ensures model iteration. This structure seamlessly connects data flow and control execution, ensuring efficiency from parameter acquisition to strategy implementation, comprehensively improving the energy efficiency and user experience of central air conditioning systems.
[0129] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "disposed," "installed," "connected," "connected," and "fixed" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; they may refer to mechanical connections or electrical connections; they may refer to direct connections or indirect connections through an intermediate medium; and they may refer to internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0130] While embodiments of the present invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A central air-conditioning intelligent control method based on user behavior prediction, characterized in that: The following steps are involved: Step S1: collecting real-time parameters during the operation of the central air conditioner, including return air temperature, supply air temperature, supply air volume, cooling capacity, heating capacity, fan speed, compressor frequency, humidifier working status, dehumidifier working status, feedback values of temperature sensors in each area, user behavior, and output signals of flow monitoring devices; Step S2: Based on user behavior, the output signal of the flow monitoring device, and the feedback value of the temperature sensor in each area, the input variable set of the behavior-density zoning and emotional load correction model is constructed to divide the human density levels in different spatial areas and determine the initial emotional load coefficient corresponding to each area; Step S3: Input the real-time operating parameters of the central air conditioning, the occupant density level, and the initial emotional load coefficient into the experience replay pool of the double-delayed deep deterministic policy gradient algorithm model. A preliminary control action sequence is generated through the actor network of the model. The preliminary control action sequence includes the compressor frequency adjustment amplitude, the fan speed change, and the air supply volume adjustment value. Step S4: Using the behavior-density zoning and emotional load correction model, the preliminary control action sequence is corrected. According to the changes in the density level of people in each area and the real-time temperature fluctuations, the emotional load coefficient is dynamically updated to obtain the corrected control action sequence. Step S5: Apply the modified control action sequence to the central air-conditioning actuator, and collect the real-time operating parameters of the central air-conditioning and user behavior feedback data after execution, and store them in the experience replay pool; Step S6: Evaluate the data in the experience replay pool through the Critic network of the double-delay deep deterministic policy gradient algorithm model, update the parameters of the Actor network and the Critic network, and repeat steps S3 to S5 to achieve closed-loop intelligent control of the central air conditioner.
2. The method according to claim 1, characterized in that When constructing the behavior-density zoning and emotional load correction model in step S2, the personnel density level classification adopts the following formula: Among them, D i is the population density level of the ith area, α and β are weight coefficients, N i is the real-time number of people in the i-th area, A i is the area of the i-th region, ΔT i is the temperature change of the i-th region within the time interval Δt; The initial emotional load coefficient is determined using the following formula: E i0 =γ×(T i -T set )+δ×D i Among them, E i0 is the initial emotional load coefficient of the i-th region, γ and δ are correction coefficients, T i is the real-time temperature of the ith region, T set Set the temperature for the area; In step S4, when dynamically updating the emotional load coefficient, the cooling capacity Q of the central air conditioner is combined with the c And heating capacity Q h Make adjustments when Q c >0, When Q h >0, Among them E i is the updated emotional load coefficient, ε and ζ are adjustment coefficients, Q cmax is the maximum cooling capacity of the central air conditioner, Q hmax This is the maximum heating capacity of the central air conditioner.
3. The method according to claim 1, characterized in that In step S3, during the experience replay pool data processing of the double-delayed deep deterministic policy gradient algorithm model, when the Actor network generates the preliminary control action sequence, the air supply volume L and return air temperature T of the central air conditioner are introduced. r As a constraint, the following formula is used to optimize the action output: a t =μ(s t ∣θ μ )×[1+η×(L-L set ) / L set ]×[1+λ×(T r -T rset ) / T rset ] Among them, a t is the optimized control action, μ(s t ∣θ μ ) is the basic action output by the Actor network, θ μ is the Actor network parameter, η and λ are adjustment coefficients, L is the real-time air supply volume, L set To set the air supply volume, T r is the real-time return air temperature, T rset To set the return air temperature; Step S3 also includes prioritizing the data in the experience replay pool, assigning corresponding priority weights to different data samples based on the temperature adjustment effect after the control action is executed and the importance of the user behavior feedback data, and preferentially extracting high-priority data samples for network parameter updating.
4. The method according to claim 1, wherein When the behavior-density partitioning and emotional load correction model is used to correct the preliminary control action sequence in step S4, the following formula is used to correct the compressor frequency adjustment amplitude: Δf i ′=Δf i ×(1+κ×ΔE i )×(1+ω×D i / D max ) Where Δf i ′ is the frequency adjustment amplitude of the compressor corresponding to the i-th region after correction, Δf i is the initial compressor frequency adjustment amplitude, κ and ω are correction parameters, ΔE i is the change in the emotional load coefficient of the i-th region, D max is the maximum personnel density level; Step S4 also includes a secondary correction of the air supply volume adjustment value based on the deviation between the real-time temperature of each area and the set temperature. When the deviation value is positive, the air supply volume adjustment value is increased proportionally. When the deviation value is negative, the air supply volume adjustment value is reduced proportionally. The correction proportional coefficient is positively correlated with the absolute value of the deviation value.
5. The method according to claim 1, wherein The real-time operating parameters of the central air conditioner collected after execution in step S5 include the compressor operating power, evaporator inlet and outlet temperatures, condenser inlet and outlet temperatures, humidifier water consumption, and temperature uniformity index of each area; The user behavior feedback data in step S5 includes changes in the time people stay in each area and changes in their location and movement trajectory; different parameters and data are associated and stored according to timestamps to form a complete data chain including control actions, execution results, and feedback data. When stored in the experience replay pool, a hierarchical storage structure is adopted, and data storage units are divided by area. Data is arranged in chronological order within each unit.
6. The method according to claim 1, characterized in that In step S6, the critic network evaluates the data in the experience replay pool using the following evaluation formula: Q(s t ,a t ∣θ Q )=Q0(s t ,a t ∣θ Q )+ξ×(Q c -Q cmin ) / (Q cmax -Q cmin )+ν×(Q h -Q hmin ) / (Q hmax -Q hmin ) Among them, Q(s t , a t ∣θ Q ) is the revised evaluation value, Q0(s t , a t ∣θ Q ) is the initial evaluation value, θ Q is the Critic network parameter, ξ and ν are evaluation coefficients, Q cmin is the minimum cooling capacity, Q cmax is the maximum cooling capacity, Q hmin is the minimum heating capacity, Q hmax is the maximum heating capacity; When updating the parameters of the Actor network and the Critic network in step S6, a dual-delay update mechanism is adopted, and different update frequencies are set. The update frequency of the Actor network is lower than that of the Critic network. A soft update strategy of the target network parameters is introduced during the update process, and the degree of approximation of the current network parameters to the target network parameters is controlled by setting the update rate parameters.
7. The method according to claim 1, characterized in that Step S3 includes the following sub-steps: S3.1: Convert the real-time central air conditioning operating parameters, including return air temperature, supply air temperature, cooling capacity, heating capacity, fan speed, and compressor frequency, into standardized feature vectors. These vectors, along with the occupant density level and initial emotional load coefficient, form a multidimensional input matrix, which is then fed into the experience replay pool of the double-delayed deep deterministic policy gradient algorithm model, ensuring that the input data dimensions match the model's required input layer dimensions. S3.2: The experience replay pool divides the input multi-dimensional input matrix into multiple data segments according to the preset time window length. Each data segment contains all parameter information within the time window. S3.3: The actor network of the double-delayed deep deterministic policy gradient algorithm model reads data fragments from the experience replay pool and, through calculations in a multi-layer neural network, generates a preliminary control action sequence, including the compressor frequency adjustment amplitude, fan speed change, and air supply volume adjustment value. The activation function of each network layer uses the ReLU function. S3.4: Limit the range of the generated preliminary control action sequence, limit the compressor frequency adjustment range to ±20% of the rated frequency, limit the fan speed change to ±15% of the rated speed, and limit the air supply volume adjustment value to ±30% of the rated air supply volume.
8. The method according to claim 1, characterized in that Step S4 includes the following sub-steps: S4.1: The behavior-density zoning and emotional load correction model receives the preliminary control action sequence and the real-time occupant density level and temperature data of each area, and calculates the rate of change of the occupant density level of each area. The rate of change is the ratio of the difference between the occupant density level at the current moment and the previous moment to the time interval; S4.2: Dynamically adjust the update rate of the emotional load coefficient based on the calculated rate of change of the occupant density level. The greater the rate of change, the higher the update rate, and vice versa. Calculate the real-time emotional load coefficient for each area based on this update rate. S4.3: Substitute the real-time emotional load coefficient into the correction formula and correct each parameter in the preliminary control action sequence one by one. First, correct the compressor frequency adjustment amplitude, then correct the fan speed change, and finally correct the air supply volume adjustment value. S4.4: After the correction is completed, the corrected control action sequence is checked for consistency to ensure that the change in the control action at two adjacent moments is within the preset smoothing threshold range. If it exceeds the range, smoothing is performed.
9. The method according to claim 1, characterized in that Step S5 includes the following sub-steps: S5.1: Convert the modified control action sequence into electrical signals recognizable by the central air-conditioning actuators, and send them to the compressor, fan, and air outlet adjustment device, respectively, to control these actuators to operate according to the control action sequence; S5.2: During the operation of the actuator, sensors installed on various components of the central air conditioner collect real-time operating parameters such as compressor operating frequency, actual fan speed, actual air supply volume, real-time temperature of each zone, return air temperature, and supply air temperature. The collection interval is 10 seconds. S5.3: Collect user behavior and mobility monitoring devices to collect feedback data on movement trajectories, dwell time, and gathering status in various areas. Associate this data with the central air conditioning operating parameters at the same time. S5.4: Organize the associated marked operating parameters and behavioral feedback data into data records in chronological order, each record including a timestamp, all operating parameter values, and all behavioral feedback data, and store them in a designated storage area of the experience replay pool.
10. Central air-conditioning intelligent control system based on user behavior prediction, characterized in that: include: The parameter acquisition and preprocessing integrated unit has its input connected to the central air conditioner's temperature sensor, pressure sensor, flow sensor, user behavior, flow monitoring device, compressor operation status monitor, and fan operation status monitor, and its output connected to the behavior-density partition parameter generation unit to collect and integrate various real-time parameters. A behavior-density partition parameter generation unit, the input end of which is connected to the output end of the parameter acquisition and preprocessing integration unit, and the output end of which is connected to the control strategy generation and correction unit, and is used to divide the personnel density level according to the input parameters and determine the initial emotional load coefficient; A double-delayed deep deterministic policy gradient algorithm operation unit, the input end of which is connected to the output end of the parameter acquisition and preprocessing integration unit and the output end of the behavior-density partition parameter generation unit, and the output end is connected to the control strategy generation and correction unit to generate a preliminary control action sequence; A control strategy generation and correction unit, the input end of which is respectively connected to the output end of the dual-delay deep deterministic policy gradient algorithm operation unit and the output end of the behavior-density partition parameter generation unit, and the output end is connected to the actuator drive unit, for correcting the preliminary control action sequence and generating the final control strategy; An actuator drive unit, whose input is connected to the output of the control strategy generation and correction unit, and whose output is connected to the compressor, fan, air outlet adjustment device, humidifier, and dehumidifier of the central air conditioner, and is used to convert the control strategy into a drive signal; The data storage and network parameter update unit has its input connected to the output of the parameter acquisition and preprocessing integrated unit and the output of the actuator drive unit, and its output connected to the dual-delay deep deterministic policy gradient algorithm operation unit, and is used to store operating data and update algorithm model parameters.
Citation Information
Cited By
Hot water kettle control method based on change of Internet of Things and hot water kettle
CN121300192A