Air conditioner control method, system, device and medium based on deep reinforcement learning

By training the air conditioning strategy scheduling model with deep reinforcement learning, the problems of slow response and insufficient applicability of air conditioning systems in complex environments are solved, enabling rapid adjustment and efficient control, and improving the flexibility and adaptability of air conditioning systems.

CN119532918BActive Publication Date: 2025-12-26WUYI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411430488.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-14
Publication Date
2025-12-26
Estimated Expiration
2044-10-14

AI Technical Summary

Technical Problem

Existing air conditioning control systems are slow to respond to complex and ever-changing indoor and outdoor environments, making it difficult to quickly adjust to the optimal state. They also lack applicability in different buildings or environments, and are inflexible and inefficient in energy management.

Method used

An air conditioning control method based on deep reinforcement learning is adopted. By acquiring environmental state data, an initial air conditioning strategy scheduling model is trained, parameters are iteratively optimized, and an efficient and highly adaptable air conditioning control strategy is generated to adjust the air conditioning operation status in real time to adapt to environmental changes.

Benefits of technology

It significantly improves the response speed and environmental adaptability of the air conditioning system, ensuring rapid adjustment to the optimal state in changing environments, and enhancing indoor comfort and the wide applicability of control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119532918B_ABST
    Figure CN119532918B_ABST
Patent Text Reader

Abstract

The application discloses an air conditioner control method and system based on deep reinforcement learning, a device and a medium. First, first environment state data is acquired and input into a pre-trained initial air conditioner strategy scheduling model. Then, a first operation strategy is obtained according to the first environment state data and a preset factor. Next, the first operation state of the air conditioner is adjusted according to the strategy, and second environment state data after adjustment is acquired again. Through a preset reward function, the first operation strategy and the second environment state data, the parameters of the initial air conditioner strategy scheduling model are optimized to obtain a target air conditioner strategy scheduling model. The current environment state data is input into the target air conditioner strategy scheduling model to generate a current operation strategy, and the current operation state of the air conditioner is accurately adjusted according to the strategy. The embodiment of the application can acquire environment state data in real time and dynamically adjust the operation strategy, thereby effectively improving the response speed and flexibility of air conditioner control.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to, but are not limited to, the technical field of artificial intelligence, and in particular to an air conditioner control method and system based on deep reinforcement learning, a device, and a medium. BACKGROUND

[0002] In the prior art, the operation mode of the air conditioner control system is relatively traditional and limited. Traditional air conditioner control systems mostly rely on fixed rules or pre-set models for temperature adjustment. This approach may be relatively slow in responding to complex and variable indoor and outdoor environments, making it difficult to quickly adjust to the optimal state, thereby leading to a decrease in indoor comfort. In addition, many systems are designed based on specific building models or hypothetical conditions, which makes them unable to achieve the expected control effect when applied to different buildings or environments. This limitation limits the flexibility of the air conditioner control system. SUMMARY

[0003] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0004] Embodiments of the present application provide an air conditioner control method based on deep reinforcement learning, which can effectively improve the response speed and flexibility of air conditioner control.

[0005] In a first aspect, embodiments of the present application provide an air conditioner control method based on deep reinforcement learning, comprising: obtaining first environment state data, inputting the first environment state data into a pre-trained initial air conditioner policy scheduling model; obtaining a first operation policy according to the first environment state data and the pre-set factor; adjusting the first operation state of the air conditioner according to the first operation policy to obtain a second operation state, and obtaining the adjusted second environment state data; optimizing the parameters of the initial air conditioner policy scheduling model according to a pre-set reward function, the first operation policy and the second environment state data to obtain a target air conditioner policy scheduling model; obtaining current environment state data, inputting the current environment state data into the target air conditioner policy scheduling model for policy generation to obtain a current operation policy; and adjusting the current operation state of the air conditioner according to the current operation policy.

[0006] With reference to the first aspect, in an embodiment of the present application, the first operation strategy is obtained according to the first environment state data and the preset factor, comprising: converting the first environment state data into an initial environment state vector, performing normalization and denoising processing on the initial environment state vector to obtain a target environment state vector; generating an initial operation strategy probability distribution according to the target environment state vector; adjusting the initial operation strategy probability distribution according to a preset factor to obtain a target operation strategy probability distribution, wherein the operation strategy corresponding to a first probability in the target operation strategy probability distribution is selected as the first operation strategy according to a preset condition.

[0007] With reference to the first aspect, in an embodiment of the present application, the adjusting the initial operation strategy probability distribution according to a preset factor to obtain a target operation strategy probability distribution, comprising: obtaining an uncertainty index of the initial operation strategy probability distribution; constructing a preset index, adjusting the preset factor according to the uncertainty index and the preset index to obtain an optimized preset factor; adjusting the initial operation strategy probability distribution according to the optimized preset factor to obtain a target operation strategy probability distribution.

[0008] With reference to the first aspect, in an embodiment of the present application, the adjusting the initial operation strategy probability distribution according to a preset factor to obtain a target operation strategy probability distribution, comprising: obtaining an uncertainty index of the initial operation strategy probability distribution; constructing a preset index, adjusting the preset factor according to the uncertainty index and the preset index to obtain an optimized preset factor; adjusting the initial operation strategy probability distribution according to the optimized preset factor to obtain a target operation strategy probability distribution.

[0009] With reference to the first aspect, in an embodiment of the present application, the preset reward function comprises a first index, a second index and a third index, the types of the first index, the second index and the third index are different, and the calculating the instant reward information according to the preset reward function and the second environment state data comprises: calculating a first performance value of the first index, a second performance value of the second index and a third performance value of the third index according to the second environment state data; assigning a first weight to the first index, a second weight to the second index, and a third weight to the third index, and performing weighted calculation on the first performance value, the second performance value and the third performance value by using the first weight, the second weight and the third weight to obtain the instant reward information.

[0010] With reference to the first aspect, in an embodiment of the present application, the first index is an energy consumption index, the second index is a comfort index, and the third index is an environmental index.

[0011] With reference to the first aspect, in an embodiment of the present application, the training step of the pre-trained initial air conditioner strategy scheduling model comprises: reading historical experience data, wherein the historical experience data comprises third environmental state data, a third running strategy, fourth environmental state data, and historical immediate reward information, the fourth environmental state data is obtained by adjusting the third environmental state data according to the third running strategy, and the historical immediate reward information is obtained according to the fourth environmental state data and the preset reward function; calculating historical actual reward information according to the third running strategy and the fourth environmental state data; constructing a second loss function, calculating a second loss value between the historical immediate reward information and the historical actual reward according to the second loss function, and updating parameters of the air conditioner strategy scheduling model according to the second loss value.

[0012] In a second aspect, an embodiment of the present application provides an air conditioner system based on deep reinforcement learning, which is applied to the air conditioner control method as described above, and comprises: a data processing module configured to obtain first environmental state data and input the first environmental state data into a pre-trained initial air conditioner strategy scheduling model; a first strategy generation module configured to obtain a first running strategy according to the first environmental state data and the preset factor; a first information processing module configured to adjust a first running state of an air conditioner according to the first running strategy to obtain a second running state and obtain second environmental state data after adjustment; a model generation module configured to optimize parameters of the initial air conditioner strategy scheduling model according to a preset reward function, the first running strategy, and the second environmental state data to obtain a target air conditioner strategy scheduling model; a second strategy generation module configured to obtain current environmental state data and input the current environmental state data into the target air conditioner strategy scheduling model to generate a strategy to obtain a current running strategy; and a second information processing module configured to adjust a current running state of the air conditioner according to the current running strategy.

[0013] In another aspect, an embodiment of the present application further provides an electronic device, which comprises:

[0014] at least one processor;

[0015] at least one memory configured to store at least one program;

[0016] when the at least one program is executed by the at least one processor, the air conditioner control method as described above is implemented.

[0017] In another aspect, the embodiments of the present application also provide a computer readable storage medium, wherein a computer program executable by a processor is stored, and the computer program executable by the processor is used to implement the air conditioner control method as above when executed by the processor.

[0018] The air conditioner control method based on deep reinforcement learning provided by the embodiments of the present application first acquires first environment state data and inputs the first environment state data into an initially trained initial air conditioner policy scheduling model. Then, the model generates a first operation policy according to the first environment state data and a preset factor. Next, the first operation state of the air conditioner is adjusted according to the policy to obtain a second operation state, and the second environment state data after adjustment is acquired again. Through a preset reward function, the first operation policy and the second environment state data, the parameters of the initial air conditioner policy scheduling model can be iteratively optimized until a target air conditioner policy scheduling model is obtained. The model can gradually master more efficient and more adaptable air conditioner control policies through deep learning and adaptation, significantly improve the response speed and environmental adaptability, and ensure that the air conditioner can quickly adjust to the optimal state when facing variable environments, thereby greatly improving the indoor comfort. The current environment state data is input into the target air conditioner policy scheduling model to generate a current operation policy, and the current operation state of the air conditioner is accurately adjusted according to the policy. The embodiments of the present application can acquire environment state data in real time and dynamically adjust the operation policy, so as to flexibly cope with various indoor and outdoor environmental changes and improve the wide applicability and flexibility of air conditioner control. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 is a flowchart of the air conditioner control method based on deep reinforcement learning provided by the embodiments of the present application;

[0020] Figure 2 is a specific flowchart of step 120 in Figure 1

[0021] Figure 3 is a specific flowchart of step 230 in Figure 2

[0022] Figure 4 is a specific flowchart of step 140 in Figure 1

[0023] Figure 5 is a schematic diagram of the training process of the initial air conditioner policy scheduling model provided by the embodiments of the present application;

[0024] Figure 6 is a structural diagram of the air conditioner system based on deep reinforcement learning provided by the embodiments of the present application;

[0025] ​​​Figure 7 is a schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0026] In order to make the purposes, technical solutions and advantages of the present application clearer, further detailed description will be made to the present application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not used to limit the present application.

[0027] It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described in the flowchart can be performed in an order different from that in the flowchart. The terms "first", "second", etc. in the specification and claims and the above drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the structure, proportion, size, etc. shown in the drawings of the specification are only used to cooperate with the content disclosed in the specification, so as to be understood and read by those skilled in the art, and do not have technical significance. Any modification of structure, change of proportion relationship or adjustment of size, without affecting the effect and purpose that can be achieved by the present application, should still fall within the scope of the technical content disclosed by the present application. At the same time, the terms such as "upper", "lower", "left", "right", "middle" and "one" used in the specification are only for the convenience of clear description, and are not intended to limit the scope of the present application. The change or adjustment of relative relationship, without substantially changing the technical content, is also considered as the implementation scope of the present application.

[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the specification is only for the purpose of describing the embodiments of the present application and is not intended to limit the present application.

[0029] In the prior art, the operation mode of the air conditioner control system is relatively traditional and limited. The traditional air conditioner control system mostly relies on fixed rules or preset models for temperature adjustment, which is not sufficient in the face of complex and variable indoor and outdoor environments, especially when dealing with non-standard or unexpected situations, its efficiency is often greatly discounted. Although some systems try to introduce optimization algorithms to improve performance, these algorithms are mostly determined at the beginning of design, lacking the ability to adjust and optimize themselves over time and with environmental changes. Therefore, when the external environment or indoor demand changes significantly, the response speed of these systems may be relatively slow, and it is difficult to quickly adjust to the optimal state, resulting in a decrease in indoor comfort.

[0030] Furthermore, existing systems have shortcomings in energy management. They may fail to fully utilize collected environmental data to optimize control strategies, resulting in insufficiently precise and efficient energy use. Many systems are often designed based on specific building models or assumptions, which may prevent them from achieving the expected control effects when applied to different buildings or environments. This limitation restricts the wide applicability and flexibility of air conditioning control systems.

[0031] Furthermore, integrating different technical components into a unified air conditioning control system is also a challenge faced by existing technologies. Compatibility and interoperability issues between different components can lead to difficulties in system integration, thereby affecting the stability and performance of the entire system.

[0032] In view of this, embodiments of this application provide an air conditioning control method, air conditioning system, electronic device, and computer-readable storage medium based on deep reinforcement learning. The method first acquires first environmental state data and inputs it into a pre-trained initial air conditioning strategy scheduling model. Then, based on the first environmental state data and preset factors, the model generates a first operating strategy. Next, the first operating state of the air conditioner is adjusted according to this strategy to obtain a second operating state, and the adjusted second environmental state data is acquired again. Through a preset reward function, the first operating strategy, and the second environmental state data, the parameters of the initial air conditioning strategy scheduling model can be iteratively optimized until a target air conditioning strategy scheduling model is obtained. This model, through deep learning and adaptation, can gradually master more efficient and adaptable air conditioning control strategies, significantly improving response speed and environmental adaptability, ensuring that the air conditioner can quickly adjust to the optimal state when facing changing environments, thereby greatly improving indoor comfort. By inputting the current environmental state data into the target air conditioning strategy scheduling model, the current operating strategy can be generated, and the current operating state of the air conditioner can be precisely adjusted according to this strategy. The embodiments of this application can acquire environmental status data in real time and dynamically adjust the operating strategy, thereby flexibly responding to various indoor and outdoor environmental changes and improving the wide applicability and flexibility of air conditioning control.

[0033] The embodiments of this application will be further described below with reference to the accompanying drawings.

[0034] like Figure 1 As shown, Figure 1 This is a flowchart of an air conditioning control method based on deep reinforcement learning provided in an embodiment of this application. The air conditioning control method includes, but is not limited to, steps 110 to 160.

[0035] Step 110: Obtain the first environmental state data and input the first environmental state data into the pre-trained initial air conditioning strategy scheduling model;

[0036] Step 120: obtaining a first operation strategy according to the first environment state data and a preset factor;

[0037] Step 130: adjusting a first operation state of the air conditioner according to the first operation strategy to obtain a second operation state, and obtaining second environment state data after adjustment;

[0038] Step 140: optimizing parameters of the initial air conditioner strategy scheduling model according to a preset reward function, the first operation strategy and the second environment state data to obtain a target air conditioner strategy scheduling model;

[0039] Step 150: obtaining current environment state data, inputting the current environment state data into the target air conditioner strategy scheduling model for strategy generation to obtain a current operation strategy;

[0040] Step 160: adjusting a current operation state of the air conditioner according to the current operation strategy.

[0041] The steps 110 to 160 will be described in detail below.

[0042] In a feasible embodiment, the first environment state data covers key parameters indoors and outdoors (for example, indoors and outdoors of a gym), including but not limited to temperature, humidity, carbon dioxide concentration, light intensity, passenger flow and activity type (such as competition, training or idle state) and the like. These data can be obtained in real time through various sensors and statistical devices, which can specifically include temperature sensors, humidity sensors, carbon dioxide concentration sensors and professional devices for counting passenger flow and the like. Through flexible selection of timing collection or event-triggered mode, the latest environment state data can be efficiently collected from the above sensors. This not only ensures the timeliness and accuracy of the data, but also provides a solid foundation for subsequent air conditioner strategy scheduling.

[0043] In a feasible embodiment, the initial air conditioner strategy scheduling model can be a model based on machine learning (such as neural network, decision tree and the like) or deep learning (such as convolutional neural network CNN, recurrent neural network RNN, deep neural network DNN and the like). The model has been pre-trained and can generate an air conditioner operation strategy based on input environment state data.

[0044] In a feasible embodiment, after inputting the first environmental state data into the pre-trained initial air conditioner strategy scheduling model, the model can formulate a first operation strategy in combination with a preset factor (e.g., a temperature factor). Specifically, the first environmental state data is analyzed and interpreted according to the preset temperature factor, which may represent the user's preference for temperature, the temperature condition of the external environment, or the performance characteristics of the air conditioning equipment, etc. By analyzing the relationship between each environmental parameter in the first environmental state data and the temperature factor, a preliminary operation strategy (i.e., the first operation strategy) can be obtained. This strategy can include the on-off state of the air conditioning equipment, the working mode, the wind speed, the temperature setting, and other aspects, aiming to provide a comfortable and energy-saving air conditioning operation scheme according to the current environmental conditions and user demand.

[0045] Referring to Figure 2 , Figure 2 is a specific flowchart of step 120 provided by the embodiments of the present application, which can include but is not limited to steps 210 to 230.

[0046] Step 210: converting the first environmental state data into an initial environmental state vector, normalizing and denoising the initial environmental state vector to obtain a target environmental state vector;

[0047] Step 220: generating an initial operation strategy probability distribution according to the target environmental state vector;

[0048] Step 230: adjusting the initial operation strategy probability distribution according to a preset factor to obtain a target operation strategy probability distribution, and selecting the operation strategy corresponding to the first probability as the first operation strategy in the target operation strategy probability distribution according to a preset condition.

[0049] It is worth noting that environmental data usually includes multiple types of information, such as temperature, humidity, air quality, etc. These data may come from different sensors or data sources, with different units and dimensions. By integrating environmental data into a state vector, the performance and accuracy of the model can be improved, and the data management and analysis process can be simplified. It should be noted that the state vector is a fixed-length vector containing all relevant information of the environmental data. Each element corresponds to a specific environmental parameter. Before integrating the data, data cleaning is usually required, including removing outliers, filling missing values, etc.

[0050] In a feasible embodiment, due to the differences in units and dimensions of different environmental parameters, in order to ensure that the model can learn and understand these data more effectively, a series of preprocessing can be performed on them, including denoising and normalization, etc. Denoising aims to eliminate outliers and noise in the data to improve the accuracy and reliability of the data. And normalization is to convert environmental parameters of different dimensions and ranges to the same scale, so that they have the same weight and influence in the model. Such preprocessing steps help the model better capture and utilize the features in the data, thereby generating more accurate air conditioning operation strategies.

[0051] In a feasible embodiment, in the process of generating an initial operation strategy probability distribution according to the target environment state vector, the model analyzes the target environment state vector, considers the influence of various environmental factors (such as temperature, humidity, air quality, etc.) on the current environment, and generates an initial operation strategy probability distribution according to these factors. This distribution represents the probability of each strategy being selected among different possible operation strategies. The size of the probability usually reflects the adaptability and effect of the strategy under the current environmental state. For example, in the air conditioning strategy scheduling model, if the target environment state vector shows that the current temperature is high and the humidity is moderate, the model may assign a higher probability to those operation strategies that can reduce the temperature, such as reducing the indoor temperature set point, increasing the air conditioning supply volume, etc. At the same time, the influence of other factors (such as air quality, energy efficiency, etc.) on strategy selection is also considered, so as to obtain a comprehensive initial operation strategy probability distribution. Next, the probability distribution is adjusted according to the preset factors (such as temperature factors) to obtain a target operation strategy that meets the actual needs. This adjustment process usually involves weighting, shifting or re-normalizing the probability distribution, etc., so that the finally selected strategy can meet the user's comfort while achieving energy saving, environmental protection, etc.

[0052] In a feasible embodiment, it is assumed that in a scenario, the environmental conditions in the greenhouse, such as temperature, humidity, light intensity, etc., need to be monitored and adjusted to ensure that the crops in the greenhouse can grow in the best environment. In this scenario, the temperature factor is the environmental factor that has a greater impact on crop growth. It is assumed that the current environmental state data of the greenhouse (i.e., the first environmental state data) includes: the temperature is 25°C, the humidity is 60%, and the light intensity is 3000 lux. At the same time, the environmental parameter range suitable for crop growth is set: the temperature is between 15-35°C, the humidity is between 40%-80%, and the light intensity is between 2000-4000 lux. After the processing of step 210, the normalized and denoised target environmental state vector can be obtained, and its value is [0.5, 0.5, 0.5]. Further, it is assumed that there is an initial running strategy probability distribution generated based on the initial air conditioning strategy scheduling model. For example, in this situation, the model predicts that the probabilities of taking the ventilation, humidification, and light supplement strategies are 0.3, 0.2, and 0.5, respectively. Given that the temperature factor has the most critical impact on crop growth, and the current temperature is 25°C, which is in the optimal temperature range (24-26°C) for crop growth, there is no need to make a substantial adjustment to the temperature. Based on this, the probability of the ventilation strategy that may cause the temperature to drop can be moderately reduced, and the probability of the light supplement strategy that may promote photosynthesis and crop growth can be increased. After the above adjustment, the probability distribution of the target running strategy is adjusted to: ventilation 0.2 (reduced), humidification 0.2 (unchanged), and light supplement 0.6 (increased). Finally, in the current target running strategy probability distribution, the probability of the light supplement strategy is the highest, reaching 0.6. Therefore, the light supplement strategy can be selected as the primary running strategy in order to further optimize the growth environment of the crops.

[0053] Referring to Figure 3 In a feasible embodiment, the process of adjusting the initial running strategy probability distribution according to the preset factor in step 230 to obtain the target running strategy probability distribution includes but is not limited to steps 310-330.

[0054] Step 310: Obtain the uncertainty index of the initial running strategy probability distribution;

[0055] Step 320: Construct a preset index, adjust the preset factor according to the uncertainty index and the preset index, and obtain an optimized preset factor;

[0056] Step 330: Adjust the initial running strategy probability distribution according to the optimized preset factor to obtain the target running strategy probability distribution.

[0057] In a feasible embodiment, the uncertainty indicator of the initial running strategy probability distribution can be used to measure the dispersion degree or uncertainty of the probabilities of each strategy in the distribution. Commonly used uncertainty indicators include entropy and variance, etc. Specifically, entropy is an indicator in information theory to measure the uncertainty of a random variable. For a probability distribution, the greater the entropy, the more uncertain the distribution, and the closer the probabilities of each strategy are. On the contrary, the smaller the entropy, the more certain the distribution, and the probability of one or more strategies is much higher than that of other strategies. Variance measures the degree of deviation of the probabilities of each strategy from the average probability in the probability distribution. The greater the variance, the more dispersed the distribution, and the greater the probability difference between strategies. The smaller the variance, the more concentrated the distribution.

[0058] In a feasible embodiment, the preset indicators are set according to actual needs and can be used to guide the adjustment direction and target of the preset factor. These indicators can include user comfort, energy consumption, environmental impact, etc. For example, if the user values comfort more, a comfort indicator based on environmental parameters such as temperature and humidity can be set; if energy saving is desired, a preset indicator based on energy consumption can be set.

[0059] In a feasible embodiment, in the process of evaluating the rationality of the current initial running strategy probability distribution according to the uncertainty indicator and the preset indicator, and adjusting the preset factor accordingly, if the uncertainty indicator is high (e.g., the entropy is large), it can indicate that the distribution is too dispersed, and the preset factor needs to be adjusted to increase the probability of certain strategies and decrease the probability of other strategies, thereby reducing the uncertainty. At the same time, the specific value of the preset factor is optimized according to the preset indicators (such as comfort, energy consumption, etc.) to make the adjusted probability distribution more in line with actual needs.

[0060] In a feasible embodiment, by steps 310 to 330, the initial running strategy probability distribution can be adjusted according to the preset factor to obtain a target running strategy probability distribution that is more in line with actual needs. This process is an iterative optimization process, which may need to adjust the preset factor and evaluate the uncertainty indicator and the preset indicator multiple times until a satisfactory result is obtained.

[0061] In a feasible embodiment, the air conditioning equipment is configured and controlled according to the first operation strategy (for example, a strategy containing specific parameter settings, such as temperature set value, air speed, working mode, etc.) generated before. The air conditioning equipment starts to operate according to the configuration of the first operation strategy, thereby changing its current first operation state. This state can include parameters such as temperature, humidity, air supply, etc. being operated. Through sensors or other monitoring equipment, real-time or periodic collection of the adjusted operation state data of the air conditioning equipment is performed, and these data constitute the second operation state of the air conditioning. The second operation state reflects the actual working condition of the air conditioning after the implementation of the first operation strategy. At the same time, the state data of the environment where the air conditioning is located also need to be collected, which reflect the influence of the operation of the air conditioning on the environment. These data can include parameters such as temperature, humidity, air quality, etc., which can change after the implementation of the first operation strategy. The adjusted second environment state data are important basis for evaluating the effect of the first operation strategy, and they can be used for subsequent strategy optimization and adjustment.

[0062] In a feasible embodiment, by comparing and analyzing the environment state data before and after adjustment with the preset reward function (such as indicators of improving comfort and reducing energy consumption, etc.), the parameter configuration of the model can be iteratively optimized until the most efficient control strategy is explored, thereby constructing the target air conditioning strategy scheduling model.

[0063] Referring to Figure 4 , Figure 4 is a specific flowchart of step 140, which can include but is not limited to steps 410 to 440.

[0064] Step 410: calculating the immediate reward information according to the preset reward function and the second environment state data;

[0065] Step 420: calculating the actual reward information according to the first operation strategy and the second environment state data;

[0066] Step 430: constructing the first loss function, and calculating the first loss value between the immediate reward information and the actual reward according to the first loss function;

[0067] Step 440: updating the parameters of the air conditioning strategy scheduling model according to the first loss value.

[0068] In an embodiment, the preset reward function is a predefined function for evaluating the reward obtained by performing an action (e.g., the first running strategy) in a given environment state. The reward function can take into account multiple factors, such as indoor temperature comfort, energy consumption, humidity control, etc. The second environment state data refers to the new environment state data collected after performing the first running strategy, such as temperature, humidity, outdoor temperature, etc. The immediate reward information refers to the immediate reward value calculated by inputting the second environment state data into the preset reward function, which can be used to represent the degree of improvement in the environment state and the efficiency of energy consumption after performing the first running strategy.

[0069] In an embodiment, the preset reward function can include a first index, a second index, and a third index, and the types of the first index, the second index, and the third index are different. The specific process of calculating the immediate reward information according to the preset reward function and the second environment state data in step 310 includes:

[0070] According to the second environment state data, calculating a first performance value of the first index, a second performance value of the second index, and a third performance value of the third index;

[0071] Assigning a first weight to the first index, a second weight to the second index, and a third weight to the third index, and performing weighted calculation on the first performance value, the second performance value, and the third performance value using the first weight, the second weight, and the third weight to obtain the immediate reward information.

[0072] In an embodiment, the first index can be an energy consumption index (e.g., kilowatt-hour consumption), the second index can be a comfort index (e.g., temperature uniformity, wind speed suitability), and the third index can be an environmental index (e.g., indoor air quality). After obtaining the second environment state data (including temperature, humidity, light intensity, PM2.5 concentration, energy consumption, etc.), the performance value of the energy consumption index can be calculated according to the energy consumption in the second environment state data. The performance value of the comfort index can be calculated according to the temperature, humidity, and other environmental data. The performance value of the environmental index can be calculated according to the light intensity, PM2.5 concentration, and other environmental data.

[0073] In an embodiment, a weight value is assigned to the energy consumption index according to business requirements. This weight reflects the importance of energy consumption in the overall reward calculation. A weight value is assigned to the comfort index, which also reflects the importance of comfort in the overall reward calculation. A weight value is assigned to the environmental index, which reflects the relative importance of environmental quality in the overall reward calculation. It is worth noting that the allocation of weights should be reasonable to ensure that the overall reward can reflect the comprehensive influence of all indexes, and the sum of the weights should usually be 1 (or some other normalized value).

[0074] In a feasible embodiment, the first performance value of the energy consumption indicator, the second performance value of the comfort indicator, and the third performance value of the environmental indicator are multiplied by their corresponding weights (the first weight, the second weight, and the third weight) respectively, and then the weighted values are added together, and the sum is the immediate reward information, which is a single value that integrates the three indicators of energy consumption, comfort, and environmental quality, used to represent the immediate reward obtained by executing a certain operation strategy under the current environmental state.

[0075] In a feasible embodiment, the actual reward information is a reward calculation based on actual business needs, for example, if the second environmental state data indicates that the indoor temperature is within the comfort range and the energy consumption is lower than a certain threshold, a higher actual reward is given.

[0076] In a feasible embodiment, the first loss function refers to a function for measuring the difference between the immediate reward information and the actual reward. It can be a mean square error, cross-entropy loss, or other types of loss functions, depending on the type of reward information and the optimization goal of the model. The first loss value refers to the loss value calculated by inputting the immediate reward information and the actual reward information into the first loss function, representing the error of the model in predicting the reward.

[0077] In a feasible embodiment, the parameters of the air conditioner strategy scheduling model can be updated according to the first loss value using optimization algorithms such as gradient descent, stochastic gradient descent, Adam, etc. By continuously iterating and optimizing the parameters of the model, the error of the model in predicting the reward is minimized, resulting in a more accurate operation strategy.

[0078] In addition, referring to Figure 5 In a feasible embodiment, the training process of the pre-trained initial air conditioner strategy scheduling model in step 110 can include but is not limited to steps 510 to 530.

[0079] Step 510: Read historical experience data, which includes third environmental state data, third operation strategy, fourth environmental state data obtained by adjusting the third environmental state data according to the third operation strategy, and historical immediate reward information obtained according to the fourth environmental state data and a preset reward function;

[0080] Step 520: Calculate the historical actual reward information according to the third operation strategy and the fourth environmental state data;

[0081] Step 530: Construct a second loss function, calculate a second loss value between the historical immediate reward information and the historical actual reward according to the second loss function, and update the parameters of the air conditioner strategy scheduling model according to the second loss value.

[0082] In a feasible embodiment, the third environmental state data records the indoor and outdoor environmental conditions at a certain time point in the past, such as temperature, humidity, light intensity, etc. The third operation strategy refers to the historical air conditioner operation strategy adopted for the above-mentioned environmental state data, including temperature setting, wind speed, working mode, etc. The fourth environmental state data is the change in the environmental state after the third operation strategy is executed, i.e., the new environmental state actually achieved. The historical immediate reward information is calculated based on the fourth environmental state data and a preset reward function, and is used to evaluate the pros and cons of the current strategy, usually considering multiple indicators such as comfort and energy efficiency.

[0083] In a feasible embodiment, based on the third operation strategy and the fourth environmental state data, the historical actual reward information can be further calculated. The purpose of this step is to quantify the actual effect of the strategy by comparing the changes in the environmental state before and after the strategy is executed, and the actual impact of these changes on user comfort.

[0084] In a feasible embodiment, in order to continuously improve the accuracy and efficiency of the strategy scheduling model, step 530 constructs a second loss function, which aims to measure the deviation between the historical immediate reward information and the historical actual reward (i.e., the second loss value). Through gradient descent or other optimization algorithms, the parameters of the model are gradually adjusted according to the calculated second loss value to minimize this deviation. This process is iterated continuously until the model can accurately predict and generate the optimal air conditioner operation strategy under a given environmental state.

[0085] Through the training process of steps 510 to 530, the initial air conditioner strategy scheduling model is gradually optimized, improving its adaptability to different environmental states and the accuracy of strategy generation, laying a solid foundation for efficient and intelligent air conditioner control.

[0086] In a feasible embodiment, after obtaining the target air conditioner strategy scheduling model, the current environmental state data can be collected in real time, and the current environmental state data is input into the target air conditioner strategy scheduling model for strategy generation to obtain the current operation strategy, and then the current operation state of the air conditioner is adjusted according to the current operation strategy. By real-time acquisition of environmental state data and dynamic adjustment of the operation strategy, various indoor and outdoor environmental changes can be flexibly responded to, thereby significantly improving the wide applicability and flexibility of the air conditioner system control. Taking the air conditioner system of a gymnasium as an example, the target air conditioner strategy scheduling model can be deployed on its control server to ensure that the model can stably and efficiently run on the server. Through a series of sensors and other devices, the environmental state data (including important information such as indoor temperature, humidity, outdoor weather conditions, and indoor personnel activity conditions) in the gymnasium can be captured in real time. Based on these data, the target air conditioner strategy scheduling model will calculate the optimal air conditioner operation strategy and apply these strategies to the air supply part of the air conditioner system. At the same time, the system will continuously monitor the environmental changes after the execution of the strategy to realize dynamic adjustment and optimization, thereby forming a closed-loop control system. In this system, the model not only receives real-time environmental feedback, but also continuously adjusts and optimizes the strategy based on these feedback. In addition, the performance of the entire system can be monitored for a long time, and the target air conditioner strategy scheduling model can be optimized according to actual needs to ensure that the operation of the air conditioner system always remains in the best state.

[0087] Referring to Figure 6 , Figure 6 is a structure diagram of an air conditioner system based on deep reinforcement learning provided by an embodiment of the present application, which can be applied to the air conditioner control method of any one of the preceding embodiments. The air conditioner system 600 comprises:

[0088] The data processing module 610 is configured to acquire first environmental state data and input the first environmental state data into the pre-trained initial air conditioner strategy scheduling model.

[0089] The first strategy generation module 620 is configured to obtain a first operation strategy according to the first environmental state data and a preset factor.

[0090] The first information processing module 630 is configured to adjust the first operation state of the air conditioner according to the first operation strategy to obtain a second operation state, and acquire second environmental state data after the adjustment.

[0091] The model generation module 640 is configured to optimize the parameters of the initial air conditioner strategy scheduling model according to a preset reward function, the first operation strategy and the second environmental state data to obtain a target air conditioner strategy scheduling model.

[0092] The second strategy generation module 650 is configured to acquire the current environment state data, input the current environment state data into the target air conditioner strategy scheduling model to generate a strategy, and obtain a current operation strategy.

[0093] The second information processing module 660 is configured to adjust a current operation state of the air conditioner according to the current operation strategy.

[0094] It should be noted that the air conditioner system 600 of the embodiment can implement the air conditioner control method of the foregoing embodiments, and therefore the air conditioner system 600 of the embodiment and the air conditioner control method of the foregoing embodiments have the same technical principles and the same beneficial effects. To avoid repetition of the content, no further description is given herein.

[0095] With reference to Figure 7 The embodiment of the present application further discloses an electronic device, which comprises:

[0096] at least one processor 701;

[0097] at least one memory 702 configured to store at least one program;

[0098] When the at least one program is executed by the at least one processor 701, the air conditioner control method as described above is implemented.

[0099] The embodiment of the present application further discloses a computer readable storage medium, wherein a computer program executable by a processor is stored, and the computer program executable by the processor is used to implement the air conditioner control method as described above when executed by the processor.

[0100] The above description of the disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for controlling an air conditioner based on deep reinforcement learning, the method comprising: The method comprises the following steps: obtaining first environment state data, inputting the first environment state data into a pre-trained initial air conditioner strategy scheduling model; obtaining a first operation strategy according to the first environment state data and a preset factor; adjusting a first operation state of an air conditioner according to the first operation strategy to obtain a second operation state, and obtaining second environment state data after adjustment; optimizing parameters of the initial air conditioner strategy scheduling model according to a preset reward function, the first operation strategy and the second environment state data to obtain a target air conditioner strategy scheduling model; obtaining current environment state data, inputting the current environment state data into the target air conditioner strategy scheduling model to generate a strategy, and obtaining a current operation strategy; adjusting a current operation state of the air conditioner according to the current operation strategy; wherein, the first operation strategy is obtained according to the first environment state data and the preset factor, which comprises: converting the first environment state data into an initial environment state vector, normalizing and denoising the initial environment state vector to obtain a target environment state vector; generating an initial operation strategy probability distribution according to the target environment state vector; adjusting the initial operation strategy probability distribution according to a preset factor to obtain a target operation strategy probability distribution, and selecting an operation strategy corresponding to a first probability in the target operation strategy probability distribution as the first operation strategy according to a preset condition. 2.The air conditioner control method based on deep reinforcement learning of claim 1, wherein, The target operation strategy probability distribution is obtained by adjusting the initial operation strategy probability distribution according to the preset factor, which comprises: obtaining an uncertainty index of the initial operation strategy probability distribution; constructing a preset index, adjusting the preset factor according to the uncertainty index and the preset index to obtain an optimized preset factor; adjusting the initial operation strategy probability distribution according to the optimized preset factor to obtain the target operation strategy probability distribution. 3.The air conditioner control method based on deep reinforcement learning of claim 1, wherein, The target air conditioner strategy scheduling model is obtained by optimizing the parameters of the initial air conditioner strategy scheduling model according to the preset reward function, the first operation strategy and the second environment state data, which comprises: calculating to obtain instant reward information according to the preset reward function and the second environment state data; calculating to obtain actual reward information according to the first operation strategy and the second environment state data; constructing a first loss function to calculate a first loss value between the instant reward information and the actual reward according to the first loss function; updating the parameters of the air conditioner strategy scheduling model according to the first loss value. 4.The air conditioner control method based on deep reinforcement learning of claim 3, wherein, The preset reward function comprises a first index, a second index and a third index, the first index is an energy consumption index, the second index is a comfort index, and the third index is an environmental index. The instant reward information is calculated according to the preset reward function and the second environment state data, and includes: calculating a first performance value of the first index, a second performance value of the second index and a third performance value of the third index according to the second environment state data; assigning a first weight to the first index, a second weight to the second index and a third weight to the third index, and performing weighted calculation on the first performance value, the second performance value and the third performance value by using the first weight, the second weight and the third weight to obtain the instant reward information. 5.The air conditioner control method based on deep reinforcement learning of claim 1, wherein, The training step of the pre-trained initial air conditioner strategy scheduling model includes: reading historical experience data, the historical experience data including third environment state data, a third running strategy, fourth environment state data and historical instant reward information, the fourth environment state data being obtained by adjusting the third environment state data according to the third running strategy, and the historical instant reward information being obtained according to the fourth environment state data and the preset reward function; calculating historical actual reward information according to the third running strategy and the fourth environment state data; constructing a second loss function, calculating a second loss value between the historical instant reward information and the historical actual reward according to the second loss function, and updating parameters of the air conditioner strategy scheduling model according to the second loss value.

6. An air conditioning system based on deep reinforcement learning, characterized by, The application is applied to the air conditioner control method in any one of claims 1 to 5, comprising: a data processing module configured to obtain first environment state data and input the first environment state data into a pre-trained initial air conditioner strategy scheduling model; a first strategy generation module configured to obtain a first running strategy according to the first environment state data and a preset factor; a first information processing module configured to adjust a first running state of an air conditioner according to the first running strategy to obtain a second running state and obtain second environment state data after adjustment; a model generation module configured to optimize parameters of an initial air conditioner strategy scheduling model according to a preset reward function, the first running strategy and the second environment state data to obtain a target air conditioner strategy scheduling model; a second strategy generation module configured to obtain current environment state data, input the current environment state data into the target air conditioner strategy scheduling model for strategy generation, and obtain a current running strategy; a second information processing module configured to adjust a current running state of an air conditioner according to the current running strategy.

7. An electronic device, comprising: comprising: at least one processor; at least one memory configured to store at least one program; when at least one of the programs is executed by at least one of the processors, the air conditioner control method in any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium, characterized in that, wherein a processor-executable computer program is stored, and the processor-executable computer program is executed by a processor to implement the air conditioner control method in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Control method and system of heating ventilation air conditioner

    CN114838470A

  • Heating ventilation air conditioner regulation and control method and device based on reinforcement learning

    CN115950080A