Air conditioning control method, device, air conditioning and storage medium based on user behavior
By performing polynomial feature processing and reward scaling function training on the historical data of the air conditioner, an accurate air conditioner control strategy is generated, which solves the problem of low control accuracy in the existing technology, and realizes the precise adjustment and energy-saving effect of the air conditioner.
Patent Information
- Application Number
- CN202310667665.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-06
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2043-06-06
AI Technical Summary
The existing air conditioner control methods cannot comprehensively consider the usage scenarios, weather and user habits, resulting in low control accuracy and difficult to determine the optimal value of the reward function parameters of deep reinforcement learning.
By obtaining the historical operation monitoring data of the target air conditioner, processing and conversion of polynomial feature data, using the reward scaling function to train the prediction model, and generating an accurate air conditioner control strategy.
It improves the control accuracy of the air conditioner, and achieves accurate adjustment and energy-saving effects of the air conditioner.
Smart Images

Figure CN116734411B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of air conditioning, and in particular to a method and device for controlling an air conditioning based on user behavior, an air conditioner, and a storage medium. Background Art
[0002] Low-temperature air source heat pump systems are widely used in the northern heating market due to their high efficiency, cleanliness, and safety. However, their control method is pre-set in the controller based on the designer's experience, and cannot comprehensively consider factors such as usage scenarios, weather, and user habits, leaving a large space for energy saving.
[0003] Existing air conditioning control methods use deep reinforcement learning as a control scheme for air conditioning systems. Different energy-saving scenarios and corresponding energy-saving strategies are set for different types of computer room air conditioners. This method can reduce the deviation between the simulation and real-world phases, improve the accuracy of energy-saving strategies, and achieve self-learning and energy conservation. However, this method's reward function is a weighted sum of parameters such as the number of air conditioner starts and stops and the energy-saving ratio. Due to the large number of weighting coefficients required, it is difficult to accurately determine the optimal value of these weighting coefficients, resulting in low air conditioning control accuracy. Summary of the Invention
[0004] Embodiments of the present invention provide a user behavior-based air conditioning control method, device, air conditioner, and storage medium to improve the control accuracy of the air conditioner.
[0005] In a first aspect, an embodiment of the present invention provides an air conditioning control method based on user behavior, which includes:
[0006] Acquiring historical operation monitoring data of a target air conditioner, and processing the historical operation monitoring data to obtain polynomial characteristic data;
[0007] Converting the polynomial feature data to obtain target data;
[0008] Training a prediction model based on the target data, and adjusting parameters of the prediction model using a reward scaling function during the model training process to obtain a target prediction model;
[0009] An air conditioning control strategy is generated based on the target prediction model, so that the target air conditioning completes adjustment control based on the air conditioning control strategy.
[0010] In a second aspect, an embodiment of the present invention provides an air conditioning control device based on user behavior, comprising:
[0011] a data acquisition unit, configured to acquire historical operation monitoring data of a target air conditioner and process the historical operation monitoring data to obtain polynomial characteristic data;
[0012] A data conversion unit, configured to convert the polynomial feature data to obtain target data;
[0013] a model training unit, configured to train a prediction model based on the target data, and, during the model training process, adjust the parameters of the prediction model using a reward scaling function to obtain a target prediction model;
[0014] An air conditioning control unit is used to generate an air conditioning control strategy based on the target prediction model, so that the target air conditioning completes adjustment control based on the air conditioning control strategy.
[0015] In a third aspect, an embodiment of the present invention provides an air conditioner, which includes a memory, a processor, and a computer program stored on the memory and runnable on the processor. When the processor executes the computer program, the air conditioner control method based on user behavior described in the first aspect above is implemented.
[0016] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the air-conditioning control method based on user behavior described in the first aspect above.
[0017] An embodiment of the present invention provides an air conditioning control method, device, air conditioner, and storage medium based on user behavior, the method comprising: obtaining historical operation monitoring data of a target air conditioner, and processing the historical operation monitoring data to obtain polynomial feature data; converting the polynomial feature data to obtain target data; training a prediction model based on the target data, and during the model training process, using a reward scaling function to adjust the parameters of the prediction model to obtain a target prediction model; generating an air conditioning control strategy based on the target prediction model, so that the target air conditioner completes adjustment control based on the air conditioning control strategy. The embodiment of the present invention processes and converts the historical operation monitoring data of the target air conditioner to obtain target data, and trains the prediction model based on the target data and the reward scaling function, so that the prediction model can generate an accurate air conditioning control strategy, thereby facilitating improving the control accuracy of the air conditioner. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1A flowchart of an air conditioning control method based on user behavior provided by an embodiment of the present invention;
[0020] Figure 2 A schematic diagram of an implementation flow of a sub-process in the air-conditioning control method based on user behavior provided by an embodiment of the present invention;
[0021] Figure 3 A schematic diagram showing data combination of historical operation monitoring data according to an embodiment of the present invention is shown;
[0022] Figure 4 A schematic diagram of another implementation flow of a sub-process in the air-conditioning control method based on user behavior provided in an embodiment of the present invention;
[0023] Figure 5 A schematic diagram of another implementation flow of a sub-process in the air-conditioning control method based on user behavior provided in an embodiment of the present invention;
[0024] Figure 6 A schematic diagram of another implementation flow of a sub-process in the air-conditioning control method based on user behavior provided in an embodiment of the present invention;
[0025] Figure 7 A schematic diagram of another implementation flow of a sub-process in the air-conditioning control method based on user behavior provided in an embodiment of the present invention;
[0026] Figure 8 A schematic block diagram of an air-conditioning control device based on user behavior provided by an embodiment of the present invention;
[0027] Figure 9 A schematic block diagram of an air conditioner provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0029] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0030] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0031] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0032] See also Figure 1 , Figure 1 The flow chart of the air-conditioning control method based on user behavior provided by the embodiment of the present invention is applied to air-conditioning.
[0033] S1. Obtain historical operation monitoring data of a target air conditioner, and process the historical operation monitoring data to obtain polynomial characteristic data.
[0034] In the embodiment of the present application, the historical operation monitoring data is the historical data of the target air conditioner during its operation. This historical operation monitoring data includes parameters such as the historical target air conditioner's set temperature and compressor frequency. In the embodiment of the present application, the historical operation monitoring data of the target air conditioner is obtained and processed to obtain polynomial feature data. This allows the feature parameters in the historical operation monitoring data to be combined, thereby facilitating the training of the prediction model and making the prediction results of the prediction model more accurate.
[0035] See also Figure 2 and Figure 3 , Figure 2 A specific implementation of step S1 is shown. Figure 3 A schematic diagram of data combination of historical operation monitoring data according to an embodiment of the present invention is shown, and is described in detail as follows:
[0036] S11: Acquire the historical operation monitoring data of the target air conditioner.
[0037] S12: Identify the time information corresponding to each data in the historical operation monitoring data, wherein the time information includes the hour and week number of each data.
[0038] S13: performing data combination processing on the historical operation monitoring data according to a preset polynomial degree and the time information to obtain the polynomial feature data.
[0039] like Figure 3As shown in the figure, the first column is the time of day, from 0 o'clock to 23 o'clock; the second column is the week number of the week, which is numbered from 1 to 7; the third column is the degree of the polynomial, which takes values 1, 2, and 3; for a given feature parameter X, X, X2, and X3 are used as extended features. In an embodiment of the present application, after obtaining the historical operation monitoring data of the target air conditioner, the time information corresponding to each data in the historical operation monitoring data is identified, and the time and week number corresponding to each historical operation monitoring data are obtained. Among them, the week number corresponds to the number of the week. Then, according to the preset polynomial degree and the time information, the historical operation monitoring data are combined to obtain the polynomial feature data corresponding to each historical operation monitoring data.
[0040] Furthermore, the embodiment of the present application adopts the grid search GridSearchCV method of Python to perform data combination processing on the historical operation monitoring data to obtain polynomial feature data.
[0041] S2. Convert the polynomial feature data to obtain target data.
[0042] See also Figure 4 , Figure 4 A specific implementation of step S2 is shown, which is described in detail as follows:
[0043] S21: Sort the polynomial feature data in order of time series to obtain time series data.
[0044] S22: using the data at adjacent moments in the time series data as independent variable data and dependent variable data respectively, so as to convert the time series data into a supervised format and obtain the target data.
[0045] In this embodiment, in order to predict the state of the next n moments based on the state of the previous m moments to generate the control strategy for the air conditioner, this embodiment of the application needs to first convert the polynomial feature data into a supervisory format. Taking m = 2 and n = 2 as an example, that is, using the state of the previous two moments to predict the state of the next two moments, the data format after the supervisory conversion is shown in Table 1:
[0046] Sample No. Independent variable X Dependent variable Y 0 S0,S1 S2,S3 1 S1,S2 S3,S4 2 S2,S3 S4,S5 … … … n Sn-4,Sn-3 Sn-2,Sn-1
[0047] Table 1
[0048] By converting the data into a supervised format, multiple target data are formed. In Table 1, n supervised format target data are formed, where X is the independent variable input and Y is the dependent variable output.
[0049] S3. Training a prediction model based on the target data, and during the model training process, adjusting the parameters of the prediction model using a reward scaling function to obtain a target prediction model.
[0050] In the embodiment of the present application, the prediction algorithm used is any one of the long short-term memory model recurrent neural network (LSTM), XGBoost and the regression moving average model (ARIMA).
[0051] See also Figure 5 , Figure 5 A specific implementation of step S3 is shown, which is described in detail as follows:
[0052] S31: performing normalization processing on the target data to obtain normalized data.
[0053] S32: Divide the normalized data into a data set according to a preset ratio to obtain a training data set and a test data set.
[0054] S33: adopting a preset prediction algorithm, and constructing a virtual environment corresponding to the prediction model based on the preset prediction algorithm.
[0055] S34: Inputting the training data set and the test data set into the prediction model for training, and during the model training process, using the reward scaling function to adjust the parameters of the prediction model to obtain the target prediction model.
[0056] In the embodiments of the present application, the target data is first normalized to convert the target data into normalized data; then, the normalized data is partitioned according to a preset ratio to obtain a training data set and a test data set. The preset ratio is set according to the actual situation and is not limited here. In one specific embodiment, the data is partitioned into a training data set and a test data set at a ratio of 3:1. The training data set is used to train the prediction model; the test data set is used to evaluate the results of the model training.
[0057] Furthermore, the main components of reinforcement learning include agents, environments, etc. The agent is the carrier of the reinforcement learning algorithm, which affects the environment through actions. The environment can be a real environment (such as a heat pump unit) or a virtual environment. Obviously, using a real unit to train a reinforcement learning model is costly and inefficient, and is not very realistic. Therefore, in the embodiment of the present application, a prediction algorithm (LSTM, ARIMA, GRU) is used to construct a virtual environment model. After the virtual environment is constructed, the training data set and the test data set are input into the prediction model for training, and during the model training process, the reward scaling function is used to adjust the parameters of the prediction model to obtain the target prediction model.
[0058] Furthermore, the reward scaling function is:
[0059]
[0060] Here, x is the input variable.
[0061] In the embodiment of the present application, after adopting the reward scaling function, x approaches the maximum value of 0.25 in the range of (-0.5, 0.5). Beyond this range, the function value decreases sharply, and after exceeding the range of (-2, 2), the function value approaches 0.
[0062] See also Figure 6 , Figure 6 A specific implementation of step S34 is shown, which is described in detail as follows:
[0063] S341: Input the training data set and the test data set into the prediction model to perform data prediction and generate action states.
[0064] S342: Calculate the reward of the prediction model based on the action state using the reward scaling function.
[0065] S343: Adjust the parameters of the prediction model based on the reward of the prediction model, and re-predict the prediction model data until the training round is reached to obtain the target prediction model.
[0066] Among them, the intelligent agent refers to the carrier of the deep reinforcement learning algorithm, which can run related algorithms according to the environmental status and rewards to generate actions. Prediction-based air-conditioning virtual environment: an environmental model that simulates the status of air-conditioning equipment is constructed through prediction methods (such as GRU, LSTM, ARIMAX and other algorithms), which can generate corresponding states and rewards for the actions of the intelligent agent. Action refers to the instructions generated by the intelligent agent, such as turning on and off the machine, adjusting the set temperature, etc. The subscript t represents a certain moment. Action state: the state of the air conditioner at a certain moment, such as temperature, pressure, power consumption, compressor frequency, indoor temperature, outdoor temperature, etc. Reward: The reward value at a certain moment is a scalar that measures the effectiveness of the action.
[0067] Furthermore, the embodiment of the present application uses a reinforcement learning algorithm to train the prediction model to obtain the optimal energy-saving strategy. Specifically, the training data set and the test data set are input into the prediction model to perform data prediction, generate future action states, and generate actions through the agent, and then generate another future action state through reinforcement learning. The reward is calculated based on the future action state and another future action state according to the reward scaling function; then the agent parameters and model parameters are adjusted based on the reward, the action state is regenerated, and then the model training is performed until the training round is reached to obtain the target prediction model. In the embodiment of the present application, the role of the reward scaling function is to encourage the parameters passed into the function to be close to 0. When the parameter difference exceeds 0.5, the function value will shrink sharply.
[0068] S4. Generate an air-conditioning control strategy based on the target prediction model, so that the target air-conditioning completes adjustment control based on the air-conditioning control strategy.
[0069] In the embodiment of the present application, since the prediction model has been trained in the above steps, the air-conditioning control strategy corresponding to the target air-conditioning to be controlled is output through the target prediction model, thereby completing the control of the target air-conditioning and achieving the purpose of air-conditioning energy saving.
[0070] See also Figure 7 , Figure 7 A specific implementation of step S4 is shown, which is described in detail as follows:
[0071] S41: Acquire current status information of the target air conditioner.
[0072] S42: Generate the air conditioning control strategy based on the target prediction model.
[0073] S43: Adjusting the current state information based on the air-conditioning control information to complete adjustment control of the target air-conditioner.
[0074] In the embodiment of the present application, since the state of the current target air conditioner needs to be changed to achieve energy saving of the air conditioner, it is necessary to obtain the current state information of the target air conditioner, then generate the air conditioner control strategy based on the target prediction model, and finally adjust the current state information based on the air conditioner control information to complete the adjustment and control of the target air conditioner. Among them, the current state information includes information such as the current set temperature and compressor frequency of the target air conditioner. The air conditioner control strategy is a strategy for adjusting the target air conditioner, which includes adjusting the temperature adjustment amount and compressor frequency of the target air conditioner, etc.
[0075] In an embodiment of the present invention, historical operation monitoring data of a target air conditioner is obtained, and the historical operation monitoring data is processed to obtain polynomial feature data; the polynomial feature data is converted to obtain target data; a prediction model is trained based on the target data, and during the model training process, a reward scaling function is used to adjust the parameters of the prediction model to obtain a target prediction model; an air conditioning control strategy is generated based on the target prediction model, so that the target air conditioner completes adjustment control based on the air conditioning control strategy. In an embodiment of the present invention, data processing and conversion processing are performed based on the historical operation monitoring data of the target air conditioner to obtain target data, and a prediction model is trained based on the target data and the reward scaling function, so that the prediction model can generate an accurate air conditioning control strategy, which is beneficial to improving the control accuracy of the air conditioner. Since the air conditioner can be accurately adjusted, the purpose of energy saving of the air conditioner is achieved.
[0076] The embodiment of the present invention further provides an air conditioning control device based on user behavior, which is used to execute any embodiment of the aforementioned air conditioning control method based on user behavior. Figure 8 , Figure 8 A schematic block diagram of an air-conditioning control device based on user behavior provided by an embodiment of the present invention.
[0077] Among them, such as Figure 8 As shown, the air-conditioning control device 5 based on user behavior includes a data acquisition unit 51 , a data conversion unit 52 , a model training unit 53 and an air-conditioning control unit 54 .
[0078] A data acquisition unit 51 is configured to acquire historical operation monitoring data of a target air conditioner and process the historical operation monitoring data to obtain polynomial characteristic data;
[0079] A data conversion unit 52 is used to convert the polynomial feature data to obtain target data;
[0080] A model training unit 53 is configured to train a prediction model based on the target data, and to adjust parameters of the prediction model using a reward scaling function during the model training process to obtain a target prediction model;
[0081] The air conditioning control unit 54 is configured to generate an air conditioning control strategy based on the target prediction model, so that the target air conditioning is adjusted and controlled based on the air conditioning control strategy.
[0082] Furthermore, the model training unit 53 includes:
[0083] A normalization processing unit, configured to perform normalization processing on the target data to obtain normalized data;
[0084] A data partitioning unit is used to partition the normalized data according to a preset ratio to obtain a training data set and a test data set;
[0085] A virtual environment construction unit, configured to adopt a preset prediction algorithm and construct a virtual environment corresponding to the prediction model based on the preset prediction algorithm;
[0086] The target prediction model generating unit is used to input the training data set and the test data set into the prediction model for training, and during the model training process, use the reward scaling function to adjust the parameters of the prediction model to obtain the target prediction model.
[0087] Furthermore, the target prediction model generation unit includes:
[0088] An action state generating unit, configured to input the training data set and the test data set into the prediction model to perform data prediction and generate an action state;
[0089] a reward calculation unit, configured to calculate a reward of the prediction model based on the action state using the reward scaling function;
[0090] A model parameter adjustment unit is used to adjust the parameters of the prediction model based on the reward of the prediction model, and re-predict the prediction model data until the training round is reached to obtain the target prediction model.
[0091] Furthermore, the reward scaling function is:
[0092]
[0093] Here, x is the input variable.
[0094] Furthermore, the data acquisition unit 51 includes:
[0095] a historical operation monitoring data acquisition unit, configured to acquire the historical operation monitoring data of the target air conditioner;
[0096] A time information identification unit, configured to identify the time information corresponding to each data in the historical operation monitoring data, wherein the time information includes the hour and week number of each data;
[0097] The data combination processing unit is used to perform data combination processing on the historical operation monitoring data according to a preset polynomial degree and the time information to obtain the polynomial characteristic data.
[0098] Furthermore, the data conversion unit 52 includes:
[0099] A data sorting unit, configured to sort the polynomial feature data in a time sequence to obtain time series data;
[0100] The target data generating unit is used to use the data at adjacent moments in the time series data as independent variable data and dependent variable data respectively, so as to convert the time series data into a supervised format to obtain the target data.
[0101] Furthermore, the air conditioning control unit 54 includes:
[0102] a current state information acquiring unit, configured to acquire current state information of the target air conditioner;
[0103] an air conditioning control strategy generating unit, configured to generate the air conditioning control strategy based on the target prediction model;
[0104] The current state information adjustment unit is configured to adjust the current state information based on the air-conditioning control information to complete adjustment control of the target air-conditioning.
[0105] In an embodiment of the present invention, historical operation monitoring data of a target air conditioner is obtained, and the historical operation monitoring data is processed to obtain polynomial feature data; the polynomial feature data is converted to obtain target data; a prediction model is trained based on the target data, and during the model training process, a reward scaling function is used to adjust the parameters of the prediction model to obtain a target prediction model; an air conditioning control strategy is generated based on the target prediction model, so that the target air conditioner completes adjustment control based on the air conditioning control strategy. In an embodiment of the present invention, data processing and conversion processing are performed based on the historical operation monitoring data of the target air conditioner to obtain target data, and a prediction model is trained based on the target data and the reward scaling function, so that the prediction model can generate an accurate air conditioning control strategy, which is beneficial to improving the control accuracy of the air conditioner. Since the air conditioner can be accurately adjusted, the purpose of energy saving of the air conditioner is achieved.
[0106] The above-mentioned air conditioning control device based on user behavior can be implemented in the form of a computer program. The computer program can be used in Figure 9 The air conditioner shown is running.
[0107] See also Figure 9 , Figure 9 5. The air conditioner 500 includes a processor 502, a memory, and a network interface 505 connected via a device bus 501. The memory may include a storage medium 503 and an internal memory 504.
[0108] The storage medium 503 may store an operating device 5031 and a computer program 5032. When the computer program 5032 is executed, the processor 502 may execute an air conditioning control method based on user behavior.
[0109] The processor 502 is used to provide computing and control capabilities to support the operation of the entire air conditioner 500.
[0110] The internal memory 504 provides an environment for the operation of the computer program 5032 in the storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute the air conditioning control method based on user behavior.
[0111] The network interface 505 is used for network communication, such as providing data information transmission, etc. Those skilled in the art will understand that the structure shown in the embodiment of the present application is merely a block diagram of a portion of the structure related to the solution of the present invention, and does not constitute a limitation on the air conditioner 500 to which the solution of the present invention is applied. The specific air conditioner 500 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0112] The processor 502 is configured to execute a computer program 5032 stored in the memory to implement the air-conditioning control method based on user behavior disclosed in an embodiment of the present invention.
[0113] Those skilled in the art will appreciate that the embodiments of the air conditioner shown in the embodiments of this application do not constitute a limitation on the specific configuration of the air conditioner. In other embodiments, the air conditioner may include more or fewer components than shown, or may combine certain components, or have different component arrangements. For example, in some embodiments, the air conditioner may only include a memory and a processor. In such embodiments, the structure and function of the memory and processor are consistent with those in the embodiments shown above and will not be further described here.
[0114] It should be understood that in the embodiment of the present invention, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0115] In another embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium may be either non-volatile or volatile. The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the user behavior-based air conditioning control method disclosed in an embodiment of the present invention.
[0116] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0117] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is merely a logical function division. In actual implementation, there may be other division methods, or units with the same function may be combined into one unit. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices or units, or may be an electrical, mechanical or other form of connection.
[0118] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected based on actual needs to achieve the objectives of the embodiments of the present invention.
[0119] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0120] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing an air conditioner to perform all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.
[0121] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. An air conditioning control method based on user behavior, characterized in that: include: Acquiring historical operation monitoring data of a target air conditioner, and processing the historical operation monitoring data to obtain polynomial characteristic data; Converting the polynomial feature data to obtain target data; Normalizing the target data to obtain normalized data; Dividing the normalized data into a training data set and a test data set according to a preset ratio; Using a preset prediction algorithm, and constructing a virtual environment corresponding to the prediction model based on the preset prediction algorithm; Inputting the training data set and the test data set into the prediction model to perform data prediction and generate action states; Calculating a reward of the prediction model based on the action state using a reward scaling function; Adjusting the parameters of the prediction model based on the reward of the prediction model, and re-predicting the prediction model data until a training round is reached to obtain a target prediction model; An air conditioning control strategy is generated based on the target prediction model, so that the target air conditioning completes adjustment control based on the air conditioning control strategy.
2. The air conditioning control method based on user behavior according to claim 1, characterized in that: The reward scaling function is: ; Here, x is the input variable.
3. The air conditioning control method based on user behavior according to claim 1, characterized in that: The acquiring of historical operation monitoring data of the target air conditioner and processing of the historical operation monitoring data to obtain polynomial characteristic data include: Acquiring the historical operation monitoring data of the target air conditioner; Identify the time information corresponding to each data in the historical operation monitoring data, wherein the time information includes the hour and week number of each data; According to the preset polynomial degree and the time information, the historical operation monitoring data is subjected to data combination processing to obtain the polynomial characteristic data.
4. The air conditioning control method based on user behavior according to any one of claims 1 to 3, characterized in that: The converting the polynomial feature data to obtain target data includes: Sorting the polynomial feature data in the order of time series to obtain time series data; The data at adjacent moments in the time series data are respectively used as independent variable data and dependent variable data, so as to convert the time series data into a supervised format and obtain the target data.
5. The air conditioning control method based on user behavior according to any one of claims 1 to 3, characterized in that: Generating an air conditioning control strategy based on the target prediction model so that the target air conditioning completes adjustment control based on the air conditioning control strategy includes: Obtaining current status information of the target air conditioner; generating the air conditioning control strategy based on the target prediction model; The current state information is adjusted based on the air-conditioning control strategy to complete adjustment control of the target air-conditioning.
6. An air conditioning control device based on user behavior, characterized in that: include: a data acquisition unit, configured to acquire historical operation monitoring data of a target air conditioner and process the historical operation monitoring data to obtain polynomial characteristic data; A data conversion unit, configured to convert the polynomial feature data to obtain target data; A normalization processing unit, configured to perform normalization processing on the target data to obtain normalized data; A data partitioning unit is used to partition the normalized data according to a preset ratio to obtain a training data set and a test data set; A virtual environment construction unit, configured to adopt a preset prediction algorithm and construct a virtual environment corresponding to the prediction model based on the preset prediction algorithm; An action state generating unit, configured to input the training data set and the test data set into the prediction model to perform data prediction and generate an action state; a reward calculation unit, configured to calculate a reward of the prediction model based on the action state using a reward scaling function; A model parameter adjustment unit, configured to adjust the parameters of the prediction model based on the reward of the prediction model, and re-predict the prediction model data until a training round is reached to obtain a target prediction model; An air conditioning control unit is used to generate an air conditioning control strategy based on the target prediction model, so that the target air conditioning completes adjustment control based on the air conditioning control strategy.
7. An air conditioner comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the air-conditioning control method based on user behavior according to any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to execute the air-conditioning control method based on user behavior according to any one of claims 1 to 5.
Citation Information
Patent Citations
Data processing method and device, terminal equipment and medium
CN114418776A
Energy-saving control method and system based on dynamic air conditioner operation data and medium
CN116085953A