Air conditioning energy-saving control method for accelerating time series confidence domain policy optimization

By combining gated cyclic units, Grover's method, and confidence domain strategy optimization, the control accuracy and stability issues of air conditioners in nonlinear time-varying systems are solved, resulting in reduced air conditioner power consumption and improved efficiency.

CN117029190BActive Publication Date: 2026-07-21GUANGXI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGXI UNIV
Filing Date
2023-08-11
Publication Date
2026-07-21

Smart Images

  • Figure CN117029190B_ABST
    Figure CN117029190B_ABST
Patent Text Reader

Abstract

The application provides an air conditioner energy-saving control method for accelerating time sequence confidence domain strategy optimization, which combines a gated recurrent unit with a confidence domain strategy optimization method and is used for energy-saving control of an air conditioner. First, a first stage of the method collects start-stop time sequences of a user air conditioner and electricity consumption data of a city where the user is located, and trains the data samples as data samples of the gated recurrent unit. Meanwhile, a Grover method is used to accelerate the training process of the gated recurrent unit. After the training of the gated recurrent unit is completed, the start-stop time sequences of the user air conditioner and the electricity consumption of the city where the user is located are predicted. Secondly, a second stage of the method uses a confidence domain strategy optimization method to control the air conditioner according to the prediction result of the gated recurrent unit, so that the electricity use efficiency of the air conditioner is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of artificial intelligence, deep learning, quantum control, quantum technology and air conditioning energy-saving control, and relates to an air conditioning energy-saving control method for accelerating time series confidence domain strategy optimization, which is applicable to the energy-saving control of air conditioners. Background Technology

[0002] Existing air conditioning energy-saving control methods can be divided into PID control, neural network control, and improved fuzzy control. PID controllers, due to their high reliability and stability, remain the most widely used. However, this control method only achieves good control results when the controlled object and the external environment are relatively stable. In reality, an air conditioning system is a typical nonlinear, time-varying system with unknown control model parameters, so PID control is not ideal, resulting in huge energy consumption. Neural network control methods achieve the desired control effect through extensive data training. Recurrent neural network control is the primary method, but it suffers from long-term dependency and gradient vanishing problems. Fuzzy control, while not requiring a precise mathematical model of the controlled object and primarily relying on its own control rules to change the magnitude of the control quantity, exhibits good control effects for air conditioners. However, fuzzy control methods have the following drawbacks:

[0003] (1) The design of fuzzy control is not systematic and it is difficult to control complex systems effectively.

[0004] (2) How to obtain fuzzy rules and membership functions, i.e. system design methods, cannot be guaranteed by relying solely on experience, as this will not ensure the system's control accuracy and stability.

[0005] (3) Simple fuzzy processing of information will lead to a deterioration in the dynamic quality and control accuracy of the system. To improve the control accuracy of the system, it is necessary to increase the number of quantization levels, which leads to an expansion of the search range of rules and a reduction in decision-making speed. Summary of the Invention

[0006] This invention proposes an air conditioning energy-saving control method that accelerates time-series confidence domain policy optimization. It combines a gated loop unit, a confidence domain policy optimization method, and the Grover method for energy-saving control of air conditioners. This method reduces energy loss during air conditioner use, improves energy efficiency and intelligence, and includes the following steps:

[0007] Step (1): Collect data on the social electricity consumption and the start-stop sequence of the user's air conditioner in the city where the user is located, and use the collected social electricity consumption and the start-stop sequence of the user's air conditioner in the city where the user is located as data samples for training the gated loop unit.

[0008] Step (2): After collecting data in step (1), the Grover method is used to accelerate the search for the optimal number of hidden units in the gated recurrent unit, thereby improving the prediction speed and accuracy of the gated recurrent unit; the Grover method can transform the search problem from the classical Step to narrow down In this step, Grover's method achieves a quadratic increase in search efficiency compared to branch-and-bound, breadth-first, and depth-first search methods, thus demonstrating quantum speedup. The speedup effect is more pronounced as N increases. Grover's method uses two registers; the first register is used to store... One qubit, the second register is used to store the Oracle's workspace;

[0009] The Grover method first... As the initial state of the quantum system, the input ground state is then subjected to a Hadamard transformation to make the initial state... Transform into an equal-weighted superposition state Secondly, use Iterative operators. Application of the Grover method. Second-rate The iterative operator completes the iteration. If we use pi (π) as the unit of measurement, then by measuring it, we can find the solution to the problem.

[0010] The Grover method generally consists of four steps: quantum state preparation, quantum unitary evolution, quantum measurement, and output of results; the basic process of the Grover method is as follows:

[0011] Step (2.1) First, initialize and prepare the equal-weighted superposition state. Perform Hadamard transformation on the input ground state to obtain the equal-weighted superposition state of all calculated ground states, as shown in Equation (1):

[0012] (1)

[0013] In the formula, It is a superposition of equal weights. For Hadamard transform, For an individual Initial state, For an indicator register;

[0014] Step (2.2) constructs the Oracle by constructing a mapping. This causes the phase of the target term to flip, but the sign of any other term orthogonal to the target term remains unchanged;

[0015] Step (2.3) involves constructing a unitary matrix. This makes the amplitude of the target state relative to the average amplitude. Flip it over, and finally apply... The iterative operator is repeatedly iterated to accelerate the training speed of the gated recurrent unit;

[0016] Step (3): Two gated loop units are used to predict the start-up and shutdown sequence of the user's air conditioner and the electricity consumption of the city where the user is located. The gated loop unit has a gating mechanism with two gates: a reset gate and an update gate. The reset gate in the gated loop unit can mine the short-term and long-term dependencies between time series data. By controlling the opening or closing of this gate, the purpose of forgetting past information can be achieved. The update gate is used to determine which information from the previous time to retain and which to forget at the current time, and can obtain the long-term dependencies between time series data. The calculation process of each gated loop unit is as follows:

[0017] (2)

[0018] (3)

[0019] (4)

[0020] (5)

[0021] In the formula, To reset the door, To update the door, To reset information, for The output of the hidden layer at all times, For the current input, This is the output of the hidden layer from the previous time step. For the Sigmoid function, it will time and The output value is controlled between [0,1], where 0 represents forgetting the information and 1 represents that all information is passed at the current moment. For activation function, , , , , and These are all training parameter matrices in the network. , and All are biased. For matrix multiplication;

[0022] Step (4): After obtaining the prediction result through step (3), the confidence region policy optimization method is used to give the policy for the next time step, thereby giving the air conditioner switch control command to control the air conditioner accordingly; the confidence region policy optimization method updates the policy by increasing the state value, adding other terms to the old reward function to represent the reward function for the next time step, and the reward function for the next time step. for:

[0023] (6)

[0024] In the formula, The old reward function; for The state at any given moment, Right now Real-time information on the user's current air conditioning power consumption and start / stop sequence; for Momentary actions Right now Monitor the user's actions when switching the air conditioner on and off; for The advantage function at time, Right now The advantages of the action-value function at time step compared to the value function of the current state. For the current moment, As a discount factor, This represents the mathematical expectation of the advantage function at the current moment;

[0025] Advantage function The calculation formula is:

[0026] (7)

[0027] In the formula, For state action value function, It is a state-value function; This is the current state. This refers to the user's current power consumption and start / stop sequence of the air conditioner; For the user's current action, That is, the current action of the user's air conditioner switch;

[0028] Meanwhile, relative entropy is used to measure the difference in policy distribution between the previous and next time steps, reducing the difference between the next and previous time steps in the gradient step. The confidence region policy optimization method ensures that the objective function of the policy in the next time step is monotonically constant and can automatically update the step size, solving the problem of step size selection in the policy gradient algorithm and ensuring that the next time step policy has better performance than the old policy.

[0029] The present invention has the following advantages and effects compared with the prior art:

[0030] (1) Compared with existing gated recurrent unit prediction methods and long short-term memory network prediction methods, combining gated recurrent units with Grover's method can predict the hourly electricity consumption of the user's city and the start-up and shutdown sequence of the user's air conditioner more quickly and accurately. At the same time, gated recurrent units have fewer training parameters and faster convergence speed than long short-term memory networks, and have better prediction performance, especially when the training data is large.

[0031] (2) The confidence region policy optimization method used can solve the problem of step size selection in the policy gradient algorithm, and the calculation process is simpler than other methods. Attached Figure Description

[0032] Figure 1 This is an overall flowchart of the method of the present invention.

[0033] Figure 2 This is a quantum circuit diagram of the method of the present invention. Detailed Implementation

[0034] This invention proposes an air conditioning energy-saving control method for accelerating time series confidence region strategy optimization, which is described in detail below with reference to the accompanying drawings:

[0035] Figure 1 This is an overall flowchart of the method of the present invention. First, in the first stage, the start-stop sequence of the user's air conditioner and the hourly electricity consumption data of the user's city are collected. This data is used as training samples for the gating loop unit, and the Grover method is used to accelerate the training of the gating loop unit. Then, after the gating loop unit is trained, the start-stop sequence of the user's air conditioner and the hourly electricity consumption of the user's city are predicted respectively. Finally, based on the prediction results of the gating loop unit, a confidence region strategy optimization method is used to give corresponding air conditioner on / off control commands for energy-saving control of the air conditioner.

[0036] Figure 2 This is a quantum circuit diagram of the method of this invention. The Grover method has two registers. The first register is used to store... One qubit, and the second register is used to store the Oracle's workspace. This method first... As the initial state of the quantum system, the input ground state is then subjected to a Hadamard transformation to make the initial state... Transform into an equal-weighted superposition state Secondly, use Iterative operators. Application of the Grover method. Second-rate The iterative operator completes the iteration. If we take pi (π), then by measuring it, we can find the solution to the problem.

[0037] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for energy-saving air conditioning control that accelerates time-series confidence region strategy optimization, characterized in that, A gated loop unit, a confidence region strategy optimization method, and the Grover method are combined for energy-saving control of air conditioning. The steps in the process are as follows: Step (1): Collect data on the social electricity consumption and the start-stop sequence of the user's air conditioner in the city where the user is located, and use the collected social electricity consumption and the start-stop sequence of the user's air conditioner in the city where the user is located as data samples for training the gated loop unit. Step (2): After collecting data in step (1), the Grover method is used to accelerate the search for the optimal number of hidden units in the gated loop unit; the Grover method has two registers; the first register is used to store One qubit, the second register is used to store the Oracle's workspace; The Grover method first will... As the initial state of the quantum system, the input ground state is then subjected to a Hadamard transformation to make the initial state... Transform into an equal-weighted superposition state Secondly, use Iterative operators; application of Grover's method Second-rate The iterative operator completes the iteration. If we use pi (π) as the mathematical constant, then by measuring it, we can find the solution to the problem. The Grover method consists of four steps: quantum state preparation, quantum unitary evolution, quantum measurement, and output results. The basic process of the Grover method is as follows: Step (2.1) First, initialize and prepare the equal-weighted superposition state. Perform Hadamard transformation on the input ground state to obtain the equal-weighted superposition state of all calculated ground states, as shown in Equation (1): (1) In the formula, It is a superposition of equal weights. For Hadamard transform, For an individual Initial state, For an indicator register; Step (2.2) constructs the Oracle by constructing a mapping. This causes the phase of the target term to flip, but the sign of any other term orthogonal to the target term remains unchanged; Step (2.3) involves constructing a unitary matrix. This makes the amplitude of the target state relative to the average amplitude. Flip it over, and finally apply... The iterative operator is repeatedly iterated to accelerate the training speed of the gated recurrent unit; Step (3): Two gated loop units are used to predict the start-up and shutdown sequence of the user's air conditioner and the electricity consumption of the city where the user is located; the gated loop unit has a gate control mechanism, with two gate control units: a reset gate and a refresh gate; the calculation process of each gated loop unit is as follows: (2) (3) (4) (5) In the formula, To reset the door, To update the door, To reset information, for The output of the hidden layer at all times, For the current input, This is the output of the hidden layer from the previous time step. For the Sigmoid function, it will time and The output value is controlled between [0,1], where 0 represents forgetting the information and 1 represents that all information is passed at the current moment. For activation function, , , , , and These are all training parameter matrices in the network. , and All are biased. For matrix multiplication; Step (4): After obtaining the prediction result through step (3), the confidence region policy optimization method is used to give the policy for the next time step, thereby giving the air conditioner switch control command to control the air conditioner to switch on and off; the confidence region policy optimization method updates the policy by increasing the state value, and the reward function of the confidence region policy optimization method for the next time step is... for: (6) In the formula, The old reward function; for The state at any given moment, Right now Real-time information on the user's current air conditioning power consumption and start / stop sequence; for Actions at any moment Right now Monitor the user's actions when switching the air conditioner on and off; for The advantage function at time, Right now The advantages of the action-value function at time step compared to the value function of the current state. For the current moment, As a discount factor, This represents the mathematical expectation of the advantage function at the current moment; Advantage function The calculation formula is: (7) In the formula, For state action value function, It is a state-value function; This is the current state. This refers to the user's current power consumption and start / stop sequence of the air conditioner; For the user's current action, This refers to the current action of the user's air conditioner switch.