A hot water system control method based on user portrait learning and deep reinforcement learning

Through the fusion of multi-source heterogeneous data features and deep reinforcement learning methods, the problems of rapid investment and inefficient energy efficiency of hot water heating devices are solved, personalized control and efficient user behavior adaptability are achieved, and user comfort and system robustness are improved.

CN116772426BActive Publication Date: 2025-08-26GUANGXI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310752083.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-25
Publication Date
2025-08-26
Estimated Expiration
2043-06-25

AI Technical Summary

Technical Problem

The existing control methods of hot water heating devices cannot be put into use quickly, and the comprehensive utilization of multi-source heterogeneous data is lacking, resulting in low energy efficiency and poor user comfort.

Method used

Combining the fusion of multi-source heterogeneous data features, user portraits, unsupervised learning and deep reinforcement learning, users can be personalized by combining agents and hot water heating devices.

Benefits of technology

It has achieved rapid input into use, improved energy utilization efficiency and user comfort, reduced the risk of misjudgment and wrong decision-making, and has good interpretability and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116772426B_ABST
    Figure CN116772426B_ABST
Patent Text Reader

Abstract

This invention proposes a hot water system control method that uses user profiling and deep reinforcement learning. The main steps of this method include collecting a large amount of historical multi-source heterogeneous data from users, fusing and extracting features from this multi-source heterogeneous data using multi-channel convolution, generating user hot water usage profiles from the fused and extracted feature data using the K-Means method, and interacting with users in real time using an online deep reinforcement learning method to continuously improve the policy model. This method can address the problem of data-driven methods in hot water systems being unable to be quickly implemented, enabling personalized hot water control, optimizing energy efficiency, and improving user comfort. It has excellent scalability and adaptability, and the overall framework offers excellent interpretability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of energy optimization operation of power system buildings, and involves multi-source heterogeneous data feature fusion extraction, user profiling, unsupervised learning and deep reinforcement learning control methods, and is suitable for the control of hot water heating devices of random behavior users. Background Art

[0002] Hot water usage is highly random, and hot water heating devices are complex and diverse. This results in some data-driven approaches requiring long installation and commissioning times, preventing immediate operational readiness, and disrupting occupant comfort. Furthermore, previous data-driven approaches lack the comprehensive utilization of heterogeneous data from multiple sources, resulting in poor interpretability, an inability to meet personalized user needs, control strategy limitations, and low energy efficiency.

[0003] Hot water usage behavior is highly random, leading to a conservative control approach for hot water heating devices, known as a two-point control method. This method turns on the heat pump when the water tank temperature falls below a lower threshold and turns off when the temperature rises above a higher threshold. While this is a simple and easy-to-use control method, it is independent of occupant behavior and consumes significant energy due to excessive hot water preparation.

[0004] Therefore, a hot water system control method based on learning user profiles and deep reinforcement learning is proposed to solve the problems of inability to quickly put into use, poor interpretability and low energy efficiency. Summary of the Invention

[0005] This paper proposes a hot water system control method that uses user profiling and deep reinforcement learning. This method combines multi-source heterogeneous data feature fusion extraction, user profiling, unsupervised learning, and deep reinforcement learning to control a hot water heating device. This method achieves personalized hot water control, optimizes energy efficiency, and improves user comfort. The steps involved in its use are:

[0006] Step (1): Collect N D The historical multi-source heterogeneous data of users; the historical multi-source heterogeneous data of the i-th user includes: hourly weather data in the past 30 days in, The weather data for the first hour of the first day; The weather data for the second hour of the first day; The weather data is for the 24th hour of the first day; The weather data for the first hour of the second day; The weather data for the second hour of the second day; The weather data is for the 24th hour of the second day; The weather data is for the first hour of the 30th day; The weather data is for the second hour of the 30th day; Weather data for the 24th hour of the 30th day; temperature data for each hour of the past 30 days in, The temperature data for the first hour of the first day; The temperature data for the second hour of the first day; The temperature data for the 24th hour of the first day; The temperature data for the first hour of the second day; The temperature data for the second hour of the second day; The temperature data for the 24th hour of the second day; The temperature data for the first hour of the 30th day; The temperature data is for the second hour on the 30th day; Temperature data for the 24th hour of the 30th day; water heater on / off status data for each hour in the past 30 days in, The water heater on / off status data for the first hour of the first day; The water heater on / off status data for the second hour of the first day; The water heater on / off status data for the 24th hour of the first day; The water heater on / off status data for the first hour of the second day; The water heater on / off status data for the second hour of the second day; The water heater on / off status data for the 24th hour of the second day; The water heater on / off status data for the first hour of the 30th day; The water heater on / off status data for the 2nd hour on the 30th day; The water heater on / off status data for the 24th hour of the 30th day; the hot water usage data for each hour in the past 30 days in, This is the hot water usage data for the first hour of the first day; This is the hot water usage data for the second hour of the first day; Hot water usage data for the 24th hour of the first day; This is the hot water usage data for the first hour of the second day; This is the hot water usage data for the second hour of the second day; Hot water usage data for the 24th hour of the second day; Hot water usage data for the first hour of the 30th day; This is the hot water usage data for the second hour on the 30th day; Hot water usage data for the 24th hour of the 30th day; date data for each hour of the past 30 days in, The date data is the first hour of the first day; The date data is the second hour of the first day; The date data is the 24th hour of the first day; The date data is the first hour of the second day; The date data is the second hour of the second day; The date data is the 24th hour of the second day; The date data is the first hour of the 30th day; The date data is the second hour of the 30th day; Date data for the 24th hour of the 30th day; day of the week data for each hour in the past 30 days in, The weekly data is the first hour of the first day; The weekly data is the second hour of the first day; The weekly data is the 24th hour of the first day; The weekly data is the first hour of the second day; The weekly data is the second hour of the second day; The weekly data is the 24th hour of the second day; The weekly data is the first hour of the 30th day; The weekly data is the second hour of the 30th day; Weekly data for the 24th hour of the 30th day; hourly data for each hour of the past 30 days in, The hourly data is for the first hour of the first day; The hourly data is for the second hour of the first day; The hourly data is the 24th hour of the first day; The hourly data is for the first hour of the second day; The hourly data is for the second hour of the second day; The hourly data is the 24th hour of the second day; The hourly data is the first hour of the 30th day; The hourly data is the second hour of the 30th day; The hourly data is the 24th hour of the 30th day;

[0007] Step (2): Splice each user’s historical multi-source heterogeneous information data to obtain a 3D tensor multi-source heterogeneous fusion sample T1 represents the 3D tensor multi-source heterogeneous fusion sample of the first user; T2 represents the 3D tensor multi-source heterogeneous fusion sample of the second user; Indicates the Nth D 3D tensor multi-source heterogeneous fusion samples of each user;

[0008] Step (3): Extract features from the 3D tensor multi-source heterogeneous fusion sample of each user to obtain the multi-source heterogeneous feature flattened sample of each user; extract features from the 3D tensor multi-source heterogeneous matrix sample T of the i-th user. i Perform multi-channel convolution operation; convolution kernel tensor M C The shape is (30,3,27); the output of the multi-channel convolution operation is a multi-source heterogeneous feature sample in, are all elements in the multi-source heterogeneous feature sample obtained by the multi-channel convolution operation of the i-th user; Z1 represents the multi-source heterogeneous feature sample of the first user; Z2 represents the multi-source heterogeneous feature sample of the second user; Indicates the Nth D The multi-source heterogeneous feature samples of each user; the multi-channel convolution operation process is as follows:

[0009] Z i =T i ⊙M C (1) Where, T i is the 3D tensor multi-source heterogeneous fusion sample of the i-th user; ⊙ is the multi-channel convolution operation; Z i is the multi-source heterogeneous feature sample of the i-th user;

[0010] Then, the multi-source heterogeneous feature samples are flattened to change the arrangement of the elements to obtain the multi-source heterogeneous feature flattened sample of the i-th user The multi-source heterogeneous feature flattening sample is a 616×1 matrix;

[0011] P i =Flatten(Z i ) (2)

[0012] Where Flatten(·) is the flattening function; P i Flatten the samples for the multi-source heterogeneous features of the i-th user;

[0013] Step (4): N D The multi-source heterogeneous feature expansion samples of each user are input into the K-Means method; the output of the K-Means method is the different types of hot water usage portraits: C1, C2, ...C k ; Among them, C1 is the first hot water usage profile; C2 is the second hot water usage profile; C k is the k-th hot water usage portrait; k is the number of types of hot water usage portraits output by the K-Means method; in addition to the k types of hot water usage portraits output by the K-Means method, the k+1-th hot water usage portrait is defined as an unknown portrait;

[0014] The operation process of the K-Means clustering method is:

[0015] (4.1) Randomly initialize k cluster labels, namely: C1, C2, ... C k ;

[0016] (4.2) Calculate the Euclidean distance between each multi-source heterogeneous feature expansion sample and all cluster labels:

[0017] d(P i ,C j )=sqrt((P i -C j ) 2 ) (3) Where, P i Expand the sample for the multi-source heterogeneous features of the i-th user; C j is the jth cluster label; d(P i ,C j ) is the Euclidean distance between the multi-source heterogeneous feature expansion sample of the i-th user and the j-th cluster label; sqrt(·) is the square root function;

[0018] (4.3) Assign each multi-source heterogeneous feature expansion sample to the cluster label with the smallest Euclidean distance;

[0019] (4.4) Update cluster labels; for each cluster label, calculate the mean of the multi-source heterogeneous feature expansion samples belonging to the cluster label as the new cluster label;

[0020]

[0021] Where C j represents the updated cluster label; N j Belongs to cluster label C j The number of samples of multi-source heterogeneous feature expansion; P m Indicates belonging to the cluster centroid C j The mth multi-source heterogeneous feature expansion sample;

[0022] (4.5) Repeat steps (4.2)-(4.4) until the maximum number of iterations is met;

[0023] (4.6) Output clustering results; K-Means method outputs k cluster labels; C1, C2, ... C k , each cluster label is a hot water usage portrait; C1 is the first hot water usage portrait; C2 is the second hot water usage portrait; C k Use portrait for the kth type of hot water;

[0024] N DFlattened samples of multi-source heterogeneous features for users With k types of hot water usage profiles C1, C2, ... C k Form the initial database; Step (5): Put the water heater into use by the new user and collect the new user's real-time multi-source heterogeneous data; the new user's real-time multi-source heterogeneous data includes: hourly weather data Hourly temperature data Hourly water heater on / off status data Hot water usage per hour Hourly date data Hourly week data Hourly data for each hour Where d is the number of days the water heater is put into operation and t is the number of hours in the day;

[0025] The water heater consists of an agent and a hot water heating device; the agent contains an initial database, a hot water image neural network discriminator, and a proximal strategy optimization method;

[0026] Before the water heater is put into use by a new user, the proximal policy optimization method is trained using offline training so that the proximal policy optimization method can control the on / off state of the water heater. The input matrix of the proximal policy optimization method is Among them, e k ∈{1,2,...,k,k+1}, indicating which hot water usage profile the new user's hot water usage behavior belongs to; e k =1 indicates that the new user's hot water usage behavior belongs to the first hot water usage profile; e k =2 means that the new user's hot water usage behavior belongs to the second hot water usage profile; k =k indicates that the new user's hot water usage behavior belongs to the kth hot water usage profile; e k =k+1 indicates that the new user's hot water usage behavior belongs to an unknown profile; the output of the proximal policy optimization method is the on / off status of the water heater;

[0027] Step (6): Splice the real-time multi-source heterogeneous data of the new user to obtain the new user's 2D tensor multi-source heterogeneous fusion sample Input R into the hot water image neural network discriminator; the output of the hot water image neural network discriminator is the probability value of R belonging to the 1st to kth category hot water usage image, which are

[0028] The input layer of the hot water image neural network discriminator is a convolutional layer; the hidden layer of the hot water image neural network discriminator consists of a width of W hw , the depth is W hdThe activation function of the fully connected layer in its hidden layer is the tanh() function; the output layer of the hot water image neural network discriminator has a width of k and a depth of W. od The activation function of the fully connected layer in its output layer is the softmax() function; the momentum stochastic gradient descent optimizer is used to train the hot water image neural network discriminator;

[0029] Step (7): Determine whether the new user's 2D tensor multi-source heterogeneous fusion sample R belongs to the hot water usage profile in the initial database;

[0030] when When both are less than 0.9, the hot water portrait neural network discriminator refuses to recognize; then the new user's hot water use behavior does not belong to the hot water use in the initial database, but belongs to the k+1th type of hot water use portrait, that is, an unknown portrait; at this time, e k =k+1;

[0031] when When it is greater than 0.9, it is determined that the new user's hot water usage behavior belongs to the i-th hot water usage profile in the initial database; at this time, e k =i;

[0032] Step (8): Online training of the proximal strategy optimization method is performed, and the control action of the hot water heating device is outputted at the same time. The online training process of the proximal strategy optimization method is as follows:

[0033] (8.1) Create a new user's hot water usage profile k and real-time multi-source heterogeneous data input into the proximal strategy optimization method;

[0034] (8.2) Generate N according to the current action function π(A|S;θ) and the evaluation function V(S;φ) P A sequence of trajectories; where θ is the network structure parameter of the action function, that is, the network weight and bias of the action function; φ is the network structure parameter of the evaluation function, that is, the network weight and bias of the evaluation function; A is the action output by the proximal policy optimization method, and there are two types of actions: one is to turn on the water heater switch, and the other is to turn off the water heater switch; S is the state input to the proximal policy optimization method, which includes the hot water usage profile of the new user and real-time multi-source heterogeneous data;

[0035] In addition, R PPO is the reward obtained when the proximal policy optimization method takes action on the heating device;

[0036] If there is a demand for hot water, reward R PPO for:

[0037] R PPO =-a reward ×Php -b reward ×max(40-T tank ,0)-c reward ×max(H time -24,0) (5)

[0038] If there is no hot water demand, reward R PPO for:

[0039] R PPO =-a reward ×P hp -c reward ×max(H time -24,0) (6) Where a reward is the energy term coefficient in the reward; b reward is the comfort factor in the reward; c reward P is the health item coefficient in the reward; hp is the energy consumed by the heating device; max(·,·) is the maximum value function; T tank is the temperature of the hot water tank; H time The duration of time when the temperature of the hot water tank last reached above 60°C;

[0040] Each trajectory sequence is:

[0041]

[0042] Where: t s Represents a certain moment, that is, the starting time of the current trajectory sequence; t s The state of the moment; t s +1 moment status; t s +N P -1 moment status; t s +N P The state of the moment; is An action in a state; is An action in a state; From the state Transition to rewards; From the state Transition to rewards;

[0043] When in state When , the probability of each action is calculated using π(A|S;θ), and the action is randomly selected according to the probability distribution At the beginning of training, t s =1, for the subsequent N P Trajectory sequence, t s ←t s +N P ;

[0044] (8.3) For t = t s +1,t s +2,...,t s +N P For these N trajectory sequences, calculate the discounted return G for each trajectory sequence t With advantage function D t ; Discounted return G t for:

[0045]

[0046] Where γ is the discount coefficient; R PPO,k From state S k Transition to S k+1 The reward value of b G To calculate G t The coefficient when ; if is the final state, then b G is 0, otherwise, b G is 1;

[0047] Advantage function D t for:

[0048]

[0049] Where λ is the smoothing factor; b D To calculate D t The coefficient of , if is the final state, then b D is 0, otherwise b D is 1;

[0050] (8.4) From this N P Learning in a sequence of trajectories:

[0051] (8.4.1) Randomly extract a trajectory of size M from the current trajectory sequence B A dataset where each element contains the corresponding discounted return and advantage function value;

[0052] (8.4.2) Minimize the evaluation loss function L by gradient descent critic(φ) is used to update the parameter φ of the evaluation function and evaluate the loss function L critic (φ) is:

[0053]

[0054] Where: G i represents the corresponding discounted return in the i-th element in the dataset;

[0055] (8.4.3) Normalize the advantage function value. The corresponding normalized discounted advantage function value of the i-th element in the data set is:

[0056]

[0057] Where i is the subscript number of each element in the data set, represents the normalized discounted advantage function value corresponding to the i-th element in the dataset; D i The advantage function value corresponding to the i-th element in the data set; the advantage function value corresponding to the first element in the D1 data set; the advantage function value corresponding to the second element in the D2 data set; M The Mth B The advantage function value corresponding to each element; mean(·,·,...,·) is the function for finding the average value; std(·,·,...,·) is the function for calculating the standard deviation;

[0058] (8.4.4) Minimize the action loss function L by gradient descent actor (θ) is used to update the parameters θ of the action function, the action loss function L actor (θ) is:

[0059]

[0060] where r i (θ) coefficient factor and entropy loss function They are:

[0061]

[0062]

[0063] Where min(·,·) is the minimum function; π)A i |S i ;θ) is in state S i Under given parameter θ, take action A i The probability of π(A i |S i θ old ) is in state S iUnder the given parameters θ before the current learning period old , take action A i The probability of r i (θ) coefficient factor; ε is the shear factor; is the entropy loss function; P z is the number of action types in the proximal strategy optimization method; π(A k |S i ;θ) is in state S i Under given parameter θ, take action A k The probability of; w is the entropy loss coefficient;

[0064] (8.5) The proximal strategy optimization method uses the updated action function to output the action for controlling the hot water heating device based on the input data in (8.1);

[0065] (8.6) Repeat (8.1)-(8.5), continuously update the proximal strategy optimization method, and output the action of controlling the hot water heating device.

[0066] The present invention has the following advantages and effects compared to the prior art:

[0067] (1) This method makes full use of multi-source heterogeneous data, including weather data, temperature data, water heater switch status data, hot water usage data, date data, week data and hourly data, to obtain information from multiple dimensions. This comprehensive consideration can more comprehensively and accurately understand the user's hot water needs and behavior patterns. In addition, compared with the patent method with an application date of June 30, 2022 and application number 202210755343.3, the method of the present patent introduces a rejection recognition mechanism for unknown portraits, which can judge and reject difficult-to-classify sample data and avoid incorrect classification results. The rejection recognition mechanism can improve the robustness of the system and reduce the risk of misjudgment and wrong decision-making.

[0068] (2) This method uses an online training method to optimize and adaptively adjust the strategy of the proximal strategy optimization method in real time; according to changes in user behavior and system status, the strategy of the proximal strategy optimization method is dynamically adjusted to adapt to the user's hot water demand in a timely manner.

[0069] (3) This method can be put into use directly without the need for debugging during equipment installation. It has good scalability and adaptability and is highly interpretable. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 This is a framework diagram of the learning user portrait and deep reinforcement learning method of the method of the present invention.

[0071] Figure 2This is a flow chart of the learning user portrait and deep reinforcement learning method of the method of the present invention. DETAILED DESCRIPTION

[0072] The present invention proposes a hot water system control method based on user profile learning and deep reinforcement learning, which is described in detail with reference to the accompanying drawings as follows:

[0073] Figure 1 This is a framework diagram for learning user profiles and deep reinforcement learning methods in the present invention. The water heater consists of an agent and a hot water heating device. The agent contains an initial database, a neural network discriminator for hot water profiles, and a proximal policy optimization method.

[0074] Figure 2 This is a flow chart of the user profile learning and deep reinforcement learning method of the present invention. First, the initial database is prepared. Then, the multi-source heterogeneous data of the new user is preliminarily collected. Secondly, the processed multi-source heterogeneous data is input into the hot water profile neural network discriminator. According to the output of the hot water profile neural network discriminator, it is judged whether the new user's hot water usage behavior belongs to the hot water profile in the initial database. If it belongs to the hot water profile in the initial database, then e k =i, where i indicates that the new user's hot water usage behavior belongs to the i-th hot water usage profile; if it does not belong to the hot water profile of the initial database, then e k =k+1, indicating that the new user's hot water usage behavior belongs to an unknown profile. Then, real-time multi-source heterogeneous data of the new user is collected. Finally, e k The real-time multi-source heterogeneous data are input into the proximal policy optimization method, the proximal policy optimization method is continuously trained, and the updated proximal policy optimization method is used to output the control action of the hot water heating device.

[0075] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A hot water system control method based on user profile learning and deep reinforcement learning, characterized in that: The system combines multi-source heterogeneous data feature fusion extraction, user profiling, unsupervised learning, and deep reinforcement learning to control hot water heating devices. This system achieves personalized hot water control, optimizes energy efficiency, and improves user comfort. The steps in the use process are as follows: Step (1): Collect N D The historical multi-source heterogeneous data of users; the historical multi-source heterogeneous data of the i-th user includes: hourly weather data in the past 30 days in, The weather data for the first hour of the first day; The weather data for the second hour of the first day; The weather data is for the 24th hour of the first day; The weather data for the first hour of the second day; The weather data for the second hour of the second day; The weather data is for the 24th hour of the second day; The weather data is for the first hour of the 30th day; The weather data is for the second hour of the 30th day; Weather data for the 24th hour of the 30th day; temperature data for each hour of the past 30 days in, The temperature data for the first hour of the first day; The temperature data for the second hour of the first day; The temperature data for the 24th hour of the first day; The temperature data for the first hour of the second day; The temperature data for the second hour of the second day; The temperature data for the 24th hour of the second day; The temperature data for the first hour of the 30th day; The temperature data is for the second hour on the 30th day; Temperature data for the 24th hour of the 30th day; water heater on / off status data for each hour in the past 30 days in, The water heater on / off status data for the first hour of the first day; The water heater on / off status data for the second hour of the first day; The water heater on / off status data for the 24th hour of the first day; The water heater on / off status data for the first hour of the second day; The water heater on / off status data for the second hour of the second day; The water heater on / off status data for the 24th hour of the second day; The water heater on / off status data for the first hour of the 30th day; The water heater on / off status data for the 2nd hour on the 30th day; The water heater on / off status data for the 24th hour of the 30th day; the hot water usage data for each hour in the past 30 days in, This is the hot water usage data for the first hour of the first day; This is the hot water usage data for the second hour of the first day; Hot water usage data for the 24th hour of the first day; This is the hot water usage data for the first hour of the second day; This is the hot water usage data for the second hour of the second day; Hot water usage data for the 24th hour of the second day; Hot water usage data for the first hour of the 30th day; This is the hot water usage data for the second hour on the 30th day; Hot water usage data for the 24th hour of the 30th day; date data for each hour of the past 30 days in, The date data is the first hour of the first day; The date data is the second hour of the first day; The date data is the 24th hour of the first day; The date data is the first hour of the second day; The date data is the second hour of the second day; The date data is the 24th hour of the second day; The date data is the first hour of the 30th day; The date data is the second hour of the 30th day; Date data for the 24th hour of the 30th day; day of the week data for each hour in the past 30 days in, The weekly data is the first hour of the first day; The weekly data is the second hour of the first day; The weekly data is the 24th hour of the first day; The weekly data is the first hour of the second day; The weekly data is the second hour of the second day; The weekly data is the 24th hour of the second day; The weekly data is the first hour of the 30th day; The weekly data is the second hour of the 30th day; Weekly data for the 24th hour of the 30th day; hourly data for each hour of the past 30 days in, i is the hourly data of the first hour of the first day; The hourly data is for the second hour of the first day; The hourly data is the 24th hour of the first day; The hourly data is for the first hour of the second day; The hourly data is for the second hour of the second day; The hourly data is the 24th hour of the second day; The hourly data is the first hour of the 30th day; The hourly data is the second hour of the 30th day; The hourly data is the 24th hour of the 30th day; Step (2): Splice each user’s historical multi-source heterogeneous information data to obtain a 3D tensor multi-source heterogeneous fusion sample T1 represents the 3D tensor multi-source heterogeneous fusion sample of the first user; T2 represents the 3D tensor multi-source heterogeneous fusion sample of the second user; Indicates the Nth D 3D tensor multi-source heterogeneous fusion samples of each user; Step (3): Extract features from the 3D tensor multi-source heterogeneous fusion sample of each user to obtain the multi-source heterogeneous feature flattened sample of each user; extract features from the 3D tensor multi-source heterogeneous matrix sample T of the i-th user. i Perform multi-channel convolution operation; convolution kernel tensor M C The shape is (30,3,27); the output of the multi-channel convolution operation is a multi-source heterogeneous feature sample in, are all elements in the multi-source heterogeneous feature sample obtained by the multi-channel convolution operation of the i-th user; Z1 represents the multi-source heterogeneous feature sample of the first user; Z2 represents the multi-source heterogeneous feature sample of the second user; Indicates the Nth D The multi-source heterogeneous feature samples of each user; the multi-channel convolution operation process is as follows: Z i =T i ⊙M C (1) Where, T i is the 3D tensor multi-source heterogeneous fusion sample of the i-th user; ⊙ is the multi-channel convolution operation; Z i is the multi-source heterogeneous feature sample of the i-th user; Then, the multi-source heterogeneous feature samples are flattened to change the arrangement of the elements to obtain the multi-source heterogeneous feature flattened sample of the i-th user The multi-source heterogeneous feature flattening sample is a 616×1 matrix; P i =Flatten(Z i ) (2) Where Flatten(·) is the flattening function; P i Flatten the samples for the multi-source heterogeneous features of the i-th user; Step (4): N D The multi-source heterogeneous feature expansion samples of each user are input into the K-Means method; the output of the K-Means method is the different types of hot water usage portraits: C1, C2, ...C k ; Among them, C1 is the first hot water usage profile; C2 is the second hot water usage profile; C k is the k-th hot water usage portrait; k is the number of types of hot water usage portraits output by the K-Means method; in addition to the k types of hot water usage portraits output by the K-Means method, the k+1-th hot water usage portrait is defined as an unknown portrait; The operation process of the K-Means clustering method is: (4.1) Randomly initialize k cluster labels, namely: C1, C2, ... C k ; (4.2) Calculate the Euclidean distance between each multi-source heterogeneous feature expansion sample and all cluster labels: d(P i ,C j )=sqrt((P i -C j ) 2 ) (3) Where, P i Expand the sample for the multi-source heterogeneous features of the i-th user; C j is the jth cluster label; d(P i ,C j ) is the Euclidean distance between the multi-source heterogeneous feature expansion sample of the i-th user and the j-th cluster label; sqrt(·) is the square root function; (4.3) Assign each multi-source heterogeneous feature expansion sample to the cluster label with the smallest Euclidean distance; (4.4) Update cluster labels; for each cluster label, calculate the mean of the multi-source heterogeneous feature expansion samples belonging to the cluster label as the new cluster label; Where C j represents the updated cluster label; N j Belongs to cluster label C j The number of samples of multi-source heterogeneous feature expansion; P m Indicates belonging to the cluster centroid C j The mth multi-source heterogeneous feature expansion sample; (4.5) Repeat steps (4.2)-(4.4) until the maximum number of iterations is met; (4.6) Output clustering results; K-Means method outputs k cluster labels; C1, C2, ... C k , each cluster label is a hot water usage portrait; C1 is the first hot water usage portrait; C2 is the second hot water usage portrait; C k Use portrait for the kth type of hot water; N D Flattened samples of multi-source heterogeneous features for users With k types of hot water usage profiles C1, C2, ... C k Form the initial database; Step (5): Put the water heater into use by the new user and collect the new user's real-time multi-source heterogeneous data; the new user's real-time multi-source heterogeneous data includes: hourly weather data Hourly temperature data Hourly water heater on / off status data Hot water usage per hour Hourly date data Hourly week data Hourly data for each hour Where d is the number of days the water heater is put into operation and t is the number of hours in the day; The water heater consists of an agent and a hot water heating device; the agent contains an initial database, a hot water image neural network discriminator, and a proximal strategy optimization method; Before the water heater is put into use by a new user, the proximal policy optimization method is trained using offline training so that the proximal policy optimization method can control the on / off state of the water heater. The input matrix of the proximal policy optimization method is Among them, e k ∈{1,2,...,k,k+1}, indicating which hot water usage profile the new user's hot water usage behavior belongs to; e k =1 indicates that the new user's hot water usage behavior belongs to the first hot water usage profile; e k =2 means that the new user's hot water usage behavior belongs to the second hot water usage profile; k =k indicates that the new user's hot water usage behavior belongs to the kth hot water usage profile; e k =k+1 indicates that the new user's hot water usage behavior belongs to an unknown profile; the output of the proximal policy optimization method is the on / off status of the water heater; Step (6): Splice the real-time multi-source heterogeneous data of the new user to obtain the new user's 2D tensor multi-source heterogeneous fusion sample Input R into the hot water image neural network discriminator; the output of the hot water image neural network discriminator is the probability value of R belonging to the 1st to kth category hot water usage image, which are The input layer of the hot water image neural network discriminator is a convolutional layer; the hidden layer of the hot water image neural network discriminator consists of a width of W hw , the depth is W hd The activation function of the fully connected layer in its hidden layer is the tanh() function; the output layer of the hot water image neural network discriminator has a width of k and a depth of W. od The activation function of the fully connected layer in its output layer is the softmax() function; the momentum stochastic gradient descent optimizer is used to train the hot water image neural network discriminator; Step (7): Determine whether the new user's 2D tensor multi-source heterogeneous fusion sample R belongs to the hot water usage profile in the initial database; when When both are less than 0.9, the hot water portrait neural network discriminator refuses to recognize; then the new user's hot water use behavior does not belong to the hot water use in the initial database, but belongs to the k+1th type of hot water use portrait, that is, an unknown portrait; at this time, e k =k+1; when When it is greater than 0.9, it is determined that the new user's hot water usage behavior belongs to the i-th hot water usage profile in the initial database; at this time, e k =i; Step (8): Online training of the proximal strategy optimization method is performed, and the control action of the hot water heating device is outputted at the same time. The online training process of the proximal strategy optimization method is as follows: (8.1) Create a new user's hot water usage profile k and real-time multi-source heterogeneous data input into the proximal strategy optimization method; (8.2) Generate N according to the current action function π(A|S;θ) and the evaluation function V(S;φ) P A sequence of trajectories; where θ is the network structure parameter of the action function, that is, the network weight and bias of the action function; φ is the network structure parameter of the evaluation function, that is, the network weight and bias of the evaluation function; A is the action output by the proximal policy optimization method, and there are two types of actions: one is to turn on the water heater switch, and the other is to turn off the water heater switch; S is the state input to the proximal policy optimization method, which includes the hot water usage profile of the new user and real-time multi-source heterogeneous data; In addition, R PPO is the reward obtained when the proximal policy optimization method takes action on the heating device; If there is a demand for hot water, reward R PPO for: R PPO =-a reward ×P hp -b reward ×max(40-T tank ,0)-c reward ×max(H time -24,0) (5) If there is no hot water demand, reward R PPO for: R PPO =-a reward ×P hp -c reward ×max(H time -24,0) (6) Where a reward is the energy term coefficient in the reward; b reward is the comfort factor in the reward; c reward P is the health item coefficient in the reward; hp is the energy consumed by the heating device; max(·,·) is the maximum value function; T tank is the temperature of the hot water tank; H time The duration of time when the temperature of the hot water tank last reached above 60°C; Each trajectory sequence is: Where: t s Represents a certain moment, that is, the starting time of the current trajectory sequence; t s The state of the moment; t s +1 moment status; t s +N P -1 moment status; t s +N P The state of the moment; is An action in a state; is An action in a state; From the state Transition to rewards; From the state Transition to rewards; When in state When , the probability of each action is calculated using π(A|S;θ), and the action is randomly selected according to the probability distribution At the beginning of training, t s =1, for the subsequent N P Trajectory sequence, t s ←t s +N P ; (8.3) For t = t s +1,t s +2,...,t s +N P For these N trajectory sequences, calculate the discounted return G for each trajectory sequence t With advantage function D t ; Discounted return G t for: Where γ is the discount coefficient; R PPO,k From state S k Transition to S k+1 The reward value of b G To calculate G t The coefficient when ; if is the final state, then b G is 0, otherwise, b G is 1; Advantage function D t for: Where λ is the smoothing factor; b D To calculate D t The coefficient of , if is the final state, then b D is 0, otherwise b D is 1; (8.4) From this N P Learning in a sequence of trajectories: (8.4.1) Randomly extract a trajectory of size M from the current trajectory sequence B A dataset where each element contains the corresponding discounted return and advantage function value; (8.4.2) Minimize the evaluation loss function L by gradient descent critic (φ) is used to update the parameter φ of the evaluation function and evaluate the loss function L critic (φ) is: Where: G i represents the corresponding discounted return in the i-th element in the dataset; (8.4.3) Normalize the advantage function value. The corresponding normalized discounted advantage function value of the i-th element in the data set is: Where i is the subscript number of each element in the data set, represents the normalized discounted advantage function value corresponding to the i-th element in the dataset; D i The advantage function value corresponding to the i-th element in the data set; the advantage function value corresponding to the first element in the D1 data set; the advantage function value corresponding to the second element in the D2 data set; M The Mth B The advantage function value corresponding to each element; mean(·,·,...,·) is the function for finding the average value; std(·,·,...,·) is the function for calculating the standard deviation; (8.4.4) Minimize the action loss function L by gradient descent actor (θ) is used to update the parameters θ of the action function, the action loss function L actor (θ) is: where r i (θ) coefficient factor and entropy loss function They are: Where min(·,·) is the minimum function; π(A i |S i ;θ) is in state S i Under given parameter θ, take action A i The probability of π(A i |S i θ old ) is in state S i Under the given parameters θ before the current learning period old , take action A i The probability of r i (θ) coefficient factor; ε is the shear factor; is the entropy loss function; P z is the number of action types in the proximal strategy optimization method; π(A k |S i ;θ) is in state S i Under given parameter θ, take action A k The probability of; w is the entropy loss coefficient; (8.5) The proximal strategy optimization method uses the updated action function to output the action for controlling the hot water heating device based on the input data in (8.1); (8.6) Repeat (8.1)-(8.5), continuously update the proximal strategy optimization method, and output the action of controlling the hot water heating device.

Citation Information

Patent Citations

  • Hot water system control method based on model prediction and deep reinforcement learning

    CN115183474A

  • Control method for comfortable bathing and water heater system

    CN111365849A

  • Control method and device for water heater, processor and water heater

    CN112556193A