Reinforcement learning device, reinforcement learning method, and reinforcement learning program

The reinforcement learning device and method address the challenge of creating significant gradients in target space states by generating task arrays and performing reinforcement learning, resulting in effective air conditioning control that enhances user comfort.

WO2025109662A1PCT designated stage expired Publication Date: 2025-05-30NT T INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/041690
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Conventional reinforcement learning methods struggle to create significant gradients in states like temperature, humidity, and illuminance within a target space, due to the vast search space for optimal actions and the difficulty in reproducing physically impossible temperature gradients.

Method used

A reinforcement learning device and method that generate a task array of target states at multiple points in the target space, using an environment simulator to reproduce state distributions, and perform reinforcement learning to determine actions that create a significant gradient in the target space.

Benefits of technology

The approach enables the determination of actions that effectively generate a significant gradient in the target space, improving comfort for multiple users with different thermal preferences by systematically controlling air conditioning equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023041690_30052025_PF_FP_ABST
    Figure JP2023041690_30052025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a reinforcement learning device for performing reinforcement learning of an action determination model for determining actions of a plurality of control targets that each affect the state of a target space. This reinforcement learning device includes: a task generation unit that generates a task array, which is composed of target states at a plurality of respective target points determined in advance in the target space, on the basis of an array composed of target space states due to predetermined actions of the plurality of control targets at the plurality of target points; and a model learning unit that performs reinforcement learning of the action determination model by using the task array.
Need to check novelty before this filing date? Find Prior Art

Description

Reinforcement learning device, reinforcement learning method, and reinforcement learning program

[0001] The disclosed technology relates to a reinforcement learning device, a reinforcement learning method, and a reinforcement learning program.

[0002] A technology has been known in the past that uses reinforcement learning using an environmental simulator of a target space to acquire a model that selects optimal air conditioning control actions to bring the target space into a target state (see, for example, Patent Document 1). In this case, the rewards used in the reinforcement learning include the predicted mean vote (PMV), which is an index of human thermal comfort, the amount of energy saved, and the temperature difference between the outside air and the room temperature.

[0003] Patent No. 7014299

[0004] In recent years, there has been a demand for control technologies that can intentionally create gradients in conditions such as temperature, humidity, and illuminance within the same space. For example, while typical office air conditioning is designed to maintain a uniform temperature within a room, if significant temperature differences could be created within a room, it is expected that the comfort of all users will be improved when multiple users with different thermal sensitivities are present together.

[0005] When creating a gradient in conditions within the same space, one possible method is to systematically control multiple control targets (e.g., air conditioners) according to the gradient of the target condition, but this control is not easy. For example, even if the set temperature of each air conditioner is simply changed using the optimal temperature for each user, the intended temperature gradient may not be achieved due to factors such as mutual interference with other air conditioners, the temperature difference with the outside air, the specifications of the air conditioner, and changes over time.

[0006] In conventional methods, a uniform target is assigned to the target space, making it impossible to achieve control that creates a gradient in the state within the target space. Furthermore, even if an attempt is made to build a model that determines control actions to create a gradient in the target state using reinforcement learning similar to conventional methods, the search space for optimal actions becomes vast due to the enormous number of patterns for states (e.g., temperature distribution) and actions (e.g., air conditioning control), making it difficult to progress in learning.

[0007] The disclosed technology has been made in consideration of the above points, and aims to provide a reinforcement learning device, a reinforcement learning method, and a reinforcement learning program that can determine actions that will produce a significant gradient in the state of a target space.

[0008] A first aspect of the present disclosure is a reinforcement learning device that performs reinforcement learning of a behavioral decision model to determine the actions of multiple control objects that each affect the state of a target space, and includes a task generation unit that generates a task array consisting of target states at each of multiple target points predetermined in the target space based on an array consisting of states of the target space obtained by predetermined actions of multiple control objects at each of multiple target points, and a model learning unit that performs reinforcement learning of the behavioral decision model using the task array.

[0009] A second aspect of the present disclosure is a reinforcement learning method for performing reinforcement learning of a behavioral decision model to determine the actions of multiple control objects that each affect the state of a target space, in which a computer generates a task array consisting of target states at each of multiple target points predetermined in the target space based on an array consisting of states of the target space obtained by predetermined actions of multiple control objects at each of multiple target points, and performs a process of reinforcement learning of the behavioral decision model using the task array.

[0010] A third aspect of the present disclosure is a reinforcement learning program that performs reinforcement learning of a behavioral decision model to determine the actions of multiple control objects that each affect the state of a target space, and is a program that causes a computer to execute a process of performing reinforcement learning of the behavioral decision model using the task array, based on an array of states of the target space obtained by predetermined actions of multiple control objects at each of multiple target points predetermined in the target space.

[0011] The disclosed techniques allow for determining actions that produce significant gradients in the state of the object space.

[0012] FIG. 1 is a block diagram showing an example of a schematic configuration of an air conditioning control system. FIG. 2 is a diagram for explaining the environment of a target space. FIG. 3 is a diagram for explaining the behavior of a controlled object. FIG. 4 is a diagram for explaining an environmental simulator. FIG. 5 is a diagram for explaining temperature distribution. FIG. 6 is a block diagram showing an example of the hardware configuration of a reinforcement learning device. FIG. 7 is a block diagram showing an example of the functional configuration of a reinforcement learning device. FIG. 8 is a diagram for explaining a difference array. FIG. 9 is a diagram for explaining reinforcement learning based on time-series data. FIG. 10 is a flowchart showing the flow of a task generation process. FIG. 11 is a flowchart showing the flow of a reinforcement learning process. FIG. 12 is a diagram showing an environment of an example. FIG. 13 is a diagram showing the results of an example. FIG. 14 is a diagram showing the results of a comparative example.

[0013] An example of an embodiment of the disclosed technology will be described below with reference to the drawings. Note that the same or equivalent components and parts in each drawing are given the same reference numerals. Also, the dimensional proportions in the drawings are exaggerated for the convenience of explanation and may differ from the actual proportions.

[0014] 1 is a schematic diagram showing an example of the configuration of an air-conditioning control system 100 to which a reinforcement learning device 10A according to this embodiment is applied. The air-conditioning control system 100 includes an air-conditioning control scenario calculation device 10, an environmental simulator 50, and multiple control targets 60A to 60C.

[0015] When the target temperature distribution in a target space is input, the air-conditioning control scenario calculation device 10 outputs optimal air-conditioning control actions for each of the multiple control targets 60A-60C. In this embodiment, this target temperature distribution may have a significant temperature gradient. For example, a target temperature distribution that maximizes overall comfort can be set, estimated based on the cold / hot sensation preferences of each user in a target space, the time of stay (e.g., arrival time at work), and the location of stay (e.g., seat location). Temperature is an example of a state in the present disclosure. Air-conditioning control actions are an example of an action in the present disclosure.

[0016] The multiple control targets 60A to 60C are air conditioning equipment that each affect the temperature of a target space, such as a variable air volume (VAV) and an air handling unit (AHU). The multiple control targets 60A to 60C can be controlled independently, and there are no particular restrictions on their type, number, or breakdown. Hereinafter, when the multiple control targets 60A to 60C are not to be distinguished from one another, they will be simply referred to as control targets 60.

[0017] An example of a target space S0 and a control target 60 will be described with reference to FIGS. 2 and 3. FIG. 2 is a plan view of an office room as an example of the target space S0, with desks, chairs, lockers, and doors arranged in the office indicated by dotted lines. Five control targets 60A-60E are installed in the target space S0. FIG. 3 shows the types of control targets 60A-60E shown in FIG. 2 and the set temperatures as an example of air conditioning control behavior. The control targets 60A-60D are VAVs with independently adjustable set temperatures, and are located at positions indicated by thick solid lines in FIG. 2. FIG. 2 also shows the positions of the air outlets BO of each VAV. The control target 60E is an AHU with an adjustable set temperature, and is located at the position indicated by thick dashed lines in FIG. 4 (the entire target space S0).

[0018] The environmental simulator 50 is constructed in advance so that it can accurately reproduce future changes in temperature distribution in a target space using current and past room temperature data, air conditioning control data, weather data, etc. Such a simulator can be realized, for example, by using a neural network and regression model that inputs various data and outputs temperature distribution.

[0019] For example, as shown in FIG. 4, the environment simulator 50 reproduces the temperature distribution shown in FIG. 5 by predicting the temperature for each region when the target space S0 is divided into 25 5x5 regions. Hereinafter, when distinguishing between the regions of the target space S0, column numbers A to E and row numbers 1 to 5 are combined and referred to as, for example, "region A1." Note that the environment simulator 50 does not need to predict the temperature for regions where no people are present because they are mostly occupied by objects, as shown in regions A2 and A3, for example. Furthermore, the granularity of the temperature distribution output by the environment simulator 50 can be determined arbitrarily.

[0020] Specifically, the air conditioning control scenario calculation device 10 includes a reinforcement learning device 10A, a behavior decision-making device 10B, a behavior decision-making model 40, and a task DB (database) 42. The behavior decision-making model 40 is a machine learning model for determining the behavior of multiple control targets 60. The reinforcement learning device 10A performs reinforcement learning using the behavior decision-making model 40 as an agent and an environmental simulator 50 as an environment. The behavior decision-making device 10B uses the learned behavior decision-making model 40 and the environmental simulator 50 to determine the optimal air conditioning control behavior for achieving a target temperature distribution.

[0021] Here, the following two examples of challenges can be raised when constructing a behavioral decision-making model 40 using reinforcement learning. The first challenge is the difficulty of preparing a temperature distribution with a temperature gradient suitable for use as a reinforcement learning task. For example, due to the specifications of multiple control objects 60 and their arrangement within the target space, each environment may have a temperature gradient that is physically impossible or difficult to reproduce. When simply setting the target space to a uniform temperature, measurements and reward calculations are performed at approximately one target point, and the temperature range that can be reproduced at the target point can be easily estimated from the specifications of the control object 60 and the steady-state room temperature, etc. Therefore, a wide range of reinforcement learning tasks can be prepared within the estimated reproducible temperature range.

[0022] On the other hand, if an agent is given a task that involves a temperature distribution with a temperature gradient that cannot be physically reproduced when attempting to create a significant temperature gradient within a target space, the agent will not receive a reward no matter what action it takes, and learning will not progress.For proper learning, it is desirable to limit the temperature gradient to those that can be reproduced and prepare a wide variety of temperature distributions, but it is difficult to manually determine whether a temperature gradient can be reproduced.

[0023] The second challenge is that the sheer number of patterns for states (temperature distribution) and actions (air conditioning control) makes the search space for optimal actions vast, making it difficult to progress in learning. For example, when training an agent model to determine optimal air conditioning control actions, one possible method is to provide the agent with the current state and the target state. If the target space is simply set to a uniform temperature, the number of dimensions of the state observed by the agent will be 1 x 2, consisting of one measured temperature and one target temperature.

[0024] On the other hand, if a two-dimensional target space is divided into 5x5 regions and the temperature distribution of the entire target space is obtained based on the state of each region, the number of dimensions of the state observed by the agent will be 5x5x2, consisting of the 5x5 measured temperature and the 5x5 target temperature. Furthermore, the number of dimensions for behavior also increases with the number of control objects 60 and their setting items. With such a large number of dimensions, similar states are unlikely to appear, making it impossible to perform the trial and error that is important in reinforcement learning, and learning does not progress. In other words, this issue arises from the large number of dimensions and the agent's inability to understand the relationship between the measured temperature and the target temperature.

[0025] Therefore, the technology of the present disclosure promotes learning of the behavior decision-making model 40 for determining, for each control object 60, an action that generates a significant gradient in the state of the target space. Hereinafter, a reinforcement learning device 10A according to an embodiment of the present disclosure will be described.

[0026] [Configuration of Reinforcement Learning Apparatus] Fig. 6 is a block diagram showing the hardware configuration of the reinforcement learning device 10A. As shown in Fig. 6, the reinforcement learning device 10A has a CPU (Central Processing Unit) 91, a ROM (Read Only Memory) 92, a RAM (Random Access Memory) 93, a storage 94, an input unit 95, a display unit 96, and a communication I / F (Interface) 97. Each component is connected to each other via a bus 99 so as to be able to communicate with each other.

[0027] The CPU 91 is a central processing unit that executes various programs and controls each part. That is, the CPU 91 reads programs from the ROM 92 or the storage 94 and executes the programs using the RAM 93 as a work area. The CPU 91 controls each of the above components and performs various arithmetic processing in accordance with the programs stored in the ROM 92 or the storage 94.

[0028] The ROM 92 stores various data and programs. The RAM 93 temporarily stores programs or data as a working area. The storage 94 is configured with a storage device such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive), and stores various programs including an operating system and various data. In this embodiment, the ROM 12 or the storage 14 stores a reinforcement learning program and a task generation program.

[0029] The input unit 95 includes, for example, a pointing device such as a mouse, a keyboard, etc., and is used to input various types of information. The display unit 96 is, for example, a liquid crystal display, and displays various types of information. The display unit 96 may also function as the input unit 95 by employing a touch panel system.

[0030] The communication I / F 97 is an interface for communicating with other devices, for example, using a wired communication standard such as Ethernet (registered trademark) or FDDI (Fiber Distributed Data Interface), or a wireless communication standard such as 4G, 5G, or Wi-Fi (registered trademark).

[0031] Next, each functional configuration of the reinforcement learning device 10A will be described. FIG. 7 is a block diagram showing the functional configuration of the reinforcement learning device 10A of this embodiment. As shown in FIG. 7, the reinforcement learning device 10A includes a task generation unit 20 and a model learning unit 30. The functional configuration of the task generation unit 20 is realized when the CPU 11 reads out a task generation program stored in the ROM 12 or the storage 14, expands it in the RAM 13, and executes it. The functional configuration of the model learning unit 30 is realized when the CPU 11 reads out a reinforcement learning program stored in the ROM 12 or the storage 14, expands it in the RAM 13, and executes it.

[0032] (Task Generation Process) First, the task generation process executed by the task generation unit 20 will be described. The task generation unit 20 generates a task array to be used for reinforcement learning based on the distribution of states in an object space obtained by predetermined actions of multiple control objects 60. In other words, the task generation unit 20 attempts to comprehensively generate task arrays corresponding to a wide variety of state distributions (temperature distributions) that can be physically reproduced by the actions of the control objects 60 in a certain object space. The task generation process is executed prior to the reinforcement learning process of the behavioral decision-making model 40 by the model learning unit 30.

[0033] 7, the task generation unit 20 includes a first control unit 21, a first initialization unit 22, a first action determination unit 23, a first sampling unit 24, and a memory determination unit 28. The task generation unit 20 stores task sequences generated by these functional components in a task DB 42. The task generation unit 20 repeats the task generation process until a predetermined number of task sequences are stored in the task DB 42 (i.e., until a wide variety of task sequences have been accumulated). The first control unit 21 performs overall control of the task generation process.

[0034] The first initialization unit 22 sets up the environment simulator 50 by providing initial conditions to the environment simulator 50. The initial conditions include, for example, the room temperature of the target space at the start, air conditioning control actions, and weather information. Note that the initial conditions are preferably such that the control target is stopped (i.e., no air conditioning control actions are being taken) and the state of the target space (e.g., room temperature) is in a steady state.

[0035] The first behavior decision unit 23 decides a predetermined behavior for each of the plurality of control objects 60 and inputs the behavior to the environment simulator 50. For example, the first behavior decision unit 23 randomly decides a set temperature for each of the plurality of control objects 60 within a range that can be set according to specifications. The environment simulator 50 reproduces the distribution of the state (e.g., temperature distribution) of the target space (environment) one step after the plurality of control objects 60 start to be controlled in accordance with the behavior input by the first behavior decision unit 23. One step is, for example, 10 minutes.

[0036] The first sampling unit 24 acquires a distribution of states of the target space at a predetermined time after the control of the plurality of control objects 60 begins in accordance with the determined action. Furthermore, the first sampling unit 24 extracts states at each of a plurality of predetermined target points in the target space from the acquired distribution of states of the target space and converts the extracted states into a 1×n-dimensional array (n is the number of target points) (see FIG. 8 ). For example, the first sampling unit 24 may acquire a temperature distribution of the target space from the environmental simulator 50 at each step and sample the temperature distribution at the eighth step.

[0037] The multiple target points used for sampling are predetermined locations in the target space by arbitrary coordinate positions and numbers, such as important locations for air conditioning control. For example, locations where users are expected to stay for long periods of time, such as entrances and exits and seats, the center of gravity of each area when the target space is divided into units larger than the granularity of the state distribution, and locations set at regular intervals within the target space can be appropriately applied. As an example, in the temperature distribution in Figure 5, the background color of the areas corresponding to the target points is changed.

[0038] The memory determination unit 28 stores an array consisting of the states at each of the extracted target points as a task array in the task DB 42 if no other arrays similar to the array are stored in the task DB 42. For example, the memory determination unit 28 compares the similarity between the generated array and all task arrays already stored in the task DB 42. Note that a known index such as inter-vector distance can be used as a method for determining the similarity.

[0039] In other words, if a task sequence similar to the generated sequence is already stored in the task DB 42, the memory determination unit 28 discards the generated sequence. In this way, by excluding similar sequences and storing novel sequences in the task DB 42, a wide variety of task sequences are accumulated in the task DB 42. The task DB 42 is an example of a memory unit in the present disclosure.

[0040] The first control unit 21 determines whether the number of task arrays stored in the task DB 42 has reached a predetermined upper limit, and if so, ends the task generation process. On the other hand, if the number of task arrays stored in the task DB 42 has not reached the upper limit, or if an array has been discarded by the memory determination unit 28, the first control unit 21 attempts to generate another task array. In other words, the task generation unit 20 repeatedly performs the task generation process while changing the behavior of multiple control targets 60.

[0041] In the second and subsequent task generation processes, the initial conditions used to set up the environment simulator 50 may be the same as or different from those used in the previous task generation process. Using different initial conditions increases the variation in task sequences, which contributes to improving the versatility of the behavioral determination model 40. However, if the initial conditions are changed too drastically (e.g., changing the weather conditions from summer to winter), it becomes impossible to ensure the reproducibility of the task sequence when different initial conditions are set during learning. Therefore, it is preferable to avoid making drastic changes to the initial conditions. If the initial conditions are changed drastically, it is preferable to separate the task DB 42 and the behavioral determination model 40 by initial condition (e.g., by season).

[0042] As described above, through the task generation process, task sequences corresponding to states (for example, temperature distributions with temperature gradients) that can be reproduced by the actions of multiple control objects 60 in the target space are comprehensively stored in the task DB 42. When performing reinforcement learning for the behavioral determination model 40, the wide variety of task sequences stored in the task DB 42 are used. In other words, while being limited to the distribution of states with reproducible gradients, learning can be performed using a wide variety of state distributions without being biased toward the distribution of states that are likely to occur, which can contribute to improving the speed of reinforcement learning and the accuracy of the agent model.

[0043] (Reinforcement Learning Process) Next, a description will be given of the reinforcement learning process executed by the model learning unit 30. The model learning unit 30 performs reinforcement learning of the behavioral determination model 40 using the task sequence stored in the task DB .

[0044] 7, the model learning unit 30 includes a second control unit 31, a second initialization unit 32, a second action determination unit 33, a second sampling unit 34, a dimension reduction unit 35, a reward calculation unit 36, and a model update unit 37. The second control unit 31 performs overall control of the reinforcement learning process in accordance with conditions based on the current number of steps.

[0045] The second initialization unit 32 acquires a task array consisting of target states at each of a plurality of target points. For example, the second initialization unit 32 randomly acquires one task array from the task DB 42 and sets the task array as the target.

[0046] The second initialization unit 32 also sets up the environment simulator 50 by providing initial conditions to the environment simulator 50. The initial conditions include, for example, the room temperature of the target space at the start, air conditioning control actions, and weather information. Note that the initial conditions are preferably such that the control target is stopped (i.e., no air conditioning control actions are being taken) and the state of the target space (e.g., room temperature) is in a steady state.

[0047] The second sampling unit 34 acquires the distribution of the state of the target space in the current step. That is, the second sampling unit 34 acquires the distribution of the state of the target space obtained by the predetermined actions of the multiple control objects 60 determined in the immediately preceding step. In the case of the initial step, the second sampling unit 34 acquires the distribution of the initial state from the environment simulator 50.

[0048] The second sampling unit 34 extracts states at each of a plurality of predetermined target points in the target space from the acquired distribution of states in the target space and converts the extracted states into a 1×n-dimensional (n is the number of target points) state array. That is, the state array is an array consisting of states at each of the plurality of target points. Here, the target points used for extracting the states are the same as the target points used for extracting the task array.

[0049] The dimension reduction unit 35 derives a difference array indicating the difference between the state array and the task array. That is, the dimension reduction unit 35 derives a difference state between the target state and the current state at each of a plurality of target points based on the state array and the task array, and converts it into a 1×n-dimensional (n is the number of target points) difference array.

[0050] FIG. 8 shows an example of a temperature array consisting of current temperatures, a task array consisting of target temperatures, and a difference array at nine predetermined target points (see FIG. 5) in the target space S0. The temperature array is an example of a state array. In this way, the number of dimensions can be reduced by sampling the 5×5-dimensional temperature distribution in the current step into a 1×9-dimensional temperature array and sampling the target 5×5-dimensional temperature distribution into a 1×9-dimensional task array. Furthermore, the number of dimensions can be halved by deriving a difference array based on the temperature array and the task array. In other words, the dimension reduction unit 35 reduces the variation of states handled in reinforcement learning by reducing the states represented in 5×5×2 dimensions to 1×9×1 dimensions.

[0051] The reward calculation unit 36 ​​calculates a reward based on the smallness of the difference between the task array and the state array (temperature array) indicated by the difference array. The reward is an index that serves as a substitute for teacher data in reinforcement learning. In this embodiment, since the relationship between the target state and the current state is made clear by the difference array, the reward calculation unit 36 ​​may increase the reward as the difference approaches "0."

[0052] The model update unit 37 updates the behavior determination model 40 by reinforcement learning based on a data set including the difference array, the action, and the reward. Specifically, the model update unit 37 updates the policy of the behavior determination model 40 so that an action that maximizes the reward (i.e., an action that makes all the differences between the target points "0") can be selected. Note that any known technology can be applied as the reinforcement learning algorithm, as appropriate.

[0053] The second behavior decision unit 33 determines the behavior of each of the multiple control objects 60 using the behavior decision model 40, which is trained to take the difference array as input and output the behavior of each of the multiple control objects 60. In other words, the second behavior decision unit 33 determines the behavior of each of the multiple control objects 60 by inputting the difference array derived in the current step into the behavior decision model 40.

[0054] Furthermore, the second action decision unit 33 inputs the decided action to the environment simulator 50. The environment simulator 50 outputs the distribution of states of the target space one step after the multiple control objects 60 start to be controlled in accordance with the action input by the second action decision unit 33. The second sampling unit 34 acquires from the environment simulator 50 the distribution of states of the target space one step after the multiple control objects 60 start to be controlled in accordance with the action decided by the second action decision unit 33, and converts it into a state array.

[0055] The second control unit 31 performs overall control so that the acquisition of state distribution, conversion to a state array, derivation of a difference array, calculation of reward, update of the behavioral determination model 40, and behavioral determination are repeated a predetermined number of M steps using the above functional configurations. M is an arbitrary integer, for example, from 18 steps (equivalent to 3 hours) to 72 steps (equivalent to 12 hours). Note that the behavioral determination model 40 may be updated every N steps, rather than every step. N is an arbitrary integer, for example, from 8 steps (equivalent to 80 minutes) to 16 steps (equivalent to 160 minutes).

[0056] Furthermore, the second control unit 31 may repeatedly perform each process with the task sequence fixed for M steps, but may perform reinforcement learning processing using a different task sequence after M steps have elapsed. Furthermore, the second control unit 31 may complete reinforcement learning of the behavioral determination model 40 when the number of steps reaches a predetermined upper limit (for example, 20,000 times).

[0057] In this embodiment, the state observed by the agent is relative information, namely, the difference array between the state array and the task array. Therefore, even if the values ​​of the state array and the task array are different, there may be multiple patterns in which the state (difference array) observed by the agent is the same. As an example, FIG. 9 shows two different examples of patterns α and β, in which the difference arrays at time t are the same. Since patterns α and β have different target task arrays and different temperature arrays at time t, the actions to be taken are essentially different. However, this difference cannot be determined from the difference array at time t alone, which can cause learning delays.

[0058] On the other hand, the difference array and the action at the past time point t-1 are different for pattern α and pattern β. Therefore, it is preferable that the model learning unit 30 also considers past information when determining the action and updating the action determination model 40. Specifically, by adding an LSTM (Long Short Term Memory) layer to the action determination model 40, the action for the current step may be determined taking into consideration the difference array and the action determined in the past step.

[0059] Furthermore, it is preferable that the model update unit 37 updates the behavioral decision-making model 40 by reinforcement learning based on a predetermined number of time-series data sets. For example, by adding an LSTM layer to the reinforcement learning algorithm, the policy of the behavioral decision-making model 40 may be updated taking into account the difference sequence, actions, and rewards in past steps.

[0060] [Operation of Reinforcement Learning Device] Next, the operation of the reinforcement learning device 10A according to this embodiment will be described.

[0061] 10 is a flowchart showing the flow of task generation processing by the reinforcement learning device 10 A. The task generation processing is performed by the CPU 91 reading a task generation program from the ROM 92 or the storage 94, expanding it into the RAM 93, and executing it.

[0062] In step S10, the CPU 91, functioning as the first initialization unit 22, sets the environment simulator 50 to an initial state by providing initial conditions to the environment simulator 50, and sets the current step number to 0. In step S12, the CPU 91, functioning as the first behavior decision unit 23, decides the behavior of each of the multiple control targets 60. For example, the CPU 91 randomly decides the set temperature within the range that can be set based on the specifications of each control target 60.

[0063] In step S14, the CPU 91, functioning as the first behavior decision unit 23, inputs the behavior of each control object 60 determined in step S12 to the environment simulator 50. In step S16, the CPU 91, functioning as the first sampling unit 24, acquires the distribution of the state of the target space one step after (e.g., 10 minutes after) the plurality of control objects 60 are controlled in accordance with the behavior input in step S14.

[0064] In step S18, the CPU 91, functioning as the first control unit 21, adds 1 to the current step number. In step S20, the CPU 91, functioning as the first control unit 21, determines whether the current step number has reached a predetermined number (e.g., 8 steps). If the current step number has not reached the predetermined number (if step S20 is N), steps S12 to S20 are repeated until the current step number reaches the predetermined number.

[0065] On the other hand, if the current number of steps has reached a predetermined number (if step S20 is Y), the process proceeds to step S22. In step S22, the CPU 91, as the first sampling unit 24, extracts (samples) the state (e.g., temperature) at each of a plurality of predetermined target points in the target space from the distribution of the state (e.g., temperature distribution) of the target space acquired in the immediately preceding step S16, and converts it into a 1×n-dimensional (n is the number of target points) array.

[0066] In step S24, the CPU 91, functioning as the memory determination unit 28, determines whether another array similar to the array converted in step S22 is stored in the task DB 42. If no other similar array is stored in the task DB 42 (if step S20 is Y), the process proceeds to step S26. In step S26, the CPU 91, functioning as the memory determination unit 28, stores the array converted in step S22 as a task array in the task DB 42. In step S28, the CPU 91, functioning as the first control unit 21, determines whether the number of task arrays stored in the task DB 42 has reached a predetermined upper limit, and if so, ends the task generation process.

[0067] On the other hand, if another sequence similar to the sequence converted in step S22 has already been stored in the task DB 42 (if step S24 is N), or if the number of task sequences stored in the task DB 42 does not reach the upper limit (if step S28 is N), the process returns to step S10. That is, another task sequence is attempted to be generated until the upper limit number of task sequences has been stored in the task DB 42.

[0068] 11 is a flowchart showing the flow of reinforcement learning processing by the reinforcement learning device 10A. The reinforcement learning processing is performed by the CPU 91 reading out a reinforcement learning program from the ROM 92 or the storage 94, expanding it into the RAM 93, and executing it. Note that at the start of the reinforcement learning processing, the current step number is 0.

[0069] In step S60, the CPU 91, functioning as the second initialization unit 32, acquires a task array from the task DB 42. In step S62, the CPU 91, functioning as the second initialization unit 32, sets the environment simulator 50 to an initial state by providing initial conditions to the environment simulator 50. In step S64, the CPU 91, functioning as the second sampling unit 34, acquires the distribution of the initial state from the environment simulator 50.

[0070] In step S66, the CPU 91, as the second sampling unit 34, extracts (samples) states (e.g., temperatures) at each of a plurality of predetermined target points in the target space from the state distribution (e.g., temperature distribution) in the target space acquired immediately before, and converts the state array (e.g., temperature array) into a 1×n-dimensional (n is the number of target points) state array. In step S68, the CPU 91, as the dimension reduction unit 35, reduces the number of dimensions of the states by deriving a difference array indicating the difference between the state array converted in step S66 and the task array acquired in step S60. In step S70, the CPU 91, as the reward calculation unit 36, calculates a reward based on the smallness of the difference between the task array and the state array indicated by the difference array derived in step S68.

[0071] In step S72, the CPU 91, functioning as the second control unit, determines whether the current number of steps is a multiple of a predetermined number, N. N is an arbitrary integer, for example, approximately 8 to 16. If the current number of steps is a multiple of N (if Y in step S72), the process proceeds to step S74. In step S74, the CPU 91, functioning as the model update unit 37, updates the behavioral determination model 40 by reinforcement learning based on a data set including time-series difference sequences, actions, and rewards for a predetermined number of past X times.

[0072] On the other hand, if the current number of steps is not a multiple of N (if step S72 is N), step S74 is skipped and the process proceeds to step S76. In step S76, the CPU 91, functioning as the second behavior decision unit 33, inputs the difference array derived in step S68 into the behavior decision model 40, thereby determining the behavior of each of the multiple control objects 60. The behavior decision model 40 is a model that learns to output the behavior of each of the multiple control objects 60 when the difference array is input.

[0073] In step S78, the CPU 91, functioning as the second behavior decision unit 33, inputs the behavior of each control object 60 determined in step S76 to the environment simulator 50. In step S80, the CPU 91, functioning as the second sampling unit 34, acquires the distribution of the state of the target space one step after (e.g., 10 minutes after) the plurality of control objects 60 are controlled in accordance with the behavior input in step S78.

[0074] In step S82, the CPU 91, functioning as the second control unit 31, adds 1 to the current number of steps. In step S84, the CPU 91, functioning as the second control unit 31, determines whether the current number of steps is a multiple of M, a predetermined number. M is an arbitrary integer, for example, approximately 18 to 72. If the current number of steps is not a multiple of M (if step S84 is N), steps S66 to S84 are repeated. Note that, in step S66, the temperature distribution acquired in the immediately preceding step S80 is used as the distribution of the state of the target space acquired immediately before.

[0075] On the other hand, if the current number of steps is a multiple of M (if step S84 is Y), the process proceeds to step S86. In step S86, the CPU 91, functioning as the second control unit 31, determines whether the current number of steps has reached a predetermined upper limit (for example, 20,000 times), and if so, ends the reinforcement learning process. On the other hand, if the current number of steps does not reach the upper limit (if step S86 is N), the process returns to step S60. In other words, reinforcement learning is performed using a different task array until the current number of steps reaches the upper limit.

[0076] [Experimental Example] An experiment was conducted to determine whether a significant temperature gradient can be achieved in a target space by controlling the air conditioning of a real space using a behavioral decision-making model 40 that underwent reinforcement learning using the method of the reinforcement learning device 10A described in the above embodiment. The experiment was conducted in a corner of an office in January 2023. FIG. 12 is a schematic diagram of the target space used in the experiment, viewed from above, illustrating the positions of desks and chairs placed in the target space. FIG. 12 also illustrates the positions of five target points P1 to P5 used in the verification. An environmental simulator that reproduced this target space was used for learning and operation of the behavioral decision-making model 40.

[0077] First, we will explain the experimental method. A task array consisting of target temperatures at each of five target points P1 to P5 in the target space was input into the behavioral decision model 40, and actual air conditioning control was performed according to the resulting air conditioning control behavior. Because the temperature distribution in the target space changes over time, air conditioning control and temperature measurements at each of target points P1 to P5 were performed over a 180-minute period. This 180-minute air conditioning control and temperature measurement over time was performed three times, with the contents of the task array (target temperature gradient) changed while the positions of target points P1 to P5 remained fixed. Table 1 shows the contents of task arrays A to C used for the three tests.

[0078] Next, we will explain the method for evaluating the measurement results. As an evaluation index, we derived the average of the absolute values ​​of the temperature error between the target value (value determined by the task array) and the actual measured value at each of the target points P1 to P5 (hereinafter referred to as the absolute temperature error average). The closer the absolute temperature error average is to 0, the more successfully the control is aligned with the target value. As mentioned above, the air conditioning control and temperature measurements were performed over time, so the absolute temperature error average was also derived multiple times at regular time intervals. Table 2 shows the progress of the absolute temperature error average for each of task arrays A to C. Figure 13 shows a line graph of the data in Table 2.

[0079] As a comparative example, we also conducted a verification in which the set temperature of the air conditioning (VAV) closest to each of the target points P1 to P5 was set to the target temperature defined in task arrays A to C. The temperature measurement method and the method for evaluating the measurement results were the same as in the above experimental example. Table 3 shows the progress of the average absolute temperature error for each of task arrays A to C in the comparative example. Figure 14 shows a line graph of the data in Table 3.

[0080]

[0081]

[0082]

[0083] As shown in Table 2 and Figure 13, when the behavioral determination model 40 according to this embodiment was used, the minimum value of the absolute temperature error average was 0.31 for task array A, 0.28 for task array B, and 0.60 for task array C. On the other hand, as shown in Table 3 and Figure 14, in the comparative example, the minimum value of the absolute temperature error average was 1.45 for task array A, 1.14 for task array B, and 0.76 for task array C. This shows that by using the behavioral determination model 40 according to this embodiment, it is possible to determine behaviors that accurately generate a significant temperature gradient within the target space.

[0084] As described above, the reinforcement learning device 10A according to this embodiment is a reinforcement learning device that performs reinforcement learning of a behavioral decision-making model 40 for determining the behaviors of a plurality of control objects 60 that each affect the state of a target space, and includes a task generation unit 20 that generates a task array consisting of target states at each of a plurality of target points predetermined in the target space, based on an array consisting of states of the target space obtained by predetermined actions of a plurality of control objects 60 at each of the plurality of target points, and a model learning unit 30 that performs reinforcement learning of the behavioral decision-making model 40 using the task array.

[0085] That is, the reinforcement learning device 10A according to this embodiment can reduce the number of dimensions of information handled by sampling only the states and tasks of target points that are important when generating a temperature gradient, such as seat positions in an office. Reducing the variation in states handled in reinforcement learning narrows the search space for actions, thereby facilitating learning. Therefore, the reinforcement learning device 10A according to this embodiment can build a behavioral decision-making model 40 that can determine actions that generate a significant gradient in the state of the target space.

[0086] In the above embodiment, the state is obtained using the environmental simulator 50, but the present invention is not limited to this. For example, a real environment having the same inputs and outputs may be used instead of the environmental simulator 50. The real environment is, for example, an environment in which there is a building with controllable air conditioning, and indoor temperature changes can be measured by a sensor or the like to collect data.

[0087] In the above embodiment, an example has been described in which temperature is used as an example of a state, an air conditioning device is used as the control target 60, and an air conditioning control action is used as an example of an action, but this is not limiting. For example, the state may be at least one of temperature, humidity, and illuminance. When humidity is used as the state, at least one of a humidifier and a dehumidifier can be used as the control target 60. When illuminance is used as the state, a lighting fixture can be used as the control target 60.

[0088] In the above embodiment, the reinforcement learning device 10A includes both the task generation unit 20 and the model learning unit 30. However, the present invention is not limited to this. For example, the processing of the task generation unit 20 may be performed by a different device.

[0089] Furthermore, the control processing executed by the CPU after reading the software (program) in the above embodiment may be executed by various processors other than the CPU. Examples of processors in this case include PLDs (Programmable Logic Devices) whose circuit configuration can be changed after manufacture, such as FPGAs (Field-Programmable Gate Arrays), and dedicated electrical circuits, such as ASICs (Application Specific Integrated Circuits), which are processors having a circuit configuration designed specifically to execute specific processing. Furthermore, the control processing may be executed by one of these various processors, or by a combination of two or more processors of the same or different types (e.g., multiple FPGAs, or a combination of a CPU and an FPGA). Furthermore, the hardware structure of these various processors is, more specifically, an electrical circuit that combines circuit elements such as semiconductor elements.

[0090] In the above embodiment, the reinforcement learning program and the task generation program are pre-stored (installed) in a storage device, but this is not limiting. The programs may be provided in a form stored in a storage medium such as a CD-ROM, DVD-ROM, Blu-ray disc, or USB memory. The programs may also be downloaded from an external device via a network.

[0091] The following additional notes are provided regarding the above-described embodiments.

[0092] (Supplementary Item 1) A reinforcement learning device that performs reinforcement learning of a behavioral decision-making model to determine the behaviors of multiple control objects that each affect the state of a target space, comprising: a memory; and at least one processor connected to the memory, wherein the processor is configured to: generate a task array consisting of target states at each of multiple target points predetermined in the target space, based on an array consisting of states of the target space obtained by predetermined actions of the multiple control objects at each of the multiple target points; and perform reinforcement learning of the behavioral decision-making model using the task array.

[0093] (Supplementary Item 2) A non-transitory storage medium storing a program executable by a computer to execute a reinforcement learning process of a behavioral decision-making model for determining the actions of a plurality of control objects that each affect the state of a target space, wherein the reinforcement learning process generates a task array consisting of target states at each of a plurality of target points predetermined in the target space based on an array consisting of states of the target space obtained by predetermined actions of the plurality of control objects at each of the plurality of target points, and performs reinforcement learning of the behavioral decision-making model using the task array.

[0094] 10 Air conditioning control scenario calculation device 10A Reinforcement learning device 10B Action decision device 21 First control unit 22 First initialization unit 23 First action decision unit 24 First sampling unit 28 Memory determination unit 31 Second control unit 32 Second initialization unit 33 Second action decision unit 34 Second sampling unit 35 Dimension reduction unit 36 ​​Reward calculation unit 37 Model update unit 40 Action decision model 42 Task DB 50 Environmental simulator 60A to 60E Control target 91 CPU 92 ROM 93 RAM 94 Storage 95 Input unit 96 Display unit 97 Communication I / F 99 Bus 100 Air conditioning control system BO Air outlet S0 Target space

Claims

1. A reinforcement learning device that performs reinforcement learning of an action decision model for determining actions of a plurality of control targets that respectively affect the state of a target space, a task generation unit that generates a task array consisting of target states at each of a plurality of target points predetermined in the target space based on an array consisting of states of the target space obtained by predetermined actions of the plurality of control targets at each of the plurality of target points; a model learning unit that performs reinforcement learning of the action decision model using the task array; A reinforcement learning device comprising:

2. The task generation unit determines the predetermined actions of each of the plurality of control targets, extracts the state at each of the plurality of target points from the distribution of the state of the target space at a time point when a predetermined period has elapsed since the plurality of control targets started being controlled according to the predetermined actions, stores the array consisting of the states at each of the plurality of extracted target points in the storage unit as the task array when no other array similar to the array is stored in the storage unit. The reinforcement learning device according to claim 1.

3. The model learning unit includes a dimensionality reduction unit that derives a difference array indicating the difference between a state array consisting of states at each of the plurality of target points extracted from the distribution of the state of the target space obtained by the actions of the plurality of control targets determined in the reinforcement learning and the task array, The action decision model takes the difference array as an input and outputs actions of each of the plurality of control targets. The reinforcement learning device according to claim 1.

4. The model learning unit further includes a reward calculation unit that calculates a reward based on the magnitude of the difference between the state array and the task array indicated by the difference array, a model update unit that updates the action decision model based on a dataset including the difference array, the action, and the reward. The reinforcement learning device according to claim 3.

5. The model update unit updates the action decision model based on a time-series dataset for a predetermined number of times. The reinforcement learning device according to claim 4.

6. The state is at least one of temperature, humidity, and illuminance. The reinforcement learning device according to claim 1.

7. A reinforcement learning method for performing reinforcement learning of an action decision model for determining actions of a plurality of control targets that respectively affect the state of a target space, the method comprising: generating a task array consisting of target states at respective ones of a plurality of target points predetermined in the target space, based on an array of states of the target space obtained by predetermined actions of the plurality of control targets at each of the plurality of target points; and performing reinforcement learning of the action decision model using the task array. A reinforcement learning method in which a computer executes the process.

8. A reinforcement learning program for causing a computer to function as the reinforcement learning apparatus according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Abnormality detecting method and abnormality detecting system

    JP2010191556A

  • Air conditioning control system, air conditioner and machine learning device

    JP2021032479A

  • Learning data collection device, learning data collection method, and learning data collection program

    WO2022269885A1