Future State Estimation Device
The future state estimation device uses a weighted probability model with basis functions to quickly estimate future states in continuous spaces, addressing the limitations of existing methods by reducing calculation time and enhancing predictive accuracy.
Patent Information
- Application Number
- JP2021187403
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-11-17
- Publication Date
- 2025-09-04
- Estimated Expiration
- 2041-11-17
AI Technical Summary
Existing methods for predicting future states of controlled objects in continuous state spaces require longer calculation times as the prediction horizon increases, limiting their effectiveness, and existing methods in discrete state spaces do not provide a clear method for estimating future states in continuous spaces.
A future state estimation device that uses a weighted probability of state transitions, represented as a linear combination of basis functions, to quickly estimate future states by performing a product-sum operation on a weight matrix, allowing for the calculation of state transition probabilities in continuous spaces.
Enables rapid estimation of future states in continuous state spaces, reducing calculation time and improving predictive accuracy by converting iterative calculations into a product-sum operation.
Smart Images

Figure 0007734051000009 
Figure 0007734051000010 
Figure 0007734051000011
Abstract
Description
[Technical Field]
[0001] The present invention relates to a future state estimation device. [Background technology]
[0002] Model predictive control, which is commonly applied in the fields of automobiles and plants (power generation and industry), tends to perform better the further into the future the state of the controlled object can be predicted. The following devices and methods exist for predicting the future state of the controlled object:
[0003] Patent Document 1 discloses a method for predicting a future state using a model that simulates the behavior of an object to be manipulated, and calculating an amount of manipulation that is appropriate for that future state.
[0004] Patent Document 2 discloses a method for predicting the current and future states of an industrial system to be controlled and optimizing a control law so as to maximize an objective function.
[0005] Patent Document 3 discloses a method of modeling a nonlinear and dynamic system such as a thermal reactor process using a regression method, and calculating optimal manipulated variables using future states predicted by the model.
[0006] Patent Document 4 relates to an automatic control parameter adjustment device that can automatically optimize control parameters according to the purpose while satisfying constraints on plant operation and can also shorten the calculation time required for optimizing the control parameters. It discloses a method for calculating a control law that takes into account future states using a plant model and a machine learning technique such as reinforcement learning.
[0007] Patent Document 5 discloses a method for quickly estimating the future state of an object to be operated at an infinite time in the future in the form of a probability density distribution within a predefined finite and discrete state space by recording a state transition model that expresses the behavior of the object to be operated as a state transition probability and performing a calculation equivalent to an infinite series of the model. [Prior art documents] [Patent documents]
[0008] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-212872 [Patent Document 2] Japanese Patent Application Laid-Open No. 2013-114666 [Patent Document 3] Japanese Patent Application Laid-Open No. 2009-076036 [Patent Document 4] Japanese Patent Application Laid-Open No. 2017-157112 [Patent Document 5] Japanese Patent Application Publication No. 2019-159876 Summary of the Invention [Problem to be solved by the invention]
[0009] The devices and methods in Patent Documents 1, 2, 3, and 4 predict future states using a model that simulates the behavior of the object to be manipulated, and calculate an optimal control method from the predicted future state. While the ability to predict future states further out tends to result in higher performance, methods that use iterative calculations require longer predictive calculations the longer the time until the desired future state. For this reason, it is common to limit calculations to a finite future state that can be calculated within an allowable time range.
[0010] The device and method of Patent Document 5 estimate the state of an operation object and its surrounding environment infinitely in the future in the form of a probability density distribution if it is in a discrete state space, but does not explicitly state a method for estimating the state of an operation object and its surrounding environment infinitely in the future in the form of a probability density distribution in a continuous state space.
[0011] Therefore, an object of the present invention is to provide a future state estimation device that can quickly estimate the future state of a prediction target within a continuous state space. [Means for solving the problem]
[0012] In order to achieve the above object, a future state estimation device of the present invention is provided, which estimates a future state by using a weighted probability of a first state transition, the probability indicating the probability that a prediction target will transition from a first state to a second state after a first time has elapsed. radius a storage device for storing a state transition model expressed by a linear combination of basis functions; radius and a calculation unit that calculates a second state transition probability indicating the probability that the prediction target will transition from the first state to the second state by the time a second time period has elapsed, by performing a product-sum operation on a weight matrix indicating a matrix having weights of basis functions as elements. the radial basis function is a normal distribution function, and the arithmetic unit stores in the storage device a matrix obtained by multiplying a transposed matrix of a transformation matrix having an integral value of the normal distribution function as an element and the weighting matrix by a decay rate, and calculates the second state transition probability based on an inverse matrix of a difference between a unit matrix and the matrix stored in the storage device. do. [Effects of the Invention]
[0013] According to the present invention, it is possible to quickly estimate the future state of a prediction target in a continuous state space. Problems, configurations, and effects other than those described above will become apparent from the following description of the embodiments. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a diagram illustrating a configuration of a processing apparatus according to a first embodiment. [Figure 2] FIG. 2 is a diagram schematically illustrating an arrangement of basis functions. [Figure 3] FIG. 10 is a diagram schematically illustrating a graph of transition probabilities. [Figure 4] FIG. 10 is a diagram illustrating the process of equation (2). [Figure 5] 2 is a diagram showing a flow of processing performed by the processing device shown in FIG. 1. FIG. [Figure 6] FIG. 10 is a diagram illustrating a configuration of a processing apparatus according to a second embodiment. [Figure 7] FIG. 7 is a diagram showing a flow of processing performed by the processing device shown in FIG. 6. [Figure 8] FIG. 10 is a diagram showing the configuration of a first display screen according to the third embodiment. [Figure 9] FIG. 10 is a diagram showing the configuration of a second display screen according to the third embodiment. [Figure 10] FIG. 10 is a diagram showing the configuration of a third display screen according to the third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0015] Examples 1 to 3 will be described below with reference to the drawings.
[0016] Example 1 1 is a configuration diagram illustrating an example of a processing device 100 (future state estimation device) according to a first embodiment of the present invention. The processing device 100 is configured with an input device 110, a data reading device 115, an output device 120, a storage device 130, and a calculation device 140 as main elements.
[0017] Of these, the input device 110 is a part that receives instructions from an operator, and is composed of buttons, a touch panel, and the like.
[0018] The data reading device 115 is a part that receives data from outside the processing device 100, and is composed of a CD drive, a USB terminal, a LAN cable terminal, a communication device, and the like.
[0019] The output device 120 is a device that outputs instruction information, scanned images, scanned results, etc. to an operator, and is composed of a display, a communication device, etc.
[0020] The above-mentioned configurations are standard, and any or all of the input device 110, data reading device 115, and output device 120 may be connected externally to the processing device 100.
[0021] The storage device 130 is a section for storing various types of data, and is composed of a model storage unit 131 and a future state prediction result storage unit 132. Of these, the model storage unit 131 is a section for storing models that simulate the behavior of objects or phenomena whose future states are to be predicted by the processing device 100. The future state prediction result storage unit 132 is a section for storing the calculation results of a future state prediction calculation unit 142, which will be described later. Details of the storage device 130 will be described later, and only an outline of its functions will be described here.
[0022] The calculation device 140 processes data input from the input device 110 and the data reading device 115 and data stored in the memory device 130, and outputs the results to the output device 120 or records them in the memory device 130. The calculation device 140 is composed of the following processing units (input control unit 141, future state prediction calculation unit 142, output control unit 143).
[0023] The input control unit 141 is a part that classifies data input from the input device 110 or the data reading device 115 into commands, models, etc., and transfers them to the storage device 130 and each part of the arithmetic device 140 .
[0024] The future state prediction calculation unit 142 calculates a decay-type state transition probability function from the model data stored in the model storage unit 131 and records the result in the future state prediction result storage unit 132 .
[0025] The output control unit 143 is a part that outputs the data stored in the storage device 130 to the output device 120. When the output destination is a screen or the like, it is preferable that the result be output each time a reading operation is performed. When the output destination is a communication destination or the like, the output process may be performed each time the state transition probability matrix is updated or the future state prediction calculation unit 142 performs a calculation, or the data may be processed by summarizing several times of data or summarizing the data at predetermined intervals.
[0026] The arithmetic device 140 is configured with, for example, a processor such as a CPU (Central Processing Unit), and the storage device 130 is configured with, for example, an HDD (Hard Disk Drive), an SSD (Solid State Drive), or a memory, etc. The processor executes a program stored in the memory, etc., causing the processor and the memory to work together to realize various functions described below.
[0027] The following describes in detail the processing executed using the processing device 100 of Fig. 1. In the following description, in this embodiment, an object or phenomenon whose future state is to be predicted will be referred to as a simulated object. Examples of simulated objects include the behavior of machines or living things, natural or physical phenomena, chemical reactions, fluctuations in money or prices, and changes in consumer demand, but this embodiment does not limit the simulated objects to these examples.
[0028] In this embodiment, the model inputs are the state of the simulated object and influencing factors such as the passage of time, operations, and disturbances, and the output is the state of the simulated object after being affected by the influencing factors, and this model is referred to as a state transition model in this embodiment. Models such as the state transition model are stored in the model storage unit 131 in Fig. 1. Furthermore, the state transition model represents the state of the simulated object at an infinite time or an infinite number of steps in a finite state space in the form of a probability density distribution.
[0029] The storage format of the state transition model in the model storage unit 131 is a linear combination of weighted functions, and may be, for example, a state transition probability matrix, a neural network, a radial basis function network, or a matrix or vector representing the weights of a neural network or a radial basis function network. However, this embodiment does not limit the storage format of the model to be simulated to these examples.
[0030] The weights of the weighting function may be set in advance according to the behavior of the object to be simulated, or may be automatically estimated from time-series data recording the behavior of the object to be simulated using an optimization technique such as a neural network.
[0031] An example of a model format stored in the model storage unit 131 in the form of a radial basis function network with uncorrelated normal distribution as the basis function is shown in the following equation (1).
[0032]
number
[0033] In equation (1), τ is the state transition probability function, s is the state before the operation is applied to the operation target (pre-transition state), s' is the state after the operation is applied to the operation target (post-transition state), M1 is the number of basis functions in the pre-transition state s direction, M2 is the number of basis functions in the post-transition state s' direction, μi (i=1,2,3,...,M1) and μ'j (j=1,2,3,...,M2) are the mean values, σi (i=1,2,3,...,M1) and σ'j (j=1,2,3,...,M2) are the variance values, λij is the weight of the basis function, Λ is a matrix that stores the weights λij of the basis functions, and G is a matrix that stores the normal distribution function, which is the basis function.
[0034] Figure 2 is a diagram showing the arrangement of the basis functions in equation (1). In the case of Figure 2, there are M1 × M2 basis functions, and the state transition probability τ(s, s') is expressed as a linear combination of the product of the basis functions and their weights.
[0035] The state transition probability function τ is a type of model that generally simulates the motion characteristics and physical phenomena of the controlled object, and is a function that stores the transition probability between all states. The output of the function τ is the state before the transition s when a preset interval Δt (or step) has elapsed. i State s' after transition from (i=1, 2, …, N) i (i=1,2,…,N) transition probability P(s'1, s'2, …, s' N |s 1, s 2, …, s N ) In the example of equation (1), the calculation formula is based on the assumption that N=1.
[0036] Figure 3 is a diagram that shows a graph of the transition probability obtained by linearly combining the product of the basis functions and their weights arranged as shown in Figure 2. The graph shows that the probability P(s'1|s1) around the post-transition state s' with the highest transition frequency from each pre-transition state s is high, and conversely, the probability P(s'1|s1) around the post-transition state s' with the lowest transition frequency is low.
[0037] For a simulated object to which this embodiment is applied, when the state of the simulated object and its surrounding environment at an infinite time or an infinite number of steps ahead is estimated in the form of a probability density distribution, the calculation time may not depend on one or more of the distance, time, and steps to the future state to be estimated. N |s1, s2, …, s N ) does not depend on time, step u, which indicates the amount or number of times that the influencing factor interferes with the simulated object, may be used instead of time t.
[0038] 1, the future state prediction result storage unit 132 is a part that stores the calculation results of the future state prediction calculation unit 142. In this embodiment, the data stored in the future state prediction result storage unit 132 is called a state transition probability series sum matrix. The state transition probability series sum matrix and its calculation method will be described later.
[0039] The future state prediction calculation unit 142 calculates a state transition probability series sum matrix from the model data recorded in the model storage unit 131, and records the result in the future state prediction result storage unit 132. An example of a method for calculating the state transition probability series sum matrix is shown in the following equation (2). Note that in the example of equation (2), it is assumed that the model storage format in the model storage unit 131 is a state transition probability function τ.
[0040]
number
[0041] In equation (2), D is a decaying state transition probability function, and γ is a constant called the decay rate, which is greater than or equal to 0 and less than 1. Also, τ(L) is a function (or matrix) that stores the transition probability between all states after a time of Δt×L has passed.
[0042] An example of a method for calculating τ(L) is shown in the following equation (3).
[0043]
number
[0044] In equation (3), kl (l=1,2,...,L-1) is the state passed through from the pre-transition state s to the post-transition state s'. The transition probability at τ(L) is the product of the results of integrating the state transition probability function τ with respect to the state kl passed through.
[0045] Figure 4 is a diagram showing the process of equation (2), in which multiple state transition probability functions τ(s, s') for each elapsed time Δt are multiplied by a weighting coefficient γ, which decays with each elapsed time Δt, and the sum is calculated.
[0046] In this way, the decaying state transition probability function D is the sum of the state transition probability function τ after Δt time has elapsed to the state transition probability function τ∞ after Δt × ∞ time has elapsed, and is also a matrix that preserves the statistical closeness between all states. Also, to lower the weight of states that transition further into the future, the decay rate γ is multiplied by a larger amount depending on the elapsed time.
[0047] Equation (2), which requires calculation from the current state transition probability function τ to the state transition probability function τ∞ after an infinite amount of time has passed, is difficult to calculate within real time. Therefore, this embodiment is characterized by converting equation (2) to the following equation (4). In essence, equation (4) performs a calculation equivalent to a series of a state transition probability matrix when estimating the state of the simulated object and its surrounding environment in the infinite time or infinite steps ahead in the form of a probability density distribution.
[0048]
number
[0049] In equation (4), E is the identity matrix, Ψ is the transformation matrix, and tΨ is the transpose matrix of the transformation matrix Ψ. Equation (4) is a calculation formula equivalent to equation (2). By converting the calculation of the sum from the state transition probability function τ to the state transition probability function τ∞ in equation (2) into the inverse matrix of (E-γΨ transpose Λ) in equation (4), the same calculation result as equation (2) can be obtained within a finite time. Here, if the transformation matrix Ψ is not linearly independent, a pseudo-inverse matrix may be used. An example of a method for calculating the transformation matrix Ψ is shown in equation (5) below.
[0050]
number
[0051] The transformation matrix Ψ is an integral value of the normal distribution, which is a basis function, and is a constant that does not depend on the pre-transition state s or the post-transition state s'.
[0052] In this way, in this embodiment, by using a state transition model as a model that simulates the behavior of the simulation target, it is possible to calculate the state transition probability after Δt×L time by calculating τ(L). In addition, by taking the sum from the state transition probability function τ after Δt time has elapsed to the state transition probability function τ(∞) after Δt×∞ time has elapsed, and weighting with a decay rate γ depending on the elapsed time, it is possible to calculate the state transition probability taking into account the time after Δt×∞ time has elapsed within a finite time.
[0053] FIG. 5 is a diagram showing the flow of processing performed by the processing device 100.
[0054] First, in the process of process step S 1201 , data relating to the model of the object to be simulated is input from data reading device 115 based on a command from input control unit 141 , and the data is recorded in model storage unit 131 .
[0055] Next, in processing step S1202, data relating to the model of the object to be simulated recorded in the model memory unit 131 is transferred to the future state prediction calculation unit 142, and the decaying state transition probability function D is calculated based on equation (4), and the result is recorded in the future state prediction result memory unit 132.
[0056] Finally, in the process of processing step S1203, the data recorded in the future state prediction result storage unit 132 is transferred to the output control unit 143 and output to the output device 120.
[0057] Example 2 6 is a configuration diagram showing an example of a processing device 101 obtained by extending the processing device 100 of the first embodiment to optimization of model-based control. The simulated object in the processing device 101 is the behavior of a control object and its surrounding environment, and the model stored in the model storage unit 131 also simulates the behavior of the control object and its surrounding environment. In this way, the second embodiment assumes a case where the simulated object includes a control object.
[0058] The processing device 101 is configured with an input device 110, a data reading device 115, an output device 120, a storage device 130, and an arithmetic device 150 as main elements.
[0059] Of these, the input device 110 is a part that receives instructions from an operator, and is composed of buttons, a touch panel, and the like.
[0060] The data reading device 115 is a part that receives data from outside the processing device 100, and is composed of a CD drive, a USB terminal, a LAN cable terminal, a communication device, and the like.
[0061] The output device 120 is a device that outputs instruction information, scanned images, scanned results, etc. to the operator, and is composed of a display, a CD drive, a USB terminal, a LAN cable terminal, a communication device, etc.
[0062] The above-mentioned configurations are standard, and any or all of the input device 110, data reading device 115, and output device 120 may be connected externally to the processing device 100.
[0063] The storage device 130 is composed of a model storage unit 131, a future state prediction result storage unit 132, a reward function storage unit 133, and a control law storage unit 134. Of these, the future state prediction result storage unit 132 has almost the same function as in the first embodiment.
[0064] The model storage unit 131 may have the same function as in the first embodiment, but in control, the behavior of the simulated object may change depending on the manipulated variable in addition to the state. When the behavior of the simulated object changes depending on the manipulated variable, the decay type state transition can be calculated in the same way as in the first embodiment by adding information about the manipulated variable to the model.
[0065] The reward function storage unit 133 is a part that stores control targets such as target position and target speed in the form of a function, table, vector, matrix, etc. In this embodiment, a function, table, vector, matrix, etc. that contains information on this control target will be called a reward function r. In this embodiment, the output value of this reward function R will be called a reward r.
[0066] An example of a reward function in functional form is shown in equation (6).
[0067]
number
[0068] Here, μr is the target state and σr is the target variance. The reward function R in equation (6) is a normal distribution for the post-transition state s', which has the characteristic that the reward r is maximized at the target state μr and the further away from the target state μr it outputs a smaller reward r. The range of states that obtain a high reward r is adjusted by the target variance σr. Note that examples of rewards in control include the desired value or objective function in reinforcement learning in AI (Artificial Intelligence).
[0069] 6, the control law storage unit 134 is a part that stores the optimum control law for the control target. An example of the control law stored in the control law storage unit 134 is shown in equation (7).
[0070]
number
[0071] Here, X is the control law, V is the value function, P is the state transition probability, and a is the manipulated variable. The value function V is a function that preserves the proximity to the target state sgoal (or a statistical index showing the ease of transition). The calculation method for the value function V will be described later. Equation (7) preserves the manipulated variable a that maximizes the value obtained by integrating the product of the value function V and the state transition probability P over the post-transition state s', among all the manipulated variables a.
[0072] Returning to FIG. 6, the arithmetic unit 150 processes data input from the input unit 110 and the data reading unit 115 and data stored in the memory unit 130, and outputs the results to the output unit 120 or records them in the memory unit 130. The arithmetic unit 150 is composed of the following processing units:
[0073] The input control unit 151 is a part that classifies data input from the input device 110 or the data reading device 115 into commands, models, etc., and transfers them to the respective parts of the storage device and arithmetic device.
[0074] The future state prediction calculation unit 152 is equivalent to the future state prediction calculation unit 142 of the first embodiment. The output control unit 153 is also equivalent to the output control unit 143 of the first embodiment.
[0075] The control law calculation unit 154 calculates the optimal control law (optimal operation amount a) from the decaying state transition probability function D recorded in the future state prediction result memory unit 132 and the reward function R recorded in the reward function memory unit 133, and records it in the control law memory unit 134.
[0076] An example of a method for calculating the optimal control law is shown below. In this example, the optimal control law is calculated in the following two stages.
[0077] Step 1: First, calculate the value function V using the decaying state transition probability function D and the reward function R. The value function V may be saved in the form of a table, vector, matrix, or other format other than a function, and the saving format is not limited in this embodiment. An example of a method for calculating the state value function V is shown in the following equation (8).
[0078]
number
[0079] As shown in equation (8), the value function V is a function obtained by integrating the product of the decaying state transition probability function D and the reward function R over the post-transition state s'. The value of the value function V is higher the easier it is for a state to transition to the target state sgoal. In this embodiment, the output of this value function V is called value. Furthermore, the value of the value function V in this embodiment is equivalent to the definition of the state value function in reinforcement learning.
[0080] Step 2: Next, the optimal manipulated variable a in the current pre-transition state s is calculated using the value function V. The optimal manipulated variable a is calculated using the above equation (7).
[0081] In this way, by calculating the value using the above equation (8), it is possible to evaluate the ease of transition to sgoal in each state, and the above equation (7) makes it possible to identify the optimal operation amount a.
[0082] Returning to Figure 6, when update data for the model data recorded in the model storage unit 131 is input from the data reading device 115, the model update unit 155 modifies the model data based on the update data and records the modified model data in the model storage unit 131.
[0083] FIG. 7 is a diagram showing the flow of processing performed by the processing device 101.
[0084] First, in processing step S1301 of FIG. 7, based on a command from the input control unit 141, data regarding the model of the object to be simulated and data regarding the reward function R are input from the data reading device 115, and the data are recorded in the model memory unit 131 and the reward function memory unit 133.
[0085] Next, in processing step S1302, data regarding the model of the object to be simulated recorded in the model memory unit 131 is transferred to the future state prediction calculation unit 142, and the decaying state transition probability function D is calculated based on equation (4), and the result is recorded in the future state prediction result memory unit 132.
[0086] Next, in processing step S1303, the decaying state transition probability function D recorded in the future state prediction result memory unit 132 and the reward function R recorded in the reward function memory unit 133 are transferred to the control law calculation unit 154, which calculates the optimal control law and records the result in the control law memory unit 134.
[0087] Next, in processing step S 1304 , the data recorded in the future state prediction result storage unit 132 and the control law storage unit 134 is transferred to the output control unit 143 and output to the output device 120 .
[0088] Next, in processing step S1305, it is determined whether or not to terminate the control of the control object. If the control is to be continued, the process proceeds to processing step S1306, and if the control is to be terminated, the flow also ends.
[0089] Next, in process step S1306, the control target calculates the manipulated variable a and executes an operation based on the control law sent to the control target from the output device 120. That is, the control target executes an operation according to the manipulated variable a.
[0090] Next, in processing step S1307, the controlled object transmits to the data reading device 115 the state of the controlled object and its surrounding environment measured before and after the execution of the operation.
[0091] Next, in processing step S1308, the input control unit 141 determines whether or not data on the state of the controlled object and its surrounding environment measured before and after the execution of the operation has been received by the data reading device 115. If data has been received, the process proceeds to processing step S1309; if data has not been received, the process returns to processing step S1305.
[0092] In processing step S1309, if the data reading device 115 receives data on the state of the controlled object and its surrounding environment measured before and after the execution of the operation in the processing of processing step S1308, the received data and the model data recorded in the model storage unit 131 are transferred to the model update unit 155, and the updated model data is recorded in the model storage unit 131. Then, the process proceeds to processing step S1302.
[0093] Example 3 8, 9, and 10 are examples of screens displayed on the output device 120 in the first and second embodiments.
[0094] 8 shows a state transition probability function τ displayed on the screen as an example of model data recorded in the model storage unit 131. In the figure, the state transition probability function τ is displayed on the screen in functional form as an example of a model storage format, and the transition probability from a pre-transition state s to a post-transition state s' is displayed. The transition probability may be updated from this screen via the input device 110.
[0095] 9 shows an example of a case where the decay-type state transition probability function D stored in the future state prediction result storage unit 132 is displayed on the screen. In the figure, the decay-type state transition probability function D is displayed on the screen in the form of a function from the pre-transition state s to the post-transition state s'.
[0096] 10 shows an example of a transition probability distribution P displayed as data obtained by processing the model data stored in the model storage unit 131. On the screen, the transition probability P is displayed with the state s' of the transition destination on the horizontal axis.
[0097] The main features of Examples 1 to 3 can be summarized as follows.
[0098] The future state estimation device (processing device 100) shown in Fig. 1 includes a storage device 130 and a calculation device 140. The storage device 130 stores a state transition model in which a first state transition probability (state transition probability function τ) indicating the probability that a prediction target will transition from a first state s to a second state s' after a first time (Δt) has elapsed is expressed as a linear combination of weighted basis functions (model storage unit 131, equation (1)). The calculation device 140 calculates a second state transition probability (decaying state transition probability function D, equation (2)) indicating the probability that a prediction target will transition from a first state s to a second state s' by the time a second time (Δt × ∞) has elapsed, by a product-sum operation of a weight matrix Λ indicating a matrix whose elements are the weights of the weighted basis functions (future state prediction calculation unit 142).
[0099] This allows the second state transition probability (decay type state transition probability function D) to be calculated by a product-sum operation of the weighting matrix Λ, rather than a multiple integral. As a result, the future state of the prediction target can be quickly estimated in the form of a transition probability distribution within a continuous state space.
[0100] In this embodiment, the product-sum operation is a series calculation of the weight matrix Λ (equations (1) and (2)), and the second time period is an infinite time period or an infinite number of steps. This allows the state of the target to be predicted after an infinite time period or an infinite number of steps has passed to be estimated quickly.
[0101] 6 calculates the optimal manipulated variable a of the device that controls the prediction target based on the second state transition probability (decaying state transition probability function D) (control law calculation unit 154). This makes it possible to statistically estimate the optimal manipulated variable for the control target.
[0102] 1 calculates the element values λij of the weighting matrix Λ from the time series of measurement values to be predicted using an optimization method such as a neural network, thereby enabling the element values λij of the weighting matrix Λ to be fed back to the state transition model.
[0103] 6 updates the state transition model using the element values λij of the weight matrix Λ (model update unit 155). The calculation device 150 may learn the element values λij of the weight matrix Λ using, for example, deep learning, and update the state transition model using the learned element values of the weight matrix. This can improve the accuracy of the state transition model.
[0104] The future state estimation device (processing device 100) shown in Fig. 1 includes an output device 120. The arithmetic device 140 may cause the output device 120 to output information indicating any two or more of the state transition model before the update, the state transition model after the update, and the difference between the state transition models before and after the update (output control unit 143). The state transition model is displayed, for example, as shown in Fig. 8.
[0105] This allows you to visually check how the state transition model has changed due to the update.
[0106] The arithmetic unit 140 may output the probability of transition from the source state to the destination state in one or more of the elapsed time, elapsed steps, time range, and step range to the output unit 120. In the example of Fig. 10, the transition probability in the source state and the specified elapsed time is displayed as a continuous function of the destination state.
[0107] This allows the transition probability distribution of the prediction target at the specified elapsed time to be visually confirmed.
[0108] In this embodiment, the basis functions are radial basis functions, which allow the first state transition probability (state transition probability function τ) to be expressed as a matrix.
[0109] In this embodiment, the radial basis functions are normal distribution functions, which makes the elements Ψij of the transformation matrix Ψ constants that do not depend on the first state s and the second state s′.
[0110] The calculation device 140 stores in the storage device 130 (e.g., memory) the probability that the prediction target will transition from the first state s to the second state s' after an integer (L) multiple of the first time (Δt) has elapsed, and calculates the second state transition probability (decaying state transition probability function D, equation (2)) from the sum of values obtained by multiplying each probability stored in the storage device 130 by the power of the decay rate γ according to the elapsed time (future state prediction calculation unit 142).
[0111] This makes it possible to calculate the second state transition probability (decay type state transition probability function D) by a product-sum operation of the weight matrix Λ.
[0112] The calculation device 140 stores in the storage device 130 a matrix γtΨΛ obtained by multiplying the product of the transposed matrix tΨ of the transformation matrix Ψ, whose elements Ψij are the integral values of the normal distribution function, and the weighting matrix Λ by a decay rate γ, and calculates a second state transition probability (decaying state transition probability function D, equation (4)) based on the inverse matrix of the difference between the unit matrix E and the matrix γtΨΛ stored in the storage device 130 (future state prediction calculation unit 142).
[0113] As a result, even if the state s is continuous, the second state transition probability (decay type state transition probability function D) can be calculated by a product-sum operation of the weight matrix Λ.
[0114] In particular, the calculation unit 140 calculates the weight matrix Λ and the inverse matrix (E-γtΨΛ) -1 and the Frobenius inner product of the Gaussian function matrix G (future state prediction calculation unit 142).
[0115] The computing device 140 is installed in a plant (for example, a power plant, a chemical plant, etc.) and calculates the operation amount of a device (for example, a steam generator, an evaporator, etc.) that controls a prediction target (for example, temperature, pressure, etc.), thereby improving the production efficiency of the plant.
[0116] In this embodiment, the weighting matrix Λ is a matrix, but it may also be a vector. Note that a 1-row, N-column or N-row, 1-column matrix can also be called a vector. The prediction target is an object (e.g., steam) controlled by a device in the plant or a physical quantity (temperature, pressure, etc.) of the surrounding environment (e.g., air) of the object (steam). This allows the distribution of the future state of the surrounding environment to be estimated quickly.
[0117] The present invention is not limited to the above-described embodiments, but includes various modifications. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, or to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace part of the configuration of each embodiment with other configurations.
[0118] Furthermore, some or all of the above-described configurations, functions, etc. may be implemented in hardware, for example, by designing them as integrated circuits. Furthermore, the above-described configurations, functions, etc. may be implemented in software by a processor interpreting and executing a program that implements each function. Information such as the programs, tables, and files that implement each function can be stored in memory, a storage device such as a hard disk or SSD, or a storage medium such as an IC card, SD card, or DVD.
[0119] The present invention may be embodied in the following manner: The following manner aims to provide a means for quickly estimating the state of an object to be manipulated or its surrounding environment at an infinite time into the future in the form of a probability density distribution, provided that the state is within a predefined finite and continuous state space.
[0120] [1]. A future state estimation device comprising a model memory unit that stores a state transition model that expresses the characteristics of the state transition probability of an object to be operated or the surrounding environment of the object to be operated as a linear combination of weighted functions, and that estimates the future state of the object to be operated or the surrounding environment of the object to be operated in the form of a probability density distribution by using a signal in which the weights of the weighted functions are converted into a matrix or vector and performing a product-sum operation on the weight matrix or vector.
[0121] [2]. A future state estimation device as described in [1], characterized in that the future state of the operation object or the surrounding environment of the operation object at an infinite time or an infinite number of steps ahead is estimated in the form of a probability density distribution by series calculation of the weight matrix or vector.
[0122] [3]. A future state estimation device according to [1] or [2], characterized in that it comprises an optimal operation amount calculation unit that calculates an optimal operation amount based on a probability density distribution of the future state of the operation object or the surrounding environment of the operation object.
[0123] [4]. A future state estimation device according to any one of [1] to [3], characterized in that it comprises a learning unit that calculates each element value of the weight matrix or vector from time series data that records the state transition characteristics of the operation object or the surrounding environment of the operation object or information including those characteristics.
[0124] [5]. A future state estimation device according to any one of [1] to [3], characterized in that it comprises a model update unit that updates the information in the model memory unit from time series data that records the state transition characteristics of the operation object or the surrounding environment of the operation object or information including those characteristics.
[0125] [6]. A future state estimation device according to the future state estimation method described in [4], characterized in that it comprises a model update unit that updates the information in the model storage unit based on the element values of the weight matrix or vector calculated by the learning unit.
[0126] [7]. A future state estimation device according to any one of [1] to [6], which is provided with a display means, characterized in that the display means outputs two or more of the following: a model before the update, a model after the update, and information regarding the difference between the model before and after the update.
[0127] [8]. A future state estimation device according to any one of [1] to [6], which is provided with a display means, characterized in that the display means displays the probability of transitioning from the source state to each state at one or more of a specified elapsed time, elapsed step, time range, or step range.
[0128] According to [1]-[8], the future state of an object of operation in an infinite time future can be calculated in the form of a probability density distribution of continuous states, regardless of the time until the future state to be predicted. Using the results of this calculation, it is possible to provide a method for calculating optimal control laws that take into account future states in an infinite time future. In addition, in the field of automated design, it is possible to provide a path optimization method that takes into account all possible paths, in the field of finance, a pricing method that takes into account distant future states, and in the field of bioengineering, a method for optimizing metabolic pathways that takes into account all pathways within the range that can be modeled. [Explanation of symbols]
[0129] 100... Processing device 101... Processing equipment 110...Input device 115...Data reading device 120...Output device 130...Storage device 131...Model memory section 132...future state prediction result storage unit 133...Reward function memory unit 134...Control law memory unit 140...Arithmetic device 141...input control unit 142...Future state prediction calculation unit 143...Output control section 150...Arithmetic device 151...input control unit 152...Future state prediction calculation unit 153...Output control section 154...Control law calculation unit 155...Model Update Section
Claims
1. a storage device that stores a state transition model in which a first state transition probability indicating the probability that a prediction target will transition from a first state to a second state after a first time has elapsed is expressed by a linear combination of weighted radial basis functions; and a calculation device that calculates a second state transition probability indicating a probability that the prediction target will transition from the first state to the second state by a second time elapse, by a product-sum operation of a weight matrix indicating a matrix having weights of the weighted radial basis functions as elements, the radial basis function is a normal distribution function, the calculation device stores in the storage device a matrix obtained by multiplying a transposed matrix of a transformation matrix having elements of the integral value of the normal distribution function and the weighting matrix by a decay rate, and calculates the second state transition probability based on an inverse matrix of a difference between an identity matrix and the matrix stored in the storage device.
2. 2. The future state estimation device according to claim 1, the multiply-and-accumulate operation is a series calculation of the weight matrix, The second time period is an infinite time or an infinite number of steps. A future state estimation device characterized by:
3. 2. The future state estimation device according to claim 1, The computing device Calculating an optimal operation amount of a device that controls the prediction target based on the second state transition probability A future state estimation device characterized by:
4. 2. The future state estimation device according to claim 1, The computing device Calculate the element values of the weight matrix from the time series of the measurement values to be predicted. A future state estimation device characterized by:
5. 5. The future state estimation device according to claim 4, The computing device The state transition model is updated using the element values of the weight matrix. A future state estimation device characterized by:
6. 5. The future state estimation device according to claim 4, The computing device learning element values of the weight matrix; The state transition model is updated using the learned element values of the weight matrix. A future state estimation device characterized by:
7. 6. The future state estimation device according to claim 5, further comprising an output device; The computing device and outputting information indicating any two or more of the state transition model before the update, the state transition model after the update, and a difference between the state transition model before the update and the state transition model after the update to the output device. A future state estimation device characterized by:
8. 2. The future state estimation device according to claim 1, further comprising an output device; The computing device The output device is caused to output the probability of transition from a source state to a destination state in one or more of an elapsed time, an elapsed step, a time range, and a step range. A future state estimation device characterized by:
9. 2. The future state estimation device according to claim 1, The computing device The second state transition probability is calculated from a Frobenius inner product of the product of the weight matrix and the inverse matrix and a Gaussian function matrix. A future state estimation device characterized by:
10. 2. The future state estimation device according to claim 1, The computing device Calculate the manipulated variable of a device that is installed in a plant and controls the target of prediction. A future state estimation device characterized by:
11. The future state estimation device according to claim 10, the weight matrix is a matrix or a vector, The prediction target is a physical quantity of an object controlled by the device or a physical quantity of the surrounding environment of the object. A future state estimation device characterized by:
Citation Information
Patent Citations
Process state prediction method
JP2009076036A
Differential calculation machine for model prediction control
JP2013114666A
Method for improving performance of method for computationally predicting future state of target object, driver assistance system, vehicle including such driver assistance system and corresponding program storage medium and program
JP2016212872A
Control parameter automatic adjustment apparatus, control parameter automatic adjustment method, and control parameter automatic adjustment apparatus network
JP2017157112A
Future state estimation device and future state estimation method
JP2019159876A