Control system, control device, and control method
Patent Information
- Application Number
- JP2022001867
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-01-07
- Publication Date
- 2026-08-27
- Estimated Expiration
- 2042-01-07
AI Technical Summary
【0009】 本発明によれば、制御対象の制御入力を適切に導出することが可能となる。
Smart Images

Figure 0007911843000007 
Figure 0007911843000008 
Figure 0007911843000009
Abstract
Description
Technical Field
[0001] The present invention relates to a control system, a control device, and a control method for controlling a control target.
Background Art
[0002] Conventionally, in order to stably operate various control targets, control devices that give control inputs according to the characteristics of each control target have been proposed. For example, a technique is known in which the supply amount of a solidifying material added to a raw material is controlled using a signal from a raw material passing sensor that detects whether or not the raw material is flowing on a belt conveyor, and the flow rate of the raw material on the belt conveyor is quantified (for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The control device can stably operate the control target by acquiring the detection results of a plurality of sensors attached to the control target and giving a control input corresponding to the detection results to the control target. However, if the detection results obtained from the sensors are not uniform or if detection results corresponding to noise are frequently included, the control system may not be stable depending on the configuration of the control device. Also, if the standard state of the control target cannot be appropriately defined, the control target may not be determined and the control input may not be accurately derived.
[0005] In view of such problems, an object of the present invention is to provide a control system, a control device, and a control method capable of appropriately deriving a control input for a control target.
Means for Solving the Problems
[0006] To solve the above problems, the control system comprises a sensor attached to the object to be controlled, and a control device including a processor and memory. The processor functions as an information acquisition unit that cooperates with a program contained in memory to acquire the detection result of the sensor, and a conversion function derivation unit that derives a conversion function that converts the sensor detection result into a state in a unimodal state space.
[0007] To solve the above problems, the control device is a control device including a processor and memory, wherein the processor functions as an information acquisition unit that cooperates with a program contained in memory to acquire detection results from sensors attached to the controlled object, and a conversion function derivation unit that derives a conversion function that converts the sensor detection results into states in a unimodal state space.
[0008] To solve the above problems, the control method involves a control device including a processor and memory, in which the processor, in cooperation with a program contained in the memory, acquires detection results from sensors attached to the controlled object and derives a transformation function that converts the sensor detection results into a state in a unimodal state space. [Effects of the Invention]
[0009] According to the present invention, it becomes possible to appropriately derive the control input of the controlled object. [Brief explanation of the drawing]
[0010] [Figure 1] This is a schematic diagram showing the general configuration of a control system according to an embodiment of the present invention. [Figure 2] This is a functional block diagram showing the schematic functions of a control device according to an embodiment of the present invention. [Figure 3] This flowchart shows the processing flow of the control method according to an embodiment of the present invention. [Figure 4] This is an explanatory diagram illustrating the image of the state distribution in the observed state space of data according to an embodiment of the present invention. [Figure 5] This is an explanatory diagram showing an image of the normalized state space according to an embodiment of the present invention. [Figure 6] This is an explanatory diagram illustrating the evaluation of maintenance input according to an embodiment of the present invention. [Figure 7] This is an explanatory diagram illustrating the evaluation of maintenance input according to an embodiment of the present invention. [Figure 8] This flowchart shows another processing flow of the control method according to an embodiment of the present invention. [Modes for carrying out the invention]
[0011] Preferred embodiments of the present invention will be described in detail below with reference to the attached drawings. The dimensions, materials, and other specific numerical values shown in these embodiments are merely examples to facilitate understanding of the invention and do not limit the present invention unless otherwise specified. In this specification and drawings, elements having substantially the same function and configuration are denoted by the same reference numerals to avoid redundant explanations, and elements not directly related to the present invention are omitted from the illustration.
[0012] <Configuration of control system 100> The configuration of the control system 100 according to an embodiment of the present invention will be described with reference to Figures 1 and 2.
[0013] Figure 1 is a schematic diagram showing the general configuration of the control system 100. The control system 100 includes the controlled object 110, the sensor 120, and the control device 130.
[0014] The controlled object 110 is an object that is controlled electrically or mechanically. For the sake of explanation, here we will use a belt conveyor that transports soil as an example of the controlled object 110. However, the controlled object 110 is not limited to a belt conveyor; it is sufficient for it to operate based on the control input, and it can be applied to various plants such as factories and wind turbines for wind power generation.
[0015] Sensor 120 is an IoT sensor with a built-in battery that detects various environmental information at its installation location. For example, if the controlled object 110 is a conveyor belt, Sensor 120 is installed at multiple different locations on the controlled object 110. Sensor 120 then detects acceleration (vibration) in three orthogonal axes at its installation location. However, the detection targets of Sensor 120 are not limited to this case; it can also detect various environmental conditions at its installation location, such as angular acceleration around the three axes, pressure, temperature, humidity, brightness, and lightness.
[0016] Here, the sensor 120 detects vibration modes at multiple installation locations on the controlled object 110 on which the sensor 120 is installed. Specifically, the sensor 120 detects the 3-axis acceleration of the controlled object 110, performs a Fourier transform on the combined 3-axis acceleration, identifies seven frequencies from its frequency components in descending order of peak value, and outputs a digital signal of these seven frequencies and their respective intensities for each sample. Therefore, the frequencies and intensities transmitted from the sensor 120 will differ for each sample.
[0017] The control device 130 acquires the detection result from the sensor 120 and provides a control input to the controlled object 110 according to the detection result, thereby ensuring stable operation of the controlled object 110. In this embodiment, as an example, maintenance is performed to keep the controlled object 110 in a normal state by providing a control input. Therefore, here we give an example of providing a maintenance input as a control input to the controlled object 110. Maintenance inputs include, for example, lubricating the drive parts of the controlled object 110, replacing parts, changing the soil transport speed, and removing foreign objects. The control device 130 provides such maintenance inputs to control the vibration mode of the controlled object 110 so that it can maintain a normal state.
[0018] To control the controlled object 110 with the control device 130, one might consider modeling the controlled object 110 and adjusting the maintenance input so that its output, for example, the detection result of the sensor 120, is stable. However, the detection results obtained from the sensor 120 may not be uniform and may be of various values. Also, if the environment in which the controlled object 110 is located is harsh and unpredictable, detection results corresponding to noise may be detected at a high frequency. In such cases, depending on the model configuration, unexpected maintenance inputs may be derived, or singularities may occur, potentially causing the control system to become unstable.
[0019] Furthermore, in the observed state space, which is a state space that simply reflects the detection results of the sensor 120, the detection results are distributed in a complex manner, and it may not be possible to properly define the standard state of the controlled object 110. For example, if the controlled object 110 is a belt conveyor, the optimal speed and load of the belt conveyor change depending on the weight and size of the object being conveyed, making it difficult to uniformly set a standard state. In that case, the control device 130 may not be able to determine what state should be the standard state that should be the control target, and may not be able to accurately derive maintenance inputs.
[0020] Therefore, in this embodiment, the observed state space is transformed into a normalized state space, and the control target and maintenance input are appropriately derived by using the state distribution in that normalized state space.
[0021] Figure 2 is a functional block diagram showing the schematic functions of the control device 130. As shown in Figure 2, the control device 130 is composed of an I / F unit 132, a data holding unit 134, and a central control unit 136.
[0022] The I / F unit 132 is an interface for bidirectional information exchange with multiple sensors 120. The data storage unit 134 consists of RAM, flash memory, HDD, etc., and stores programs and various information necessary for processing each of the functional modules shown below.
[0023] The central control unit 136 is composed of a semiconductor integrated circuit including a processor, a ROM containing programs, a RAM as a work area, and controls the I / F unit 132, data holding unit 134, etc., via the system bus 138. In this embodiment, the processor in the central control unit 136 works in cooperation with the programs contained in the ROM to function as functional modules such as an information acquisition unit 140, a conversion function derivation unit 142, an evaluation unit 144, a reinforcement learning unit 146, and a control input derivation unit 148.
[0024] <Operation of control system 100> The operation of the control system 100 according to an embodiment of the present invention will be described with reference to Figures 3 to 8.
[0025] Figure 3 is a flowchart showing the processing flow of the control method. First, when the control method is started (S100), the information acquisition unit 140 of the central control unit 136 acquires the detection results of multiple sensors 120 (S110). The conversion function derivation unit 142 uses Normalizing Flow to derive a conversion function f that converts the detection results of multiple sensors 120 into states in the normalized state space (S120). The evaluation unit 144 uses the normalized state space to evaluate a predetermined maintenance input (S130). Thus, the control method is completed (S140). Below, the control method for appropriately deriving the maintenance input of the controlled object 110, which is characteristic of this embodiment, will be described in detail, taking into account the operation of each functional module of the central control unit 136.
[0026] In step S110 of Figure 3, the information acquisition unit 140 receives the data x represented by the following formula 1 as the detection result of sensor 120. i (t) is acquired and stored in the data holding unit 134. Here, i is an identifier that identifies the sensor 120. Sensor 120 outputs seven frequency data and intensity data for each frequency, i.e., n=14 data in one sampling.
number
[0027] Furthermore, the information acquisition unit 140 can acquire detection results from multiple sensors 120 (for example, h sensors) at once in a single sampling. Therefore, the total data X that the information acquisition unit 140 can acquire during a predetermined sampling period is expressed by equation 2. Here, the sampling count r is obtained by dividing the sampling period by the sampling interval, which is set to 10000, for example.
number
[0028] Figure 4 shows data x i This is an explanatory diagram illustrating the image of the observed state space O of (t). When the total data X described above is plotted on the observed state space O, the state distribution is as shown in Figure 4. In Figure 4, the data is biased and complex due to various environmental differences such as the weight and size of the cargo transported by the belt conveyor (the controlled object 110), and the weather, season, and temperature at the time the belt conveyor is being used. As shown in Figure 4, if the state distribution does not have regularity, the control device 130 cannot identify which state in the observed state space O should be the standard state, which is the control target. Therefore, the control device 130 defines the standard state, which is the control target, by appropriately transforming the observed state space O onto a normalized state space S that is easier to control.
[0029] In step S120 of Figure 3, the conversion function derivation unit 142 derives a conversion function f that converts the detection results of the multiple sensors 120 into states on the normalized state space S. Specifically, the conversion function derivation unit 142 first provisionally sets the conversion function f. Then, the data x acquired by the sensor 120 in the first sampling... i (t0) is transformed by the transformation function f (x i (t0)) is a state on the normalized state space S as shown in equation 3. i Let (t0) be the case.
number
[0030] The conversion function derivation unit 142 determines that such a normalized state space S is a unimodal state space, here represented by a Gaussian distribution, that is, the distribution function P(S i (t j ))⊃N(0, σ), and derives the conversion function f using Normalizing Flow, which is an example of a neural network.
[0031] Similarly, the conversion function derivation unit 142 sequentially converts the data x i (t1),…,x i (t r ) obtained during a predetermined sampling period using the conversion function f, and adjusts the conversion function f using Normalizing Flow so that the states s i (t1),…,s i (t r ) on the normalized state space S, which are the conversion results, form a state space represented by a Gaussian distribution. Normalizing Flow is a method of learning such that the distribution transformed by the conversion function f matches a target distribution, for example, a Gaussian distribution. The match between the transformed distribution and the Gaussian distribution can be evaluated using, for example, the Kullback-Leibler information measure, which is a measure of the difference between two probability distributions. At this time, Normalizing Flow is executed so that data close to the standard state is transformed to the center of the normalized state space S, and irregular states different from the standard state are transformed to the periphery of the normalized state space S. In this way, the conversion function derivation unit 142 can derive the conversion function f that converts the detection results of a plurality of sensors into states on the normalized state space S represented by a Gaussian distribution.
[0032] Here, a Lyapunov function V is given as the target Gaussian distribution. The Lyapunov function V is a convex-down unimodal function with an equilibrium point at the origin "0 (zero)". Therefore, the conversion function derivation unit 142 makes the normalized state space S a state space represented by the Lyapunov function V, that is, the distribution function P(S iWe derive the transformation function f using Normalizing Flow such that the relationship between (t) and the Lyapunov function V is as shown in equation 4.
number
[0033] Data x for the number of sampling steps r i (t0), ..., x i (t r Using this, once the transformation function f is derived, this transformation function f is used as the final value, and thereafter, the data x, which is the detection result of sensor 120, is used. i (t) is the state s as shown in equation 5 by its transformation function f. i It is converted to (t).
number
[0034] Note that, in this example, the adjustment of the transformation function f was limited to the number of samples r. However, even after the adjustment for the number of samples r is complete, the transformation function f continues to be applied to newly input data x. i It may also be said that it is updated by (t). Also, once the effectiveness of the fixed transformation function f decreases, for example, every year, the most recent data x i (t) may be used to revise the transformation function f.
[0035] Figure 5 is an explanatory diagram illustrating the image of the normalized state space S. The normalized state space S of the total data X shown in Figure 4 can be expressed by equation 6 using the transformation function f derived by the transformation function derivation unit 142.
number
[0036] Comparing Figure 4 and Figure 5, it can be seen that the normalized state space S is smoother compared to the total data X. Furthermore, since the equilibrium point (origin) in the normalized state space S can be defined as the standard state, the control target can be easily assumed. Moreover, since the normalized state space S is a unimodal space with the standard state as its equilibrium point, it is possible to determine whether the system is approaching the standard state due to the maintenance input m using a simple criterion such as whether it is moving towards the equilibrium point.
[0037] In step S130 of Figure 3, the evaluation unit 144 determines the maintenance input m and evaluates the value (good or bad) of the maintenance input m based on the normalized state space S.
[0038] Figure 6 is an explanatory diagram illustrating the evaluation of the maintenance input m. First, the evaluation unit 144 evaluates the data x, which is the detection result of the sensor 120 before the maintenance input m is given to the controlled object 110. i (t o The evaluation unit 144 then obtains the acquired data x as shown in Figure 6. i (t o The state s on the normalized state space S obtained by transforming ) with the transformation function f. i (t o Derive the (prior state).
[0039] Here, based on the state in the normalized state space S, a predetermined maintenance input m, for example, the operation of applying oil to the drive part of the belt conveyor, is given to the controlled object 110. The evaluation unit 144 receives data x, which is the detection result of the sensor 120 after the maintenance input m has been given to the controlled object 110. i (t n The evaluation unit 144 then obtains the acquired data x as shown in Figure 6. i (t n The state s on the normalized state space S obtained by transforming ) with the transformation function f. i (t n The (post-state) is derived. In this way, two states (the prior state and the post-state) are identified on the normalized state space S in Figure 6.
[0040] In step S120 of Figure 3, the evaluation unit 144 evaluates a predetermined maintenance input m based on the relative relationship between the prior state and the subsequent state in the normalized state space S. For example, when referring to a solid line segment from the prior state to the subsequent state as shown in Figure 6, the longer the distance and the closer the angle between that direction and the line segment from the prior state to the equilibrium point is to 0 degrees, the more the evaluation unit 144 evaluates that the maintenance input m is effective or that the maintenance is correct.
[0041] Figure 7 is an explanatory diagram illustrating the evaluation of maintenance input m. The maintenance input m is defined as follows on the normalized state space S in Figure 7. For example, for the controlled object 110, the maintenance input m is m1, m2, ..., m k Assume the following maintenance inputs are given in this order: m1, m2, ..., m k Therefore, data X is x i (t0)x i (t1), ..., x i (t k ) changes as follows, and the normalized state space S becomes s i (t0), s i (t1), ..., s i (t k It changes like this.
[0042] At this time, the distribution function P(S i The log-likelihood of (t) ΣlogP(S i (t)) is maximized, for example, in a given state s i The gradient dV / ds of the Lyapunov function V at (t) i According to data x i By increasing or decreasing (t) to select the maintenance input m, the evaluation of the maintenance input m is improved, and the controlled object 110 can be efficiently transitioned to the standard state.
[0043] Furthermore, such maintenance inputs m can also be obtained by reinforcement learning on a normalized state space S.
[0044] Figure 8 is a flowchart showing another processing flow of the control method. Instead of, or in addition to, the evaluation of the evaluation unit 144 (S130) in Figure 3, the reinforcement learning unit 146 performs the data x detection result of the sensor 120. i (t) or state s i Based on (t), an action model that derives the maintenance input m is reinforced through learning (S230). Then, the control input derivation unit 148 uses the reinforced learning action model to derive the maintenance input m to be executed (S240).
[0045] Here, given a normalized state space S and a control objective, we assume that in reinforcement learning, rewards are defined for the distance to the control objective and for reaching the control objective.
[0046] In step S230 of Figure 8, the reinforcement learning unit 146 uses the behavioral model to determine the data x, which is the detection result of the sensor 120. i (t o ) or state s i (t o The maintenance input m is derived from ) and the maintenance input m is given to the controlled object 110. The reinforcement learning unit 146 then uses the data x, which is the detection result of the sensor 120 after the maintenance input m has been given to the controlled object 110. i (t n ) is obtained, and the state s on the normalized state space S is transformed by the transformation function f. i (t n The reinforcement learning unit 146 then derives the following: i (t o ) from state s i (t n The reinforcement learning unit 146 learns the behavioral model so that the reward for transitioning to ) is maximized. By repeating this learning process, the reinforcement learning unit 146 learns the log-likelihood ΣlogP(S i We can construct an action model that maximizes (t).
[0047] In step S240 of Figure 8, the control input derivation unit 148 uses the reinforcement-learned behavior model to obtain the data x, which is the detection result of the sensor 120. i (t o) or state s i (t o From this, we derive the maintenance input m that should be executed.
[0048] Here, reinforcement learning makes it possible to appropriately derive the maintenance input m of the controlled system.
[0049] <Effects of Control System 100> The effects of the control system 100 according to an embodiment of the present invention will be described.
[0050] In the control system 100, the information acquisition unit 140 acquires detection results from multiple sensors 120, and the conversion function derivation unit 142 uses Normalizing Flow to derive a conversion function f that converts the detection results from the multiple sensors 120 into states on a unimodal state space (normalized state space S). This allows the system to generate a normalized state space S that is normalized compared to the total data X. Furthermore, since the equilibrium point in the normalized state space S can be defined as the standard state, the control objective can be easily assumed.
[0051] Furthermore, in the control system 100, the evaluation unit 144 evaluates a predetermined control input (maintenance input m) based on the relative relationship between the pre-state, which is obtained by transforming the detection result of the sensor 120 before applying a predetermined control input (maintenance input m) using a transformation function f, and the post-state, which is obtained by transforming the detection result of the sensor after applying a predetermined control input using the transformation function f. Here, since the normalized state space S is a unimodal space with the standard state as the equilibrium point, it is possible to determine whether or not the control input has brought the state closer to the standard state using a simple criterion such as whether or not it is moving towards the equilibrium point.
[0052] Furthermore, in the control system 100, the reinforcement learning unit 146 performs reinforcement learning to develop an action model that derives control inputs based on the detection results of the sensor 120, and the control input derivation unit 148 derives control inputs using the reinforcement-learned action model. Through such reinforcement learning, it becomes possible to appropriately derive the control input of the controlled object.
[0053] Preferred embodiments of the present invention have been described above with reference to the attached drawings, but it goes without saying that the present invention is not limited to these embodiments. It will be obvious to those skilled in the art that various modifications or alterations can be conceived within the scope of the claims, and these will naturally also fall within the technical scope of the present invention.
[0054] For example, the processes described using flowcharts in this specification do not necessarily have to be executed in the order shown in the flowcharts. Some processing steps may be executed in parallel. Additional processing steps may be adopted, and some processing steps may be omitted.
[0055] Furthermore, for example, the series of control processes performed by the control system 100 and the control device 130 described above may be implemented using software, hardware, or a combination of software and hardware. The programs constituting the software are pre-stored in a storage medium provided inside or outside the information processing device, for example. [Explanation of symbols]
[0056] 100 control systems 110 Controlled object 120 sensors 130 Control device 140 Information Acquisition Department 142 Derivation of the Transformation Function 144 Evaluation Department 146 Reinforcement Learning Department 148 Control Input Derivation Section
Claims
1. A sensor (120) attached to the controlled object (110), A control device (130) including a processor and memory, Equipped with, The processor cooperates with the program contained in the memory, An information acquisition unit (140) that acquires the detection result of the sensor (120), A conversion function derivation unit (142) derives a conversion function that converts the detection result of the sensor (120) into a state in a unimodal state space, It functions, The aforementioned unimodal state space is a control system represented by a Lyapunov function.
2. The aforementioned processor, The control system according to claim 1, which functions as an evaluation unit (144) that evaluates the predetermined control input based on the relative relationship between a pre-state obtained by converting the detection result of the sensor (120) before a predetermined control input is applied using the conversion function, and a post-state obtained by converting the detection result of the sensor (120) after a predetermined control input is applied using the conversion function.
3. The aforementioned processor, A reinforcement learning unit (146) performs reinforcement learning to develop an action model that derives a control input based on the detection results of the sensor (120), A control input derivation unit (148) that derives a control input using the reinforcement-learned behavior model, The control system according to claim 1, which functions as follows.
4. A control device (130) including a processor and memory, The processor cooperates with the program contained in the memory, An information acquisition unit (140) that acquires the detection result of a sensor (120) attached to the controlled object (110), A conversion function derivation unit (142) derives a conversion function that converts the detection result of the sensor (120) into a state in a unimodal state space, It functions, The aforementioned unimodal state space is a control device represented by a Lyapunov function.
5. In a control device (130) including a processor and memory, the processor cooperates with a program contained in the memory, The detection result of the sensor (120) attached to the controlled object (110) is obtained (S110), A transformation function is derived (S120) that converts the detection result of the sensor (120) into a state in a unimodal state space. The aforementioned monomodal state space is a control method represented by a Lyapunov function.
Citation Information
Patent Citations
A machine learnable system with normalizing flow
EP3767533A1
Soil mixing and improving machine
JP2000328598A
Optimization method, control device, and robot
JP2019113985A