Determination device, determination method, and recording medium having determination program recorded thereon
By obtaining device status and operation quantity data, using machine learning to generate control models and simulate device status, the control accuracy and stability problems of robots under motor interference are solved, and more reliable AI control is achieved.
Patent Information
- Application Number
- CN202210190756.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-03
- Filing Date
- 2022-02-28
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2042-02-28
AI Technical Summary
In the prior art, it is difficult for robots to effectively correct the teaching position when facing motor interference, resulting in insufficient control accuracy and stability.
By obtaining device status and operation quantity data, using machine learning to generate a control model, simulate the device status under the output of the model, and determine whether AI control can be performed based on the simulation results, including reinforcement learning and the use of simulation models.
It improves the control accuracy and stability of the robot when facing motor interference, avoids equipment abnormalities caused by AI control, and ensures the normal operation of the equipment.
Smart Images

Figure CN115032950B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a determination device, a determination method, and a recording medium having a determination program recorded thereon. Background Art
[0002] Patent document 1 states that "machine learning is performed on the correction amount of the robot's teaching position in response to the interference applied to the motors driving the joints of the robot, and based on the results of the machine learning, the robot is controlled while correcting the teaching position in a manner that suppresses the interference as it moves toward the teaching position."
[0003] Patent Document 1: Japanese Patent Application Laid-Open No. 2018-202564 Summary of the Invention
[0004] (Item 1)
[0005] In the first aspect of the present invention, a determination device is provided. The determination device may include a state data acquisition unit that acquires state data representing the state of a device provided with a control object. The determination device may include an operation quantity data acquisition unit that acquires operation quantity data representing the operation quantity of the above-mentioned control object. The determination device may include a control model generation unit that uses the above-mentioned state data and the above-mentioned operation quantity data to generate a control model that outputs the above-mentioned operation quantity corresponding to the state of the above-mentioned device through machine learning. The determination device may include a simulation unit that uses a simulation model to simulate the state of the above-mentioned device when the above-mentioned operation quantity output by the above-mentioned control model is given to the above-mentioned control object. The determination device may include a determination unit that determines whether the above-mentioned control model can control the above-mentioned control object based on the simulation result.
[0006] (Item 2)
[0007] The determination unit may determine that the control model can control the control target when the period during which the device is determined to be able to operate normally exceeds a predetermined threshold value based on the simulation result.
[0008] (Item 3)
[0009] The determination unit may determine that the control model can control the control target when the number of times the device is determined to be able to operate normally exceeds a predetermined threshold based on the simulation result.
[0010] (Item 4)
[0011] The determination device may further include an output unit that outputs the simulation result, and the determination unit may determine that control of the control object by the control model is possible when an instruction to permit control is obtained in response to the output of the simulation result.
[0012] (Item 5)
[0013] The determination device may further include an instruction unit configured to instruct the start of control of the controlled object based on the control model when determining that the controlled object can be controlled by the control model.
[0014] (Item 6)
[0015] The control model generation unit may regenerate the control model through the machine learning when it is determined that the control model cannot control the control object.
[0016] (Item 7)
[0017] The determination device may further include an end determination unit for determining the end of the machine learning, and the simulation unit may simulate the state of the device when it is determined that the machine learning has ended.
[0018] (Item 8)
[0019] The termination determination unit may determine the termination of the machine learning based on an elapsed time from the start of the machine learning.
[0020] (Item 9)
[0021] The termination determination unit may determine the termination of the machine learning based on a value of an evaluation function of the machine learning.
[0022] (Item 10)
[0023] The control model generation unit may generate the control model by performing reinforcement learning in response to the input of the state data so as to output a more recommended operation amount for an operation amount having a higher reward value defined by a predetermined reward function.
[0024] (Item 11)
[0025] In a second aspect of the present invention, a determination method is provided. The determination method may include the following steps: obtaining status data indicating the status of a device provided with a control object. The determination method may include the following steps: obtaining operation quantity data indicating the operation quantity of the above-mentioned control object. The determination method may include the following steps: using the above-mentioned status data and the above-mentioned operation quantity data, through machine learning, to generate a control model that outputs the above-mentioned operation quantity corresponding to the status of the above-mentioned device. The determination method may include the following steps: using a simulation model to simulate the status of the above-mentioned device when the above-mentioned operation quantity output by the above-mentioned control model is given to the above-mentioned control object. The determination method may include the following steps: determining whether the above-mentioned control model can control the above-mentioned control object based on the simulation result.
[0026] (Item 12)
[0027] In a third aspect of the present invention, a recording medium is provided, which records a determination program. The determination program can be executed by a computer. The determination program can cause the computer to function as a state data acquisition unit that acquires state data representing the state of a device provided with a control object. The determination program can cause the computer to function as an operation quantity data acquisition unit that acquires operation quantity data representing the operation quantity of the control object. The determination program can cause the computer to function as a control model generation unit that uses the state data and the operation quantity data to generate a control model that outputs the operation quantity corresponding to the state of the device through machine learning. The determination program can cause the computer to function as a simulation unit that uses a simulation model to simulate the state of the device when the operation quantity output by the control model is given to the control object. The determination program can cause the computer to function as a determination unit that determines whether the control model can control the control object based on the simulation result.
[0028] The above summary of the invention does not list all the features of the present invention. In addition, sub-components of the above-mentioned feature groups can also constitute the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 An example of a block diagram of the device 10 in which the control target 20 is installed is shown together with the determination device 100 according to this embodiment.
[0030] Figure 2 An example of a flow in which the determination device 100 according to the present embodiment determines whether AI control is possible is shown.
[0031] Figure 3An example of a block diagram of a device 10 in which a control target 20 is installed is shown together with a determination device 100 according to a modification of the present embodiment.
[0032] Figure 4 An example of a flow in which the determination device 100 according to the modification of the present embodiment determines whether AI control is possible is shown.
[0033] Figure 5 An example of a block diagram of a device 10 provided with a control target 20 is shown together with a determination device 100 according to another modified example of the present embodiment.
[0034] Figure 6 An example of a flow in which the determination device 100 according to another modified example of the present embodiment determines whether AI control is possible is shown.
[0035] Figure 7 This figure shows an example of a computer 9900 that can embody all or part of the various embodiments of the present invention. DETAILED DESCRIPTION
[0036] The present invention will be described below by way of embodiments of the invention, but the following embodiments do not limit the invention as defined in the claims. Furthermore, all combinations of features described in the embodiments are not necessarily essential for the solution to the problem.
[0037] Figure 1 An example of a block diagram of a device 10 equipped with a control target 20 is shown together with a determination device 100 according to this embodiment. Before starting control of the control target 20 based on a learning model generated through machine learning (also known as AI control), the determination device 100 according to this embodiment simulates the state of the device 10 when the output of the learning model is applied to the control target 20. Based on the simulation results, the determination device 100 according to this embodiment determines whether AI control is possible.
[0038] Equipment 10 is a facility, device, or the like equipped with a controlled object 20. For example, equipment 10 may be a factory or a complex device composed of multiple instruments. Examples of factories include, in addition to chemical and biological industrial plants, plants that manage and control the wellheads of gas and oil fields and their surrounding areas; plants that manage and control hydropower, thermal power, and nuclear power generation; plants that manage and control environmental resource generation such as solar and wind power; and plants that manage and control water supply and drainage systems, dams, and the like.
[0039] The device 10 is provided with a control target 20. In this figure, as an example, only one control target 20 is provided for the device 10, but the present invention is not limited thereto. A plurality of control targets 20 may be provided for the device 10.
[0040] Furthermore, the device 10 may be provided with one or more sensors (not shown) for measuring various states (physical quantities) inside and outside the device 10. Such sensors measure, for example, operating data, consumption data, and external environment data.
[0041] Here, the operating data represents the operating state resulting from controlling the controlled object 20. For example, the operating data may represent a measured value PV (Process Variable) obtained by measuring the controlled object 20. For example, the operating data may represent the output (controlled variable) of the controlled object 20, or various values that vary depending on the output of the controlled object 20.
[0042] The consumption data indicates the consumption of at least one of energy and raw materials by the equipment 10. For example, the consumption data may indicate the consumption of electricity or fuel (for example, LPG: Liquefied Petroleum Gas) as energy consumption.
[0043] External environment data represents physical quantities that can act as interference with the control of control object 20. For example, external environment data may represent the temperature, humidity, sunshine, wind direction, wind volume, precipitation, and other physical quantities of the atmosphere outside device 10 that vary with the control of other devices installed in device 10.
[0044] Controlled object 20 is an instrument or device that is controlled. For example, controlled object 20 may be an actuator such as a valve, pump, heater, fan, motor, or switch that controls at least one physical quantity of the process in equipment 10, such as pressure, temperature, pH, velocity, or flow rate. It receives a manipulated variable (MV) as input and outputs a controlled variable.
[0045] Furthermore, the controlled object 20 can be switched between feedback control based on a manipulated variable MV (FB) assigned from a controller (not shown) and AI control based on a manipulated variable MV (AI) assigned from a control model. This FB control can be, for example, at least one of proportional control (P control), integral control (I control), or differential control (D control), and, as an example, PID control. Furthermore, the controller can be integrally formed as part of the determination device 100 of this embodiment, or it can be configured as a separate structure independent of the determination device 100.
[0046] For example, before starting AI control of the controlled object 20 (e.g., switching from FB control to AI control and starting AI control), the determination device 100 according to this embodiment simulates the state of the device 10 when the output of the learning model is applied to the controlled object 20. Based on the simulation results, the determination device 100 according to this embodiment determines whether AI control is possible.
[0047] The determination device 100 can be a computer such as a PC (personal computer), a tablet computer, a smart phone, a workstation, a server computer, or a general-purpose computer, or a computer system composed of multiple computers connected together. This computer system is also a computer in a broad sense. In addition, the determination device 100 can be installed in a computer using one or more executable virtual computer environments. Alternatively, the determination device 100 can be a dedicated computer designed to determine whether control is possible, or it can be dedicated hardware implemented using dedicated circuits. In addition, if the determination device 100 can be connected to the Internet, the determination device 100 can be implemented through cloud computing.
[0048] The determination device 100 includes a state data acquisition unit 110, an operation variable data acquisition unit 120, a control model generation unit 130, a control model 135, a simulation unit 140, a simulation model 145, a determination unit 150, and an instruction unit 160. Furthermore, the above modules may be functionally separate modules and may not correspond to the actual device structure. That is, although shown as a single module in this figure, they do not need to be composed of a single device. Furthermore, although shown as separate modules in this figure, they do not need to be composed of separate devices.
[0049] The status data acquisition unit 110 acquires status data representing the status of the device 10 in which the control object 20 is installed. For example, the status data acquisition unit 110 acquires operating data, consumption data, external environment data, etc. measured by sensors installed in the device 10 via a network. However, this is not limited to this. The status data acquisition unit 110 can acquire this status data from an operator or from various memory devices. The status data acquisition unit 110 supplies the acquired status data to the control model generation unit 130. In addition, the status data acquisition unit 110 supplies the acquired status data to the control model 135.
[0050] The manipulated variable data acquisition unit 120 acquires manipulated variable data representing the manipulated variable of the controlled object 20. For example, the manipulated variable data acquisition unit 120 acquires data representing the manipulated variable MV(FB) assigned to the controlled object 20 by the controller (not shown) via a network. However, this is not a limitation. The manipulated variable data acquisition unit 120 may acquire such manipulated variable data from an operator or from various memory devices. The manipulated variable data acquisition unit 120 supplies the acquired manipulated variable data to the control model generation unit 130.
[0051] The control model generation unit 130 uses state data and manipulated variable data through machine learning to generate a control model 135 that outputs manipulated variables corresponding to the state of the device 10. For example, the control model generation unit 130 performs reinforcement learning using the state data supplied by the state data acquisition unit 110 and the data representing the manipulated variable MV(FB) supplied by the manipulated variable data acquisition unit 120 as learning data, thereby generating a control model 135 that outputs the manipulated variable MV(AI) corresponding to the state of the device 10. Specifically, the control model generation unit 130 generates the control model 135 by performing reinforcement learning based on the input state data, such that the manipulated variable with a higher reward value specified by a predetermined reward function is output as a more recommended manipulated variable. This will be described in detail later.
[0052] The control model 135 is a learning model generated by the control model generation unit 130 through reinforcement learning. It outputs the manipulated variable MV (AI) corresponding to the state of the device 10. For example, the control model 135 inputs the state data supplied by the state data acquisition unit 110 and outputs the recommended manipulated variable MV (AI) to be assigned to the controlled object 20 based on the state of the device 10. The control model 135 supplies the output manipulated variable MV (AI) to the controlled object 20. Furthermore, the control model 135 supplies the output manipulated variable MV (AI) to the simulation unit 140. While this figure illustrates an example in which the control model 135 is built into the determination device 100, this is not limiting. The control model 135 may be stored on a device separate from the determination device 100 (e.g., a cloud server). While this figure illustrates an example in which the control model 135 outputs the manipulated variable MV (AI), the output manipulated variable MV (AI) is always supplied to the controlled object 20, this is not limiting. The control model 135 may supply the output manipulated variable MV(AI) to the controlled object 20 only when the result of determination described later indicates that AI control is possible.
[0053] The simulation unit 140 uses a simulation model 145 to simulate the state of the device 10 when the manipulated variable MV (AI) output by the control model 135 is applied to the controlled object 20. The term "simulation" here encompasses not only the simulation unit 140 itself simulating the state of the device 10, but also the simulation unit 140 causing another device (e.g., a simulator (not shown)) to simulate the state of the device 10, and obtaining the state of the device 10 simulated by another device from another device. For example, the simulation unit 140 inputs the manipulated variable MV (AI) output by the control model 135 into the simulation model 145 and obtains multiple output values from the simulation model 145 as simulation results. The simulation unit 140 supplies the obtained simulation results to the determination unit 150.
[0054] The simulation model 145 is a model (e.g., a plant model) constructed in a manner that simulates the operation of the device 10. For example, the simulation model 145 inputs the manipulated variable MV (AI) and simulates the operation of the device 10 when the manipulated variable MV (AI) is given to the control object 20. Moreover, the simulation model 145 outputs a plurality of output values representing the state of the simulated device 10. As an example, the simulation model 145 can be a simple physical model or a low-dimensional linear model that can simulate the state of the device 10 with a cycle that is the same as or shorter than the control cycle of the device 10 and has a lighter processing load. In addition, in this figure, as an example, a case where the simulation model 145 is built into the determination device 100 is shown, but it is not limited to this. Similar to the control model 135, the simulation model 145 can be stored in a device different from the determination device 100 (e.g., a cloud server). In addition, the above-mentioned simulator can be equipped with a device different from the determination device 100.
[0055] Based on the simulation results, the determination unit 150 determines whether the control model 135 can control the controlled object 20. For example, the determination unit 150 determines whether the simulation results supplied from the simulation unit 140 satisfy predetermined conditions (e.g., abnormality diagnosis conditions). If the simulation results do not satisfy the predetermined conditions, the determination unit 150 determines that the control model 135 can control the controlled object 20. If the simulation results satisfy the predetermined conditions, the determination unit 150 determines that the control model 135 cannot control the controlled object 20. The determination unit 150 supplies the determination results to the instruction unit 160.
[0056] If the control model 135 determines that the controlled object 20 can be controlled, the instruction unit 160 instructs the controlled object 20 to begin control based on the control model 135. In this case, the instruction unit 160 can directly output the instruction to the controlled object 20. This allows, for example, the controlled object 20 to switch from FB control based on the manipulated variable MV(FB) supplied by the controller to AI control based on the manipulated variable MV(AI) supplied by the control model 135, thereby initiating AI control. Alternatively, if another controller, such as a PID controller, is integrated as part of the determination device 100, the instruction unit 160 can issue an instruction to a switch that switches between outputting the manipulated variable MV(FB) and the manipulated variable MV(AI) to the controlled object 20. This allows the manipulated variable MV output from the determination device 100 to the controlled object 20 to be switched from the manipulated variable MV(FB) to the manipulated variable MV(AI), thereby initiating AI control of the controlled object 20.
[0057] Figure 2 An example of a flow in which the determination device 100 according to the present embodiment determines whether AI control is possible is shown.
[0058] In step 210, the determination device 100 acquires status data. For example, the status data acquisition unit 110 acquires status data indicating the status of the device 10 in which the control target 20 is installed. As an example, the status data acquisition unit 110 acquires operating data, consumption data, and external environment data measured by sensors installed in the device 10 via a network as status data. The status data acquisition unit 110 supplies the acquired status data to the control model generation unit 130 and the control model 135.
[0059] In step 220, the determination device 100 acquires the operation amount data. For example, the operation amount data acquisition unit 120 acquires the operation amount data representing the operation amount of the control object 20. As an example, the operation amount data acquisition unit 120 acquires the operation amount MV (FB) assigned from the controller to the control object 20 when the control object 20 is subjected to FB control from the controller via the network. The operation amount data acquisition unit 120 supplies the acquired operation amount data to the control model generation unit 130. In addition, in this figure, as an example, a case where the determination device 100 acquires the operation amount data after acquiring the state data is shown, but it is not limited to this. The determination device 100 may acquire the state data after acquiring the operation amount data, or may acquire the state data and the operation amount data at the same time.
[0060] In step 230, the determination device 100 generates a control model 135. For example, the control model generation unit 130 uses the state data and the manipulated variable data to generate the control model 135 through machine learning, thereby outputting a manipulated variable corresponding to the state of the device 10. As an example, the control model generation unit 130 uses the state data acquired in step 210 and the data representing the manipulated variable MV(FB) acquired in step 220 as learning data to perform reinforcement learning to generate the control model 135 that outputs the manipulated variable MV(AI) corresponding to the state of the device 10.
[0061] Typically, if an agent observes the state of the environment and chooses an action, the environment changes based on that action. In reinforcement learning, by giving a certain reward along with changes in the environment, the agent learns to choose a better action (willingness decision). In contrast to learning with a teacher, which can obtain completely correct answers, reinforcement learning gives rewards as discontinuous values based on changes in part of the environment. Therefore, the agent learns to choose the action that will maximize the total reward in the future. In this way, in reinforcement learning, the agent learns appropriate actions based on the interaction of learning actions and taking actions on the environment, that is, learning actions to maximize the reward received in the future.
[0062] In this embodiment, the reward for reinforcement learning can be an indicator used to evaluate the operation of the device 10, or a value determined by a predetermined reward function. Here, a function refers to a mapping with a rule that establishes a one-to-one correspondence between elements of another set and elements of a particular set. For example, it can be a mathematical formula or a table.
[0063] The reward function outputs a value (reward value) that evaluates the state of the device 10 represented by the state data based on the input of the state data. As described above, for example, the state data includes the measured value PV measured for the control object 20. Therefore, the reward function can be defined as a function in which the closer the measured value PV is to the target value SV (Setting Variable), the higher the reward value. Here, the evaluation function can be defined as a function whose variable is the absolute value of the difference between the measured value PV and the target value SV. That is, as an example, when the control object 20 is a valve, the evaluation function can be a function whose variable is the absolute value of the difference between the valve opening actually measured by the sensor, that is, the measured value PV, and the valve opening set as the target, that is, the target value SV. Furthermore, the reward function can be a function whose variable is the value of the evaluation function obtained by this evaluation function.
[0064] Furthermore, as described above, in addition to the measured value PV, the state data also includes, for example, various values that change depending on the output of the controlled object 20, consumption data, external environmental data, and the like. Therefore, the reward function may be a function that increases or decreases the reward value based on such various values, consumption data, and external environmental data. As an example, when constraints are imposed on such various values or consumption data that must be observed, the reward function may be a function that minimizes the reward value when, with reference to the external environmental data, such various values or consumption data do not meet the constraints. Furthermore, when targets are imposed on such various values or consumption data that must be achieved, the reward function may be a function that increases the reward value as such various values or consumption data approach the targets, and decreases the reward value as such various values or consumption data move further away from the targets, with reference to the external environmental data.
[0065] The control model generation unit 130 obtains reward values for various learning data based on this reward function. Furthermore, the control model generation unit 130 performs reinforcement learning using each set of learning data and reward values. In this case, the control model generation unit 130 can perform learning processing based on well-known methods such as the steepest descent method, neural networks, DQN (Deep Q-Network), Gaussian processes, and deep learning. Furthermore, the control model generation unit 130 performs learning in such a way that the higher the reward value, the more recommended the operation amount is output preferentially. Specifically, the control model generation unit 130 performs reinforcement learning based on the input state data in such a way that the higher the reward value, as determined by a predetermined reward function, the more recommended the operation amount is output, thereby generating the control model 135. This updates the model and generates the control model 135.
[0066] In step 240, the determination device 100 performs a simulation. For example, the simulation unit 140 uses the simulation model 145 to simulate the state of the device 10 when the manipulated variable MV (AI) output by the control model 135 is applied to the controlled object 20. As an example, the simulation unit 140 inputs the manipulated variable MV (AI) output by the control model 135 generated in step 230 into the simulation model 145 and obtains multiple output values from the simulation model as simulation results. The simulation unit 140 supplies the obtained simulation results to the determination unit 150.
[0067] In step 250, the determination device 100 determines whether AI control can be performed. For example, based on the simulation results, the determination unit 150 determines whether the control model 135 can control the controlled object 20. As an example, the determination unit 150 determines whether the simulation results in step 240 meet predetermined conditions. For example, the determination unit 150 may pre-store abnormality diagnosis conditions for diagnosing abnormalities in the device 10. Furthermore, if none of the multiple output values output by the simulation model 145 meet the abnormality diagnosis conditions, the determination unit 150 may infer that the device 10 is operating normally. Alternatively, if at least one of the multiple output values output by the simulation model 145 meets the abnormality diagnosis conditions, the determination unit 150 may infer that the device 10 is not operating normally (an abnormality has occurred in the device). Furthermore, if the period of time during which the device 10 is determined to be operating normally based on the simulation results exceeds a predetermined threshold, the determination unit 150 may determine that the control model 135 can control the controlled object 20. In other words, if the determination unit 150 determines that the device 10 is operating normally for more than a predetermined period P, the determination unit 150 may determine that AI control is possible. In addition, the determination unit 150 may determine that the control model 135 can control the control object 20 when it is determined based on the simulation results that the number of times the device 10 can operate normally exceeds a predetermined threshold. That is, the determination unit 150 may determine that AI control can be performed when it is determined that the number of times the device 10 can operate normally exceeds a predetermined number, that is, N times. In addition, the determination unit 150 may use period-based determination and number-based determination at the same time. For example, the determination unit 150 may determine that AI control can be performed when it is determined that the period of normal operation exceeds period P and the number of times the device can operate normally exceeds N times. In addition, the determination unit 150 may determine that AI control can be performed when it is determined that the number of times the device 10 can operate normally exceeds period P exceeds N times. The determination unit 150 supplies the determination result to the indication unit 160.
[0068] If it is determined in step 250 that AI control is not possible (No), the determination device 100 returns the process to step 210 to continue the flow. In other words, if it is determined that the control model 135 cannot control the controlled object 20, the control model generation unit 130 regenerates the control model 135 through machine learning.
[0069] If it is determined in step 250 that AI control is possible (Yes), the determination device 100 advances the process to step 260, instructing the control object 20 to start AI control. For example, if it is determined that the control model 135 can control the control object 20, the instruction unit 160 instructs the control object 20 to start control based on the control model 135. Thus, for example, the control object 20 switches from FB control based on the manipulated variable MV(FB) supplied from the controller to AI control based on the manipulated variable MV(AI) supplied from the control model 135, and AI control begins.
[0070] Typically, in machine learning, input data is used to determine the parameters of the learning model. These parameters are subject to probability and cannot be theoretically guaranteed. Therefore, the learning model may output abnormal inference data. Therefore, before commencing AI control, the determination device 100 in this embodiment uses a simulation model 145 to simulate the state of the device 10 when the manipulated variable MV(AI) output by the control model 135 is applied to the controlled object 20. Furthermore, the determination device 100 determines whether AI control is feasible based on the simulation results. Thus, the determination device 100 in this embodiment can prevent the device 10 from performing abnormal operations associated with AI control after the actual machine is put into AI control, i.e., after the device 10 begins operating under AI control. Alternatively, the determination of whether AI control is feasible may be based on whether the manipulated variable MV(AI) output by the control model 135 meets a predetermined benchmark. However, such benchmarks are artificially assigned based on experience, and even if the manipulated variable MV(AI) meets such benchmarks, it is not certain that the device 10 will not exhibit abnormalities. Similarly, even if the manipulated variable MV(AI) does not meet this criterion, the device 10 may not necessarily be malfunctioning. In contrast, the determination device 100 according to this embodiment determines whether AI control can be performed based not on the manipulated variable MV(AI) itself but on the results of simulating the state of the device 10 when the manipulated variable MV(AI) is applied to the controlled object 20. Therefore, it is possible to determine whether to stop AI control based on a more accurate basis than actual operation.
[0071] Furthermore, if the period during which the device 10 is determined to be capable of normal operation based on the simulation results exceeds a threshold, or if the number of times it is determined to be capable of normal operation exceeds a threshold, the determination device 100 according to this embodiment determines that AI control is possible. Thus, the determination device 100 according to this embodiment determines that AI control is possible after temporarily observing that normal operation is possible. Therefore, when the manipulated variable MV(AI) is assigned to the controlled object 20, it is possible to avoid erroneously determining that AI control is possible even if it is accidentally inferred that the device 10 is not abnormal.
[0072] Furthermore, if it is determined that AI control is possible, the determination device 100 according to this embodiment instructs the controlled object 20 to switch to control based on the control model 135. Thus, according to the determination device 100 according to this embodiment, it is possible to instruct the controlled object 20 to start AI control using the determination based on the simulation results as a trigger condition. Furthermore, if it is determined that AI control is not possible, the determination device 100 according to this embodiment regenerates the control model 135 through machine learning. Thus, according to the determination device 100 according to this embodiment, even if it is temporarily determined that AI control is not possible, it is possible to regenerate the control model through relearning and repeatedly determine whether AI control can be performed using the regenerated control model 135.
[0073] Figure 3 An example of a block diagram of a device 10 provided with a control target 20 is shown together with a determination device 100 according to a modification of this embodiment. Figure 3 In the Figure 1 Components with the same functions and structures are designated by the same reference numerals, and description thereof will be omitted except for the following differences. In the determination device 100 of the above embodiment, an example is shown in which the determination of whether AI control is possible is automatically determined based on simulation results. However, in the determination device 100 of this modified example, the simulation results are output, and the determination of whether AI control is possible is made based on permission from an operator, etc., who has reviewed the simulation results. The determination device 100 of this modified example further includes an output unit 310 and an input unit 320.
[0074] In the determination device 100 according to this variation, the simulation unit 140 supplies the simulation results to the output unit 310 in addition to the determination unit 150. Furthermore, the output unit 310 outputs the simulation results. For example, the output unit 310 may output the simulation results by displaying them on a monitor, printing them, or transmitting the data to another device.
[0075] The input unit 320 receives user input from an operator who has examined the simulation results, etc., in response to the output of the simulation results. The input unit 320 supplies the instruction input by the user from the operator to the determination unit 150 .
[0076] If the instruction supplied from the input unit 320 indicates that AI control is permitted, the determination unit 150 determines that the control model 135 can control the controlled object 20. In other words, if an instruction indicating that control is permitted is obtained in accordance with the output of the simulation results, the determination unit 150 determines that the control model 135 can control the controlled object 20.
[0077] Figure 4 An example of a flow in which the determination device 100 according to a modification of this embodiment determines whether AI control is possible is shown. Figure 4 In, with Figure 2 The same processing is denoted by the same reference numerals, and description thereof will be omitted except for the following differences: In this flow, step 250 is replaced with steps 410 and 420 .
[0078] In step 410, the determination device 100 outputs the simulation result. For example, the output unit 310 obtains the simulation result of the simulation unit 140 in step 240 and outputs the simulation result by displaying it on a monitor.
[0079] In step 420, the determination device 100 determines whether AI control is permitted. For example, the determination unit 150 determines whether an instruction to permit AI control is obtained via the input unit 320 by an operator who has studied the simulation results. In the case where an instruction to permit AI control is not obtained in step 420 (in the case of No), the determination unit 150 determines that AI control cannot be performed. Furthermore, the determination device 100 returns the processing to step 210 and continues the process. In the case where an instruction to permit AI control is obtained in step 420 (in the case of Yes), the determination unit 150 determines that AI control can be performed. Furthermore, the determination device 100 enters the processing into step 260. That is, in the case where an instruction to permit control is obtained in accordance with the output of the simulation results (in the case of Yes), the determination unit 150 determines that the control model 135 can control the control object 20.
[0080] Thus, the determination device 100 according to this variation outputs simulation results and determines whether AI control can be performed based on instructions from an operator who has studied the simulation results. Thus, the determination device 100 according to this variation can reflect the operator's intentions when AI control is implemented in the actual machine.
[0081] In addition, in the above description, the case where the determination device 100 performs steps 410 and 420 instead of step 250 is shown as an example, but it is not limited to this. On the basis of step 250, the determination device 100 involved in this modification can perform the processing of steps 410 and 420. At this time, the determination device 100 can determine that AI control can be performed when at least any one of the permission instructions of the operator, etc. and the automatic judgment based on the computer (for example, the judgment based on the period and number of times of normal operation) is met. Alternatively, the determination device 100 can determine that AI control can be performed only when both the permission instructions of the operator, etc. and the automatic judgment based on the computer are met. Thus, according to the determination device 100 involved in this modification, the actual machine investment of AI control can be more carefully determined by simultaneously utilizing the automatic judgment based on the computer and the manual judgment based on the operator.
[0082] Figure 5 An example of a block diagram of a device 10 provided with a control target 20 is shown together with a determination device 100 according to another modified example of the present embodiment. Figure 5 In the Figure 1 Components with the same functions and structures are designated by the same reference numerals, and description thereof will be omitted except for the following differences. In the determination device 100 according to the above embodiment, as an example, when the control model 135 outputs the manipulated variable MV (AI), the state of the device 10 is always simulated when the manipulated variable MV (AI) is applied to the controlled object 20. However, in the determination device 100 according to this other variation, a trigger condition for simulating the state of the device 10 is provided based on the progress of the machine learning used to generate the control model 135. The determination device 100 according to this other variation further includes an end determination unit 510.
[0083] The end determination unit 510 monitors the progress of the machine learning performed by the control model generation unit 130 to generate the control model 135. Furthermore, the end determination unit 510 determines the end of the machine learning performed to generate the control model 135. If the end determination unit 510 determines that the machine learning has ended, it instructs the simulation unit 140 to simulate the state of the device 10. In response, the simulation unit 140 simulates the state of the device 10. That is, if the simulation unit 140 determines that the machine learning has ended, it simulates the state of the device 10.
[0084] Figure 6 FIG. 1 shows an example of a process for determining whether AI control can be performed by the determination device 100 according to another modified example of the present embodiment. Figure 6 In, with Figure 2The same processes are marked with the same reference numerals, and description thereof will be omitted except for the following differences.
[0085] In step 610, the determination device 100 determines the end of machine learning. For example, the end determination unit 510 monitors the progress of the machine learning used by the control model generation unit 130 to generate the control model 135, and determines the end of the machine learning used to generate the control model 135. At this time, the end determination unit 510 can determine the end of machine learning based on the time that has passed since the start of machine learning. Alternatively or in addition to this, the end determination unit 510 can determine the end of machine learning based on the value of the evaluation function of machine learning. For example, the end determination unit 510 can determine that machine learning has ended when at least any one of the minimum value, maximum value, average value or median value of the evaluation function, which is a function with the absolute value of the difference between the measured value PV and the target value SV as a variable, is lower than a predetermined threshold.
[0086] If it is determined in step 610 that machine learning has not yet been completed (if the answer is No), the determination device 100 returns the process to step 210 to continue the flow. If it is determined in step 610 that machine learning has been completed (if the answer is Yes), the determination device 100 proceeds to step 240. In other words, the simulation unit 140 simulates the state of the device 10 when it is determined that machine learning has been completed.
[0087] In this manner, the determination device 100 according to this other variation provides a trigger condition for simulating the state of the device 10 based on the progress of machine learning for generating the control model 135. Thus, the determination device 100 according to this other variation avoids simulating the state of the device 10 until machine learning is complete, thereby reducing the processing load on the determination device 100. Furthermore, the determination device 100 according to this other variation simulates the state of the device 10 after learning is complete, thereby improving the reliability of the simulation results used to determine whether AI control is possible.
[0088] At this time, the determination device 100 involved in this other variant example determines the end of machine learning based on, for example, the elapsed time from the start of machine learning. Thus, according to the determination device 100 involved in this other variant example, the end of machine learning can be tentatively determined based on the elapsed time. In addition, the determination device 100 involved in this other variant example determines the end of machine learning based on, for example, the value function of machine learning. Thus, according to the determination device 100 involved in this other variant example, the end of machine learning can be determined based on an objective value. In addition, the determination device 100 involved in this other variant example can use both the judgment based on the above-mentioned elapsed time and the judgment based on the value function when determining the end of machine learning. Thus, according to the determination device 100 involved in this other variant example, the simulation is triggered only when machine learning is carried out for a long time and the result of the machine learning meets a predetermined benchmark, thereby further reducing the processing load of the determination device 100.
[0089] Various embodiments of the present invention may be described with reference to flowcharts and block diagrams, where a module may represent (1) a stage of a process for performing an operation, or (2) a portion of a device having the function of performing an operation. Specific stages and portions may be implemented using dedicated circuits, programmable circuits supplied with computer-readable instructions stored on a computer-readable medium, and / or processors supplied with computer-readable instructions stored on a computer-readable medium. Dedicated circuits may include digital and / or analog hardware circuits, as well as integrated circuits (ICs) and / or discrete circuits. Programmable circuits may include reconfigurable hardware circuits such as logical AND, logical OR, logical XOR, logical NAND, logical NOR, and other logical operations, trigger circuits, registers, memory elements such as field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), and the like.
[0090] The computer-readable medium may include any tangible device capable of storing instructions for execution using an appropriate device. As a result, the computer-readable medium having the instructions stored therein has a product containing executable instructions for making a means for performing the operations specified in the flowchart or block diagram. Examples of computer-readable media include electronic storage media, magnetic storage media, optical storage media, electromagnetic storage media, semiconductor storage media, etc. More specific examples of computer-readable media include Floppy (registered trademark) floppy disks, floppy disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), electrically erasable programmable read-only memory (EEPROM), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disk (DVD), Blu-ray (RTM) disks, memory sticks, integrated circuit cards, etc.
[0091] Computer readable instructions may include any source code or object code described in assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or any combination of one or more programming languages including object-oriented programming languages such as Smalltalk, JAVA (registered trademark), C++, and procedural programming languages such as the "C" programming language or similar programming languages.
[0092] Computer-readable instructions may be provided to a processor or programmable circuit of a general-purpose computer, a special-purpose computer, or other programmable data processing device via a local area network (LAN), a wide area network (WAN), or the like, and the computer-readable instructions may be executed to create a means for performing the operations specified in the flowchart or block diagram. Examples of processors include computer processors, processing units, microprocessors, digital signal processors, controllers, and microcontrollers.
[0093] Figure 7 This represents an example of a computer 9900 that can embody all or part of the various aspects of the present invention. The program installed in computer 9900 can cause computer 9900 to function as an operation associated with a device according to an embodiment of the present invention or one or more components of such a device, or to execute such operation or such one or more components, and / or can cause computer 9900 to perform a process or a stage of such a process according to an embodiment of the present invention. This program can be executed by CPU 9912 to cause computer 9900 to perform specific operations associated with some or all of the modules in the flowcharts and block diagrams described in this specification.
[0094] The computer 9900 according to this embodiment includes a CPU 9912, a RAM 9914, a graphics controller 9916, and a display device 9918, and these components are interconnected via a main controller 9910. Furthermore, the computer 9900 includes input / output units such as a communication interface 9922, a hard disk drive 9924, a DVD-ROM drive 9926, and an IC card drive, and these components are connected to the main controller 9910 via an input / output controller 9920. Furthermore, the computer includes conventional input / output units such as a ROM 9930 and a keyboard 9942, and these components are connected to the input / output controller 9920 via an input / output chip 9940.
[0095] The CPU 9912 controls each unit by executing operations according to programs stored in the ROM 9930 and the RAM 9914. The graphics controller 9916 obtains image data generated by the CPU 9912 in a frame buffer or the like or in the graphics controller itself and supplies it to the RAM 9914, and displays the image data on the display device 9918.
[0096] The communication interface 9922 communicates with other electronic devices via a network. The hard disk drive 9924 stores programs and data used by the CPU 9912 in the computer 9900. The DVD-ROM drive 9926 reads programs and data from the DVD-ROM 9901 and provides the programs and data to the hard disk drive 9924 via the RAM 9914. The IC card driver reads programs and data from an IC card and / or writes programs and data to an IC card.
[0097] The ROM 9930 stores therein a startup program and the like executed by the computer 9900 upon activation, and / or programs that depend on the hardware of the computer 9900. In addition, the input / output chip 9940 connects various input / output units to the input / output controller 9920 via a parallel port, a serial port, a keyboard port, a mouse port, and the like.
[0098] The program is provided on a computer-readable medium such as a DVD-ROM 9901 or an IC card. The program is read from the computer-readable medium and installed in a hard disk drive 9924, a RAM 9914, or a ROM 9930, also an example of a computer-readable medium, and then executed by the CPU 9912. The information processing described in the program is read by the computer 9900, and the program and the various types of hardware resources described above are coordinated. By using the computer 9900 to implement information manipulation or processing, a device or method can be constructed.
[0099] For example, when communication is performed between the computer 9900 and an external device, the CPU 9912 can execute a communication program loaded in the RAM 9914 and, based on the processing described in the communication program, issue communication processing instructions to the communication interface 9922. Under the control of the CPU 9912, the communication interface 9922 reads transmission data stored in a transmission buffer area provided in the RAM 9914, the hard disk drive 9924, the DVD-ROM 9901, or a recording medium such as an IC card, and transmits the read transmission data to the network or writes reception data received from the network to a reception buffer area provided on the recording medium.
[0100] Furthermore, the CPU 9912 can read all or a required portion of a file or database stored in an external recording medium such as the hard disk drive 9924, DVD-ROM drive 9926 (DVD-ROM 9901), or IC card into the RAM 9914, and perform various types of processing on the data in the RAM 9914. The CPU 9912 then writes the processed data back to the external recording medium.
[0101] Various types of information such as various types of programs, data, tables, and databases can be stored in a recording medium and subjected to information processing. The CPU 9912 can perform various types of processing including various types of operations, information processing, conditional judgments, conditional branches, unconditional branches, information retrieval / replacement, etc., which are recorded at any location in the present disclosure and specified by the instruction sequence of the program, on the data read from the RAM 9914, and write the results back to the RAM 9914. In addition, the CPU 9912 can retrieve information in files, databases, etc. in the recording medium. For example, in the case where a plurality of entries each having an attribute value of a first attribute associated with an attribute value of a second attribute are stored in the recording medium, the CPU 9912 can retrieve an entry that specifies an attribute value of the first attribute that is consistent with the condition from the plurality of records, read the attribute value of the second attribute stored in the entry, and obtain the attribute value of the second attribute associated with the first attribute that satisfies the condition predetermined thereby.
[0102] The program or software module described above can be stored in a computer-readable medium on or near the computer 9900. In addition, a recording medium such as a hard disk or RAM provided in a server system connected to a dedicated communication network or the Internet can be used as a computer-readable medium, thereby providing the program to the computer 9900 via the network.
[0103] The present invention has been described above using the embodiments. However, the technical scope of the present invention is not limited to the scope described in the above embodiments. Those skilled in the art will appreciate that various modifications or improvements may be made to the above embodiments. As is clear from the claims, embodiments incorporating such modifications or improvements are also within the technical scope of the present invention.
[0104] Regarding the order in which actions, sequences, steps, and stages, etc., of the apparatuses, systems, programs, and methods described in the claims, specifications, and drawings are executed, it should be noted that unless otherwise expressly indicated as "before," "before," or the like, and unless the output of a previous process is used in a subsequent process, the execution order may be any order. Even if the flow of actions in the claims, specifications, and drawings is described using the phrases "first," "next," or the like for convenience, it does not necessarily mean that the actions must be executed in that order.
[0105] Description of the label
[0106] 10 Equipment
[0107] 20 Control Objects
[0108] 100 Determination Device
[0109] 110 Status data acquisition unit
[0110] 120 Operation volume data acquisition unit
[0111] 130 Control Model Generation Unit
[0112] 135 Control Model
[0113] 140 Simulation Department
[0114] 145 simulation model
[0115] 150 Judgment Department
[0116] 160 Instruction Department
[0117] 310 Output
[0118] 320 Input
[0119] 510 End Judgment Unit
[0120] 9900 Computer
[0121] 9901 DVD-ROM
[0122] 9910 Master Controller
[0123] 9912 CPU
[0124] 9914 RAM
[0125] 9916 Graphics Controller
[0126] 9918 Display Device
[0127] 9920 Input / Output Controller
[0128] 9922 Communication Interface
[0129] 9924 Hard Drive
[0130] 9926 DVD drive
[0131] 9930 ROM
[0132] 9940 Input / Output Chip
[0133] 9942 Keyboard
Claims
1. A determination device, wherein: The determination device comprises: a status data acquisition unit that acquires status data indicating a status of a device provided with a control target; an operation amount data acquisition unit that acquires operation amount data representing an operation amount of the control object; a control model generating unit that generates, by machine learning, a control model that outputs the operation amount corresponding to the state of the device using the state data and the operation amount data; a simulation unit that simulates, using a simulation model, a state of the device when the manipulated variable output by the control model is applied to the controlled object; a determination unit that determines whether the control object can be controlled by the control model based on a simulation result; as well as The instructing unit instructs, when it is determined that the control of the controlled object can be performed by the control model, to start the control of the controlled object based on the manipulated variable supplied from the control model to the controlled object.
2. The determination device according to claim 1, wherein: The determination unit determines that the control target can be controlled by the control model when it is determined based on the simulation result that the period during which the device can operate normally exceeds a predetermined threshold.
3. The determination device according to claim 1 or 2, wherein: The determination unit determines that the control target can be controlled by the control model when the number of times the device is determined to be able to operate normally exceeds a predetermined threshold based on the simulation result.
4. The determination device according to claim 1 or 2, wherein: The determination device further includes an output unit for outputting the simulation result. The determination unit determines that the control object can be controlled by the control model when an instruction to permit control is acquired in response to the output of the simulation result.
5. The determination device according to claim 1 or 2, wherein: The control model generation unit regenerates the control model through the machine learning when it is determined that the control model cannot control the controlled object.
6. The determination device according to claim 1 or 2, wherein: The determination device further includes an end determination unit for determining the end of the machine learning. The simulation unit simulates the state of the device when it is determined that the machine learning has been completed.
7. The determination device according to claim 6, wherein: The end determination unit determines the end of the machine learning based on an elapsed time from the start of the machine learning.
8. The determination device according to claim 6, wherein: The termination determination unit determines the termination of the machine learning based on a value of an evaluation function of the machine learning.
9. The determination device according to claim 1 or 2, wherein: The control model generation unit generates the control model by performing reinforcement learning in response to the input of the state data so as to output a more recommended operation amount for an operation amount having a higher reward value defined by a predetermined reward function.
10. A determination method, wherein: The determination method has the following steps: acquiring status data indicating a status of a device in which a control target is provided; acquiring operation amount data representing an operation amount of the control object; Using the state data and the operation amount data, a control model is generated by machine learning to output the operation amount corresponding to the state of the device; simulating, using a simulation model, a state of the device when the manipulated variable output by the control model is applied to the controlled object; as well as Determining whether the control object can be controlled by the control model based on the simulation result, When it is determined that the control model can control the controlled object, an instruction is given to start the control of the controlled object based on the manipulated variable supplied from the control model to the controlled object.
11. A recording medium having a determination program recorded thereon, wherein: The determination program is executed by a computer, causing the computer to function as the following functional unit: a status data acquisition unit that acquires status data indicating a status of a device provided with a control target; an operation amount data acquisition unit that acquires operation amount data representing an operation amount of the control object; a control model generating unit that generates, by machine learning, a control model that outputs the operation amount corresponding to the state of the device using the state data and the operation amount data; a simulation unit that simulates, using a simulation model, a state of the device when the manipulated variable output by the control model is applied to the controlled object; a determination unit that determines whether the control object can be controlled by the control model based on a simulation result; as well as The instructing unit instructs, when it is determined that the control of the controlled object can be performed by the control model, to start the control of the controlled object based on the manipulated variable supplied from the control model to the controlled object.
Citation Information
Patent Citations
Controller and machine learning device
JP2018202564A
Unmanned plane control model training method and system based on AI (artificial intelligence)
CN107479368A
Industrial plant controller
US20200192340A1