Machine learning device, inference device, machine learning method, inference method, machine learning program, and inference program
Patent Information
- Application Number
- PCT/JP2026/011949
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2026-03-25
- Publication Date
- 2026-10-01
Smart Images

Figure JP2026011949_01102026_PF_FP_ABST
Abstract
Description
Machine learning device, inference device, machine learning method, inference method, machine learning program and inference program
[0001] (Cross-Reference to Related Application) This application claims the benefit of priority from the specification of Japanese Patent Application No. 2025-053093 filed on March 27, 2025, and the entire content of said specification is incorporated herein by reference. (Technical Field) The present disclosure relates to a machine learning device, an inference device, a machine learning method, an inference method, a machine learning program, and an inference program.
[0002] Various methods have been proposed to solve problems. For example, it has been devised to learn a model through machine learning and infer a solution to a problem.
[0003] For example, Patent Document 1 describes exploring actions using a mask.
[0004] Japanese Unexamined Patent Application Publication No. 2019-73271
[0005] For example, as in Patent Document 1, there are cases where action decision is performed while restricting actions using a mask. For example, corresponding to a discrete probability distribution for discrete actions (variables), individual restrictions may be set for each discrete action. When individual restrictions are set for each action, simulation is performed for each action to verify feasibility, and restrictions are set based on the results, which tends to impose a processing burden on information processing devices such as machine learning devices and inference devices. Reducing the processing burden has been a technical problem.
[0006] In view of the above problems, an object of the present disclosure is to provide a machine learning device, an inference device, a machine learning method, an inference method, a machine learning program, and an inference program that can reduce processing burden.
[0007] To solve the above problems, a machine learning device according to one aspect of the present disclosure includes: an acquisition unit that acquires a state corresponding to a predetermined environment; an output unit that uses a learning model to output a discrete probability distribution relating to an action on the environment corresponding to the state, which is also a discrete action; a restriction unit that sets a continuous restriction range for the action on the discrete probability distribution; a decision unit that determines the action on the environment based on the discrete probability distribution on which the continuous restriction range has been set; and an update unit that updates the learning model based on the result of performing the action on the environment.
[0008] An inference device according to another aspect of the present disclosure includes: a receiving unit for receiving input data; an output unit that uses a trained model to output a discrete probability distribution corresponding to the input data and relating to discrete variables; a limiting unit that sets a continuous limiting range for the variables in the discrete probability distribution; an estimation unit that estimates output data corresponding to the input data based on the discrete probability distribution with the continuous limiting range set; and a data output unit that outputs the estimated output data.
[0009] A machine learning method according to another aspect of the present disclosure includes the steps of: acquiring a state corresponding to a predetermined environment; using a learning model to output a discrete probability distribution relating to a discrete action that is an action on the environment corresponding to the state; setting a continuous limit range for the action on the discrete probability distribution; determining the action on the environment based on the discrete probability distribution on which the continuous limit range has been set; and updating the learning model based on the result of performing the action on the environment.
[0010] An inference method according to another aspect of the present disclosure includes the steps of: receiving input data; using a trained model to output a discrete probability distribution relating to the input data and discrete variables; setting a continuous limit range for the discrete probability distribution with respect to the variables; estimating output data corresponding to the input data based on the discrete probability distribution with the continuous limit range set; and outputting the estimated output data.
[0011] A machine learning program according to another aspect of the present disclosure causes a computer to function as: an acquisition unit that acquires a state corresponding to a predetermined environment; an output unit that uses a learning model to output an action for the environment corresponding to the state, which is also a discrete probability distribution relating to the discrete action; a restriction unit that sets a continuous restriction range for the discrete probability distribution relating to the action; a decision unit that determines the action for the environment based on the discrete probability distribution for which the continuous restriction range has been set; and an update unit that updates the learning model based on the result of performing the action for the environment.
[0012] An inference program according to another aspect of the present disclosure causes a computer to function as: a receiving unit for receiving input data; an output unit that uses a trained model to output a discrete probability distribution corresponding to the input data and relating to discrete variables; a limiting unit that sets a continuous limiting range for the discrete probability distribution relating to the variables; an estimation unit that estimates output data corresponding to the input data based on the discrete probability distribution with the continuous limiting range set; and a data output unit that outputs the estimated output data.
[0013] Furthermore, this disclosure may be implemented as a semiconductor integrated circuit that implements part or all of the program, as an information processing device, or as a system including an information processing device.
[0014] The machine learning apparatus, inference apparatus, machine learning method, inference method, machine learning program, and inference program related to this disclosure can reduce the processing load.
[0015] This figure schematically shows an example of the configuration of a planning system according to the embodiment of this disclosure. This figure shows an example of raw material loading and unloading to and from a tank base. This figure shows the relationship between the environment and the agent related to machine learning. This figure schematically shows an example of the hardware configuration of a server device. This block diagram shows an example of various functions in the server device. This figure shows an example of loading and unloading plan information. This figure shows an example of a probability distribution corresponding to the amount of material unloaded. This figure shows an example of constraint conditions in the limiting unit. This figure shows an example of a case where a limit is set on the probability distribution. This figure shows an example of processing when multiple actions are determined step by step. This figure shows an example of the processing following the processing in Figure 10A when multiple actions are determined step by step. This figure shows an example of the processing following the processing in Figure 10B when multiple actions are determined step by step. This figure shows an example of processing using a trained model. This flowchart shows an example of the learning process flow. This flowchart shows an example of the inference process flow.
[0016] The following descriptions illustrate some aspects of this disclosure.
[0017] A machine learning device according to a first aspect of this disclosure includes: an acquisition unit that acquires a state corresponding to a predetermined environment; an output unit that uses a learning model to output a discrete probability distribution relating to an action on the environment corresponding to the state, which is also a discrete action; a restriction unit that sets a continuous restriction range for the action on the discrete probability distribution; a decision unit that determines the action on the environment based on the discrete probability distribution on which the continuous restriction range has been set; and an update unit that updates the learning model based on the result of performing the action on the environment.
[0018] In the machine learning apparatus according to the second aspect of this disclosure, the continuous limit range is set as a continuous range that includes a plurality of discrete actions in the discrete probability distribution, in relation to the first aspect.
[0019] In the machine learning apparatus according to the third aspect of this disclosure, relating to the first and second aspects, the limiting unit sets a continuous limiting range using predetermined constraint conditions.
[0020] In the machine learning apparatus according to the fourth aspect of this disclosure, the actions toward the environment, as per the first and second aspects, are extensive actions.
[0021] In the machine learning apparatus according to the fifth aspect of this disclosure, relating to the first to second aspects, the decision unit determines the action for the environment from among the actions that are not included in the continuous limit range in the discrete probability distribution.
[0022] In the machine learning apparatus according to the sixth aspect of this disclosure, as per the fifth aspect, the decision unit probabilistically determines the action in relation to the environment using the probability distribution.
[0023] In the machine learning apparatus according to the seventh aspect of this disclosure, according to the first to second aspects, the limiting unit sets a limit on the probability distribution relating to the non-extensive behavior based on the result of setting the limiting range on the probability distribution relating to the extensive behavior that is related to the non-extensive behavior.
[0024] A machine learning apparatus according to the eighth aspect of this disclosure further comprises a calculation unit that calculates a loss based on the reward for the determined action, according to the first to second aspects, and the update unit updates the learning model using the loss as a result.
[0025] In the machine learning apparatus according to the ninth aspect of this disclosure, as per the eighth aspect, the calculation unit calculates the loss based on the probability of the determined action and the reward for the determined action.
[0026] In the machine learning apparatus according to the tenth aspect of this disclosure, relating to the first to second aspects, the learning model takes as input data relating to at least one of the plans for loading and unloading of multiple types of raw materials to and from a tank base, and outputs data relating to at least one of the plans for loading and unloading of each of the multiple tanks owned by the tank base.
[0027] In the machine learning apparatus according to the eleventh aspect of this disclosure, relating to the first to second aspects, the environment corresponds to a raw material tank base, and the action is an action relating to at least one of the loading and unloading of the raw material at the tank base.
[0028] An inference device according to a twelfth aspect of this disclosure uses the learned model, which has been learned by the above-described machine learning device, as a trained model to infer output data corresponding to input data.
[0029] An inference device according to a thirteenth aspect of this disclosure includes: a receiving unit for receiving input data; an output unit that uses a trained model to output a discrete probability distribution corresponding to the input data and relating to discrete variables; a limiting unit that sets a continuous limiting range for the variables in the discrete probability distribution; an estimation unit that estimates output data corresponding to the input data based on the discrete probability distribution with the continuous limiting range set; and a data output unit that outputs the estimated output data.
[0030] A machine learning method according to a fourteenth aspect of this disclosure includes the steps of: acquiring a state corresponding to a predetermined environment; using a learning model to output a discrete probability distribution relating to an action on the environment corresponding to the state, which is also a discrete action; setting a continuous limit range for the action on the discrete probability distribution; determining the action on the environment based on the discrete probability distribution on which the continuous limit range has been set; and updating the learning model based on the result of performing the action on the environment.
[0031] An inference method according to a 15th aspect of this disclosure includes the steps of: receiving input data; using a trained model to output a discrete probability distribution relating to the input data and discrete variables; setting a continuous limit range for the discrete probability distribution with respect to the variables; estimating output data corresponding to the input data based on the discrete probability distribution with the continuous limit range set; and outputting the estimated output data.
[0032] A machine learning program according to the sixteenth aspect of this disclosure causes a computer to function as: an acquisition unit that acquires a state corresponding to a predetermined environment; an output unit that uses a learning model to output a discrete probability distribution that is an action toward the environment corresponding to the state and relates to the discrete action; a restriction unit that sets a continuous restriction range for the discrete probability distribution and relates to the action; a decision unit that determines the action toward the environment based on the discrete probability distribution with the continuous restriction range set; and an update unit that updates the learning model based on the result of performing the action toward the environment.
[0033] The inference program according to the 17th aspect of this disclosure causes the computer to function as: a receiving unit that receives input data; an output unit that uses a trained model to output a discrete probability distribution corresponding to the input data and relating to discrete variables; a limiting unit that sets a continuous limiting range for the discrete probability distribution with respect to the variables; an estimation unit that estimates output data corresponding to the input data based on the discrete probability distribution with the continuous limiting range set; and a data output unit that outputs the estimated output data.
[0034] The following description illustrates embodiments of the present disclosure. To facilitate understanding of the description, the same reference numerals are used for identical components and steps in each drawing whenever possible, and redundant descriptions are omitted.
[0035] <Overall Configuration> Figure 1 is a schematic diagram showing an example of the configuration of a planning system 1 according to one embodiment. The planning system 1 is a system that creates a plan as a solution corresponding to a pre-set problem.
[0036] As shown in Figure 1, the planning system 1 consists of a user terminal 2 and a server device 3. The server device 3 and the user terminal 2 can communicate with each other via the network NT.
[0037] User terminal 2 is a terminal device, an information processing device (computer) used by the user. User terminal 2 is, for example, a personal computer. The user can input various information using user terminal 2 and instruct server device 3 to create a plan.
[0038] Server device 3 is an information processing device (computer) that creates a plan to address a problem using information input by user terminal 2. In this embodiment, one example of a problem is a problem related to the loading (inbound) and unloading (shipping) of raw materials to and from the tank base 21, which will be described later. In this embodiment, the problem of planning the loading and unloading of raw materials to and from the tank base 21 is referred to as the "loading and unloading problem." Note that the problem to be solved is not limited to the loading and unloading problem to and from the tank base 21.
[0039] Figure 2 shows an example of raw material being brought into and out of the tank base 21.
[0040] As shown in Figure 2, the tank base 21 is a base for temporarily storing raw materials. For example, raw materials are brought into the tank base 21 using a ship F1, and raw materials are unloaded from the tank base 21 using a ship F2, etc. The means for bringing in and unloading are not limited to ships F1 and F2. The tank base 21 is equipped with multiple tanks 24. In the example shown in Figure 2, the tank base 21 has tanks 24a, tank 24b, and tank 24c. The number of tanks 24 in the tank base 21 is not limited. The tank base 21 can store the brought-in raw materials in each of the tanks 24. The tank base 21 can also unload the raw materials stored in each of the tanks 24. Multiple types of raw materials may be brought into the tank base 21 and stored in each of the tanks 24. Furthermore, the tank base 21 may have multiple tanks 24 into which different types of raw materials are delivered, or multiple types of raw materials may be delivered to a single tank 24. Each of the multiple types of raw materials stored in each tank 24 can be delivered from each tank 24. In this embodiment, the problem of delivery and delivery is given as an example to the problem of both the delivery and delivery of raw materials to and from the tank base 21, but it may also be a problem concerning only one of the delivery or delivery of raw materials to and from the tank base 21.
[0041] In this embodiment, the server device 3 creates a general plan (broad plan) for the loading and unloading of multiple types of raw materials to and from the tank base 21, and then creates a specific plan for the loading and unloading of each raw material to and from each tank 24 in the tank base 21. In other words, the broad plan is a prerequisite for creating the specific plan. To put it another way, the broad plan is a plan for the entire tank base 21, and the specific plan is a detailed plan within the tank base 21.
[0042] Server device 3 creates a plan using machine learning. Therefore, server device 3 corresponds to a machine learning device when performing learning and to an inference device when performing inference (plan creation).
[0043] Figure 3 is a diagram illustrating the overview of machine learning.
[0044] As shown in FIG. 3, in the present embodiment, the server device 3 performs reinforcement learning. For example, the server device 3 performs deep reinforcement learning. Reinforcement learning is a learning method in which an environment E1 and an agent A1 interact with each other to complete a task (a solution to a problem), whereby the agent A1 learns what kind of action should be taken to obtain more rewards (evaluations). The agent A1 is the subject of actions. An action is the behavior of the agent A1. The environment E1 is the object on which the agent A1 takes actions, and is also a precondition for the agent A1. That is, the state of the environment E1 is the state in which the agent A1 is placed within the environment E1. The state of the environment E1 changes according to the action taken by the agent A1 on the environment E1. That is, an action can be said to be a factor that changes the environment E1. The agent A1 is provided with a reward corresponding to the action (an evaluation of the action). The agent A1 receives a larger reward when it executes a favorable action in the environment E1, compared with when it executes an unfavorable action in the environment E1. Note that the reward evaluation method can be appropriately set according to the learning objective and the like.
[0045] As shown in operation S1, the agent A1 takes an action on the environment E1 according to the state of the environment E1. Thereby, the state of the environment E1 changes to a new state according to the action taken by the agent A1. Further, as shown in operation S2, the agent A1 is provided with the new state of the environment E1 and a reward corresponding to the action taken by the agent A1.
[0046] Then, in reinforcement learning, learning is performed based on the interaction between the environment E1 and the agent A1 generated by repeating operation S1 and operation S2. For example, when learning is performed using an episode including a plurality of steps, interaction between the environment E1 and the agent A1 such as operation S1 and operation S2 is executed corresponding to each step. Note that an episode is a flow (period) from the start to the end of a task to be solved in reinforcement learning.
[0047] Furthermore, the action performed by agent A1 in operation S1 is determined based on policy W1. Policy W1 is a rule (policy) serving as an index for determining the action of agent A1 according to the state of environment E1 before the action. For example, policy W1 is information in which a plurality of patterns of actions that can be taken by agent A1 corresponding to the state of environment E1 are associated with the probability (selection probability) that each action is executed. In the present embodiment, information in which a plurality of patterns of actions are associated with the probability that each action is executed is referred to as a "probability distribution". For example, reinforcement learning is adjusting the probability distribution as policy W1 with the aim of enabling agent A1 to execute more preferable actions in the state of environment E1.
[0048] In deep reinforcement learning, a policy W1 corresponding to the state of environment E1 is obtained using a neural network. That is, the neural network outputs information related to the action of agent A1 from the state of environment E1. Then, the neural network is updated (trained) such that, for example, the reward is maximized (the loss is minimized). That is, the neural network is an object of learning and is an example of a "learning model". Note that the learning model is not limited to a neural network, and can be variously applied according to the reinforcement learning method selected by a user, the type of problem, and the like.
[0049] <Hardware Configuration> FIG. 4 is a diagram schematically illustrating an example of the hardware configuration of server device 3.
[0050] As illustrated in FIG. 4, the server device 3 includes a control device 40, a communication device 41, and a storage device 42. The control device 40 is mainly configured to include a processor 43 and a memory 44.
[0051] In the control device 40, the processor 43 executes a predetermined program stored in the memory 44, the storage device 42, or the like, thereby functioning as various functional configurations described later.
[0052] The processor 43 is, for example, a CPU (Central Processing Unit). However, the processor 43 is not limited to a CPU. The processor 43 may be, for example, a GPU (Graphics Processing Unit) or an NPU (Neural Processing Unit). In a specific example, the processor 43 is a multi-core processor. The processor 43 may be a single-core processor. The processor 43 may include multiple processors or cores and be capable of performing parallel processing. The processor 43 is configured to execute computer programs. The processor 43 may, for example, include an ASIC (Application Specific Integrated Circuit) in part, or it may include programmable hardware such as an FPGA (Field Programmable Gate Array) or a CPLD (Complex Programmable Logic Device) in part.
[0053] The memory 44 is a computer-readable storage medium and may consist of at least one of the following: RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), etc. The memory 44 can store various types of data, including programs necessary for executing processing in the server device 3.
[0054] The communication device 41 consists of a communication interface and the like for communicating with an external device. For example, the communication device 41 can communicate with the user terminal 2.
[0055] The storage device 42 is, for example, a computer-readable instruction recording medium (a non-temporary computer-readable instruction recording medium) and is composed of a hard disk, a solid-state drive, or the like. The storage device 42 stores various programs and information necessary for executing processing in the control device 40, as well as information on the processing results. Other examples of non-temporary computer-readable instruction recording media include portable recording media such as magnetic tapes, flexible disks, optical disks, digital versatile disks, Blu-ray discs, magneto-optical disks, memory cards, and USB memory.
[0056] The server device 3 may consist of a single information processing device or multiple information processing devices. Furthermore, Figure 4 only shows a part of the main hardware configuration of the server device 3, and the server device 3 may have other configurations. For example, the server device 3 may further include an input device (not shown) and a display device (not shown). The input device is an input device that receives input from the outside (e.g., a keyboard, mouse, etc.). The input device receives user operations and inputs those operations to the server device 3. The display device is a display device that performs output to the outside (e.g., a display, etc.). The display device outputs characters and images. The server device 3 may have the input device and output device integrated (e.g., a touch panel). The user terminal 2 may also have a configuration similar to the server device 3 as an information processing device, including a control device (processor and memory), a communication device, a storage device, etc.
[0057] <Functional Configuration> Figure 5 is a block diagram showing an example of various functions in the server device 3. Various processes are executed according to the functions in each block. Computer programs that implement the functions of at least some of the functional blocks shown in Figure 5 may be installed in the storage of one or more computers. The processors of one or more computers may perform the functions of multiple functional blocks shown in Figure 5 by reading the computer programs installed on their own machines into main memory and executing them.
[0058] Furthermore, the functions of each functional block shown in Figure 5 may be executed by a single computer, or they may be executed in a distributed manner across multiple computers. When the functions of each functional block shown in Figure 5 are executed in a distributed manner across multiple computers, these multiple computers may send and receive data via a communication network including a LAN (Local Area Network), a WAN (Wide Area Network), or the Internet.
[0059] As shown in Figure 5, the server device 3 has a functional configuration that mainly consists of a learning unit 51 and an inference unit 52. In the server device 3, the learning unit 51 functions as a machine learning device, and the inference unit 52 functions as an inference device.
[0060] The learning unit 51 performs learning (deep reinforcement learning) on the learning model. The learning unit 51 mainly comprises a simulation unit 60, a first acquisition unit 61, a setting unit 62, an output unit 63, a limiting unit 64, a decision unit 65, a second acquisition unit 66, a calculation unit 67, and an update unit 68.
[0061] The simulation unit 60 executes a simulation related to the loading and unloading problem. For example, with respect to the loading and unloading problem, environment E1 becomes a predetermined environment corresponding to the simulation model of the tank base 21. Agent A1 is a virtual entity that performs actions such as loading and unloading at the tank base 21. Agent A1 executes an action determined by the state of the tank base 21. An episode is the entire loading and unloading problem (the whole process), and the loading and unloading of raw materials to and from the tank base 21 each corresponds to a step (stage). That is, loading included in the loading and unloading problem corresponds to one step, unloading included in the loading and unloading problem corresponds to another step, and unloading and loading each correspond to different steps. In this way, the loading and unloading problem includes multiple steps arranged in chronological order. The learning model is trained so that it can determine the best action (or the closest to the best action) for the loading and unloading problem.
[0062] Specifically, the learning model can accept input data related to the planning of loading and unloading multiple types of raw materials to and from the tank base 21. In other words, the input data is information related to the overall plan for the loading and unloading problem. For example, the input data includes at least one of initial inventory information, loading and unloading plan information, and constraint information. The initial inventory information, loading and unloading plan information, and constraint information are preconditions for the loading and unloading problem.
[0063] The initial inventory information shows the inventory status of each tank 24 in the initial state of the loading and unloading problem. The inventory status is the type and quantity of raw materials stored in each tank 24. In other words, the initial inventory information is information that associates tank 24 (name, identification information, etc.) with the type of raw materials stored and the quantity (weight or percentage) stored.
[0064] The loading and unloading plan information consists of the loading and unloading plans included in the loading and unloading problem. In other words, the loading and unloading plan information, in particular, constitutes the overall plan related to the loading and unloading problem.
[0065] Figure 6 shows an example of loading and unloading plan information.
[0066] As shown in Figure 6, the loading and unloading plan information is associated with attributes, dates (or sequences), and loading or unloading quantities for each type of raw material. Attributes indicate loading or unloading. Loading or unloading quantities for each type of raw material indicate the amount of raw material loaded into or unloaded from the tank base 21. In other words, the loading and unloading plan information does not limit individual loading or unloading to tanks 24, but shows the loading and unloading plan for the entire tank base 21. In Figure 6, loading or unloading quantities for each type of raw material are shown by the quantity of each type of raw material corresponding to the attribute (loading or unloading). The example in Figure 6 shows a case where six types of raw materials, G1, G2, G3, G4, G5, and G6, are handled in the loading and unloading problem. Note that the loading and unloading plan information may also be associated with identification information (name, etc.) of the vessel involved in the loading or unloading.
[0067] In the loading and unloading problem, loading and unloading in the loading and unloading plan information correspond to steps, forming an episode as a whole.
[0068] The constraint information indicates the constraints in the loading and unloading problem. For example, the constraint information includes at least one of the following: upper and lower limit constraints on the inventory in tank 24, constraints on the number of tanks 24 used during loading and unloading, minimum quantities related to loading and unloading, concentration constraints, and tank internal layer separation constraints. The concentration constraint is a constraint on the combination and proportion (concentration) of multiple types of raw materials stored in one tank 24. The tank internal layer separation constraint is information that restricts the combination and order in which multiple types of raw materials are stored to prevent separation of multiple types of raw materials within tank 24. The constraint information may also include upper and lower limit constraints on the API (American Petroleum Institute) specific gravity of tank 24. Here, API specific gravity refers to the specific gravity of crude oil as defined by the American Petroleum Institute. API specific gravity is a value that can be measured, for example, in accordance with ASTM D1298. In this embodiment, API specific gravity may be simply referred to as "API".
[0069] Furthermore, when training a learning model, it is preferable to perform the training using various patterns of input data (prerequisites).
[0070] The learning model then provides output data regarding the loading and unloading plans for each of the multiple tanks 24 located within the tank base 21. In other words, the output data represents a specific plan within the tank base 21 that corresponds to the preconditions (overall plan, etc.).
[0071] For example, the output data includes information such as tank selection information, raw material quantity information, and type information. The tank selection information is information about the tank 24 that is operated on in relation to loading or unloading. That is, the tank 24 that is the loading destination or the unloading source is selected in the tank selection information. The tank selection information may also be information about the tank 24 involved in the transfer of raw materials from one tank 24 to another tank 24 within the tank base 21 (raw material shift). The raw material quantity information is information about the amount of raw material to be loaded into or unloaded from the selected tank 24, corresponding to the tank 24. The type information is information about the type of raw material to be loaded into or unloaded from the selected tank 24, corresponding to the tank 24. For example, the output data is shown as loading 100 kl of raw material G3 into tank 24a. If the loading / unloading problem involves multiple steps (loading or unloading), the tank selection information, raw material quantity information, and type information are output corresponding to each step.
[0072] Furthermore, the output data may include evaluation indicators corresponding to the output tank selection information, raw material quantity information, and type information. Examples of evaluation indicators include raw material agreement rate, raw material group agreement rate, and API error. The raw material agreement rate is the ratio of the planned quantity (actual quantity) to the required quantity (ideal quantity) of raw material for each type. For example, in the shipment of raw materials, the raw material agreement rate is the ratio of the quantity of raw material to be shipped to the quantity of raw material required for shipment. When calculating the raw material agreement rate for multiple types, for example, the average of the raw material agreement rates for each type of raw material may be used.
[0073] Furthermore, the raw material group agreement rate is the ratio of the planned quantity (actual quantity) of raw materials to the required quantity (ideal quantity) for each group. A group is, for example, a group of several types of raw materials with similar properties. For example, in the shipment of raw materials, the raw material group agreement rate is the ratio of the total quantity of raw materials of the same group that are scheduled to be shipped to the total quantity of raw materials of the same group that are required at the time of shipment. When calculating the raw material group agreement rate for multiple groups, for example, the average of the raw material agreement rates for each group may be used. The API error is the difference between the API of the required raw material (for example, the average of the APIs of various raw materials) and the API of the scheduled raw material (for example, the average of the APIs of various raw materials). For example, in the shipment of raw materials, the API error represents the difference between the API of the raw material required at the time of shipment and the API of the raw material that is scheduled to be shipped.
[0074] In this way, the simulation unit 60 is capable of executing simulations related to the loading and unloading problem in accordance with the learned model.
[0075] Returning to Figure 5, the first acquisition unit 61 has the function of an acquisition unit and acquires the current state in environment E1 as the "first state". That is, the first acquisition unit 61 acquires the first state of environment E1 in which agent A1 is located before agent A1 takes action. Specifically, the first acquisition unit 61 observes the state of the tank base 21 as environment E1 and sets it as the first state.
[0076] In the loading and unloading problem, the first state includes, for example, the state of the tanks 24 in the tank base 21 as environment E1. The state of the tanks 24 refers to information such as the amount of raw materials stored in each of the tanks 24 and the types of raw materials stored.
[0077] When learning is performed through an episode, the first acquisition unit 61 acquires a first state corresponding to each of the multiple steps included in the episode.
[0078] The setting unit 62 sets candidate actions. Specifically, the setting unit 62 sets candidate actions corresponding to the steps. For example, in the case of a step related to removal, the setting unit 62 sets actions related to removal as candidate actions. For example, actions related to removal include selecting the source tank 24 and the amount to be removed. The setting unit 62 then sets an effective range for the candidate actions corresponding to the steps. The effective range is the range of actions that can be taken in relation to the environment E1. The effective range may be set, for example, in relation to the steps (receiving or removing), or in relation to the first state. If the candidate action corresponding to the step is the amount to be removed, the effective range may be set, for example, from 30 kl to 100 kl.
[0079] In this way, the setting unit 62 sets the candidate actions and their effective range.
[0080] The output unit 63 uses a learning model to output a probability distribution related to actions in environment E1 based on the state of environment E1. In other words, the output unit 63 uses a learning model to output a probability distribution of actions corresponding to the first state.
[0081] In particular, the output unit 63 uses a learning model to output a probability distribution corresponding to the candidate actions set in the setting unit 62. The probability distribution corresponds to actions within the effective range set in the setting unit 62. For example, the output unit 63 uses a learning model to output a probability distribution corresponding to the amount of material dispensed, as well as a probability distribution corresponding to the effective range (e.g., from 30 kl to 100 kl).
[0082] A probability distribution is probability information that associates multiple actions of agent A1 with their probabilities (selection probabilities) corresponding to the first state. A probability distribution can be represented, for example, by a function that shows probability. A probability distribution is a function that associates the probability of a random variable taking a particular value. The function that shows probability may also be a probability density function. In a probability distribution, actions are represented as variables. These variables are also called random variables.
[0083] Behavior can be divided into descriptive behavior and non-descriptive behavior. Descriptive behavior can also be called descriptive behavior or descriptive variable (a variable that possesses descriptive properties). Non-descriptive behavior is descriptive behavior, and can also be called descriptive behavior or descriptive variable.
[0084] Extensive behaviors are those related to variables that are proportional to changes in the size of the system. Extensive behaviors can also be described as those related to variables that depend on quantity. In other words, extensive behaviors can be described as actions that change the quantity (state of quantity) in environment E1. Furthermore, extensive behaviors may also be considered as actions that can take on continuous values.
[0085] In this embodiment, actions include, for example, selecting a tank 24 related to tank selection information, selecting a quantity of raw material related to raw material quantity information, and selecting a type of raw material related to type information. Therefore, raw material quantity information (selection of raw material quantity) is an extensive action. Specifically, the amount of raw material brought into tank 24 (input amount) and the amount of raw material discharged from tank 24 (discharge amount) are examples of extensive actions. In contrast, tank selection information and type information are examples of actions that do not have extensive properties.
[0086] In this embodiment, the action having extensive properties is defined as "discharge amount." The discharge amount indicates the amount of raw material discharged from a certain tank 24 in the tank base 21.
[0087] In this embodiment, the output unit 63 outputs a discrete probability distribution (discrete probability distribution) related to discrete actions using a learning model. In particular, the output unit 63 discretizes extensive actions using the learning model and outputs a discrete probability distribution. For example, if the effective range of the output amount is set from 30 kl to 100 kl in the setting unit 62, the learning model discretizes the output amount within the effective range and outputs a discrete probability distribution corresponding to the discretized output amount. That is, the learning model outputs a probability distribution while discretizing output amounts that can take continuous values (extensive actions). Discretization is set in advance in the learning model. For example, information such as the number of divisions related to discretization is set in advance in the learning model. Output amounts that can take continuous values are converted to take discrete values through discretization.
[0088] Figure 7 shows an example of a probability distribution corresponding to the output quantity as an example of an action. In Figure 7, the vertical axis represents probability, and the horizontal axis represents the output quantity (an example of a random variable). In a probability distribution, actions are discrete. That is, the probability distribution is represented as a discrete space for actions. Specifically, the output quantity has discrete values. In the example in Figure 7, the output quantity takes eight discrete values, specifically output quantity X1, X2, X3, X4, X5, X6, X7, and X8. Note that the number of discrete values (patterns) is not limited to eight and can be set as appropriate. For example, the bar graph extends along the vertical axis with each output quantity value as the center of the horizontal axis. In this way, the output quantity is a discrete variable in the probability distribution.
[0089] Since the output quantities as variables are discrete, the probability distribution is a discrete probability distribution. Specifically, in Figure 7, the probability distribution is shown by the probabilities corresponding to each output quantity from output quantity X1 to output quantity X8.
[0090] For example, the probability distribution may have a shape that is nearly symmetrical with respect to the output volume, as shown in Figure 7, or it may have an asymmetrical shape. In other words, the shape of the distribution in the probability distribution is not limited.
[0091] In the above example, an example of a probability distribution for an extensive action is shown in Figure 7, where the output quantity is used. However, the probability distribution for extensive actions other than the output quantity may be a discrete probability distribution. Examples of extensive actions other than the output quantity include the input quantity. The input quantity represents the amount of raw material input into a tank 24 at the tank base 21. The probability distribution for actions that do not have extensive properties may also be a discrete probability distribution. Furthermore, discrete and continuous probability distributions may be applied depending on the type of action, such as one type of action and another type of action, or an extensive action and an action that does not have extensive properties.
[0092] In this way, the output unit 63 outputs a discrete probability distribution related to the behavior using the learning model.
[0093] When learning is performed through episodes, the output unit 63 outputs a probability distribution relating to the actions of agent A1 for each of the first states of each step acquired by the first acquisition unit 61.
[0094] The restriction unit 64 sets restrictions on the action selection based on the probability distribution in the decision unit 65, which will be described later. The restriction unit 64 sets restrictions on actions with respect to the probability distribution. That is, the restriction unit 64 sets a restriction range with respect to the probability distribution. The restriction range is also called a mask.
[0095] Specifically, the restriction unit 64 sets restrictions based on predetermined constraint conditions (constraint information) to prevent inappropriate actions corresponding to the first state from being determined as actions of agent A1.
[0096] Figure 8 shows an example of the constraint conditions used in the limiting section 64.
[0097] As shown in Figure 8, for example, the constraints are upper and lower limits on the inventory in tank 24, upper and lower limits on the API of the inventory in tank 24, upper limit limit on the number of tanks 24 used during unloading, and upper limit limit on the number of tanks 24 used during loading. Also, for example, the constraints are lower limit limits on the amount loaded per tank, lower limit limits on the amount unloaded per tank, upper limit limits on the amount transported during shifts between tanks, and lower limit limits on the amount transported during shifts between tanks. Also, for example, the constraints are upper limit limits on the number of shifts between tanks per predetermined period (e.g., one month) and upper limit limits on the number of shifts between tanks per predetermined period (e.g., one day). A shift is an operation to move raw materials from one tank 24 to another tank 24. The limiting unit 64 may use at least one of the above constraints.
[0098] Furthermore, each constraint is assigned a type of action to which it applies. The types of actions are the selection of tank 24 during discharge, the amount of raw material during discharge, the selection of tank 24 during loading, and the amount of raw material during loading. Additionally, the types of actions corresponding to shifts are whether or not the tank is used, the selection of tank 24 to discharge, the selection of tank 24 to load, and the amount of raw material. Figure 8 shows an example of the correspondence between constraints and the types of actions to which those constraints can be applied. In Figure 8, "○" indicates that a particular constraint is applicable to a particular action. The limiting unit 64 sets limits using constraints that are applicable to the probability distribution (applicable to actions).
[0099] Furthermore, each constraint may or may not be assigned a priority. In the example in Figure 8, an example is shown where three levels of priority (high, medium, and low) are set. In the example in Figure 8, "high" is the highest priority, "medium" is the next highest, and "low" is the next highest. More specifically, when there are multiple constraints imposed on the same action, the constraint with high priority is applied preferentially over the constraints with medium and low priority. Also, when there are multiple constraints imposed on the same action, the constraint with medium priority is applied preferentially over the constraint with low priority. Note that "preferential application" may also mean, for example, when there are multiple constraints imposed on the same action, the degree of influence of each constraint on the action is corrected using a weighting according to the order of priority of the multiple constraints. Alternatively, "preferential application" may also mean, for example, when there are multiple constraints imposed on the same action, only the constraint with the highest priority among the multiple constraints is applied to that action. Furthermore, "preferential application" may mean, for example, that when multiple constraints are imposed on the same action, the constraint with the highest priority will always be satisfied, while the other constraints of lower priority will be satisfied as much as possible. Note that the number of priority levels is not limited to three, and can be changed as appropriate by the user, for example, to two levels. The priority of each constraint can also be changed as appropriate by the user.
[0100] The limiting unit 64 sets a continuous limiting range for actions with respect to discrete probability distributions. In this way, the limiting unit 64 sets a continuous limiting range for the relationship between discretely represented actions and probabilities.
[0101] A continuous limit range is defined as a continuous range of behavior, representing the range that restricts behavior. A continuous limit range is a continuous range for a variable representing behavior in a probability distribution. Specifically, if the behavior is an output quantity, the limit range is set as a continuous range of the output quantity being restricted. In other words, a limit range corresponding to an extensive behavior is set as a continuous range relating to the quantity.
[0102] Figure 9 shows an example of a case where restrictions are set on the behavior of a probability distribution. Specifically, Figure 9 shows an example where the restriction unit 64 sets restrictions on the probability distribution shown in Figure 7.
[0103] As shown in Figure 9, for example, the lower limit constraint on the amount of material to be removed per tank, as shown in Figure 8, can be applied as a constraint condition. In this case, the limiting unit 64 sets the limiting range for the amount of material to be removed based on the lower limit constraint on the amount of material to be removed per tank. In Figure 9, the limiting range corresponding to the lower limit constraint on the amount of material to be removed per tank is shown as the limiting range MS1. The limiting range MS1 indicates that the amount of material to be removed from X1 to X3 is subject to the limiting.
[0104] For example, the lower limit constraint on the amount discharged per tank is defined as a numerical range related to the amount discharged. Therefore, the limiting unit 64 can set a continuous limiting range MS1 as shown in Figure 9 by using the lower limit constraint on the amount discharged per tank. Other constraints may also be defined as numerical ranges.
[0105] Furthermore, regarding the amount of material to be discharged, it is also possible to apply upper and lower limit constraints on the inventory of tank 24, as shown in Figure 8. For example, the limiting unit 64 converts the numerical range related to the upper and lower limits of the inventory of tank 24, indicated by the upper and lower limit constraints on the inventory of tank 24, into a numerical range related to the amount of material to be discharged, for example, using a function. For example, the amount of raw materials stored in tank 24 before discharge is used to calculate the numerical range related to the amount of material to be discharged that corresponds to the upper and lower limit constraints on the inventory of tank 24. As a result, the limiting unit 64 sets a continuous limiting range that corresponds to the upper and lower limit constraints on the inventory of tank 24. In Figure 9, the limiting ranges corresponding to the upper and lower limit constraints on the inventory of tank 24 are shown as limiting range MS2 and limiting range MS3. Limiting range MS2 indicates that the amount of material to be discharged from X1 to X2 is subject to restriction. Limiting range MS3 indicates that the amount of material to be discharged from X7 to X8 is subject to restriction.
[0106] Thus, the limiting unit 64 sets a continuous limiting range for a discrete probability distribution, for example, by using the numerical range indicated by the constraint conditions, or by transforming the numerical range. That is, the limiting unit 64 may set the numerical range indicated by the constraint conditions as the limiting range, or it may set the limiting range by transforming the numerical range indicated by the constraint conditions using a function or the like. The function can be represented, for example, by arithmetic operations.
[0107] Furthermore, the restriction range is not set as an individual restriction corresponding to each discrete action in the probability distribution, but rather as a continuous range. In other words, the continuous restriction range is set as a continuous range that includes multiple discrete actions in the discrete probability distribution. Note that, depending on the range indicated by the restriction range, it may include only one discrete action.
[0108] Among the discrete actions (variables) shown in the probability distribution, the actions that fall within the restricted range are restricted in their selection when selecting an action in the decision unit 65, which will be described later.
[0109] Furthermore, the constraints used by the limiting unit 64 are not limited to those described above in the loading / unloading problem, and the user may change them as appropriate. Also, when applying a problem other than the loading / unloading problem to the planning system 1, the constraints used by the limiting unit 64 may be set as appropriate for the problem being applied. Moreover, in this case, when multiple constraints are set, each of the multiple constraints may be set as an appropriate priority.
[0110] The decision unit 65 determines the action of agent A1 based on a probability distribution. Specifically, the decision unit 65 determines the specific content of the action based on a discrete probability distribution with a continuous limit range.
[0111] The decision unit 65 determines an action by selecting one action from a group of actions based on a probability distribution. Specifically, the decision unit 65 makes an action decision by selecting one action from among several discrete actions in a probability distribution using probability. For example, the decision unit 65 probabilistically selects one action based on a probability distribution. "Probabilistically selecting an action" means that an action is selected randomly according to the probabilities shown in the probability distribution, rather than by a decisive rule. In other words, because an action is selected probabilistically, an action with a high probability may be selected, or an action with a low probability may be selected, making the selection probabilistic. The decision unit 65 may also make a decisive choice of action. "Decisively selecting an action" means that an action is selected not probabilistically, but based on predetermined (set) rules using a probability distribution. Decisively selecting an action could mean, for example, selecting an action whose probability in the probability distribution is greater than or equal to a predetermined value, or selecting the action with the highest probability in the probability distribution.
[0112] Furthermore, when the decision unit 65 performs learning (updating) of the learning model, it is preferable to determine the action probabilistically using a probability distribution. That is, during the learning phase of the learning model, by probabilistically determining the action and performing action exploration while considering not only high-probability actions but also low-probability actions as options, it is possible to perform learning that takes into account various action patterns. In this way, the decision unit 65 determines the action of agent A1 corresponding to the first state.
[0113] In this embodiment, a continuous limit range is set for a discrete probability distribution. Therefore, the decision unit 65 determines an action for environment E1 from among actions that are not included in the continuous limit range of the discrete probability distribution. For example, as shown in Figure 9, if limit ranges MS1, MS2, and MS3 are set for the probability distribution, the decision unit 65 excludes the output amounts X1, X2, X3, X7, and X8, which have limit ranges set, from the options. In other words, the decision unit 65 excludes actions included in the discrete probability distribution that are included in the continuous limit range. Then, the decision unit 65 determines the output amount as an action from among the output amounts X4, X5, and X6 that are included in the range of output amounts for which no limit range is set. In other words, the decision unit 65 determines an action for environment E1 from among actions included in the discrete probability distribution that are not included in the continuous limit range. For example, if the targets are output quantities X4, X5, and X6, the decision unit 65 uses probability to select output quantity X5 and determines the specific content of the action as output quantity X5.
[0114] In this way, the decision unit 65 determines an action that complies with the constraints. Since a continuous limit range is set for the discrete probability distribution, the decision unit 65 can efficiently determine an action that complies with the constraints by determining whether or not each action in the discrete probability distribution falls within the limit range.
[0115] When learning is performed through episodes, the decision unit 65 determines the action of agent A1 using a probability distribution for each of the first states of each step acquired by the first acquisition unit 61.
[0116] Furthermore, when the decision unit 65 determines multiple actions such as tank selection information, raw material quantity information, and type information, it is preferable to determine each action in stages. For example, the decision unit 65 determines the tank selection information using a probability distribution related to the tank selection information. Then, the decision unit 65 determines the type information using a probability distribution related to the type information obtained corresponding to the selected tank 24. Then, the decision unit 65 determines the raw material quantity information using a probability distribution related to the raw material quantity information obtained corresponding to the selected tank 24 and the type of raw material. This allows the decision unit 65 to determine a specific action, for example, to transport 100 kl of raw material G3 into tank 24a. When determining actions in stages in this way, for example, the output unit 63 may output a probability distribution related to the tank selection information, and the decision unit 65 may determine the tank selection information using the probability distribution related to the tank selection information. Subsequently, the output unit 63 may be configured to output a probability distribution related to the type information corresponding to the selected tank 24.
[0117] In this manner, when determining each action step by step in response to multiple actions such as tank selection information, raw material quantity information, and type information, the limiting unit 64 sets a limit based on constraint conditions for each probability distribution.
[0118] Furthermore, the limiting unit 64 may set a limit on the probability distribution related to the non-extensive behavior based on the result of setting a limit range for the probability distribution related to the extensive behavior associated with the non-extensive behavior. In other words, the result of the limit may be fed back when the behavior is decided step by step.
[0119] Figures 10A, 10B, and 10C illustrate an example of processing when multiple actions are decided in stages. Figures 10A to 10C show an example where a decision is made regarding tank selection information as an action, followed by a decision regarding raw material quantity information as another action. Tank selection is an example of an action that does not have extensive quantitative properties, while raw material quantity selection is an example of an action that does have extensive quantitative properties. Furthermore, raw material quantity selection is an action that has extensive quantitative properties related to tank selection.
[0120] First, the output unit 63 outputs a discrete probability distribution related to tank selection, as shown in Figure 10A. Then, the decision unit 65 determines the tank selection information based on the discrete probability distribution related to tank selection. The restriction unit 64 may set restrictions on the discrete probability distribution related to tank selection based on constraint information, etc. For example, the output unit 63 outputs discrete probability distributions corresponding to tanks 24a, 24b, and 24c, as shown in Figure 10A. Then, the decision unit 65 selects tank 24c based on the discrete probability distribution.
[0121] Next, the output unit 63 outputs a discrete probability distribution relating to the amount of raw material (e.g., discharge amount) corresponding to the tank 24c, as shown in Figure 10B. For example, the discrete actions corresponding to the discrete probability distribution relating to the selection of the discharge amount from tank 24c are discharge amounts T1, T2, T3, and T4. Then, the limiting unit 64 sets a continuous limiting range for the discrete probability distribution relating to the amount of raw material using constraint conditions. For example, if a continuous limiting range MS5 is set as shown in Figure 10B, there are cases where restrictions are set on all actions in the discrete probability distribution. That is, restrictions are set on all of the discharge amounts T1, T2, T3, and T4. In such a case, it is assumed that tank 24c should not be selected in the discrete probability distribution relating to tank selection. For this reason, the limiting unit 64 uses the result of setting a limiting range for the discrete probability distribution relating to the amount of raw material to set a restriction on the discrete probability distribution relating to tank selection.
[0122] Specifically, as shown in Figure 10C, the limiting unit 64 returns to the discrete probability distribution related to tank selection and sets a limit on tank 24c. This limit is shown, for example, as the limiting range MS6 in Figure 10C. As a result, the decision unit 65 can select a tank 24 other than tank 24c based on this discrete probability distribution. For example, the decision unit 65 selects tank 24b. The output unit 63 then outputs a discrete probability distribution related to the amount of raw material corresponding to tank 24b, and the limiting unit 64 sets the limit and the decision unit 65 determines the amount of raw material, etc.
[0123] In the above example, we used the case where restrictions are set on all actions and feedback is given regarding those restrictions as an example, but the feedback is not limited to cases where there are no action options left. For example, feedback may be given when setting restrictions reduces the number of action options to less than a predetermined number, or when setting restrictions leaves only low-probability actions as options.
[0124] Returning to Figure 5, the second acquisition unit 66 functions as an acquisition unit and acquires the state of the environment E1, which has changed as a result of the action determined by the decision unit 65, as the "second state". The second acquisition unit 66 also acquires the reward for the action determined by the decision unit 65.
[0125] Specifically, the second acquisition unit 66 observes the state of the tank base 21 as environment E1 after the action and sets it to the second state. The second acquisition unit 66 also acquires an evaluation of the action that changed the state of the tank base 21 from the first state to the second state as a reward.
[0126] In the loading and unloading problem, the second state includes, for example, the state of the tank 24 in the tank base 21 as environment E1.
[0127] The reward is an evaluation of the decided action, and in the loading / unloading problem, it is an evaluation of the action that changes the state of the tank base 21 from the first state to the second state. For example, the action is evaluated using a predetermined evaluation index, and a reward is set. The predetermined evaluation index is, for example, the raw material matching rate. However, as long as the goodness or badness of the action can be evaluated, the specific method of setting the reward is not limited. The more desirable the decided action, the larger the reward value.
[0128] When learning is performed through episodes, the second acquisition unit 66 acquires a second state and a reward corresponding to each step's action determined by the decision unit 65.
[0129] The calculation unit 67 calculates the loss based on the reward for the decided action. The calculation unit 67 uses the probability of the decided action and the reward for the decided action to calculate the loss. Specifically, the calculation unit 67 calculates the loss by multiplying the log probability related to the probability of the decided action and the advantage related to the reward for the decided action.
[0130] The logarithmic probability is determined by the probability of the decided action. For example, the logarithmic probability is the value obtained by transforming the probability of the decided action using a predetermined logarithmic function. For example, the calculation unit 67 performs the transformation such that when a probability less than a predetermined value is transformed into a logarithmic probability, the logarithmic probability becomes a negative value, and when a probability greater than or equal to the predetermined value is transformed into a logarithmic probability, the logarithmic probability becomes a positive value. As a result, the probability of the decided action becomes logarithmic and its sign (positive or negative) is set. For example, the lower the probability of the decided action is compared to the predetermined value, the more negative the logarithmic probability becomes and the larger its absolute value becomes. Also, for example, the higher the probability of the decided action is compared to the predetermined value, the more positive the logarithmic probability becomes and the larger its absolute value becomes.
[0131] The advantage is determined by the reward for the decided action. Specifically, the advantage is the difference between the reward for the action decided using the learning model and the reward for the action decided using the baseline model. The baseline model is a baseline model corresponding to the learning model. For example, the baseline model is the learning model of the initial state before the start of learning. The reward for the action decided using the baseline model is, like the learning model, the reward for the action decided corresponding to the first state. For example, the reward for the action decided using the baseline model is obtained using a simulation related to the loading and unloading problem, just like the learning model. Although the action decided using the learning model and the action decided using the baseline model are based on the same first state, the specific content of the decided actions may be the same or different.
[0132] The advantage is the degree of improvement (reward difference) between the behavior determined using the learning model and the behavior determined using the baseline model. For example, the advantage is set by subtracting the reward for the behavior determined using the baseline model from the reward for the behavior determined using the learning model. In this way, the advantage is an evaluation that corresponds to the learning model's progress from its initial state through learning. In other words, the higher the reward of the learning model compared to the reward of the baseline model, the higher the advantage.
[0133] The calculation unit 67 then calculates the loss as the product of the log probability, the advantage, and a constant. The constant is set to a negative value, for example, -1. As a result, the higher the probability of the action (positive log probability) and the higher the advantage, the larger the negative value of the loss, and the smaller the loss. This indicates that the desirable action is associated with a high probability, and the learning model is in a favorable state. Conversely, the lower the probability of the action (negative log probability) and the higher the advantage, the larger the positive value of the loss, and the larger the loss. This indicates that the desirable action is associated with a low probability, and the learning model is not in a favorable state. Thus, even if an action is desirable, the lower its probability, the larger the loss.
[0134] In the example above, the loss was calculated using the product of the log probability, the advantage, and a constant, but the method of calculating the loss is not limited to this.
[0135] The update unit 68 updates the learning model based on the results of actions taken in response to environment E1. The results of actions taken in response to environment E1 are the losses calculated by the calculation unit 67. In other words, the update unit 68 updates the learning model based on the losses. Specifically, the update unit 68 updates the learning model so that the loss is reduced. More specifically, the update unit 68 updates the weights and other components included in the neural network as a learning model so that the loss value of the updated learning model is smaller than the loss value of the learning model before the update. In other words, the update unit 68 updates the learning model so that the probability of favorable actions (actions with high rewards or advantages) is increased in the probability distribution output from the learning model.
[0136] In this embodiment, a limit range is set for the probability distribution. Therefore, it becomes possible to proceed with learning in response to actions that comply with the constraints in the probability distribution. That is, the learning model learns in response to actions that comply with the constraints, and also learns in a way that increases the probability of desirable actions among the actions that comply with the constraints.
[0137] For example, the update unit 68 updates the neural network as a learning model by backpropagation using loss. Note that the method for updating the learning model is not limited to backpropagation, and other methods may be used.
[0138] In this way, reinforcement learning (deep reinforcement learning) is performed in the learning unit 51. Using the learned model in this manner, inference is performed by the inference unit 52, which will be described later.
[0139] The inference unit 52 performs inference using the trained model. The trained model is the trained model that has been trained in the learning unit 51. In other words, the inference unit 52 uses the trained model trained by the learning unit 51 as the trained model to infer output data corresponding to the input data.
[0140] The inference unit 52 mainly comprises a reception unit 71, an output unit 72, a limiting unit 73, an estimation unit 74, and a data output unit 75.
[0141] The reception unit 71 receives input data. The input data is data that serves as a prerequisite for the loading and unloading problem. For example, the input data is data related to the planning of loading and unloading multiple types of raw materials to and from the tank base 21. Specifically, the input data is data related to initial inventory information, loading and unloading plan information, and constraint information. For example, the input data is set by the user as a prerequisite for the problem to be applied and is entered using the user terminal 2.
[0142] The output unit 72 uses the trained model to output a discrete probability distribution corresponding to the input data and relating to discrete variables. The output unit 72 is a functional unit that corresponds to (and has similar functionality to) the output unit 63 in the learning unit 51.
[0143] Figure 11 shows an example of processing using a trained model. Specifically, the output unit 72 inputs the input data into the trained model. The input data may be preprocessed to make it suitable for input into the trained model. The trained model then outputs a probability distribution. That is, the output unit 72 outputs a discrete probability distribution relating to discrete actions (variables). For example, the probability distribution will be a discrete probability distribution, similar to that in the learning stage, as shown in Figure 7. The output unit 72 may also set candidate actions and their effective ranges before outputting the probability distribution, similar to the setting unit 62 in the learning unit 51.
[0144] The restriction unit 73 sets a continuous restriction range for actions (variables) in a discrete probability distribution. The restriction unit 73 is a functional unit that corresponds to (has a similar function to) the restriction unit 64 in the learning unit 51. Specifically, the restriction unit 73 sets a continuous restriction range for a discrete probability distribution using constraint conditions. The restriction unit 73 sets a continuous restriction range for a discrete probability distribution, for example, using or by transforming the numerical range indicated by the constraint conditions. The restriction range is set as a continuous range, not as an individual restriction corresponding to each discrete action in the probability distribution. That is, the continuous restriction range is set as a continuous range that includes multiple discrete actions in the discrete probability distribution. Depending on the range indicated by the restriction range, it may also be a range that includes only one discrete action.
[0145] Among the discrete actions (variables) shown in the probability distribution, those that fall within the restricted range are restricted in their selection when selecting an action in the estimation unit 74, which will be described later.
[0146] The estimation unit 74 estimates output data corresponding to input data based on a discrete probability distribution with a continuous limit range. The estimation unit 74 is a functional unit that corresponds to (has similar functionality to) the decision unit 65 in the learning unit 51. Specifically, the estimation unit 74 estimates output data by determining the specific content of actions (variables) from a discrete probability distribution. The output data includes, for example, information on tank selection, raw material quantity, and type. In other words, the output data is information corresponding to the actions of agent A1 during the learning phase. That is, the actions (variables) determined based on the probability distribution become the solution to the problem. For example, corresponding to a probability distribution like the one in Figure 7, a specific value of the amount of material to be shipped is estimated as the specific content of the output data. The value of the amount of material to be shipped thus estimated becomes part of the plan corresponding to the loading and unloading problem. Other parts of the plan are similarly estimated using the trained model.
[0147] In this way, the estimation unit 74 estimates, in the loading and unloading problem, information regarding the loading and unloading plans (specific plans) for each of the multiple tanks 24 owned by the tank base 21, corresponding to the preconditions (overall plan), as output data. The output data may be subjected to post-processing.
[0148] When the estimation unit 74 performs inference using a trained model, it is preferable to definitively determine its actions using a probability distribution. That is, the estimation unit 74 prefers, for example, to select a variable (such as the amount of goods to be shipped) whose probability in the probability distribution is greater than or equal to a predetermined value, or to select the variable (such as the amount of goods to be shipped) with the highest probability in the probability distribution. However, if the problem involves multiple time-series stages, such as an import / export problem, and a variable with a high probability is not necessarily good in relation to other stages, the estimation unit 74 may also determine its actions probabilistically using a probability distribution.
[0149] Furthermore, when the estimation unit 74 performs inference in response to a predetermined episode, it estimates output data corresponding to each step. For example, for each step corresponding to the loading and unloading of raw materials, the loading and unloading plans for each raw material corresponding to the tank 24 are estimated. Then, the plans corresponding to each step are combined to form the overall plan. Note that the plans corresponding to each step are estimated in such a way that consistency is ensured between each step. For example, considering the loading plan for the tank 24 estimated in response to the loading step (considering the state of environment E1 when that loading plan is executed), the unloading plan for the tank 24 is estimated in response to the unloading step of the next step.
[0150] In this way, output data corresponding to the input data is inferred using the trained model.
[0151] The data output unit 75 outputs the output data estimated by the estimation unit 74. For example, the data output unit 75 outputs the output data as an inference result to a predetermined device such as a user terminal 2, providing the inference result to the user. This allows the user to recognize the loading and unloading plans for each of the multiple tanks 24 owned by the tank base 21.
[0152] The destination of the output data from the data output unit 75 is not limited. For example, if further processing is performed using the output data, the output data may be output to other functional units within the server device, or it may be output to a device different from the server device 3.
[0153] <Processing Flow> Figure 12 is a flowchart showing an example of the learning process flow according to this embodiment. Each of the following processes is started, for example, according to a user's instruction to start learning. In the learning process, each episode consists of M steps, and each step is associated with a number i (an integer from 1 to M). The order and content of each of the following steps can be changed as appropriate.
[0154] (Step SP10) The simulation unit 60 initializes various information related to the loading and unloading problem. For example, the state of the environment E1 and the weights of the neural network as a learning model are initialized. That is, each parameter related to the loading and unloading problem is set to its initial state. Alternatively, each hyperparameter related to learning may be set. Then, the process moves on to step SP11.
[0155] (Step SP11) The simulation unit 60 sets the step number i included in the episode to 1 (i=1). Then, the process moves on to step SP12.
[0156] (Step SP12) The simulation unit 60 starts the episode and executes step number i. Then, the process moves on to step SP13.
[0157] (Step SP13) The first acquisition unit 61 acquires the current state in environment E1 as the first state. That is, the first acquisition unit 61 acquires the state of environment E1 before agent A1's action in step number i as the first state. Then the process moves on to step SP14.
[0158] (Step SP14) The setting unit 62 sets the candidate actions. Specifically, the setting unit 62 sets the candidate actions along with the effective range. Then, the process moves on to step SP15.
[0159] (Step SP15) The output unit 63 inputs the first state acquired by the first acquisition unit 61 into the learning model and outputs a probability distribution corresponding to the first state. Specifically, the output unit 63 outputs a discrete probability distribution. Then, the process moves on to step SP16.
[0160] (Step SP16) The restriction unit 64 sets a restriction on the probability distribution based on the constraint conditions. Specifically, the restriction unit 64 sets a continuous restriction range for a discrete probability distribution. Then, the process moves on to step SP17.
[0161] (Step SP17) The decision unit 65 determines the action of agent A1 according to a probability distribution with set restrictions. That is, the decision unit 65 determines the action of agent A1 in step number i. Then the process moves on to step SP18.
[0162] (Step SP18) The simulation unit 60 reflects the determined actions of agent A1 into environment E1. Then, the process moves on to step SP19.
[0163] (Step SP19) The second acquisition unit 66 acquires the state that has changed due to agent A1's actions in environment E1 as the second state, and also acquires the reward for the action. That is, the second acquisition unit 66 acquires the second state and the reward corresponding to environment E1 after agent A1's action in step i. Then the process moves on to step SP20.
[0164] (Step SP20) The calculation unit 67 calculates the loss using the probability of the decided action and the reward of the decided action. That is, the calculation unit 67 calculates the loss corresponding to step number i. Then the process moves on to step SP21.
[0165] (Step SP21) The update unit 68 updates the learning model based on the loss. Then, the process moves on to step SP22.
[0166] (Step SP22) The simulation unit 60 determines whether the episode has finished. Specifically, the simulation unit 60 determines that the episode has finished if the step number i is the final number (i = M). If the episode has not finished, the process proceeds to step SP23. If the episode has finished, the process ends. Note that if another episode is to be executed, learning may be performed again in accordance with that other episode.
[0167] (Step SP23) The simulation unit 60 adds 1 to the number i. Then, the process moves to step SP12 and is executed again. In other words, each step is repeatedly executed until the episode ends.
[0168] In this way, deep reinforcement learning is performed. That is, a trained model is generated by training (updating) a neural network, which is an example of a learning model. The trained model updated in step SP20 may be evaluated for its learning status using test episodes, etc., and if the learning objective has not been achieved, relearning may be performed. Reearning may also be performed until a predetermined number of trials are reached.
[0169] Figure 13 is a flowchart showing an example of the inference process flow according to this embodiment. Each of the following processes is started, for example, in accordance with a user's instruction to start inference. Note that the order and content of each of the following steps can be changed as appropriate.
[0170] (Step SP30) The reception unit 71 receives input data from the user as preconditions for the problem. Then, the process moves on to step SP31.
[0171] (Step SP31) The output unit 72 inputs the input data into the trained model and outputs a probability distribution. This probability distribution is a discrete probability distribution. Then the process moves on to step SP32.
[0172] (Step SP32) The limiting unit 73 sets a continuous limiting range for the behavior (variable) with respect to the probability distribution. Then, the process moves on to step SP33.
[0173] (Step SP33) The estimation unit 74 estimates output data corresponding to the input data using a probability distribution. Then, the process moves on to step SP34.
[0174] (Step SP34) The data output unit 75 outputs the output data to the user terminal 2. Then the process ends.
[0175] <Effects> As described above, the server device 3 according to this embodiment includes: a first acquisition unit 61 that acquires a state corresponding to a predetermined environment E1; an output unit 63 that uses a learning model to output a discrete probability distribution that is an action for the environment E1 corresponding to the state and relates to discrete actions; a limiting unit 64 that sets a continuous limiting range for actions for the discrete probability distribution; a decision unit 65 that determines an action for the environment E1 based on the discrete probability distribution for which a continuous limiting range has been set; and an update unit 68 that updates the learning model based on the result of the action taken for the environment E1.
[0176] This configuration sets a continuous restriction range for discrete probability distributions. In other words, instead of a discrete restriction, a continuous restriction range is set for discrete probability distributions. For example, when setting discrete restrictions for discrete actions, it is necessary to perform simulations for each action to verify the feasibility of each action against the restriction. That is, setting discrete restrictions for discrete actions requires processing load for verification. In contrast, when setting a continuous restriction range for discrete actions, it is possible to omit the need to perform simulations for each action to verify the feasibility of each action. That is, in this case, the feasibility of each action can be determined by comparing each discrete action with the continuous restriction range. As a result, the processing load (computational load) when deciding on actions can be reduced compared to when discrete restrictions are set, while still setting restrictions.
[0177] Furthermore, using discrete probability distributions reduces the range of the action space to be explored, thereby improving processing efficiency.
[0178] In this way, reducing the processing load allows for more efficient decision-making. This enables improved processing efficiency in the computer (server device 3). In other words, the performance of the processing in the computer is enhanced.
[0179] Furthermore, in the server device 3, the continuous limit range is set as a continuous range that includes multiple discrete actions in a discrete probability distribution.
[0180] This configuration allows for more effective setting of constraints for discrete probability distributions by making the constraint range a continuous range that includes multiple discrete actions. In other words, by comparing each discrete action with the continuous constraint range, it is possible to increase the number of actions that can be excluded from the candidates. This further reduces the processing burden when making action decisions while setting constraints.
[0181] Furthermore, in the server device 3, the limiting unit 64 sets a continuous limiting range using predetermined constraint conditions.
[0182] This configuration allows for setting a limit range for discrete probability distributions while suppressing the processing load, by using predetermined constraints.
[0183] Furthermore, in server device 3, the actions taken in relation to environment E1 are extensive actions.
[0184] This configuration allows for the setting of a continuous limit range for discrete probability distributions corresponding to extensive actions. Although extensive actions can take continuous values, using them as discrete actions reduces the number of action options and thus the processing burden. Furthermore, even when the probability distribution corresponding to an extensive action is represented as a discrete probability distribution, a limit range can be set as a continuous range corresponding to that action, reducing the processing burden associated with setting the limit.
[0185] Furthermore, in the server device 3, the decision unit 65 determines an action for the environment E1 from among actions that are not included in a continuous limit range in a discrete probability distribution.
[0186] This configuration allows us to determine an action from among actions that are not included in the continuous limit range of a discrete probability distribution. In other words, we can determine an appropriate action by excluding actions that are included in the continuous limit range.
[0187] Furthermore, in the server device 3, the decision unit 65 determines the action to take in response to the environment E1 probabilistically using a probability distribution.
[0188] With this configuration, actions are determined probabilistically during the update (learning) phase, allowing for the learning of a wide range of patterns, including actions with low probability.
[0189] Furthermore, in the server device 3, the limiting unit 64 sets a limit on the probability distribution related to actions that do not have extensive properties, based on the result of setting a limit range for the probability distribution related to actions that do have extensive properties and actions that do not have extensive properties.
[0190] This configuration allows the results of restrictions on the probability distribution of actions with extensive properties to be reflected in restrictions on the probability distribution of actions without extensive properties, making it possible to make appropriate action decisions for actions without extensive properties.
[0191] Furthermore, the server device 3 includes a calculation unit 67 that calculates the loss based on the reward for the determined action, and the update unit 68 uses the resulting loss to update the learning model.
[0192] With this configuration, the learning model is updated to facilitate more desirable behaviors.
[0193] Furthermore, in the server device 3, the calculation unit 67 calculates the loss based on the probability of the decided action and the reward for the decided action.
[0194] With this configuration, the learning model is updated to increase the probability of more favorable behavior.
[0195] Furthermore, in the server device 3, the learning model takes data relating to at least one of the plans for loading and unloading multiple types of raw materials to and from the tank base 21 as input, and outputs data relating to at least one of the plans for loading and unloading each of the multiple tanks 24 that the tank base 21 has.
[0196] This configuration allows for training a learning model to address the issue of raw material loading and unloading.
[0197] Furthermore, the server device 3 uses the trained model as a trained model to infer output data corresponding to the input data.
[0198] With this configuration, inference can be performed using the trained model.
[0199] Furthermore, the server device 3 includes a reception unit 71 that receives input data, an output unit 72 that uses a trained model to output a discrete probability distribution corresponding to the input data and relating to discrete variables, a restriction unit 73 that sets a continuous restriction range for the variables in the discrete probability distribution, an estimation unit 74 that estimates output data corresponding to the input data based on the discrete probability distribution with the set continuous restriction range, and a data output unit 75 that outputs the estimated output data.
[0200] This configuration sets a continuous restriction range for discrete probability distributions. In other words, instead of a discrete restriction, a continuous restriction range is set for discrete probability distributions. For example, when setting discrete restrictions for discrete actions, it is necessary to perform simulations for each action to verify the feasibility of each action against the restriction. That is, setting discrete restrictions for discrete actions requires processing load for verification. In contrast, when setting a continuous restriction range for discrete actions, it is possible to omit the need to perform simulations for each action to verify the feasibility of each action. That is, in this case, the feasibility of each action can be determined by comparing each discrete action with the continuous restriction range. As a result, the processing load (computational load) when deciding on actions can be reduced compared to when discrete restrictions are set, while still setting restrictions.
[0201] In this way, by reducing the processing load, actions (variables) can be determined more efficiently. This makes it possible to improve the processing efficiency of the computer (server device 3). In other words, the performance of the processing in the computer performing the processing is improved.
[0202] <Modifications> This disclosure is not limited to the embodiments described above. That is, any modifications made to the embodiments described above by a person skilled in the art are also included in the scope of this disclosure, as long as they retain the features of this disclosure. Furthermore, the elements of the embodiments described above and the modifications described later can be combined to the extent that it is technically possible, and any combination thereof is also included in the scope of this disclosure, as long as it retains the features of this disclosure.
[0203] In the above embodiment, one example is that each function is provided by the server device 3, but each function may also be provided by the user terminal 2. Alternatively, each function may be distributed between the server device 3 and the user terminal 2. For example, the user terminal 2 may function as both a machine learning device and an inference device. Alternatively, one of the server device 3 and the user terminal 2 may function as a machine learning device, and the other of the server device 3 and the user terminal 2 may function as an inference device. Alternatively, two server devices 3 that can communicate with each other may be installed, with one server device 3 functioning as a machine learning device and the other server device 3 functioning as an inference device. Alternatively, multiple server devices 3 that can communicate with each other may be installed, with each function constituting the machine learning device and each function constituting the inference device distributed among multiple server devices 3, and these may function together as a machine learning device or an inference device.
[0204] Furthermore, while the above embodiment described the application of the problem of loading and unloading raw materials at the tank base 21 to the planning system 1 as an example, the problems that can be applied to the planning system 1 are not limited to those described above. For example, the problem of loading and unloading vehicles at a parking facility may be applied to the planning system 1. The problem of loading and unloading ships at a port may also be applied to the planning system 1. The problem of receiving and shipping goods and products at a warehouse or store may also be applied to the planning system 1. It should be noted that the problems that can be applied to the planning system 1 are not limited to those described above, and a variety of problems can be applied. In addition, while the planning system 1 can be applied to problems that involve multiple chronological stages, such as loading and unloading problems, it is also possible to apply problems that involve multiple non-chronological stages or problems that do not involve multiple stages (for example, problems with only one stage).
[0205] Furthermore, while the above embodiment illustrates a case where a discrete probability distribution is output for extensive actions and a continuous restriction range is set, it is also possible to output a discrete probability distribution for non-extensive actions and set a continuous restriction range. In this case as well, it is possible to omit the verification of the feasibility of each action by performing a simulation for each action. In other words, while setting restrictions, the processing burden (computational burden) when deciding on an action can be reduced compared to the case where discrete restrictions are set.
[0206] Furthermore, in the above embodiment, one example was the case where the learning model discretizes actions that have extensive properties and outputs a discrete probability distribution, but the timing of discretization is not limited to the above. For example, if the setting unit 62 sets the output amount as a candidate action and also sets the effective range, it may discretize the output amount with respect to the effective range. The output unit 63 may then output a discrete probability distribution corresponding to the discretized output amount. Alternatively, the output unit 63 may output a continuous probability distribution corresponding to the output amount and then convert this continuous probability distribution to a discrete probability distribution according to the output amount (variable). Thus, the timing of discretization is not limited. Similarly, the timing of discretization in the inference unit 52 is not limited.
[0207] Furthermore, in the above embodiment, the limiting unit 64 exemplified a case where it sets a continuous limiting range for a discrete probability distribution using constraint conditions and numerical ranges, or by transforming numerical ranges. However, the method of setting the limiting range is not limited to the above. For example, the limiting unit 64 may set the limiting range using simulation or the like. In other words, the limiting unit 64 may switch the method of setting the limiting range according to the constraint conditions. Similarly, the method of setting the limiting range in the inference unit 52 is not limited to the above.
[0208] Furthermore, in the above embodiment, the restriction unit 64 is shown as restricting actions included in the restriction range so that they are not selected when selecting an action, but it is not limited to this. For example, actions included in a discrete probability distribution may be modified so that they are not included in the restriction range. For example, suppose there are six output amounts as actions included in a discrete probability distribution: 10kl, 20kl, 30kl, 40kl, 50kl, and 60kl. Suppose the restriction range is set to 0kl to 35kl. In this case, the output amounts included in the restriction range may be modified so that they are not included in the restriction range and do not overlap with output amounts that were not originally included in the restriction range. For example, since the output amounts of 10kl, 20kl, and 30kl are included in the restriction range (0kl to 35kl), the output amount of 30kl may be modified to 36kl so that it is not included in the restriction range. This makes it possible to set a restriction range while suppressing a reduction in the number of action options. Similarly, the inference unit 52 is not limited to restricting actions included in the restriction range so that they are not selected.
[0209] The various types of information described in this disclosure (e.g., status, reward, etc.) may be expressed using absolute values, relative values from a given value, or other corresponding information.
[0210] In this disclosure, expressions such as "based on," "using," and "by" (including equivalent expressions) do not mean "based solely on," "using only," or "by" unless otherwise specified. In other words, the phrase "based on" means both "based solely on" and "at least on," and the same applies to equivalent expressions such as "using" and "by."
[0211] The term “decision” in this disclosure may encompass a wide variety of actions. “Decision” may include, for example, judgment, calculation, calculation, processing, derivation, investigation, exploration, and confirmation. Furthermore, “decision” may include, for example, considering something to have been “decided,” such as resolving, selecting, choosing, establishing, or comparing. In short, “decision” may include considering any action to have been “decided.”
[0212] In this disclosure, where expressions such as "obtain / set / use / based on" (including similar expressions) are used, unless otherwise specified, this includes cases where the information itself is used, or where the information has been processed in some way (e.g., noise-added, normalized, features extracted from the information, intermediate representation of the information, etc.). Furthermore, where it is stated that some result is obtained by "obtaining / setting / using / based on" (including similar expressions) (including similar expressions), unless otherwise specified, this includes cases where the result is obtained solely based on the information in question, or where the result is influenced by other information, factors, conditions, and / or states other than the information in question. Furthermore, where it is stated that "output" (including similar expressions), unless otherwise specified, this includes cases where the information itself is used as output, or where the information has been processed in some way (e.g., noise-added, normalized, features extracted from the information, intermediate representation of various types of information, etc.) is used as output.
[0213] Where terms meaning "containing" or "including" are used in this disclosure (e.g., "containing" or "including"), "having," etc.), they are intended as open-ended terms, including cases where the object of such term contains or possesses something other than the object indicated by the object of the term. Where the object of such terms meaning "containing" or "possessing" is an expression that does not specify a quantity or suggests a singular number (an expression with the article "a" or "an"), such expression should be interpreted as not being limited to a specific number.
[0214] In this disclosure, even if expressions such as "one or more" or "at least one" are used in some places, and expressions that do not specify a quantity or suggest singularity (expressions using the articles a or an) are used in other places, the latter expressions are not intended to mean "one." In general, expressions that do not specify a quantity or suggest singularity (expressions using the articles a or an) should be interpreted as not necessarily being limited to a specific number.
Claims
1. A machine learning device comprising: an acquisition unit that acquires a state corresponding to a predetermined environment; an output unit that uses a learning model to output a discrete probability distribution relating to an action on the environment corresponding to the state, which is also a discrete action; a restriction unit that sets a continuous restriction range for the action on the discrete probability distribution; a decision unit that determines the action on the environment based on the discrete probability distribution on which the continuous restriction range has been set; and an update unit that updates the learning model based on the result of performing the action on the environment.
2. The machine learning apparatus according to claim 1, wherein the continuous limit range is set as a continuous range that includes a plurality of discrete actions in the discrete probability distribution.
3. The machine learning apparatus according to claim 1 or 2, wherein the limiting unit sets a continuous limiting range using predetermined constraint conditions.
4. The machine learning apparatus according to claim 1 or 2, wherein the action taken in response to the environment is an extensive action.
5. The machine learning apparatus according to claim 1 or 2, wherein the decision unit determines the action for the environment from among the actions that are not included in the continuous limit range in the discrete probability distribution.
6. The machine learning apparatus according to claim 5, wherein the decision unit determines the action in relation to the environment probabilistically using the probability distribution.
7. The machine learning apparatus according to claim 1 or 2, wherein the limiting unit sets a limit on the probability distribution relating to the non-extensive behavior based on the result of setting the limiting range on the probability distribution relating to the extensive behavior that is related to the non-extensive behavior.
8. A machine learning device according to claim 1 or 2, further comprising: a calculation unit that calculates a loss based on the reward for the determined action, wherein the update unit updates the learning model using the loss as a result.
9. The machine learning apparatus according to claim 8, wherein the calculation unit calculates the loss based on the probability of the determined action and the reward for the determined action.
10. The machine learning device according to claim 1 or 2, wherein the learning model takes data relating to at least one of the plans for loading and unloading of multiple types of raw materials to and from a tank base as input, and outputs data relating to at least one of the plans for loading and unloading of each of the multiple tanks owned by the tank base.
11. The machine learning apparatus according to claim 1 or 2, wherein the environment corresponds to a raw material tank base, and the action is an action relating to at least one of loading and unloading the raw material at the tank base.
12. An inference device that uses the learned model, which has been trained by the machine learning device described in claim 1 or 2, as a trained model to infer output data corresponding to input data.
13. An inference device comprising: a receiving unit for receiving input data; an output unit that uses a trained model to output a discrete probability distribution corresponding to the input data and relating to discrete variables; a limiting unit that sets a continuous limiting range for the discrete probability distribution with respect to the variables; an estimation unit that estimates output data corresponding to the input data based on the discrete probability distribution with the continuous limiting range set; and a data output unit that outputs the estimated output data.
14. A machine learning method comprising: a step of acquiring a state corresponding to a predetermined environment; a step of using a learning model to output a discrete probability distribution relating to a discrete action that is an action for the environment corresponding to the state; a step of setting a continuous limit range for the action for the discrete probability distribution; a step of determining the action for the environment based on the discrete probability distribution for which the continuous limit range has been set; and a step of updating the learning model based on the result of performing the action for the environment.
15. An inference method comprising: a step of receiving input data; a step of using a trained model to output a discrete probability distribution corresponding to the input data and relating to discrete variables; a step of setting a continuous limit range for the discrete probability distribution with respect to the variables; a step of estimating output data corresponding to the input data based on the discrete probability distribution with the continuous limit range set; and a step of outputting the estimated output data.
16. A machine learning program that causes a computer to function as: an acquisition unit that acquires a state corresponding to a predetermined environment; an output unit that uses a learning model to output a discrete probability distribution that is an action for the environment corresponding to the state and relates to the discrete action; a restriction unit that sets a continuous restriction range for the discrete probability distribution and relates to the action; a decision unit that determines the action for the environment based on the discrete probability distribution with the continuous restriction range set; and an update unit that updates the learning model based on the result of performing the action for the environment.
17. An inference program that causes a computer to function as: a receiving unit that receives input data; an output unit that uses a trained model to output a discrete probability distribution corresponding to the input data and relating to discrete variables; a limiting unit that sets a continuous limiting range for the discrete probability distribution with respect to the variables; an estimation unit that estimates output data corresponding to the input data based on the discrete probability distribution with the continuous limiting range set; and a data output unit that outputs the estimated output data.