Machine learning device, inference device, machine learning method, inference method, machine learning program, and inference program

WO2026205144A1PCT designated stage Publication Date: 2026-10-01ENEOS HLDG INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/011946
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-27
Filing Date
2026-03-25
Publication Date
2026-10-01

Smart Images

  • Figure JP2026011946_01102026_PF_FP_ABST
    Figure JP2026011946_01102026_PF_FP_ABST
Patent Text Reader

Abstract

This machine learning device is provided with: an acquisition unit that acquires a state corresponding to a prescribed environment; an output unit that uses a learning model to output, from the state, a probability distribution (P1) relating to actions for the environment; a determination unit that determines an action for the environment on the basis of the probability distribution (P1); and an update unit that updates the learning model on the basis of the result of performing the action for the environment. The probability distribution (P1) has an asymmetric probability distribution shape respect to an action having a quantitative property.
Need to check novelty before this filing date? Find Prior Art

Description

Machine learning device, inference device, machine learning method, inference method, machine learning program and inference program

[0001] (CROSS-REFERENCE TO RELATED APPLICATION) The present application claims the benefit of priority from Japanese Patent Application No. 2025-053087, filed on March 27, 2025, the entire content of which is incorporated herein by reference. (Technical Field) The present disclosure relates to a machine learning device, an inference device, a machine learning method, an inference method, a machine learning program, and an inference program.

[0002] Various methods have been proposed to solve problems. For example, learning a model via machine learning to infer a solution to a problem has been devised.

[0003] For example, Patent Document 1 describes that a probability distribution is output from a state of an environment, and an action is selected using the probability distribution.

[0004] Japanese Unexamined Patent Publication No. 2023-37385

[0005] However, for example, when an action is an unloading amount, a loading amount, or the like, attempting to determine an action or the like using a probability distribution obtained from a model may result in excess or deficiency of the unloading amount, loading amount, or the like. As described above, for example, an action determined using a probability distribution obtained from a model may not be an appropriate action. Thus, determining an appropriate action has been a technical problem.

[0006] In view of the above problem, an object of the present disclosure is to provide a machine learning device, an inference device, a machine learning method, an inference method, a machine learning program, and an inference program that can facilitate determination of an appropriate action.

[0007] In order to solve the above problem, a machine learning device according to an aspect of the present disclosure includes: an acquisition unit that acquires a state corresponding to a predetermined environment; an output unit that outputs a probability distribution related to an action for the environment from the state using a learning model; a determination unit that determines the action for the environment based on the probability distribution; and an update unit that updates the learning model based on a result of performing the action on the environment, wherein the probability distribution has an asymmetric probability distribution shape with respect to the quantitative action.

[0008] An inference device according to another aspect of the present disclosure comprises: a receiving unit for receiving input data; an estimation unit for outputting a probability distribution corresponding to the input data using a trained model and estimating output data corresponding to the input data based on the probability distribution; and an output unit for outputting the estimated output data, wherein the probability distribution has a probability distribution shape that is asymmetric with respect to extensive variables.

[0009] A machine learning method according to another aspect of the present disclosure includes the steps of: acquiring a state corresponding to a predetermined environment; using a learning model to output a probability distribution relating to an action on the environment from the state; determining the action on the environment based on the probability distribution; and updating the learning model based on the result of performing the action on the environment, wherein the probability distribution has an extensive nature and an asymmetric probability distribution shape with respect to the action.

[0010] An inference method according to another aspect of the present disclosure includes the steps of: receiving input data; outputting a probability distribution corresponding to the input data using a trained model; estimating output data corresponding to the input data based on the probability distribution; and outputting the estimated output data, wherein the probability distribution has a probability distribution shape that is asymmetric with respect to extensive variables.

[0011] A machine learning program according to another aspect of this disclosure causes a computer to function as an acquisition unit that acquires a state corresponding to a predetermined environment, an output unit that uses a learning model to output a probability distribution relating to an action on the environment from the state, a decision unit that determines the action on the environment based on the probability distribution, and an update unit that updates the learning model based on the result of performing the action on the environment, wherein the probability distribution has an extensive nature and an asymmetric probability distribution shape with respect to the action.

[0012] An inference program according to another aspect of the present disclosure causes a computer to function as a receiving unit that receives input data, an estimation unit that outputs a probability distribution corresponding to the input data using a trained model and estimates output data corresponding to the input data based on the probability distribution, and an output unit that outputs the estimated output data, wherein the probability distribution has a probability distribution shape that is asymmetric with respect to extensive variables.

[0013] Furthermore, this disclosure may be implemented as a semiconductor integrated circuit that implements part or all of the program, as an information processing device, or as a system including an information processing device.

[0014] The machine learning apparatus, inference apparatus, machine learning method, inference method, machine learning program, and inference program related to this disclosure make it easier to determine appropriate actions.

[0015] This diagram schematically shows an example of the configuration of a planning system according to the embodiment of this disclosure. This diagram shows an example of raw material loading and unloading to and from a tank base. This diagram shows the relationship between the environment and the agent related to machine learning. This diagram schematically shows an example of the hardware configuration of a server device. This block diagram shows an example of various functions in the server device. This diagram shows an example of loading and unloading plan information. This diagram shows an example of a probability distribution corresponding to the amount of material unloaded. This diagram shows an example of constraint conditions in the limiting unit. This diagram shows an example of a case where a limit is set on the probability distribution. This diagram shows an example of processing in the estimation unit. This flowchart shows an example of the learning process flow. This flowchart shows an example of the inference process flow.

[0016] The following descriptions illustrate some aspects of this disclosure.

[0017] A machine learning apparatus according to a first aspect of this disclosure includes: an acquisition unit that acquires a state corresponding to a predetermined environment; an output unit that uses a learning model to output a probability distribution relating to an action in relation to the environment from the state; a decision unit that determines the action in relation to the environment based on the probability distribution; and an update unit that updates the learning model based on the result of performing the action in relation to the environment, wherein the probability distribution has an extensive nature and an asymmetric probability distribution shape with respect to the action.

[0018] In the machine learning apparatus according to the second aspect of this disclosure, the extensive behavior according to the first aspect is a behavior that changes the quantity in the environment.

[0019] In the machine learning apparatus according to the third aspect of this disclosure, in the first and second aspects, the probability distribution is a continuous probability distribution for a continuous variable relating to the behavior having extensive properties.

[0020] In the machine learning apparatus according to the fourth aspect of this disclosure, in the first to second aspects, the probability distribution has different probability distribution shapes in the region where it is smaller than the mode of the extensive behavior and in the region where it is larger than the mode.

[0021] In the machine learning apparatus according to the fifth aspect of this disclosure, as per the fourth aspect, the area of ​​the region smaller than the mode of the probability distribution is larger than the area of ​​the region larger than the mode of the probability distribution.

[0022] In the machine learning apparatus according to the sixth aspect of this disclosure, as per the fourth aspect, the probability distribution increases in the region smaller than the mode, and decreases in the region larger than the mode, at a rate of change greater than the increase.

[0023] In the machine learning apparatus according to the seventh aspect of this disclosure, relating to the first to second aspects, the decision unit probabilistically determines the action in relation to the environment using the probability distribution.

[0024] A machine learning apparatus according to the eighth aspect of this disclosure further comprises a calculation unit that calculates a loss based on the reward for the determined action, according to the first to second aspects, and the update unit updates the learning model using the loss as a result.

[0025] In the machine learning apparatus according to the ninth aspect of this disclosure, as per the eighth aspect, the update unit updates the asymmetry of the probability distribution output from the learning model when updating the learning model with the loss.

[0026] In the machine learning apparatus according to the tenth aspect of this disclosure, as per the eighth aspect, the calculation unit calculates the loss based on the probability of the determined action and the reward for the determined action.

[0027] A machine learning apparatus according to the eleventh aspect of this disclosure further comprises a restriction unit that sets restrictions on the probability distribution relating to the action, relating to the first to second aspects, and the determination unit determines the action from within the range of the probability distribution for which no restrictions are set.

[0028] In the machine learning apparatus according to the twelfth aspect of this disclosure, according to the first to second aspects, the learning model takes as input data relating to at least one of the plans for loading and unloading of multiple types of raw materials to and from a tank base, and outputs data relating to at least one of the plans for loading and unloading of each of the multiple tanks owned by the tank base.

[0029] In the machine learning apparatus according to the thirteenth aspect of this disclosure, relating to the first to second aspects, the environment corresponds to a raw material tank base, and the action is an action relating to at least one of the loading and unloading of the raw material at the tank base.

[0030] The inference device according to the 14th aspect of this disclosure uses the learned model, which has been learned by the above-described machine learning device, as a trained model to infer output data corresponding to input data.

[0031] An inference device according to a 15th aspect of this disclosure comprises: a receiving unit for receiving input data; an estimation unit for outputting a probability distribution corresponding to the input data using a trained model and estimating output data corresponding to the input data based on the probability distribution; and an output unit for outputting the estimated output data, wherein the probability distribution has a probability distribution shape that is asymmetric with respect to an extensive variable.

[0032] A machine learning method according to a sixteenth aspect of this disclosure includes the steps of: acquiring a state corresponding to a predetermined environment; using a learning model to output a probability distribution relating to an action in relation to the environment from the state; determining the action in relation to the environment based on the probability distribution; and updating the learning model based on the result of performing the action in relation to the environment, wherein the probability distribution has an extensive nature and an asymmetric probability distribution shape with respect to the action.

[0033] An inference method according to a 17th aspect of this disclosure comprises the steps of: receiving input data; outputting a probability distribution corresponding to the input data using a trained model; estimating output data corresponding to the input data based on the probability distribution; and outputting the estimated output data, wherein the probability distribution has a probability distribution shape that is asymmetric with respect to extensive variables.

[0034] A machine learning program according to the 18th aspect of this disclosure causes a computer to function as an acquisition unit that acquires a state corresponding to a predetermined environment, an output unit that uses a learning model to output a probability distribution relating to an action on the environment from the state, a decision unit that determines the action on the environment based on the probability distribution, and an update unit that updates the learning model based on the result of performing the action on the environment, wherein the probability distribution has an extensive nature and an asymmetric probability distribution shape with respect to the action.

[0035] An inference program according to a 19th aspect of this disclosure causes a computer to function as a receiving unit that receives input data, an estimation unit that outputs a probability distribution corresponding to the input data using a trained model and estimates output data corresponding to the input data based on the probability distribution, and an output unit that outputs the estimated output data, wherein the probability distribution has a probability distribution shape that is asymmetric with respect to an extensive variable.

[0036] The following description illustrates embodiments of the present disclosure. To facilitate understanding of the description, the same components and steps are denoted by the same reference numerals in each drawing whenever possible, and redundant descriptions are omitted.

[0037] <Overall Configuration> Figure 1 is a schematic diagram showing an example of the configuration of a planning system 1 according to one embodiment. The planning system 1 is a system that creates a plan as a solution corresponding to a pre-set problem.

[0038] As shown in Figure 1, the planning system 1 consists of a user terminal 2 and a server device 3. The server device 3 and the user terminal 2 can communicate with each other via the network NT.

[0039] User terminal 2 is a terminal device, an information processing device (computer) used by the user. User terminal 2 is, for example, a personal computer. The user can input various information using user terminal 2 and instruct server device 3 to create a plan.

[0040] Server device 3 is an information processing device (computer) that creates a plan to address a problem using information input by user terminal 2. In this embodiment, one example of a problem is a problem related to the loading (inbound) and unloading (shipping) of raw materials to and from the tank base 21, which will be described later. In this embodiment, the problem of planning the loading and unloading of raw materials to and from the tank base 21 is referred to as the "loading and unloading problem." Note that the problem to be solved is not limited to the loading and unloading problem to and from the tank base 21.

[0041] Figure 2 shows an example of raw material being brought into and out of the tank base 21.

[0042] As shown in Figure 2, the tank terminal 21 is a terminal that temporarily stores raw materials. For example, raw materials are carried into the tank terminal 21 by using a ship F1 or the like, and raw materials are carried out from the tank terminal 21 by using a ship F2 or the like. Note that the means for carrying in and carrying out are not limited to the ship F1 and the ship F2. The tank terminal 21 is provided with a plurality of tanks 24. In the example shown in Figure 2, the tank terminal 21 includes, as the tanks 24, a tank 24a, a tank 24b, and a tank 24c. Note that the number of tanks 24 that the tank terminal 21 has is not limited. The tank terminal 21 can store carried-in raw materials in each tank 24. In addition, the tank terminal 21 can carry out raw materials stored in each tank 24 from each tank 24. A plurality of types of raw materials may be carried into the tank terminal 21 and stored in each tank 24. Note that different types of raw materials may be carried into the plurality of tanks 24 respectively in the tank terminal 21, or a plurality of types of raw materials may be carried into one tank 24. Each of the plurality of types of raw materials stored in each tank 24 can be carried out from each tank 24. Note that, in the present embodiment, as an example of the carry-in / carry-out problem, the problem relates to both carrying in and carrying out raw materials with respect to the tank terminal 21, but the problem may also relate to one of carrying in raw materials and carrying out raw materials with respect to the tank terminal 21.

[0043] In the present embodiment, the server device 3 creates a specific plan for carrying in and carrying out each raw material to and from each tank 24 in the tank terminal 21, based on a general plan (outline plan) for carrying in and carrying out a plurality of types of raw materials with respect to the tank terminal 21. That is, the outline plan serves as a precondition for creating the specific plan. In other words, the outline plan is a plan for the entire tank terminal 21, and the specific plan is a detailed plan within the tank terminal 21.

[0044] The server device 3 creates the plan by using machine learning. Therefore, the server device 3 corresponds to a machine learning device when performing learning, and corresponds to an inference device when performing inference (creating a plan).

[0045] Figure 3 is a diagram showing an outline of machine learning.

[0046] As shown in FIG. 3, in the present embodiment, the server device 3 performs reinforcement learning. The server device 3 performs, for example, deep reinforcement learning. Reinforcement learning is a learning method in which an environment E1 and an agent A1 interact with each other to complete a task (a solution to a problem), thereby learning what kind of action the agent A1 should take to receive more rewards (evaluations). The agent A1 is the subject of actions. An action is the behavior of the agent A1. The environment E1 is not only the object on which the agent A1 takes actions, but also a prerequisite for the agent A1. That is, the state of the environment E1 is the state where the agent A1 is placed in the environment E1. The state of the environment E1 changes according to the action that the agent A1 takes on the environment E1. In other words, an action can be said to be a factor that changes the environment E1. The agent A1 is presented with a reward corresponding to the action (an evaluation of the action). When the agent A1 executes a preferable action in the environment E1, more rewards are presented to the agent A1 compared to when the agent A1 executes an unpreferable action in the environment E1. Note that the reward evaluation method can be appropriately set according to the learning objective and the like.

[0047] As shown in operation S1, the agent A1 takes an action on the environment E1 according to the state of the environment E1. As a result, the state of the environment E1 changes to a new state according to the action taken by the agent A1. Further, as shown in operation S2, the agent A1 is presented with the new state of the environment E1 and a reward corresponding to the action taken by the agent A1.

[0048] Then, in reinforcement learning, learning is performed based on the interaction between the environment E1 and the agent A1 generated by repeating operation S1 and operation S2. For example, when learning is performed by an episode having a plurality of steps, interaction between the environment E1 and the agent A1 such as operation S1 and operation S2 is executed corresponding to each step. Note that an episode is a flow (period) from the start to the end of a task to be solved in reinforcement learning.

[0049] Furthermore, the action that agent A1 takes in operation S1 is determined based on policy W1. Policy W1 is a rule (policy) that serves as an indicator for determining agent A1's action based on the state of environment E1 before the action. For example, policy W1 is information that associates multiple patterns of action that agent A1 can take in response to the state of environment E1 with the probability (selection probability) that each action will be performed. In this embodiment, the information that associates multiple patterns of action with the probability that each action will be performed is referred to as a "probability distribution". For example, reinforcement learning aims to adjust the probability distribution as policy W1 so that agent A1 can perform a more favorable action in the state of environment E1.

[0050] In deep reinforcement learning, a neural network is used to obtain a policy W1 that corresponds to the state of environment E1. That is, the neural network outputs information about the action of agent A1 based on the state of environment E1. The neural network is then updated (learned) to maximize rewards (minimize losses), for example. In other words, the neural network is the object of learning and is an example of a "learning model." It should be noted that learning models are not limited to neural networks and can be applied in various ways depending on the reinforcement learning method and type of problem chosen by the user.

[0051] <Hardware Configuration> Figure 4 is a schematic diagram showing an example of the hardware configuration of server device 3.

[0052] As shown in Figure 4, the server device 3 comprises a control device 40, a communication device 41, and a storage device 42. The control device 40 is mainly composed of a processor 43 and a memory 44.

[0053] In the control device 40, the processor 43 executes a predetermined program stored in the memory 44 or storage device 42, thereby functioning as one of the various functional configurations described later.

[0054] The processor 43 is, for example, a CPU (Central Processing Unit). However, the processor 43 is not limited to a CPU. The processor 43 may be, for example, a GPU (Graphics Processing Unit) or an NPU (Neural Processing Unit). In a specific example, the processor 43 is a multi-core processor. The processor 43 may be a single-core processor. The processor 43 may include multiple processors or cores and be capable of performing parallel processing. The processor 43 is configured to execute computer programs. The processor 43 may, for example, include an ASIC (Application Specific Integrated Circuit) in part, or it may include programmable hardware such as an FPGA (Field Programmable Gate Array) or a CPLD (Complex Programmable Logic Device) in part.

[0055] The memory 44 is a computer-readable storage medium and may consist of at least one of the following: RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), etc. The memory 44 can store various types of data, including programs necessary for executing processing in the server device 3.

[0056] The communication device 41 consists of a communication interface and the like for communicating with an external device. For example, the communication device 41 is capable of communicating with the user terminal 2.

[0057] The storage device 42 is, for example, a computer-readable instruction recording medium (a non-temporary computer-readable instruction recording medium) and is composed of a hard disk, a solid-state drive, or the like. The storage device 42 stores various programs and information necessary for executing processing in the control device 40, as well as information on the processing results. Other examples of non-temporary computer-readable instruction recording media include portable recording media such as magnetic tapes, flexible disks, optical disks, digital versatile disks, Blu-ray discs, magneto-optical disks, memory cards, and USB memory.

[0058] The server device 3 may consist of a single information processing device or multiple information processing devices. Furthermore, Figure 4 only shows a part of the main hardware configuration of the server device 3, and the server device 3 may have other configurations. For example, the server device 3 may further include an input device (not shown) and a display device (not shown). The input device is an input device that receives input from the outside (e.g., a keyboard, mouse, etc.). The input device receives user operations and inputs those operations to the server device 3. The display device is a display device that performs output to the outside (e.g., a display, etc.). The display device outputs characters and images. The server device 3 may have the input device and output device integrated (e.g., a touch panel). The user terminal 2 may also have a configuration similar to the server device 3 as an information processing device, including a control device (processor and memory), a communication device, a storage device, etc.

[0059] <Functional Configuration> Figure 5 is a block diagram showing an example of various functions in the server device 3. Various processes are executed according to the functions in each block. Computer programs that implement the functions of at least some of the functional blocks shown in Figure 5 may be installed in the storage of one or more computers. The processors of one or more computers may perform the functions of multiple functional blocks shown in Figure 5 by reading the computer programs installed on their own machines into main memory and executing them.

[0060] Furthermore, the functions of each functional block shown in Figure 5 may be executed by a single computer, or they may be executed in a distributed manner across multiple computers. When the functions of each functional block shown in Figure 5 are executed in a distributed manner across multiple computers, these multiple computers may send and receive data via a communication network including a LAN (Local Area Network), a WAN (Wide Area Network), or the Internet.

[0061] As shown in Figure 5, the server device 3 has a functional configuration that mainly consists of a learning unit 51 and an inference unit 52. In the server device 3, the learning unit 51 functions as a machine learning device, and the inference unit 52 functions as an inference device.

[0062] The learning unit 51 performs learning (deep reinforcement learning) on ​​the learning model. The learning unit 51 mainly comprises a simulation unit 60, a first acquisition unit 61, an output unit 62, a limiting unit 63, a decision unit 64, a second acquisition unit 65, a calculation unit 66, and an update unit 67.

[0063] The simulation unit 60 executes a simulation related to the loading and unloading problem. For example, with respect to the loading and unloading problem, environment E1 becomes a predetermined environment corresponding to the simulation model of the tank base 21. Agent A1 is a virtual entity that performs actions such as loading and unloading at the tank base 21. Agent A1 executes an action determined by the state of the tank base 21. An episode is the entire loading and unloading problem (the whole process), and the loading and unloading of raw materials to and from the tank base 21 each corresponds to a step (stage). That is, loading included in the loading and unloading problem corresponds to one step, unloading included in the loading and unloading problem corresponds to another step, and unloading and loading each correspond to different steps. In this way, the loading and unloading problem includes multiple steps arranged in chronological order. The learning model is trained so that it can determine the best action (or the closest to the best action) for the loading and unloading problem.

[0064] Specifically, the learning model can accept input data related to the planning of loading and unloading multiple types of raw materials to and from the tank base 21. In other words, the input data is information related to the overall plan for the loading and unloading problem. For example, the input data includes at least one of initial inventory information, loading and unloading plan information, and constraint information. The initial inventory information, loading and unloading plan information, and constraint information are preconditions for the loading and unloading problem.

[0065] The initial inventory information shows the inventory status of each tank 24 in the initial state of the loading and unloading problem. The inventory status is the type and quantity of raw materials stored in each tank 24. In other words, the initial inventory information is information that associates tank 24 (name, identification information, etc.) with the type of raw materials stored and the quantity (weight or percentage) stored.

[0066] The loading and unloading plan information consists of the loading and unloading plans included in the loading and unloading problem. In other words, the loading and unloading plan information, in particular, constitutes the overall plan related to the loading and unloading problem.

[0067] Figure 6 shows an example of loading and unloading plan information.

[0068] As shown in Figure 6, the loading and unloading plan information is associated with attributes, dates (or sequences), and loading or unloading quantities for each type of raw material. Attributes indicate loading or unloading. Loading or unloading quantities for each type of raw material indicate the amount of raw material loaded into or unloaded from the tank base 21. In other words, the loading and unloading plan information does not limit individual loading or unloading to tanks 24, but shows the loading and unloading plan for the entire tank base 21. In Figure 6, loading or unloading quantities for each type of raw material are shown by the quantity of each type of raw material corresponding to the attribute (loading or unloading). The example in Figure 6 shows a case where six types of raw materials, G1, G2, G3, G4, G5, and G6, are handled in the loading and unloading problem. Note that the loading and unloading plan information may also be associated with identification information (name, etc.) of the vessel involved in the loading or unloading.

[0069] In the loading and unloading problem, loading and unloading in the loading and unloading plan information correspond to steps, forming an episode as a whole.

[0070] The constraint information indicates the constraints in the loading and unloading problem. For example, the constraint information includes at least one of the following: upper and lower limit constraints on the inventory in tank 24, constraints on the number of tanks 24 used during loading and unloading, minimum quantities related to loading and unloading, concentration constraints, and tank internal layer separation constraints. The concentration constraint is a constraint on the combination and proportion (concentration) of multiple types of raw materials stored in one tank 24. The tank internal layer separation constraint is information that restricts the combination and order in which multiple types of raw materials are stored to prevent separation of multiple types of raw materials within tank 24. The constraint information may also include upper and lower limit constraints on the API (American Petroleum Institute) specific gravity of tank 24. Here, API specific gravity refers to the specific gravity of crude oil as defined by the American Petroleum Institute. API specific gravity is a value that can be measured, for example, in accordance with ASTM D1298. In this embodiment, API specific gravity may be simply referred to as "API".

[0071] Furthermore, when training a learning model, it is preferable to perform the training using various patterns of input data (prerequisites).

[0072] The learning model then provides output data regarding the loading and unloading plans for each of the multiple tanks 24 located within the tank base 21. In other words, the output data represents a specific plan within the tank base 21 that corresponds to the preconditions (overall plan, etc.).

[0073] For example, the output data includes information such as tank selection information, raw material quantity information, and type information. Tank selection information is information about the tank 24 that is operated on in relation to loading or unloading. That is, the tank 24 that is the loading destination or unloading source is selected in the tank selection information. Note that the tank selection information may also be information about the tank 24 involved in the transfer of raw materials from one tank 24 to another tank 24 within the tank base 21 (raw material shift). Raw material quantity information is information about the amount of raw material to be loaded into or unloaded from the selected tank 24, corresponding to the selected tank 24. Type information is information about the type of raw material to be loaded into or unloaded from the selected tank 24, corresponding to the selected tank 24. For example, the output data is shown as loading 100 kl of raw material G3 into tank 24a. Note that if the loading / unloading problem involves multiple steps (loading or unloading), tank selection information, raw material quantity information, and type information are output corresponding to each step.

[0074] Furthermore, the output data may include evaluation indicators corresponding to the output tank selection information, raw material quantity information, and type information. Examples of evaluation indicators include raw material agreement rate, raw material group agreement rate, and API error. The raw material agreement rate is the ratio of the planned quantity (actual quantity) of raw material to the required quantity (ideal quantity) for each type. For example, in raw material shipment, the raw material agreement rate is the ratio of the quantity of raw material to be shipped to the quantity of raw material required for shipment. When calculating the raw material agreement rate for multiple types, for example, the average of the raw material agreement rates for each type of raw material may be used. The raw material group agreement rate is the ratio of the planned quantity (actual quantity) to the required quantity (ideal quantity) of raw material for each group. A group is, for example, a group of multiple types of raw materials with similar properties. For example, in raw material shipment, the raw material group agreement rate is the ratio of the total quantity of raw material from the same group to the total quantity of raw material from that group to the total quantity of raw material from that group required for shipment. When calculating the raw material group agreement rate for multiple groups, for example, the average of the raw material agreement rates for each group may be used. API error is the difference between the API of the required raw material (e.g., the average of the APIs of various raw materials) and the API of the planned raw material (e.g., the average of the APIs of various raw materials). For example, in the shipment of raw materials, the API error represents the difference between the API of the raw material required for shipment and the API of the raw material scheduled to be shipped.

[0075] In this way, the simulation unit 60 is capable of executing simulations related to the loading and unloading problem in accordance with the learned model.

[0076] Returning to Figure 5, the first acquisition unit 61 has the function of an acquisition unit and acquires the current state in environment E1 as the "first state". That is, the first acquisition unit 61 acquires the first state of environment E1 in which agent A1 is located before agent A1 takes action. Specifically, the first acquisition unit 61 observes the state of the tank base 21 as environment E1 and sets it as the first state.

[0077] In the loading and unloading problem, the first state includes, for example, the state of the tanks 24 in the tank base 21 as environment E1. The state of the tanks 24 refers to information such as the amount of raw materials stored in each of the tanks 24 and the types of raw materials stored.

[0078] When learning is performed through an episode, the first acquisition unit 61 acquires a first state corresponding to each of the multiple steps included in the episode.

[0079] The output unit 62 uses a learning model to output a probability distribution related to actions in environment E1 based on the state of environment E1. In other words, the output unit 62 uses a learning model to output a probability distribution of actions corresponding to the first state.

[0080] A probability distribution is probability information that associates multiple actions that agent A1 can perform with a given first state, along with their probabilities (selection probabilities). A probability distribution can be represented, for example, by a probability function. This probability function may also be a probability density function. In a probability distribution, actions are represented, for example, by variables. These variables are also called random variables.

[0081] Behavior can be divided into descriptive behavior and non-descriptive behavior. Descriptive behavior can also be called descriptive behavior or descriptive variable (a variable that possesses descriptive properties). Non-descriptive behavior is descriptive behavior, and can also be called descriptive behavior or descriptive variable.

[0082] Extensive behaviors are those related to variables that are proportional to changes in the size of the system. Furthermore, extensive behaviors can also be described as those related to variables that depend on quantity. In other words, extensive behaviors can be described as actions that change the quantity (state of quantity) in environment E1. Moreover, extensive behaviors can be described as actions that can take on continuous values.

[0083] In this embodiment, actions include, for example, selecting a tank 24 related to tank selection information, selecting a quantity of raw material related to raw material quantity information, and selecting a type of raw material related to type information. Therefore, raw material quantity information (selection of raw material quantity) is an extensive action. Specifically, the amount of raw material brought into tank 24 (input amount) and the amount of raw material discharged from tank 24 (discharge amount) are examples of extensive actions. In contrast, tank selection information and type information are examples of actions that do not have extensive properties.

[0084] In this embodiment, an extensive action will be explained using "discharge amount" as an example. A probability distribution is a function that associates the probability of a random variable taking a particular value with the value of that random variable. Discharge amount indicates the amount of raw material discharged from a certain tank 24 in the tank base 21. The probability distribution corresponding to the discharge amount will be specifically referred to as "probability distribution P1".

[0085] Figure 7 shows an example of a probability distribution P1 corresponding to the output quantity as an example of an action. In Figure 7, the vertical axis represents probability and the horizontal axis represents the output quantity (an example of a random variable). As shown in equation (1) described later, the probability corresponds to "P" and the output quantity corresponds to "x". The probability distribution P1 is a continuous probability distribution for a continuous variable. That is, the output quantity is a continuous variable, and correspondingly the probability is also a continuous value. Note that since the variable corresponding to an extensive action is continuous, the probability distribution corresponding to an extensive action is a continuous probability distribution.

[0086] In this embodiment, for example, the probability distribution P1 corresponding to the output amount, as shown in probability distribution P1, has an asymmetrical probability distribution shape. Specifically, the probability distribution P1 does not have a symmetrical distribution shape with respect to the mode X1 of the output amount, but rather an asymmetrical distribution shape with respect to the mode X1. The mode X1 is the value of the variable (output amount) for which the probability is maximized. The mode X1 is learned to approach, for example, the value of the variable (output amount) for which the reward obtained by agent A1 is maximized. Specifically, the probability distribution P1 has a region R1 in which the output amount is smaller than the mode X1, and a region R2 in which the output amount is larger than the mode X1. The probability distribution P1 has an asymmetry such that the probability distribution shapes are different in region R1 and region R2. In particular, in region R1, the probability of the probability distribution P1 increases as the output amount increases toward the mode X1. In region R2, the probability of the probability distribution P1 decreases as the output amount increases. The probability distribution P1 has a point B1 where, as the output volume increases, the probability decreases at a larger rate of change than the increase in probability in region R1. In the example in Figure 7, the probability distribution P1 shows that at point B1 in region R2, immediately after the output volume exceeds the mode X1, the probability decreases sharply as the output volume increases. Furthermore, in region R2, the probability distribution P1 has a point (point B2) where the probability decreases more slowly as the output volume increases compared to point B1.

[0087] In the probability distribution P1 illustrated in Figure 7, the output quantity value corresponding to location B1 is closer to the mode X1 than, for example, the output quantity value corresponding to location B2. In other words, the absolute difference between the output quantity value corresponding to location B1 and the mode X1 is smaller than the absolute difference between the output quantity value corresponding to location B2 and the mode X1.

[0088] In the probability distribution P1 illustrated in Figure 7, location B1 of probability distribution P1 is closer to region R1 than location B2 of probability distribution P1. In other words, location B1 of probability distribution P1, which is a continuous probability distribution, is continuous with the probability corresponding to the mode X1. Location B2 of probability distribution P1, which is a continuous probability distribution, is continuous with location B1.

[0089] In the probability distribution P1 illustrated in Figure 7, it is preferable that the area M1 of region R1 where the output amount is less than the mode X1 is greater than the area M2 of region R2 where the output amount is greater than the mode X1. The area corresponds to the area between a predetermined value of probability (e.g., 0) and the curve of the probability distribution. In other words, areas M1 and M2 can be obtained, for example, by using a definite integral with respect to the probability distribution P1 expressed by equation (1) described later, based on the value of the random variable.

[0090] The probability distribution P1 illustrated in Figure 7 is not symmetric with respect to the mode X1 (the line x = X1) in a two-dimensional space represented by P as a probability and x as a random variable (output). Thus, the probability distribution P1 illustrated in Figure 7 is asymmetric with respect to the output (an example of a random variable). Furthermore, the probability distribution shape of the probability distribution P1 illustrated in Figure 7 differs between the region R1, which is smaller than the mode X1, and the region R2, which is larger than the mode X1. In addition, the probability distribution P1 illustrated in Figure 7 has a probability distribution shape that is asymmetric with respect to extensive behavior. Furthermore, the rate of change (absolute value) of the decrease in probability at point B1 in region R2 is greater than the rate of change (absolute value) of the increase in probability in region R1. Furthermore, the rate of change (absolute value) of the decrease in probability at point B2 in region R2 is greater than the rate of change (absolute value) of the decrease in probability at point B1 in region R2.

[0091] In the probability distribution P1 illustrated in Figure 7, when agent A1's action is probabilistically determined (selected), for example, the probability of selecting the mode X1 (e.g., the amount of material dispensed that maximizes the reward) is the highest. Also, for example, the probability of selecting an action (amount of material dispensed) belonging to region R1 is higher than the probability of selecting an action (amount of material dispensed) belonging to region R2. That is, in probability distribution P1, the area M1 of region R1, which is smaller than the mode X1, is larger than the area M2 of region R2, which is larger than the mode X1. As a result, the probability of selecting an action that exceeds the amount of raw materials stored in tank 24 (e.g., the amount of material dispensed belonging to location B2), and immediately failing to answer, can be reduced compared to using a probability distribution symmetrical with respect to the mode X1.

[0092] The asymmetrical probability distribution P1 illustrated in Figure 7 ensures that the probability P3 of selecting an action that agent A1 may choose (e.g., output amount X3) is higher than the probability P2 of selecting an action that agent A1 should not choose (e.g., output amount X2 belonging to location B2) in answering the problem. Note that output amount X3 is smaller than the mode X1 by the absolute value of the difference between the mode X1 and output amount X2. This reduces the probability of selecting an action that exceeds the amount of raw materials stored in tank 24 (e.g., output amount belonging to location B2), resulting in immediate failure to answer the problem, compared to using a probability distribution symmetrical with respect to the mode X1.

[0093] The asymmetrical probability distribution P1 illustrated in Figure 7 shows that the probability P5 of selecting an action that agent A1 may choose (e.g., output quantity X5) is higher than the probability P4 of selecting an action that agent A1 should not choose in answering the problem (e.g., output quantity X4 belonging to location B1). Note that output quantity X5 is an output quantity that is smaller than the mode X1 by the absolute value of the difference between the mode X1 and output quantity X4. For example, if the problem in question is one that has multiple time-series stages, such as an input / output plan, the action decided in the earlier stage of the input / output plan may affect the range of actions that are permissible in the later stage of the input / output plan. Therefore, when deriving a solution to such a problem, it is important not only to reduce the possibility of deciding on an action that immediately results in a failed (inappropriate) solution to the problem (an example of an inappropriate action), but also to reduce the probability of deciding on an action that should not be chosen when considering the impact on the range of actions that are permissible in the later stage (e.g., output quantity X4). For example, consider a scenario where raw materials stored in tank 24 are removed twice, in the first (first) and second (second) stages, and the amount of raw materials in tank 24 is 2 x 1 (twice the mode x 1). If the removal amount x 4 is selected for the first removal and the mode x 1 is selected for the second removal, the amount of inventory in tank 24 will be insufficient, making the operation impossible and resulting in a failed solution. Therefore, removing amount x 4 for the first removal is an appropriate action when considering only the first removal, but it is not necessarily an appropriate action when considering the problem as a whole (an example of an action that should not be selected). On the other hand, if the removal amount x 5 is selected for the first removal and the mode x 1 is selected for the second removal, the amount of inventory in tank 24 will not be insufficient, making the operation possible and resulting in a successful solution. Therefore, especially in problems with multiple time-series stages, using a probability distribution with an asymmetric probability distribution shape for extensive actions can make it easier to determine the appropriate action more effectively. In the above explanation, for the sake of simplicity, we have illustrated the case where the same probability distribution P1 is used to decide on the action for both the first and second removals (actions). However, for the decision on the second action, a different probability distribution P1 may be used depending on the change in the environment E1 caused by the first action.

[0094] Thus, the mode X1 (for example, the amount of output that maximizes the reward) and an inappropriate action for agent A1 (an action that results in an excess or shortage) may be adjacent in the continuous action space of the probability distribution P1. For example, in the probability distribution P1 of Figure 7, an inappropriate action for agent A1 is included in a region R2 (locations B1 and B2) that is larger than the mode X1. On the other hand, the region that is not an inappropriate action from the perspective of the mode X1 (for example, region R1) is not an inappropriate action that would immediately result in a failed answer. Therefore, if a symmetrical probability distribution is set with respect to the mode X1, the probabilities of an inappropriate action (region R2, which is larger than the mode X1) and a non-inappropriate action (region R1, which is smaller than the mode X1) may be about the same. Therefore, by applying an asymmetrical probability distribution P1 as shown in Figure 7, the probability of an inappropriate action in region R2 that results in an excess or shortage and a failed answer can be relatively reduced compared to the probability of a non-inappropriate action in region R1. In other words, the selection of an inappropriate action can be effectively suppressed.

[0095] The above explanation using Figure 7 is merely an example for conceptually explaining this embodiment, and the output amount X2 belonging to location B2 is not necessarily an action that agent A1 should not select. Similarly, the output amount X4 belonging to location B1 is not necessarily an action that agent A1 should not select. For example, the action that agent A1 should not select may be output amount X2. For example, the action that agent A1 should not select may be an output amount larger than output amount X2. For example, there may not even be an action that agent A1 should not select.

[0096] The probability distribution P1 illustrated in Figure 7 is an example where the output quantity is the random variable, and the shape of the probability distribution P1 is not limited to this. For example, if the random variable is changed from the output quantity, the shape of the probability distribution in region R1 and region R2 may be the inverse of the example in Figure 7 (the probabilities corresponding to each random variable may be inverted with respect to the line x = X1). In this case, the probability of selecting a value significantly lower than the mode X1 (for example, a value that, given the characteristics of the random variable, would immediately result in failure to answer) can be reduced compared to the case where a probability distribution symmetrical with respect to the mode X1 is used. That is, the shape of the probability distribution P1 that is asymmetrical with respect to the mode X1 can be changed according to the characteristics of the random variable.

[0097] A probability distribution P1 having asymmetry as shown in Figure 7 can be represented, for example, by the following equation (1).

[0098]

[0099] Equation (1) is an example of a function that represents an asymmetric probability distribution P1 based on a normal distribution (Gaussian distribution). In equation (1), x is a variable (random variable) corresponding to the output quantity (an extensive action), and P(x) is a function with respect to the output quantity as the variable and represents the probability. Also in equation (1), μ represents the mean and σ represents the variance. And in equation (1), α is the adjustment parameter. The adjustment parameter α is a parameter that adjusts the asymmetry of the probability distribution P1. For example, when the adjustment parameter α is a value less than 1, the probability distribution P1 has the asymmetry shown in Figure 7.

[0100] For example, a function like that in equation (1) is incorporated into the learning model. The learning model learns the appropriate asymmetry by changing the tuning parameter α according to the state of environment E1. The learning model is trained so that the mode X1 in Figure 7 matches or approaches the ideal output amount according to the state of environment E1. That is, the learning model learns the probability distribution P1 so that the reward obtained by agent A1 is maximized or approaches the maximum.

[0101] In the above example, an example of probability distribution P1 was shown with the output quantity as an extensive action. However, probability distributions related to extensive actions other than output quantities also exhibit asymmetry, as shown in probability distribution P1. Examples of extensive actions other than output quantities include input quantities. Input quantities represent the amount of raw materials delivered to a tank 24 in the tank base 21. In other words, any probability distribution corresponding to an extensive action is not limited to output quantities and has an asymmetrical probability distribution shape.

[0102] In the example above, a probability distribution P1 based on a normal distribution was used as an example, but the probability distribution P1 is not limited to this. The probability distribution P1 may be based on at least one of the following: the normal distribution, the log-normal distribution, the Weibull distribution, the Gumbel distribution, the half-normal distribution, the Rayleigh distribution, the exponential distribution, the gamma distribution, the Pareto distribution, the Fréchet distribution, and the inverse gamma distribution. In other words, the probability distribution is not limited to a probability distribution based on a normal distribution as shown in Figure 7, but can also be an asymmetrical probability distribution based on the above distribution.

[0103] Furthermore, for actions that do not have extensive properties (such as tank selection information), the probability distribution may be asymmetrical or symmetrical. In other words, the probability distribution may be switched between having a symmetrical or asymmetrical shape depending on the type of action. For example, the output unit 62 may output a probability distribution with an asymmetrical shape when the action has extensive properties, and output a probability distribution with a symmetrical shape when the action does not have extensive properties.

[0104] In this way, the output unit 62 outputs a probability distribution related to the action using the learning model.

[0105] When learning is performed through episodes, the output unit 62 outputs a probability distribution relating to the actions of agent A1 for each of the first states of each step acquired by the first acquisition unit 61.

[0106] The restriction unit 63 sets restrictions on the action selection based on the probability distribution in the decision unit 64, which will be described later. Specifically, the restriction unit 63 sets restrictions on the actions in the probability distribution. That is, the restriction unit 63 sets a mask on the actions included in the probability distribution.

[0107] Specifically, the restriction unit 63 sets restrictions based on predetermined constraint conditions (constraint information) to prevent actions that are inappropriate as actions for agent A1 corresponding to the first state from being determined as actions for agent A1.

[0108] Figure 8 shows an example of the constraint conditions used in the limiting section 63.

[0109] As shown in Figure 8, for example, the constraints are upper and lower limits on the inventory in tank 24, upper and lower limits on the API of the inventory in tank 24, upper limit limit on the number of tanks 24 used during unloading, and upper limit limit on the number of tanks 24 used during loading. Also, for example, the constraints are lower limit limits on the amount loaded per tank, lower limit limits on the amount unloaded per tank, upper limit limits on the amount transported during shifts between tanks, and lower limit limits on the amount transported during shifts between tanks. Also, for example, the constraints are upper limit limits on the number of shifts between tanks per predetermined period (e.g., one month) and upper limit limits on the number of shifts between tanks per predetermined period (e.g., one day). A shift is an operation to move raw materials from one tank 24 to another tank 24. The limiting unit 63 may use at least one of the above constraints.

[0110] Furthermore, each constraint is associated with a set type of action to which it applies. The types of actions are the selection of tank 24 during unloading, the amount of raw material during unloading, the selection of tank 24 during loading, and the amount of raw material during loading. Additionally, the types of actions corresponding to shifts are whether or not the tank is used, the selection of tank 24 to unload, the selection of tank 24 to load, and the amount of raw material. Figure 8 shows an example of the correspondence between constraints and the types of actions to which those constraints can be applied. In Figure 8, "○" indicates that a particular constraint is applicable to a particular action.

[0111] Furthermore, each constraint may or may not be assigned a priority. In the example in Figure 8, an example is shown where three levels of priority (high, medium, and low) are set. In the example in Figure 8, "high" is the highest priority, "medium" is the next highest, and "low" is the next highest. More specifically, when there are multiple constraints imposed on the same action, the constraint with high priority is applied preferentially over the constraints with medium and low priority. Also, when there are multiple constraints imposed on the same action, the constraint with medium priority is applied preferentially over the constraint with low priority. Note that "preferential application" may also mean, for example, when there are multiple constraints imposed on the same action, the degree of influence of each constraint on the action is corrected using a weighting according to the order of priority of the multiple constraints. Alternatively, "preferential application" may also mean, for example, when there are multiple constraints imposed on the same action, only the constraint with the highest priority among the multiple constraints is applied to that action. Furthermore, "preferential application" may mean, for example, that when multiple constraints are imposed on the same action, the constraint with the highest priority will always be satisfied, while the other constraints of lower priority will be satisfied as much as possible. Note that the number of priority levels is not limited to three, and can be changed as appropriate by the user, for example, to two levels. The priority of each constraint can also be changed as appropriate by the user.

[0112] Figure 9 shows an example of a case where restrictions are set on the behavior of a probability distribution. Specifically, Figure 9 shows an example where the restriction unit 63 sets restrictions on the probability distribution P1 in Figure 7.

[0113] As shown in Figure 9, for example, the lower limit constraint on the amount of material to be removed per tank, as shown in Figure 8, may be applied to the amount of material to be removed. In this case, the limiting unit 63 sets a lower limit value X6 for the amount of material to be removed based on the lower limit constraint on the amount of material to be removed per tank, and sets a mask MS for the range of material to be removed relative to the lower limit value X6.

[0114] In this way, the restriction unit 63 sets a restriction on the probability distribution. In this embodiment, the case in which the restriction unit 63 is provided in the learning unit 51 is given as an example, but the restriction unit 63 may be omitted. Furthermore, the constraints used by the restriction unit 63 are not limited to the above in the loading and unloading problem and may be changed by the user as appropriate. Also, when applying a problem other than the loading and unloading problem to the planning creation system 1, the constraints used by the restriction unit 63 may be set as appropriate constraints that are effective for the applied problem. Furthermore, in this case, when multiple constraints are set, each of the multiple constraints may be set as an appropriate priority.

[0115] The decision unit 64 determines the action of agent A1 based on the probability distribution. Specifically, the decision unit 64 determines the specific content of agent A1's action based on the probability distribution obtained using the learning model in response to the first state.

[0116] The decision unit 64 determines an action by selecting one of several actions based on a probability distribution. For example, the decision unit 64 determines a specific amount to be shipped based on a probability distribution P1 related to the amount to be shipped. For example, the decision unit 64 probabilistically selects one action based on a probability distribution. "Probabilistically selecting an action" means that an action is selected randomly according to the probabilities shown by the probability distribution, rather than by a decisive rule. That is, because an action is selected probabilistically, an action with a high probability may be selected, or an action with a low probability may be selected, making the selection probabilistic. The decision unit 64 may also choose to make a decisive decision to select an action. "Decisively selecting an action" means that an action is selected not probabilistically, but based on predetermined (set) rules using a probability distribution. Decisively selecting an action means, for example, selecting an action whose probability in the probability distribution is greater than or equal to a predetermined value, or selecting the action with the highest probability in the probability distribution.

[0117] Furthermore, when the decision unit 64 performs learning (updating) of the learning model, it is preferable to determine actions probabilistically using a probability distribution. That is, during the learning phase of the learning model, by probabilistically determining actions and performing action exploration while considering not only high-probability actions but also low-probability actions as options, it is possible to perform learning that takes into account various action patterns. In this way, the decision unit 64 determines the action of agent A1 corresponding to the first state.

[0118] The decision unit 64 determines an action considering the restriction set by the restriction unit 63 on the probability distribution. For example, if a probability distribution P1 and a mask MS are set as shown in Figure 9, the decision unit 64 determines an action from within the range of the probability distribution P1 for which the mask MS is not set as a restriction.

[0119] Furthermore, when the decision unit 64 determines multiple actions such as tank selection information, raw material quantity information, and type information, it is preferable to determine each action in stages. For example, the decision unit 64 determines the tank selection information using a probability distribution related to the tank selection information. Then, the decision unit 64 determines the type information using a probability distribution related to the type information obtained corresponding to the selected tank 24. Then, the decision unit 64 determines the raw material quantity information using a probability distribution related to the raw material quantity information obtained corresponding to the selected tank 24 and the type of raw material. This allows the decision unit 64 to determine a specific action, for example, to transport 100 kl of raw material G3 into tank 24a. When determining actions in stages in this way, for example, the output unit 62 may output a probability distribution related to the tank selection information, and the decision unit 64 may determine the tank selection information using the probability distribution related to the tank selection information. After that, the output unit 62 may be configured to output a probability distribution related to the type information corresponding to the selected tank 24.

[0120] When learning is performed through episodes, the decision unit 64 determines the action of agent A1 using a probability distribution for each of the first states of each step acquired by the first acquisition unit 61.

[0121] Returning to Figure 5, the second acquisition unit 65 functions as an acquisition unit and acquires the state of the environment E1 that has changed as a result of the action determined by the decision unit 64 as the "second state". The second acquisition unit 65 also acquires the reward for the action determined by the decision unit 64.

[0122] Specifically, the second acquisition unit 65 observes the state of the tank base 21 as environment E1 after the action and sets it to the second state. The second acquisition unit 65 also acquires an evaluation of the action that changed the state of the tank base 21 from the first state to the second state as a reward.

[0123] In the loading and unloading problem, the second state includes, for example, the state of the tank 24 in the tank base 21 as environment E1.

[0124] The reward is an evaluation of the decided action, and in the loading / unloading problem, it is an evaluation of the action that changes the state of the tank base 21 from the first state to the second state. For example, the action is evaluated using a predetermined evaluation index, and a reward is set. The predetermined evaluation index is, for example, the raw material matching rate. For example, the higher the raw material matching rate, the higher the reward. However, as long as the quality of the action can be evaluated, the specific method of setting the reward is not limited. The more desirable the decided action, the larger the reward value.

[0125] When learning is performed through episodes, the second acquisition unit 65 acquires a second state and a reward corresponding to each step's action determined by the decision unit 64.

[0126] The calculation unit 66 calculates the loss based on the reward for the decided action. The calculation unit 66 uses the probability of the decided action and the reward for the decided action to calculate the loss. Specifically, the calculation unit 66 calculates the loss by multiplying the log probability related to the probability of the decided action and the advantage related to the reward for the decided action.

[0127] The logarithmic probability is determined by the probability of the decided action. For example, the logarithmic probability is the value obtained by transforming the probability of the decided action using a predetermined logarithmic function. For example, the calculation unit 66 performs the transformation such that when a probability less than a predetermined value is transformed into a logarithmic probability, the logarithmic probability becomes a negative value, and when a probability greater than or equal to the predetermined value is transformed into a logarithmic probability, the logarithmic probability becomes a positive value. As a result, the probability of the decided action becomes logarithmic and its sign (positive or negative) is set. For example, the lower the probability of the decided action is compared to the predetermined value, the more negative the logarithmic probability becomes and the larger its absolute value becomes. Also, for example, the higher the probability of the decided action is compared to the predetermined value, the more positive the logarithmic probability becomes and the larger its absolute value becomes.

[0128] The advantage is determined by the reward for the decided action. Specifically, the advantage is the difference between the reward for the action decided using the learning model and the reward for the action decided using the baseline model. The baseline model is a baseline model corresponding to the learning model. For example, the baseline model is the learning model of the initial state before the start of learning. The reward for the action decided using the baseline model is, like the learning model, the reward for the action decided corresponding to the first state. For example, the reward for the action decided using the baseline model is obtained using a simulation related to the loading and unloading problem, just like the learning model. Although the action decided using the learning model and the action decided using the baseline model are based on the same first state, the specific content of the decided actions may be the same or different.

[0129] The advantage is the degree of improvement (reward difference) between the behavior determined using the learning model and the behavior determined using the baseline model. For example, the advantage is set by subtracting the reward for the behavior determined using the baseline model from the reward for the behavior determined using the learning model. In this way, the advantage is an evaluation that corresponds to the learning model's progress from its initial state through learning. In other words, the higher the reward of the learning model compared to the reward of the baseline model, the higher the advantage.

[0130] The calculation unit 66 then calculates the loss as the product of the log probability, the advantage, and a constant. The constant is set to a negative value, for example, -1. As a result, the higher the probability of the action (positive log probability) and the higher the advantage, the larger the negative value of the loss, and the smaller the loss. This indicates that the desirable action is associated with a high probability, and the learning model is in a favorable state. Conversely, the lower the probability of the action (negative log probability) and the higher the advantage, the larger the positive value of the loss, and the larger the loss. This indicates that the desirable action is associated with a low probability, and the learning model is not in a favorable state. Thus, even if an action is desirable, the lower its probability, the larger the loss.

[0131] In the example above, the loss was calculated using the product of the log probability, the advantage, and a constant, but the method of calculating the loss is not limited to this.

[0132] The update unit 67 updates the learning model based on the results of actions taken in response to environment E1. The results of actions taken in response to environment E1 are the losses calculated by the calculation unit 66. In other words, the update unit 67 updates the learning model based on the losses. Specifically, the update unit 67 updates the learning model so that the loss is reduced. More specifically, the update unit 67 updates the weights and other components included in the neural network as a learning model so that the loss value of the updated learning model is smaller than the loss value of the learning model before the update. In other words, the update unit 67 updates the learning model so that the probability of favorable actions (actions with high rewards or advantages) is increased in the probability distribution output from the learning model.

[0133] In particular, the learning model learns the asymmetry of the probability distribution by updating it using loss. Specifically, the adjustment parameter α in equation (1) is updated and the asymmetry shape of the probability distribution is optimized.

[0134] For example, the update unit 67 updates the neural network as a learning model by backpropagation using loss. Note that the method for updating the learning model is not limited to backpropagation, and other methods may be used.

[0135] In this way, reinforcement learning (deep reinforcement learning) is performed in the learning unit 51. Using the learned model in this manner, inference is performed by the inference unit 52, which will be described later.

[0136] The inference unit 52 performs inference using the trained model. The trained model is the trained model that has been trained in the learning unit 51. In other words, the inference unit 52 uses the trained model trained by the learning unit 51 as the trained model to infer output data corresponding to the input data.

[0137] The inference unit 52 mainly comprises a reception unit 71, an estimation unit 72, and an output unit 73.

[0138] The reception unit 71 receives input data. The input data is data that serves as a prerequisite for the loading and unloading problem. For example, the input data is data related to the planning of loading and unloading multiple types of raw materials to and from the tank base 21. Specifically, the input data is data related to initial inventory information, loading and unloading plan information, and constraint information. For example, the input data is set by the user as a prerequisite for the problem to be applied and is entered using the user terminal 2.

[0139] The estimation unit 72 estimates output data using a trained model. The estimation unit 72 has the same functions as the output unit 62, restriction unit 63, and decision unit 64 in the learning unit 51. Specifically, the estimation unit 72 outputs a probability distribution corresponding to the input data using the trained model, and estimates output data corresponding to the input data based on the probability distribution. The output data is, for example, information about tank selection, raw material quantity, and type. In other words, the output data is information corresponding to the actions of agent A1 during the learning phase.

[0140] Figure 10 shows an example of processing in the estimation unit 72.

[0141] As shown in Figure 10, the estimation unit 72 inputs the input data into the trained model. The input data may be preprocessed to make it suitable for input into the trained model. The trained model then outputs a probability distribution, for example, as shown in Figure 7. Furthermore, the probability distribution has an asymmetrical probability distribution shape with respect to extensive actions (variables with extensive properties), such as the amount of material removed, as shown in Figure 7. That is, the probability distribution has an asymmetrical probability distribution shape with respect to extensive actions, as shown in probability distribution P1, similar to the learning stage. Since the adjustment parameter α in equation (1) is determined by learning, the trained model outputs a probability distribution with an asymmetry associated with the adjustment parameter α. Similarly, for actions (variables without extensive properties), the probability distribution may be asymmetrical or symmetrical, and the distribution shape may be switched depending on the type of action (variable).

[0142] The estimation unit 72 then determines output data corresponding to the input data based on the probability distribution. Specifically, the estimation unit 72 determines the specific content of the actions (variables) based on the probability distribution, as in the determination unit 64 described above, and uses this as output data. In other words, the actions (variables) determined based on the probability distribution become the solution to the problem. For example, corresponding to the probability distribution shown in Figure 7, the specific value of the output quantity is estimated as the specific content of the output data. The output quantity value estimated in this way becomes part of the plan corresponding to the input / output problem. Other parts of the plan are similarly estimated using the trained model.

[0143] In this way, the estimation unit 72 estimates, as output data, information regarding the loading and unloading plans (specific plans) for each of the multiple tanks 24 owned by the tank base 21, corresponding to the preconditions (overall plan) for the loading and unloading problem. The output data may be subjected to post-processing.

[0144] When the estimation unit 72 performs inference using a trained model, it is preferable to definitively determine the action using a probability distribution. That is, the estimation unit 72 is preferable to select a variable (such as the amount of goods to be shipped) whose probability in the probability distribution is greater than or equal to a predetermined value, or to select the variable (such as the amount of goods to be shipped) with the highest probability in the probability distribution. However, if the problem involves multiple time-series stages, such as an import / export problem, and a variable with a high probability is not necessarily good in relation to other stages, the estimation unit 72 may also determine the action probabilistically using a probability distribution.

[0145] Furthermore, when the estimation unit 72 performs inference in response to a predetermined episode, it estimates output data corresponding to each step. For example, for each step corresponding to the loading and unloading of raw materials, the loading and unloading plans for each raw material corresponding to the tank 24 are estimated. Then, the plans corresponding to each step are combined to form the overall plan. Note that the plans corresponding to each step are estimated in such a way that consistency is ensured between each step. For example, considering the loading plan for the tank 24 estimated in response to the loading step (considering the state of environment E1 when that loading plan is executed), the unloading plan for the tank 24 is estimated in response to the unloading step of the next step.

[0146] Alternatively, the estimation unit 72 may set constraints on the probability distribution based on constraint conditions, as in the restriction unit 63, and then estimate the output data.

[0147] In this way, output data corresponding to the input data is inferred using the trained model.

[0148] The output unit 73 outputs the output data estimated by the estimation unit 72. For example, the output unit 73 outputs the output data as an inference result to a predetermined device such as a user terminal 2, providing the inference result to the user. This allows the user to recognize the loading and unloading plans for each of the multiple tanks 24 owned by the tank base 21.

[0149] The destination of the output data in the output unit 73 is not limited. For example, if further processing is performed using the output data, the output data may be output to other functional units within the server device, or it may be output to a device different from the server device 3.

[0150] <Processing Flow> Figure 11 is a flowchart showing an example of the learning process flow according to this embodiment. Each of the following processes is started, for example, according to a user's instruction to start learning. In the learning process, each episode consists of M steps, and each step is associated with a number i (an integer from 1 to M). The order and content of each of the following steps can be changed as appropriate.

[0151] (Step SP10) The simulation unit 60 initializes various information related to the loading and unloading problem. For example, the state of the environment E1 and the weights of the neural network as a learning model are initialized. That is, each parameter related to the loading and unloading problem is set to its initial state. Alternatively, each hyperparameter related to learning may be set. Then, the process moves on to step SP11.

[0152] (Step SP11) The simulation unit 60 sets the step number i included in the episode to 1 (i=1). Then, the process moves on to step SP12.

[0153] (Step SP12) The simulation unit 60 starts the episode and executes step number i. Then, the process moves on to step SP13.

[0154] (Step SP13) The first acquisition unit 61 acquires the current state in environment E1 as the first state. That is, the first acquisition unit 61 acquires the state of environment E1 before agent A1's action in step number i as the first state. Then the process moves on to step SP14.

[0155] (Step SP14) The output unit 62 inputs the first state acquired by the first acquisition unit 61 into the learning model and outputs a probability distribution corresponding to the first state. The probability distribution P1 relating to the extensive behavior has an asymmetric probability distribution shape. Then the process moves on to step SP15.

[0156] (Step SP15) The restriction unit 63 sets a restriction on the probability distribution based on the constraint conditions. Then, the process moves on to step SP16.

[0157] (Step SP16) The decision unit 64 determines the action of agent A1 according to a probability distribution with set restrictions. That is, the decision unit 64 determines the action of agent A1 in step number i. Then the process moves on to step SP17.

[0158] (Step SP17) The simulation unit 60 reflects the determined actions of agent A1 into environment E1. Then, the process moves on to step SP18.

[0159] (Step SP18) The second acquisition unit 65 acquires the state that has changed due to agent A1's actions in environment E1 as the second state, and also acquires the reward for the action. That is, the second acquisition unit 65 acquires the second state and the reward corresponding to environment E1 after agent A1's action in step i. Then the process moves on to step SP19.

[0160] (Step SP19) The calculation unit 66 calculates the loss using the probability of the decided action and the reward of the decided action. That is, the calculation unit 66 calculates the loss corresponding to step number i. Then the process moves on to step SP20.

[0161] (Step SP20) The update unit 67 updates the learning model based on the loss. Then, the process moves on to step SP21.

[0162] (Step SP21) The simulation unit 60 determines whether the episode has finished. Specifically, the simulation unit 60 determines that the episode has finished if the step number i is the final number (i = M). If the episode has not finished, the process proceeds to step SP22. If the episode has finished, the process ends. If another episode is to be executed, learning may be performed again in accordance with that other episode.

[0163] (Step SP22) The simulation unit 60 adds 1 to the number i. Then, the process moves to step SP12 and is executed again. In other words, each step is repeatedly executed until the episode ends.

[0164] In this way, deep reinforcement learning is performed. That is, a trained model is generated by training (updating) a neural network, which is an example of a learning model. The trained model updated in step SP20 may be evaluated for its learning status using test episodes, etc., and if the learning objective has not been achieved, relearning may be performed. Reearning may also be performed until a predetermined number of trials are reached.

[0165] Figure 12 is a flowchart showing an example of the inference process flow according to this embodiment. Each of the following processes is started, for example, in accordance with a user's instruction to start inference. Note that the order and content of each of the following steps can be changed as appropriate.

[0166] (Step SP30) The reception unit 71 receives input data from the user as preconditions for the problem. Then, the process moves on to step SP31.

[0167] (Step SP31) The estimation unit 72 inputs the input data into the trained model and outputs the probability distribution. Then, the process moves on to step SP32.

[0168] (Step SP32) The estimation unit 72 estimates output data corresponding to the input data using a probability distribution. Then, the process moves on to step SP33.

[0169] (Step SP33) The output unit 73 outputs the output data to the user terminal 2. Then the process ends.

[0170] <Effects> As described above, the server device 3 according to this embodiment includes a first acquisition unit 61 that acquires a state corresponding to a predetermined environment E1, an output unit 62 that uses a learning model to output a probability distribution related to actions for environment E1 from the state, a decision unit 64 that determines actions for environment E1 based on the probability distribution, and an update unit 67 that updates the learning model based on the results of actions taken for environment E1, wherein the probability distribution has a probability distribution shape that is asymmetric with respect to extensive actions.

[0171] This configuration makes it easier to determine appropriate actions compared to using a symmetric probability distribution, because it uses a probability distribution with an asymmetric probability distribution shape to determine actions with extensive properties. For example, in actions with extensive properties, if the amount of material to be discharged is too large relative to the amount of material stored in tank 24 (an example of inappropriate action), it is expected that the amount of material stored in tank 24 will be insufficient (short). Furthermore, if the amount of material to be brought in is too large relative to the empty capacity of tank 24 (an example of inappropriate action), it is expected that the empty capacity of tank 24 will be insufficient (short) (the amount of inventory in tank 24 will be excessive). Therefore, by determining actions with extensive properties using a probability distribution with an asymmetric probability distribution shape, it is possible to reduce the probability corresponding to inappropriate actions compared to using a probability distribution with a symmetric probability distribution shape, thereby making it easier to determine appropriate actions. In this way, it is possible to make it easier to determine appropriate actions with extensive properties.

[0172] In particular, when the problem in question has multiple time-series stages, such as an inbound / outbound plan, the actions decided in the earlier stages of the plan may affect the range of actions that are permissible in the later stages. For example, it is possible to reduce the likelihood of deciding on an action that immediately leads to failure in solving the problem (an example of an inappropriate action), such as a stock shortage as described above, or to reduce the probability of deciding on an action that should be avoided when considering its impact on the range of actions that are permissible in the later stages of the plan. For this reason, in problems with multiple time-series stages, for example, by deciding on an action using a probability distribution whose probability distribution shape is asymmetric with respect to extensive actions, it is possible to make it easier to decide on an appropriate action in a particularly effective way.

[0173] Furthermore, the above configuration allows for efficient determination of appropriate actions. This enables improved processing efficiency in the computer (server device 3). In other words, the performance of the processing in the computer is improved.

[0174] Furthermore, in server device 3, extensive actions are those that change the quantity in environment E1.

[0175] This configuration allows for action decisions to be made using a probability distribution with an asymmetrical probability distribution shape for actions that change quantities in environment E1. In other words, it makes it easier to determine the appropriate action for actions that change quantities.

[0176] Furthermore, in server device 3, the probability distribution is a continuous probability distribution for continuous variables relating to extensive behavior.

[0177] This configuration allows for an improvement in the degree of freedom in decision-making for extensive actions by using a continuous probability distribution.

[0178] Furthermore, in server device 3, the probability distribution has different shapes in the region R1, which is smaller than the mode X1 for extensive behaviors, and in the region R2, which is larger than the mode X1.

[0179] This configuration allows for setting a probability distribution shape that is asymmetric with respect to the mode X1.

[0180] Furthermore, in server device 3, the area M1 of region R1, where the probability distribution is smaller than the mode X1, is larger than the area M2 of region R2, where the probability distribution is larger than the mode X1.

[0181] With this configuration, actions corresponding to the region R2, which is larger than the mode X1, are less likely to be selected than actions corresponding to the region R1, which is smaller than the mode X1. Therefore, for example, excessive amounts of incoming or outgoing goods are suppressed.

[0182] Furthermore, in server device 3, the probability distribution increases towards the mode X1 in region R1, which is smaller than the mode X1, while decreasing at a rate of change greater than the increase in region R2, which is larger than the mode X1.

[0183] This configuration makes it less likely that actions corresponding to the region R2, which is larger than the mode X1, will be effectively selected. As a result, for example, excessive amounts of incoming or outgoing data are suppressed.

[0184] Furthermore, in the server device 3, the decision unit 64 determines the action to take in response to the environment E1 probabilistically using a probability distribution.

[0185] With this configuration, actions are determined probabilistically during the update (learning) phase, allowing for the learning of a wide range of patterns, including actions with low probability.

[0186] Furthermore, the server device 3 includes a calculation unit 66 that calculates the loss based on the reward for the determined action, and the update unit 67 uses the resulting loss to update the learning model.

[0187] With this configuration, the learning model is updated to facilitate more desirable behaviors.

[0188] Furthermore, in the server device 3, the update unit 67 updates the asymmetry of the probability distribution output from the learning model when updating the learning model due to loss.

[0189] This configuration updates the probability distribution so that its asymmetry is optimal or approaches optimal.

[0190] Furthermore, in the server device 3, the calculation unit 66 calculates the loss based on the probability of the decided action and the reward for the decided action.

[0191] With this configuration, the learning model is updated to increase the probability of more favorable behavior.

[0192] Furthermore, the server device 3 is further equipped with a restriction unit 63 that sets restrictions on the probability distribution related to actions, and the decision unit 64 determines an action from within the range of the probability distribution for which no restrictions have been set.

[0193] This configuration allows for the effective exclusion of inappropriate behavior.

[0194] Furthermore, in the server device 3, the learning model takes data relating to at least one of the plans for loading and unloading multiple types of raw materials to and from the tank base 21 as input, and outputs data relating to at least one of the plans for loading and unloading each of the multiple tanks 24 that the tank base 21 has.

[0195] This configuration allows for training a learning model to address the issue of raw material loading and unloading.

[0196] Furthermore, the server device 3 uses the learned model acquired by the learning unit 51 as a trained model to infer output data corresponding to the input data.

[0197] With this configuration, inference can be performed using the trained model.

[0198] Furthermore, the server device 3 includes a reception unit 71 that receives input data, an estimation unit 72 that outputs a probability distribution corresponding to the input data using a trained model and estimates output data corresponding to the input data based on the probability distribution, and an output unit 73 that outputs the estimated output data. The probability distribution has a probability distribution shape that is asymmetric with respect to the extensive variable.

[0199] This configuration allows for inference of output data using a probability distribution with an asymmetric probability distribution shape for extensive variables (behaviors). For extensive variables, depending on the content of the variable, there are possibilities such as, for example, a shortage of raw materials due to excessive output, or an excess of raw materials due to excessive input. For example, when performing inference using a probability distribution with a symmetric probability distribution shape, such inappropriate variable content may be determined. Therefore, by using an asymmetric probability distribution for extensive variables, it becomes easier to determine appropriate variables that suppress the above-mentioned excesses and shortages. In this way, it becomes easier to determine appropriate content for extensive variables (behaviors).

[0200] In particular, when the problem in question involves multiple time-series stages, such as loading and unloading plans, the aforementioned surpluses or deficiencies can affect later stages. Therefore, this method makes it easier to effectively determine appropriate output data, especially for problems involving multiple time-series stages.

[0201] Furthermore, the above configuration allows for efficient inference of appropriate output data. This enables improved processing efficiency in the computer (server device 3). In other words, the performance of the processing in the computer is enhanced.

[0202] <Modifications> This disclosure is not limited to the embodiments described above. That is, any modifications made to the embodiments described above by a person skilled in the art are also included in the scope of this disclosure, as long as they retain the features of this disclosure. Furthermore, the elements of the embodiments described above and the modifications described later can be combined to the extent that it is technically possible, and any combination thereof is also included in the scope of this disclosure, as long as it retains the features of this disclosure.

[0203] In the above embodiment, one example is that each function is provided by the server device 3, but each function may also be provided by the user terminal 2. Alternatively, each function may be distributed between the server device 3 and the user terminal 2. For example, the user terminal 2 may function as both a machine learning device and an inference device. Alternatively, one of the server device 3 and the user terminal 2 may function as a machine learning device, and the other of the server device 3 and the user terminal 2 may function as an inference device. Alternatively, two server devices 3 that can communicate with each other may be installed, with one server device 3 functioning as a machine learning device and the other server device 3 functioning as an inference device. Alternatively, multiple server devices 3 that can communicate with each other may be installed, with each function constituting the machine learning device and each function constituting the inference device distributed among multiple server devices 3, and these may function together as a machine learning device or an inference device.

[0204] Furthermore, while the above embodiment described the application of the problem of loading and unloading raw materials at the tank base 21 to the planning system 1 as an example, the problems that can be applied to the planning system 1 are not limited to those described above. For example, the problem of loading and unloading vehicles at a parking facility may be applied to the planning system 1. The problem of loading and unloading ships at a port may also be applied to the planning system 1. The problem of receiving and shipping goods and products at a warehouse or store may also be applied to the planning system 1. It should be noted that the problems that can be applied to the planning system 1 are not limited to those described above, and a variety of problems can be applied. In addition, while the planning system 1 can be applied to problems that involve multiple chronological stages, such as loading and unloading problems, it is also possible to apply problems that involve multiple non-chronological stages or problems that do not involve multiple stages (for example, problems with only one stage).

[0205] Furthermore, in the above embodiment, as shown in Figure 7, for example, the probability distribution P1 is given as an example in which the area M1 of the region R1 smaller than the mode X1 is larger than the area M2 of the region R2 larger than the mode X1. However, it is not limited to the above. For example, the probability distribution P1 may have an asymmetry such that the area M2 of the region R2 larger than the mode X1 is larger than the area M1 of the region R1 smaller than the mode X1. That is, in the region R2 larger than the mode X1, the probability of the probability distribution P1 may increase toward the mode X1, while in the region R1 smaller than the mode X1, the probability may decrease at a rate of change larger than that increase.

[0206] Furthermore, while the above embodiment illustrates a case where the learning model outputs a probability distribution P1 having asymmetric properties, the model is not limited to the above as long as it can output a probability distribution P1 having asymmetric properties for extensive actions. For example, the learning model may output a probability distribution P1 having symmetry corresponding to a state, and this symmetric probability distribution P1 may be modified to have an asymmetric shape.

[0207] Furthermore, in the above embodiment, the asymmetric probability distribution shape is constructed based on a normal distribution as shown in equation (1), but the specific type of distribution is not limited. For example, the asymmetric probability distribution shape may be constructed based on a function that results in a symmetric probability distribution shape other than a normal distribution. Alternatively, for example, a function that results in an asymmetric probability distribution shape may be used.

[0208] The various types of information described in this disclosure (e.g., status, reward, etc.) may be expressed using absolute values, relative values ​​from a given value, or other corresponding information.

[0209] In this disclosure, expressions such as "based on," "using," and "by" (including equivalent expressions) do not mean "based solely on," "using only," or "by" unless otherwise specified. In other words, the phrase "based on" means both "based solely on" and "at least on," and the same applies to equivalent expressions such as "using" and "by."

[0210] The term “decision” in this disclosure may encompass a wide variety of actions. “Decision” may include, for example, judgment, calculation, calculation, processing, derivation, investigation, exploration, and confirmation. Furthermore, “decision” may include, for example, considering something to have been “decided,” such as resolving, selecting, choosing, establishing, or comparing. In short, “decision” may include considering any action to have been “decided.”

[0211] In this disclosure, where expressions such as "obtain / set / use / based on" (including similar expressions) are used, unless otherwise specified, this includes cases where the information itself is used, or where the information has been processed in some way (e.g., noise-added, normalized, features extracted from the information, intermediate representation of the information, etc.). Furthermore, where it is stated that some result is obtained by "obtaining / setting / using / based on" (including similar expressions) (including similar expressions), unless otherwise specified, this includes cases where the result is obtained solely based on the information in question, or where the result is influenced by other information, factors, conditions, and / or states other than the information in question. Furthermore, where it is stated that "output" (including similar expressions), unless otherwise specified, this includes cases where the information itself is used as output, or where the information has been processed in some way (e.g., noise-added, normalized, features extracted from the information, intermediate representation of various types of information, etc.) is used as output.

[0212] Where terms meaning "containing" or "including" are used in this disclosure (e.g., "containing" or "including"), "having," etc.), they are intended as open-ended terms, including cases where the object of such term contains or possesses something other than the object indicated by the object of the term. Where the object of such terms meaning "containing" or "possessing" is an expression that does not specify a quantity or suggests a singular number (an expression with the article "a" or "an"), such expression should be interpreted as not being limited to a specific number.

[0213] In this disclosure, even if expressions such as "one or more" or "at least one" are used in some places, and expressions that do not specify a quantity or suggest singularity (expressions using the articles a or an) are used in other places, the latter expressions are not intended to mean "one." In general, expressions that do not specify a quantity or suggest singularity (expressions using the articles a or an) should be interpreted as not necessarily being limited to a specific number.

Claims

1. A machine learning device comprising: an acquisition unit that acquires a state corresponding to a predetermined environment; an output unit that uses a learning model to output a probability distribution relating to an action in relation to the environment from the state; a decision unit that determines the action in relation to the environment based on the probability distribution; and an update unit that updates the learning model based on the result of performing the action in relation to the environment, wherein the probability distribution has an extensive nature and an asymmetric probability distribution shape with respect to the action.

2. The machine learning apparatus according to claim 1, wherein the extensive behavior is a behavior that changes the quantity in the environment.

3. The machine learning apparatus according to claim 1 or 2, wherein the probability distribution is a continuous probability distribution for a continuous variable relating to the behavior having extensive properties.

4. The machine learning apparatus according to claim 1 or 2, wherein the probability distribution has different probability distribution shapes in the region where it is smaller than the mode of the extensive behavior and in the region where it is larger than the mode.

5. The machine learning apparatus according to claim 4, wherein the area of ​​the region smaller than the mode in the probability distribution is larger than the area of ​​the region larger than the mode.

6. The machine learning apparatus according to claim 4, wherein the probability distribution increases toward the mode in the region smaller than the mode, and decreases at a rate of change greater than the increase in the region larger than the mode.

7. The machine learning apparatus according to claim 1 or 2, wherein the decision unit determines the action in relation to the environment probabilistically using the probability distribution.

8. A machine learning device according to claim 1 or 2, further comprising: a calculation unit that calculates a loss based on the reward for the determined action, wherein the update unit updates the learning model using the loss as a result.

9. The machine learning apparatus according to claim 8, wherein the update unit updates the asymmetry of the probability distribution output from the learning model when updating the learning model with the loss.

10. The machine learning apparatus according to claim 8, wherein the calculation unit calculates the loss based on the probability of the determined action and the reward for the determined action.

11. A machine learning device according to claim 1 or 2, further comprising a limiting unit that sets a limit on the probability distribution relating to the action, wherein the determination unit determines the action from within the range of the probability distribution for which no limit is set.

12. The machine learning device according to claim 1 or 2, wherein the learning model takes data relating to at least one of the plans for loading and unloading of multiple types of raw materials to and from a tank base as input, and outputs data relating to at least one of the plans for loading and unloading of each of the multiple tanks owned by the tank base.

13. The machine learning apparatus according to claim 1 or 2, wherein the environment corresponds to a raw material tank base, and the action is an action relating to at least one of loading and unloading the raw material at the tank base.

14. An inference device that uses the learned model, which has been trained by the machine learning device described in claim 1 or 2, as a trained model to infer output data corresponding to input data.

15. An inference device comprising: a receiving unit for receiving input data; an estimation unit for outputting a probability distribution corresponding to the input data using a trained model and estimating output data corresponding to the input data based on the probability distribution; and an output unit for outputting the estimated output data, wherein the probability distribution has a probability distribution shape that is asymmetric with respect to an extensive variable.

16. A machine learning method comprising: a step of acquiring a state corresponding to a predetermined environment; a step of using a learning model to output a probability distribution relating to an action on the environment from the state; a step of determining the action on the environment based on the probability distribution; and a step of updating the learning model based on the result of performing the action on the environment, wherein the probability distribution has an extensive nature and an asymmetric probability distribution shape with respect to the action.

17. An inference method comprising: a step of receiving input data; a step of outputting a probability distribution corresponding to the input data using a trained model, estimating output data corresponding to the input data based on the probability distribution; and a step of outputting the estimated output data, wherein the probability distribution has a probability distribution shape that is asymmetric with respect to extensive variables.

18. A machine learning program that enables a computer to function as: an acquisition unit that acquires a state corresponding to a predetermined environment; an output unit that uses a learning model to output a probability distribution relating to an action on the environment from the state; a decision unit that determines the action on the environment based on the probability distribution; and an update unit that updates the learning model based on the result of performing the action on the environment, wherein the probability distribution has an extensive nature and an asymmetric probability distribution shape with respect to the action.

19. An inference program in which a computer functions as a receiving unit that receives input data, an estimation unit that outputs a probability distribution corresponding to the input data using a trained model and estimates output data corresponding to the input data based on the probability distribution, and an output unit that outputs the estimated output data, wherein the probability distribution has a probability distribution shape that is asymmetric with respect to an extensive variable.