Information processing device, information processing method, information processing system, and program

The information processing apparatus enhances decision-making by deriving a second optimal solution based on the reliability of first derivation processes, addressing limitations in existing sequential decision-making techniques.

WO2025126326A1PCT designated stage expired Publication Date: 2025-06-19NEC CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/044465
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing sequential decision-making techniques struggle to derive more appropriate optimal solutions due to limitations in handling uncertainty and reliability in decision-making processes.

Method used

An information processing apparatus and method that acquire output values from multiple models, perform multiple first derivation processes to derive a first optimal solution, and a second derivation process to derive a second optimal solution based on the reliability of each first derivation process.

Benefits of technology

This approach enables the derivation of a more appropriate decision-making result (optimal solution) by incorporating the reliability of each derivation process, thereby improving the accuracy and effectiveness of decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023044465_19062025_PF_FP_ABST
    Figure JP2023044465_19062025_PF_FP_ABST
Patent Text Reader

Abstract

To be able to derive a more suitable decision-making result (optimal solution), an information processing device (1) is provided with: an acquisition means for acquiring an output value obtained from each of one or a plurality of models; and a derivation means for executing a plurality of first derivation processes for deriving a first optimal solution by referring to the output value acquired by the acquisition means, and a second derivation process for deriving a second optimal solution in accordance with the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, information processing system, and program

[0001] The present invention relates to an information processing device, an information processing method, an information processing system, and a program.

[0002] A sequential decision-making technique is known in which a process of deriving a prediction (decision-making result) regarding demand or supply, etc., observing the execution results of the derived prediction, and deriving a further prediction based on the observed results is sequentially repeated.

[0003] For example, Patent Document 1 describes an optimal decision-making method for formulating and evaluating a capital investment plan that includes uncertain factors.

[0004] Japanese Patent Application Laid-Open No. 2005-108147

[0005] Generally, sequential decision-making techniques are required to derive more appropriate decision-making results (optimal solutions), but the technique described in Patent Document 1 leaves room for improvement in this regard.

[0006] One aspect of the present invention has been made in consideration of the above-mentioned problems, and one of its objectives is to provide a technique that can derive more appropriate decision-making results (optimal solutions).

[0007] An information processing device according to one aspect of the present invention includes an acquisition means for acquiring output values ​​obtained from each of one or more models, a plurality of first derivation processes for deriving a first optimal solution by referring to the output values ​​acquired by the acquisition means, and a derivation means for executing a second derivation process for deriving a second optimal solution in accordance with the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes.

[0008] An information processing method according to one aspect of the present invention includes an information processing device that acquires output values ​​from each of one or more models and executes a plurality of first derivation processes to derive a first optimal solution by referring to the acquired output values, and a second derivation process that derives a second optimal solution based on the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes.

[0009] A program according to one aspect of the present invention is a program that causes a computer to function as an information processing device, and the program causes the computer to acquire output values ​​obtained from each of one or more models, and to execute a plurality of first derivation processes that derive a first optimal solution by referring to the acquired output values, and a second derivation process that derives a second optimal solution depending on the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes.

[0010] An information processing system according to one aspect of the present invention is an information processing system including an information processing device and a terminal device, wherein the information processing device comprises an acquisition means for acquiring output values ​​obtained from each of one or more models, a plurality of first derivation processes for deriving a first optimal solution by referring to the output values ​​acquired by the acquisition means, and a derivation means for executing a second derivation process for deriving a second optimal solution in accordance with the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes, and the terminal device comprises an execution means for executing the second optimal solution derived by the information processing device.

[0011] According to one aspect of the present invention, it is possible to derive a more appropriate decision-making result (optimal solution).

[0012] FIG. 1 is a block diagram showing the configuration of an information processing device according to an exemplary embodiment. FIG. 2 is a flow diagram showing the flow of an information processing method according to an exemplary embodiment. FIG. 3 is a diagram for explaining processing by an information processing device according to an exemplary embodiment. FIG. 4 is a block diagram showing the configuration of an information processing system according to an exemplary embodiment. FIG. 5 is a flow diagram showing the flow of processing by an information processing system according to an exemplary embodiment. FIG. 6 is a block diagram showing the configuration of an information processing device according to an exemplary embodiment. FIG. 7 is a diagram for explaining processing by an information processing device according to an exemplary embodiment. FIG. 8 is a diagram for explaining effects of an information processing device according to an exemplary embodiment. FIG. 9 is a diagram for explaining processing by an information processing device according to an application example of an exemplary embodiment. FIG. 10 is a diagram showing an example of information referenced by an information processing device according to an application example of an exemplary embodiment. FIG. 11 is a block diagram showing the configuration of a computer functioning as an information processing device according to each exemplary embodiment.

[0013] The following are examples of embodiments of the present invention. However, the present invention is not limited to the exemplary embodiments shown below, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, embodiments obtained by appropriately omitting some of the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, the effects mentioned in the exemplary embodiments shown below are examples of effects expected in the exemplary embodiments, and do not define the scope of the present invention. In other words, embodiments that do not exhibit the effects mentioned in the exemplary embodiments shown below may also be included in the scope of the present invention.

[0014] [First Exemplary Embodiment] A first exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. This exemplary embodiment is a basic form for each of the exemplary embodiments described below. Note that the scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in the drawings referenced to explain this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise.

[0015] <Overview of Information Processing Device 1> First, an overview of the information processing device 1 according to this exemplary embodiment will be described. The information processing device 1 according to this exemplary embodiment is an information processing device that acquires output values ​​from multiple experts (models) and makes decisions by referring to the acquired output values. Furthermore, the information processing device 1 according to this exemplary embodiment sequentially acquires output values ​​and makes decisions. For example, in round t, it acquires output value P1t from expert 1 and output value P2t from expert 2, and derives a decision-making result (t) for round t by referring to the acquired output values ​​P1t and P2t. The information processing device 1 then updates parameters for decision-making, acquires output values ​​for the next round, and derives a decision-making result for the next round. Note that t is an index representing the number of repetitions and can also be interpreted as an index indicating timing.

[0016] In this exemplary embodiment, the "expert" may be any of hardware, software, or a living organism that outputs some kind of output value. For example, the "expert" may be hardware such as a predicted value derivation device that outputs a predicted value as an output value, software such as a predicted value derivation algorithm that outputs a predicted value as an output value, or a person who outputs a predicted value as an output value using some method. Furthermore, the "expert" is not limited to an entity that outputs a predicted value, but may also be a generative model that outputs some kind of generation result as an output value, or a control model that outputs some kind of control value as an output value. Furthermore, the information processing device 1 according to this exemplary embodiment may be configured to include an "expert" or may be configured to acquire output values ​​from an external "expert." The "expert" may also be referred to as a "model," "agent," or the like.

[0017] Furthermore, in this exemplary embodiment, the "output value" may be anything. As an example, the "output value" may be, for example, a predicted value related to supply or demand, or a predicted value related to other cases (events). Furthermore, the "output value" does not have to be related to prediction. For example, the "output value" in this exemplary embodiment may be an output value related to some parameter referenced by the information processing device 1. The information processing device 1 according to this exemplary embodiment can be applied to the general process of making decisions regarding a target event by referring to output values ​​from each of a plurality of experts.

[0018] Furthermore, in this exemplary embodiment, "intention" refers to some information related to a target event, and is not limited to the intention of a living organism (person). For example, in an application scenario in which demand for a target product is predicted, a predicted value of future demand is an example of an "intention" determined by the information processing device 1 according to this exemplary embodiment, or an example of a "decision-making result" derived by the information processing device 1. The information processing device 1 according to this exemplary embodiment can also be expressed as a decision-making device or a decision-making result derivation device. The "decision-making result" is also called an "optimization solution" or an "optimization result."

[0019] Furthermore, in this exemplary embodiment, a loss value may be provided (acquired) corresponding to an output value provided by each of the multiple experts. Here, the loss value may be acquired by observation depending on the "decision-making result" by the information processing device 1, or may not be acquired by observation. The information processing device 1 may acquire a loss value that cannot be acquired by observation by derivation. As an example, what loss value can be "observed" depending on what "decision-making result" can be expressed by a directed graph structure called a feedback graph, but this does not limit the present exemplary embodiment.

[0020] In this exemplary embodiment, the loss value can be expressed as the difference between the output value (predicted value) and the observed value (actually measured value), for example, but this does not limit the present exemplary embodiment. The loss value may also be the difference between the predicted value and another predetermined value. Furthermore, the loss value may be an estimated value related to the loss. Furthermore, the term "loss value" may include the concept of "reward." For example, the loss value may be expressed as the reward value with the sign reversed (the reward value multiplied by a negative constant). Therefore, the loss value according to this exemplary embodiment may be read as the reward value.

[0021] <Configuration of information processing device 1> Next, the configuration of the information processing device 1 will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of the information processing device 1. As shown in Fig. 1, the information processing device 1 includes an acquisition unit 11 and a derivation unit 12.

[0022] (Acquisition Unit 11) The acquisition unit 11 acquires output values ​​obtained from each of the multiple models (experts). Here, as described above, the "output value" may be, for example, a predicted value related to supply and demand, or an output value related to other cases (events). The "output value" may also be an output value related to some parameter referenced by the information processing device 1. The acquisition unit 11 may be configured to acquire the output value for each round in the sequential decision-making process, but this does not limit the present exemplary embodiment.

[0023] (Derivation Unit 12) The derivation unit 12 executes: a plurality of first derivation processes that derive a first optimal solution by referring to the output values ​​acquired by the acquisition means; and a second derivation process that derives a second optimal solution according to the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes. Here, the first derivation process is sometimes referred to as a base algorithm, and the second derivation process is sometimes referred to as a master algorithm, but these terms do not limit this exemplary embodiment.

[0024] Furthermore, the reliability is an index indicating the degree to which the output value of each expert is reflected in the decision-making process. The reliability may also be expressed as an index indicating the degree to which the first optimal solution obtained by each first derivation process is reflected in the decision-making process. For example, the reliability can be expressed as a relative weight calculated on the predicted value of each expert, or as a relative weight calculated on each of the first optimal solutions.

[0025] As described above, the information processing device 1 according to this exemplary embodiment acquires output values ​​from one or more models (experts) and executes a plurality of first derivation processes to derive a first optimal solution by referring to the acquired output values, and a second derivation process to derive a second optimal solution based on the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes. In other words, the information processing device 1 according to this exemplary embodiment derives an optimal solution (decision-making) through hierarchical processing using reliability. Therefore, the information processing device 1 according to this exemplary embodiment can derive a more appropriate decision-making result (optimal solution).

[0026] <Flow of Information Processing Method S1> Next, the flow of the information processing method S1 according to the present exemplary embodiment 1 will be described with reference to Fig. 2. Fig. 2 is a flow chart showing the flow of the information processing method S1.

[0027] (Step S11) In step S11, the acquisition unit 11 acquires output values ​​obtained from one or more models (experts).

[0028] (Step S12) In step S12, the derivation unit 12 executes: a plurality of first derivation processes that derive a first optimal solution by referring to the output values ​​acquired in step S11; and a second derivation process that derives a second optimal solution in accordance with the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes.

[0029] The information processing device 1 performs the processes of steps S11 and S12 in a certain round, and then performs the processes of steps S11 and S12 in the next round.

[0030] FIG. 3 is a diagram for schematically illustrating sequential decision-making processing by the information processing method S1 according to this exemplary embodiment. As shown in FIG. 3 , the information processing device 1 acquires output values ​​provided by each expert for multiple experts in a certain round. Then, the information processing device 1 derives a decision-making result (optimal solution) by referring to each acquired output value. The derived decision-making result (optimal solution) is then executed, and each expert provides an output value for the next round. A loss value corresponding to the next round may also be observed. As described above, the loss value may be acquired by observation or may not be acquired by observation, depending on the decision-making result (optimal solution) by the information processing device 1. The information processing device 1 may acquire a loss value that cannot be acquired by observation by derivation. In this manner, the information processing device 1 sequentially derives decision-making results.

[0031] As described above, the information processing method S1 according to this exemplary embodiment executes a plurality of first derivation processes that acquire output values ​​from one or more models (experts) and derive a first optimal solution by referring to the acquired output values, and a second derivation process that derives a second optimal solution based on the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes. In other words, the information processing method S1 according to this exemplary embodiment derives an optimal solution (decision-making) through hierarchical processing using reliability. Therefore, the information processing method S1 according to this exemplary embodiment can derive a more appropriate decision-making result (optimal solution).

[0032] <Configuration of Information Processing System 100> Next, the configuration of the information processing system 100 according to this exemplary embodiment will be described with reference to Fig. 4. Fig. 4 is a block diagram showing the configuration of the information processing system 100. As shown in Fig. 4, the information processing device 100 includes an information processing device 1 and a terminal device 2 that are communicably connected to each other. Since the components of the information processing device 1 have been described above, description thereof will be omitted here.

[0033] 4, the terminal device 2 includes an execution unit 21 and a loss value acquisition unit 22. The execution unit 21 executes the decision-making result (optimal solution) derived by the information processing device 1, or a process corresponding to the decision-making result (optimal solution). As an example, if the decision-making result predicts X units of product A as today's demand, the execution unit 21 places an order for X units of product A.

[0034] <Flow of Information Processing Method S100> Next, the flow of the information processing method S100 according to the present exemplary embodiment 1 will be described with reference to Fig. 5. Fig. 5 is a flow diagram showing the flow of the information processing method S100 executed by the information processing system 100.

[0035] (Steps S11-1, S12-1) As shown in FIG. 4, in step S11-1, the acquisition unit 11 acquires output values ​​obtained from each of a plurality of experts (models).

[0036] In step S12-1, the derivation unit 12 executes: a plurality of first derivation processes that derive a first optimal solution by referring to the output values ​​acquired in step S11-1; and a second derivation process that derives a second optimal solution in accordance with the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes.

[0037] (Steps S21-1, S22-1) In step S21-1, the execution unit 21 of the terminal device 2 executes the decision-making result (more specifically, the second optimal solution) derived in step S12-1, or a process corresponding to the decision-making result. In step S22-1, the terminal device 2 provides the result of the execution to the information processing device 1. If a loss value is obtained as a result of the executed decision-making, the result of the execution may include the loss value.

[0038] (Steps S11-2, S12-2) In step S11-2, the acquisition unit 11 acquires output values ​​for the current round (round t=2) obtained from each of the multiple experts (models). Here, each predicted value may be, for example, the loss value acquired by the loss value acquisition unit 22 in step S22-1 or the loss value derived by the information processing device 1.

[0039] In step S12-2, the derivation unit 12 executes: a plurality of first derivation processes that derive a first optimal solution by referring to the output values ​​acquired in step S11-2; and a second derivation process that derives a second optimal solution in accordance with the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes. The decision-making result derived in step S12-2 (more specifically, the second optimal solution) is provided to the terminal device 2 and executed in step S21-2.

[0040] As described above, the information processing method S100 according to this exemplary embodiment executes a plurality of first derivation processes that acquire output values ​​from one or more models (experts) and derive a first optimal solution by referring to the acquired output values, and a second derivation process that derives a second optimal solution based on the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes. In other words, the information processing method S100 according to this exemplary embodiment derives an optimal solution (decision-making) through hierarchical processing using reliability. Therefore, the information processing method S100 according to this exemplary embodiment can derive a more appropriate decision-making result (optimal solution).

[0041] [Exemplary Embodiment 2] A second exemplary embodiment, which is one example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiment will be assigned the same reference numerals, and their description will be omitted as appropriate. The scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in each drawing referenced to describe this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise.

[0042] <Configuration of Information Processing System 100A> The configuration of the information processing system 100A according to this exemplary embodiment will be described with reference to Fig. 6. Fig. 6 is a block diagram showing the configuration of the information processing system 100A. As shown in Fig. 6, the information processing system 100A includes an information processing device 1A and a terminal device 2A. Also, as shown in Fig. 6, the information processing device 1A and the terminal device 2A are configured to be able to communicate with each other via a network N. Here, the specific configuration of the network N does not limit this exemplary embodiment, but as an example, a wireless local area network (LAN), a wired LAN, a wide area network (WAN), a public line network, a mobile data communication network, or a combination of these networks can be used.

[0043] <Configuration of Information Processing Apparatus 1A> The configuration of the information processing apparatus 1A according to this exemplary embodiment will be described with reference to Fig. 6. Fig. 6 is a block diagram showing the configuration of the information processing apparatus 1A.

[0044] As shown in FIG. 6, the information processing device 1A includes a control unit 10A, a storage unit 15A, and a communication unit 16A.

[0045] The communication unit 16A communicates with devices external to the information processing device 1A. As an example, the communication unit 16A communicates with the terminal device 2A. The communication unit 16A transmits data supplied from the control unit 10A to the terminal device 2A, and supplies data received from the terminal device 2A to the control unit 10A.

[0046] (Storage Unit 15A) The storage unit 15A stores various types of information referenced by the control unit 10A and various types of information derived by the control unit 10A. As an example, the storage unit 15A stores: output value information PI including output values ​​from each of the multiple experts; loss value information LI including loss values ​​corresponding to the output values ​​from each of the multiple experts; reliability information CI indicating at least one of the reliability of each of the multiple experts and the reliability of each of multiple first derivation units 121-1, 121-2, ... described below; first optimal solutions (first decision-making results) DR1-1, DR1-2, ... derived by each of multiple first derivation units 121-1, 121-2, ... described below; and a second optimal solution (second decision-making result) DR2 derived by a second derivation unit 122 described below.

[0047] As described in exemplary embodiment 1, the information processing device system 100A according to this exemplary embodiment may be configured to include an "expert (model)" or may be configured to obtain the output values ​​from an "expert (model)" outside the system.

[0048] (Control Unit 10A) As shown in FIG. 6, the control unit 10A includes an acquisition unit 11 and a derivation unit 12.

[0049] (Acquisition Unit 11) As in the first exemplary embodiment, the acquisition unit 11 acquires output values ​​obtained from each of a plurality of experts (models). Here, if a loss value corresponding to the output value can be acquired by observation, the acquisition unit 11 may further acquire the loss value. The output values ​​and loss values ​​have been explained in the first exemplary embodiment, and therefore similar explanations will be omitted. Specific examples of the output values ​​and loss values ​​will be described later.

[0050] 6 , the derivation unit 12 includes a plurality of first derivation units 12-1, 121-2, ..., and a second derivation unit 122. As in the first exemplary embodiment, the derivation unit 12 according to this exemplary embodiment performs the following: a plurality of first derivation processes that derive a first optimal solution by referring to the output values ​​acquired by the acquisition unit 11, and a second derivation process that derives a second optimal solution in accordance with the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes. In this exemplary embodiment, each of the plurality of first derivation processes is performed by each of the plurality of first derivation units 121-1, 121-2, ..., and the second derivation process is performed by the second derivation unit 122.

[0051] In other words, each first derivation unit 121-j (j is an index for distinguishing the first derivation units from each other) obtains each output value (m t ) to obtain the first optimal solution (first decision-making result) (p t The second derivation unit 122 derives the first optimal solution (first decision-making result) (p t (j)) and the reliability (w t (j)) and derive a second optimal solution (second decision-making result).

[0052] The first derivation process may be referred to as a base algorithm, and the second derivation process may be referred to as a master algorithm, but these terms do not limit this exemplary embodiment. Each of the first derivation units 121-1, 121-2, ... may be simply referred to as the first derivation unit 121. The first derivation process and the second derivation process may also be referred to as online machine learning processes (online learning algorithms) that refer to the output values ​​sequentially acquired.

[0053] Furthermore, as described in the first exemplary embodiment, the reliability is an index indicating the extent to which the output value of each expert (model) is reflected in the decision-making process. The reliability may also be expressed as an index indicating the extent to which the first optimal solution obtained by each first derivation process is reflected in the decision-making process. For example, the reliability can be expressed as a relative weight calculated for the predicted value of each expert, or as a relative weight calculated for each of the first optimal solutions.

[0054] Furthermore, the derivation unit 12 may derive the reliability by referring to at least one of the output value acquired by the acquisition unit 11 and the first optimal solutions derived by each of the plurality of first derivation processes. Furthermore, the reliability derivation process may be executed as part of the second derivation process. According to this configuration, the reliability is derived by referring to at least one of the output value and the first optimal solution, so that a suitable reliability can be derived. Furthermore, by referring to such a suitable reliability, a suitable second optimal solution can be derived. A more specific example of the reliability derivation process will be described later.

[0055] The derivation unit 12 may also be configured to initialize the reliability at a predetermined timing. As an example, the derivation unit 12 may be configured to initialize the reliability (set the reliability to an initial value) every time the second optimal solution is derived a predetermined number of times (e.g., 100 times). In this way, the derivation unit 12 can appropriately respond to changes in the environment by initializing the reliability at a predetermined timing. Note that the initialization of the reliability may be executed as part of a process for restarting at least one of the first derivation process and the second derivation process described above.

[0056] Furthermore, the derivation unit 12 may determine the timing for initializing the reliability depending on the period of the environmental change. As an example, when the decision-making process by the information processing device system 100A according to this exemplary embodiment is applied to a situation in which the environmental change occurs approximately once a month, the derivation unit 12 may be configured to initialize the reliability every month.

[0057] The derivation unit 12 may also perform a process of estimating a loss value corresponding to the first optimal solution. Here, the process of estimating the loss value may be executed as part of the second derivation process described above. Furthermore, the derivation unit 12 may also perform a process of updating the first optimal solution by referring to the derived loss value in the first derivation process. As described above, depending on the result of decision-making, the loss value corresponding to the output value may or may not be obtainable by observation. According to the above configuration, the derivation unit 12 estimates a loss value corresponding to the first optimal solution and updates the first optimal solution by referring to the estimated loss value. Therefore, even if the loss value corresponding to the output value cannot be obtained by observation, a suitable optimal solution (decision-making result) can be derived.

[0058] The derivation unit 12 may be configured to calculate the loss value as an unbiased estimator. With this configuration, the optimization is derived by referring to the loss value calculated as the unbiased estimator, so that a more suitable optimal solution (decision-making result) can be derived.

[0059] Furthermore, the derivation unit 12 may derive the decision-making result by further referring to a parameter for correcting the loss value or the reliability. With this configuration, at least one of the reliability of each expert (model), the reliability of each first derivation process, and the loss value can be appropriately corrected using the parameter, thereby enabling the information processing device 1 to be operated with a minimum level of accuracy guaranteed, i.e., so-called conservative operation.

[0060] 6, the terminal device 2A includes a control unit 20A, an execution unit 21, and a communication unit 26. As an example, the terminal device 2A can be specifically realized as a checkout terminal located in a store, an inventory management terminal located in a warehouse, or the like, but this does not limit the present exemplary embodiment.

[0061] The communication unit 26 communicates with devices external to the terminal device 2A. As an example, the communication unit 26 communicates with the information processing device 1A. The communication unit 26 transmits data supplied from the control unit 20A to the information processing device 1A, and supplies data received from the information processing device 1A to the control unit 20A.

[0062] The execution unit 21 performs processing corresponding to the decision-making result derived by the derivation unit 12 or the decision-making result. As an example, the execution unit 21 executes the second decision-making result (second optimal solution) derived by the derivation unit 12. The execution unit 21 may be configured to execute at least a portion of the multiple first decision-making results derived by the derivation unit 12, instead of or together with the second decision-making result derived by the derivation unit 12. Furthermore, the execution unit 21 may be configured to display the second decision-making result derived by the derivation unit 12 and at least a portion of the multiple first decision-making results derived by the derivation unit 12, or may be configured to display the second decision-making result and the multiple first decision-making results so that they can be compared with each other.

[0063] 6, the control unit 20A includes a loss value acquisition unit 22 and a loss value providing unit 23. When a loss value for each expert (model) can be acquired by observation (measurement) as a result of execution by the execution unit 21, the loss value acquisition unit 22 acquires the loss value. The loss value providing unit 23 provides the loss value acquired by the loss value acquisition unit 22 to the information processing device 1A via the communication unit 26.

[0064] FIG. 7 is a diagram for schematically illustrating sequential decision-making processing by the information processing system 100A according to this exemplary embodiment. As shown in FIG. 7 , in a certain round, the information processing device 1A acquires output values ​​associated with each expert (model) for multiple experts and derives a decision-making result by referring to the acquired output values. Here, if a loss value associated with the power value can be acquired through observation, the loss value may be further acquired and the acquired loss value may be further referenced to derive a decision-making result. Furthermore, as an example, the decision-making processing performed by the information processing device 1A is a hierarchical decision-making processing performed by multiple first derivation units 12-1, 121-2, ..., and a second derivation unit 122, as described above.

[0065] Then, the derived decision-making result (optimal solution) is executed, and a loss value corresponding to the next round can be observed. If a loss value can be obtained by observation, the loss value and an output value corresponding to the loss value are provided to the information processing device 1A and are referenced in the decision-making process in the next round. In this way, the information processing system 100A sequentially derives decision-making results.

[0066] (Specific Processing Example by Information Processing Device 1A) A specific processing example by information processing device 1A will be described below. The information processing device 1A according to this example executes the following algorithm 1 and algorithm 2. Algorithm 1 is mainly executed by each of the multiple first derivation units 121-1, 121-2, ..., and algorithm 2 is mainly executed by second derivation unit 122. (Processing of Algorithm 1) First, the processing of Algorithm 1 will be described. As shown at the beginning of Algorithm 1, the acquisition unit 11 acquires a graph G, a parameter η, and a parameter T. Here, as shown at the beginning of Algorithm 1, the graph G is defined by a vertex V and an edge (side, link) E. In this exemplary embodiment, the graph G is, as an example, a directed graph called a feedback graph.

[0067] In the feedback graph, each vertex V corresponds to a possible option of the decision-making result, and each edge indicates the observability of the loss. As an example, a directed edge E(i,j) from vertex V(i) to vertex V(j) indicates that if the decision-making result indicates option i (if the player selects option i), the loss value associated with option j is observable. Here, the case of i=j is also called a self-loop.

[0068] The upper part of Figure 8 shows Example 1 of a feedback graph. This example is a feedback graph corresponding to the problem statement, "If the apples are tasty, it is better to ship them, but if they are not tasty, it is better not to ship them." As shown in the upper part of Figure 8, the feedback graph of this example includes an option V(1) of not shipping the apples and an option V(2) of shipping the apples. If option V(1) of not shipping the apples is selected, the user can taste the apples and determine whether they are tasty. Therefore, it is possible to determine whether it would have been better to ship the apples. This indicates that if option V(1) of not shipping the apples is selected, the following occurs: - the loss value associated with option V(1) of not shipping the apples is observable (corresponding to edge E(1,1)), and - the loss value associated with option V(2) of shipping the apples is also observable (corresponding to edge E(1,2)).

[0069] On the other hand, if option V(2) of shipping the apples is selected, the apples cannot be tasted and therefore it is not known whether they are delicious or not. Therefore, if option V(2) of shipping the apples is selected, no loss value can be observed. This corresponds to the absence of an outward arrow (outward edge) originating from option V(2) in the feedback graph.

[0070] The lower part of Figure 8 shows Example 2 of a feedback graph. This example is a feedback graph corresponding to the problem statement, "We want to order the appropriate quantity of a certain product." As shown in the lower part of Figure 8, the feedback graph of this example has option V(1) of ordering 300 units, option V(2) of ordering 200 units, and option V(3) of ordering 100 units. This shows that when a larger quantity is ordered, the loss value when a smaller quantity is ordered can be observed, but when a smaller quantity is ordered, the loss value when a larger quantity is ordered cannot be observed.

[0071] In this way, feedback graphs can be used to express various problem settings. For example, bandit feedback corresponds to the case where each of multiple vertices has only self-loops, while full feedback (full information feedback) corresponds to the case where all pairs included in multiple vertices are bidirectionally connected and all vertices have self-loops.

[0072] Since Algorithm 1 according to this exemplary embodiment is configured to be applicable to any feedback graph, the information processing system 100A according to this exemplary embodiment can be applied to a very wide range of problem settings.

[0073] Returning to the explanation of Algorithm 1, the parameter η acquired by the acquisition unit 11 serves as a learning rate referenced for deriving the reliability, as will be described later. The parameter T acquired by the acquisition unit 11 is a natural number that defines the total number of rounds.

[0074] Next, as shown in “1:” of Algorithm 1, the first derivation unit 121 derives a first optimal solution (first decision-making result) p t The parameter (weight factor) p' is used to derive t The initial value of Here, K represents the total number of options that can be taken as the first optimal solution (first decision-making result), and is also the total number of vertices |V| of the feedback graph described above. As will be understood from the explanation below, p tThe parameter (weight factor) p' is used to derive t is the output value of each expert (model) (m t ) can be regarded as a parameter representing the reliability of each of the above. However, this interpretation does not limit the present exemplary embodiment.

[0075] In addition, in the above formula (1), the numerator on the right side is where [K] is a K-dimensional vector (array) defined by (i.e., a set whose elements are the natural numbers 1 to K).

[0076] Next, the first derivation unit 121 executes the loop processing specified by "2:" to "7:" in Algorithm 1. The loop processing is executed while incrementing an index t indicating the round until t = T (T is the total number of rounds). In other words, the first derivation unit 121 repeats the processing specified by "3:" to "7:" in Algorithm 1 for each round.

[0077] More specifically, as shown in “3:” of Algorithm 1, the first derivation unit 121 first derives the function ψ(p) as follows: The function ψ(p) serves as a convex function that defines the Bregman divergence, which will be described later. In the above formula, p(i) is defined as In addition, N in the definition of the function ψ(p) in (i) is (in other words, for a vertex V(i) with vertex number i, a set whose element is the vertex number j of a vertex V(j) that is the starting point (starting point) of an edge entering the vertex V(i)), and |N in (i) | indicates the number of elements in the set. is satisfied if the condition in [ ] is satisfied, i.e., |N in(i)|<K is satisfied, and returns 0 otherwise. Therefore, the second term on the right-hand side of the definition equation (Equation 4) of the function ψ(p) has a non-zero value proportional to log p(i) when, for a certain vertex V(i), the number of vertices V(j) that are the origins (starting points) of edges entering the vertex V(i) is less than the total number of vertices |V|=K. This second term is also called the logarithmic barrier term.

[0078] In this manner, the Bregman information referred to by the first derivation unit 121 according to this exemplary embodiment to derive the first optimal solution is described by a convex function ψ(p), which includes the logarithmic barrier term described above. By using such a logarithmic barrier term, it is possible to perform decision-making processing that is flexibly applicable to various environments (various problem settings (various feedback graphs)).

[0079] Next, as shown in “4:” of Algorithm 1, the acquisition unit 11 acquires the output values ​​(m t The output values ​​output by each expert (model) have been described above, so a detailed description will be omitted here.

[0080] Next, as shown in “5:” of Algorithm 1, the first derivation unit 121 The first optimal solution (first decision-making result) p t Here, the first term on the right side of (Equation 8) is m t and p, and the second term on the right side is the Bregman divergence defined using the function ψ. As arguments of t In addition, Δ' in (Equation 8) is used. K teeth, is the set defined by

[0081] As shown in (Equation 8), the first derivation part 121 is t and p's inner product < m t , p> and p, p' t , and the Bregman information D defined by ψψ (p, p' t ) and p that minimizes the linear sum of p is derived as the first optimal solution (first decision-making result).

[0082] In this way, the first derivation unit 121 calculates the output values ​​(m t ) to find the first optimal solution (p t ) is derived.

[0083] Next, as shown in “6:” of Algorithm 1, the information processing system 100A calculates the first optimal solution p t Execute the calculation to obtain the loss value l t The acquisition unit 11 also observes a parameter α t Here, the first optimal solution p t The execution of the above may be performed by the execution unit 21 of the terminal device 2A, for example. The acquisition of the loss value may be performed by the loss value acquisition unit 22, for example. However, the derived first optimal solution p t At least a part of the first optimal solution p may be supplied to the second derivation unit 122 that executes the algorithm 2 described later without being executed. t Depending on what is indicated, the loss value l t As mentioned above, there are cases where the loss value l t , and a correction parameter α t At least a part of the above may be derived using a value derived by Algorithm 2 described later.

[0084] Next, as shown in “7:” of Algorithm 1, the first derivation unit 121 By this, the parameter p' t In this way, the first derivation unit 121 updates the loss value l t The first optimal solution (p t ) to derive the parameter (p' t ) is updated, where a t teeth, More specifically, at is the loss l t (i) Expert (model) output value m t (i) Correction parameter α t , and the learning rate η. As is clear from (Equation 11) and (Equation 12), the correction parameter α t is the loss value l t , output value m t , parameter p′ t , and the first optimal solution p t The parameter a defined by (Equation 12) can be regarded as a parameter for correcting at least one of the above. t can also be considered as a parameter for correction in the same sense.

[0085] In this way, the first derivation unit 121 calculates the correction parameter α t or a t Further referring to the first optimal solution (p t ) can be considered to be a configuration that updates the

[0086] (Processing of Algorithm 2) Next, the processing of Algorithm 2 will be described. As shown at the beginning of Algorithm 2, the acquisition unit 11 acquires a graph G(V, E) and a parameter T. The graph G(V, E) and the parameter T have been described above, and therefore will not be described here.

[0087] Furthermore, as indicated by "Input" at the beginning of Algorithm 2, Algorithm 2 is executed with reference to Base Algorithm B. Here, Base Algorithm B refers to Algorithm 1 described above, as an example. Algorithm 2 is executed with reference to a plurality of Base Algorithms B. In other words, the second derivation unit 122 executes the second algorithm in cooperation with each of the first derivation units 121 that execute each of the plurality of Algorithms 1.

[0088] Next, as shown in “1:” of Algorithm 2, the second derivation unit 122 uses the total number of rounds T to derive the parameter M as follows: where log 2The symbols on both sides of T indicate ceiling functions. Therefore, the parameter M is calculated by the second derivation unit 122 using a real argument (log 2 T) is set to the smallest integer value (e.g., log 2 (If T=3.3..., then M=4 is set.) Here, the parameter M has the meaning of the total number of base algorithms (Algorithm 1) that Algorithm 2 refers to, but this does not limit the present exemplary embodiment.

[0089] As shown in "1:" of Algorithm 2, the second derivation unit 122 calculates the parameter η(j) as Here, the index j is an index for distinguishing between multiple base algorithms, and as shown in "1:" of algorithm 2, it takes an integer value from 1 to M, for example. Note that the parameter η(j) is the reliability (reliability of each base algorithm) w t Therefore, the second derivation unit 122 determines the reliability w t This can be expressed as setting the learning rate η(j) referred to in order to derive (j) to a value proportional to the −1 / 2 power of the total number M of base algorithms (first derivation process). According to the findings of the inventors, by setting the learning rate η(j) as described above, the reliability w t In the following description, each of the first derivation units 121-1, 121-2, ... may be expressed as 121-j using the index j.

[0090] Next, as shown in “2:” of Algorithm 2, the second derivation unit 122 calculates the weight factor w 1 ' is the parameter p 1 ' for each j Here, the weight factor w 1 ' is as shown in "2:" of Algorithm 2, where Δ M denotes the set obtained by replacing K with M in the following definition. Next, as shown in “3:” of Algorithm 2, the second derivation unit 122 calculates the Here, in the initialization process, as shown in “3:” of Algorithm 2, the second derivation unit 122 initializes the base algorithm B j This includes a process of passing the values ​​of G, η(j), and T to

[0091] Next, the second derivation unit 122 executes the loop processing specified by "4:" to "13:" in Algorithm 2. The loop processing is executed while incrementing an index t indicating the round until t = T. In other words, the second derivation unit 122 repeats the processing specified by "5:" to "13:" in Algorithm 1 for each round.

[0092] More specifically, as shown in "5:" of Algorithm 2, the acquisition unit 11 first calculates the predicted value m t The second derivation unit 122 obtains the predicted value m t supply.

[0093] Next, as shown in “6:” of Algorithm 2, the acquisition unit 11 acquires each first optimal solution (first decision-making result) p by each base algorithm (each algorithm 1). t,j Then, the second derivation unit 122 obtains the parameter h t (j) Here, < , > indicates the dot product.

[0094] Next, as shown in “7:” of Algorithm 2, the second derivation unit 122 The reliability vector w indicating the reliability of each base algorithm (each algorithm 1) (each first derivation means 121-j) is calculated by t Here, D φrepresents the Bregman divergence described above. However, the convex function φ that defines the Bregman divergence is In this way, and as partly mentioned above, w' t Is the reliability w t In this way, the second derivation unit 122 calculates the output value (m t ) and the first optimal solution (p t ) and the reliability (w t (j)) is derived.

[0095] Next, as shown in “8:” of Algorithm 2, the second derivation unit 122 derives the first optimal solution (first decision-making result) p by each base algorithm (each algorithm 1). t,j , a reliability vector w indicating the reliability of each base algorithm t Each component of the reliability w t (j) Using The second optimal solution (second decision-making result) p t is derived.

[0096] In this way, the second derivation unit 122 calculates the first optimal solution (p t,j ) and the reliability (w t (j)) and the second optimal solution (p t More specifically, as expressed by the above formula, the second derivation unit 122 executes a second derivation process (master algorithm) to derive each of the first optimal solutions (first decision-making results) p t,j and the reliability w of each of the plurality of first derivation parts 121-j is a weighted sum of t The second optimal solution (second decision-making result) p t is derived.

[0097] As described above, the second derivation unit 122 calculates the first decision-making result (p t,j) and the reliability (w t (j)) and the second decision-making result p t Here, the first decision-making result is the output value (m t ) is referred to. Therefore, with the above configuration, it is possible to derive a suitable decision-making result by hierarchical processing by the first derivation unit 121 and the second derivation unit 122 with reference to the output values ​​provided by each expert (model). Next, as shown in "9:" of Algorithm 2, the second derivation unit 122 derives the second optimal solution p t Depending on the option i t More specifically, the second derivation unit 122 identifies the second optimal solution p t With the probability according to the probability distribution shown by t where i is and i t indicates the i in the t step. out (i) is In other words, N out (i) is a set whose elements are the vertex number j of the vertex V(j) that is the end point of the edge going out from the vertex V(i) with the vertex number i. Next, as shown in "10:" of Algorithm 2, the derived second optimal solution (second decision-making result) p t Option i selected according to t is executed by the execution unit 21 of the terminal device 2A, for example. Then, the second optimal solution p t The loss value l corresponding to t (In other words, the selected option i t The loss value l corresponding to t ) is acquired by the loss value acquisition unit 22 of the terminal device 2A and provided to the information processing device 1A.

[0098] As shown in "10:" of Algorithm 2, the option i t The execution of This is performed for all options i that satisfy out(i t ) is, as mentioned above, the vertex number i t The vertex V(i t ) for the vertex V(i t ) is a set whose elements are the vertex number j of the vertex V(j) that is the end point of the edge going out from t The loss value l corresponding to t is an observable loss value. Next, as shown in “11:” of Algorithm 2, the second derivation unit 122 calculates the loss value l obtained by observation in “10:” of Algorithm 2. t , and the output value m of the expert (model) t See, The loss value with a hat (^l t ) is derived. Here, P t (i) is the first optimal solution p derived by each of the first derivation processes (base algorithms). t Using (j) In addition, in the first term on the right side of (Equation 26), is satisfied if the condition in [ ] is satisfied, i.e., index i is in set N out (i t ) element (in other words, the loss value l t is obtained by observation), and returns 0 otherwise. In this way, the loss value with a hat (^l t )teeth, (where E t [ ] represents an expected value). That is, the second derivation unit 122 calculates the loss value with a hat (^l t ) as an unbiased estimate. In this way, the second derivation unit 122 derives the loss value l t Whether or not is obtained by observation, the hatted loss value (^l t ) can be suitably derived, so that the loss value l tWhether or not the loss value l is obtained by observation, the information processing system 100A according to this exemplary embodiment can perform decision-making (derive an optimal solution) in an appropriate manner. In other words, the information processing system 100A according to this exemplary embodiment can perform decision-making processing that is flexibly applicable to various environments (various problem settings (various feedback graphs)). Note that in this specification, the loss value l obtained by observation is t , and the hatched loss value (^l t ) are sometimes simply referred to as loss values.

[0099] Next, as shown in “12:” of Algorithm 2, the second derivation unit 122 calculates the loss value (^l t ), and the correction parameter α t Base algorithm B j Here, the correction parameter α t As shown in "12:" of Algorithm 2, for example, In other words, the second derivation unit 122 derives the correction parameter α t The loss value ( t ) and optimal solution (decision result) p t The loss value (^l) provided in this step is derived as the inner product of t ) is obtained in "6:" of Algorithm 1 as an example. t In addition, the α provided in this step corresponds to t (j) is the correction parameter α obtained in “7:” of Algorithm 1. t It corresponds to.

[0100] Next, as shown in "13:" of Algorithm 2, the second derivation unit 122 calculates the weight factor w' in round t. t the weight factor w′ in round t+1 t+1 More specifically, the second derivation unit 122 updates the weight factor w′ in round t+1 to t+1 of, Here, g t teeth, is defined by b t teeth, is defined by

[0101] In addition, reliability w t The parameter (weight factor) w' is used to derive t can also be considered as a parameter representing the reliability of each base algorithm (the reliability of each first optimal solution), although this interpretation does not limit the present exemplary embodiment.

[0102] As described above, in this exemplary embodiment, the information processing device 1 executes a plurality of first derivation processes that acquire output values ​​from one or more models (experts) and derive a first optimal solution by referring to the acquired output values, and a second derivation process that derives a second optimal solution based on the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes. In other words, the information processing device 1 according to this exemplary embodiment derives an optimal solution (decision-making) through hierarchical processing using reliability. Therefore, the information processing device 1 according to this exemplary embodiment can derive a more appropriate decision-making result (optimal solution).

[0103] The reliability may be initialized at a predetermined timing. Such a configuration allows the system to respond appropriately to changes in the environment. As described above, the information processing system 100A according to this exemplary embodiment is applicable to any feedback graph G(V, E). In other words, the information processing system 100A according to this exemplary embodiment is applicable whether or not the loss value can be observed. More specifically, as described above, in this exemplary embodiment, the loss value (^l t ) may be derived by estimation (more specifically, as an unbiased estimator), thereby allowing the information processing system 100A to be suitably applied to any feedback graph G(V, E). In other words, the information processing system 100A is capable of handling any feedback graph G(V, E) and is also capable of suitably responding to changes in the environment.

[0104] (Explanation of More Specific Effects of Information Processing System 100A) More specific effects of the information processing system 100A will be described below with reference to FIG. 9. FIG. 9 shows an example in which the information processing system 100A is applied to a decision-making problem regarding the shipping of apples. In this problem setting, there are two options: to ship the apples or not to ship the apples (corresponding to K=2 in the above-mentioned algorithm 1), and the value of the parameter T is set to T=200. In other words, the total number of decision-making times is set to 200. In addition, each loss is set as follows: Loss when the apples are tasty: 1 if they are not shipped, 0 if they are shipped Loss when the apples are not tasty: 0 if they are not shipped, 1 if they are shipped

[0105] As an environmental change, the probability that the apple is delicious is set to be 0.9 on average in the first 100 times, and 0.1 on average in the last 100 times. In addition, the information processing system 100A according to this exemplary embodiment is configured to initialize the reliability (initialize the algorithm) every 50 times. In addition, the output value (m t ) was always set to 0.

[0106] Figure 9 shows the transition of loss values ​​for the information processing system 100A according to the present exemplary embodiment and the transition of loss values ​​for the configuration according to the comparative example, based on the problem set described above. As shown in Figure 9, the loss values ​​for the information processing system 100A are significantly smaller than the loss values ​​for the configuration according to the comparative example, and it can be seen that the system is able to quickly adapt to environmental changes that occur on the 100th iteration. It can also be seen that the loss values ​​converge quickly even after the reliability is initialized every 50 iterations. As such, it can be seen that the information processing system 100A according to the present exemplary embodiment has significantly higher adaptability to environmental changes than the configuration according to the comparative example.

[0107] [Application Example] A more specific application example of the information processing system 100A according to this exemplary embodiment will be described below. Fig. 10 is a diagram schematically showing processing by the information processing device 1 according to this example.

[0108] 8, the information processing device 1 according to this example makes a decision regarding matching between a plurality of medical professionals and a plurality of hospitals (derives a decision-making result). Here, as described above, the information processing device 1 according to this example acquires output values ​​associated with each expert (model) for a plurality of experts in a certain round, and derives a decision-making result by referring to the acquired output values.

[0109] The specific configuration of the information processing device 1 in this example does not limit this application example, but may be the same as the information processing device 1 described in exemplary embodiment 1, or may be the same as the information processing device 1A described in exemplary embodiment 2.

[0110] Furthermore, the decision-making process performed by the information processing device 1 in this example can be, as an example, a hierarchical decision-making process using multiple first derivation units 121-1, 121-2, ... and a second derivation unit 122, as described with respect to the information processing device 1A, but this does not limit this application example.

[0111] Then, the derived decision-making result is executed, and a loss value corresponding to the next round is obtained by observation or by derivation. The loss value and an output value corresponding to the loss value are provided to the information processing device 1 according to this example and are referred to in the decision-making process in the next round.

[0112] Examples of inputs to each expert and output values ​​of each expert according to this example are as follows: As described in the first exemplary embodiment, the information processing system according to this example may be configured to include an "expert (model)" or may be configured to obtain predicted values ​​from an external "expert (model)."

[0113] (Example of input to the expert) - Features of each hospital and each medical worker observed at time t (round t) - Number of patients visiting each hospital in the previous round (round t-1) (Example of output from the expert) - Hospital assigned to each medical worker at time t (round t).

[0114] (Processing Flow According to This Example) An example of processing by the information processing system 100 according to this example will be described below.

[0115] First, a person in charge of inputting information at the hospital inputs information (also referred to as hospital data) such as diagnosis status, availability of hospital rooms, medical departments, and consultation hours via a terminal device 2A or the like into the information processing system 100. The input information is stored in the storage unit 15A, for example, and referenced by the control unit 10A. The lower part of Fig. 11 shows an example of hospital data managed by the information processing system 100 according to this example.

[0116] Next, each medical worker inputs their own data (specialty, years of service, preferred hospital, etc.) (also referred to as medical worker data) into the information processing system 100. The input information is stored in the storage unit 15A, for example, and referenced by the control unit 10A. The upper part of Fig. 11 is an example of medical worker data managed by the information processing system 100 according to this example.

[0117] Next, the information processing device 1 according to this example derives a decision-making result regarding the optimal matching of hospitals and medical professionals by referring to the hospital data and the medical professional data. As an example, each of the multiple experts (models) according to this example calculates an output value by referring to the hospital data and the medical professional data. Then, the information processing device 1 according to this example derives a decision-making result regarding the optimal matching of hospitals and medical professionals by referring to these output values.

[0118] The information processing system 100 according to the present example then proposes optimal hospital candidates to the medical worker via the terminal device 2A, etc. As an example, the information processing system 100 according to the present example presents optimal hospital candidates to the medical worker via a display panel, etc., provided in the terminal device 2A. Furthermore, the information processing system 100 according to the present example performs work-related registration for each medical worker.

[0119] For example, the information processing system 100 according to this example records the number of patients visiting each hospital each month (each round). Then, the information processing system 100 according to this example determines the assignment destination for the next round (next round) based on a loss value corresponding to the number of patients visiting each hospital. While specific examples of loss values ​​are not limited to this example, a loss value corresponding to the degree of congestion at a hospital can be used as an example.

[0120] More specifically, the loss value can be calculated by (the actual number of patients visiting each hospital) - (the number of medical professionals assigned to each hospital x the number of patients that each medical professional can examine).

[0121] However, some hospitals may not provide the number of patients visiting the hospital due to privacy considerations, etc. In such cases, the loss value for that hospital (the decision-making result) is not observable.

[0122] As described above, the information processing system 100 according to this example can make appropriate decisions whether or not the loss value is observable, and therefore can appropriately match hospitals and medical professionals even in the above-mentioned cases. Furthermore, hospital data and medical professional data may change over time. The information processing system 100 according to this example can appropriately match hospitals and medical professionals even when such environmental changes occur.

[0123] [Example of implementation by software] The control blocks (particularly the acquisition unit 11 and the derivation unit 12) of the information processing device 1, 1A and the terminal device 2, 2A may be implemented by a logic circuit (hardware) formed on an integrated circuit (IC chip) or the like, or may be implemented by software.

[0124] In the latter case, the information processing device 1, 1A, and the terminal device 2, 2A each include a computer that executes instructions from a program, which is software that realizes each function. This computer includes, for example, at least one processor (control device) and at least one computer-readable recording medium storing the program. The object of the present invention is achieved when the processor in the computer reads and executes the program from the recording medium. The processor may be, for example, a CPU (Central Processing Unit). The recording medium may be a "non-transitory tangible medium," such as a ROM (Read Only Memory), tape, disk, card, semiconductor memory, or programmable logic circuit. The computer may also include a RAM (Random Access Memory) for loading the program. The program may also be supplied to the computer via any transmission medium capable of transmitting the program (such as a communication network or broadcast waves). Note that one aspect of the present invention may also be realized in the form of a data signal embedded in a carrier wave, in which the program is embodied by electronic transmission.

[0125] [Appendix A] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0126] (Appendix A1) An information processing device comprising: an acquisition means for acquiring an output value obtained from each of one or more models; a plurality of first derivation processes for deriving a first optimal solution by referring to the output value acquired by the acquisition means; and a derivation means for executing a second derivation process for deriving a second optimal solution in accordance with the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes.

[0127] (Supplementary Note A2) The information processing device according to Supplementary Note A1, wherein the derivation means derives the reliability by referring to at least one of the output value and the first optimal solution, and initializes the reliability at a predetermined timing.

[0128] (Supplementary Note A3) The information processing device according to Supplementary Note A1 or A2, wherein the derivation means estimates a loss value corresponding to the first optimal solution, and updates a parameter for deriving the first optimal solution by referring to the loss value.

[0129] (Supplementary Note A4) The information processing device according to Supplementary Note A3, wherein the derivation means calculates the loss value as an unbiased estimator.

[0130] (Supplementary Note A5) The information processing device according to Supplementary Note A3 or A4, wherein the derivation means further refers to a parameter for correction to update the first optimal solution.

[0131] (Supplementary Note A6) The information processing device according to any one of Supplementary Notes A1 to A5, wherein the derivation means sets a learning rate referred to in order to derive the reliability to a value proportional to the −½ power of a total number of the first derivation processes.

[0132] (Appendix A7) The information processing device according to any one of Appendices A1 to A5, wherein the Bregman divergence referred to by the derivation means to derive the first optimal solution is defined by a convex function including a logarithmic barrier term.

[0133] (Supplementary Note A8) The information processing device according to any one of Supplementary Notes A1 to A7, wherein the first derivation process and the second derivation process are online machine learning processes that refer to the output values ​​that are sequentially acquired.

[0134] [Appendix B] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0135] (Appendix B1) An information processing method including a derivation process, the derivation process including: an acquisition process in which at least one processor acquires output values ​​obtained from each of one or more models; a plurality of first derivation processes in which a first optimal solution is derived by referring to the output values ​​acquired by the acquisition process; and a second derivation process in which a second optimal solution is derived depending on the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes.

[0136] (Supplementary Note B2) The information processing method according to Supplementary Note B1, wherein in the derivation process, the at least one processor derives the reliability by referring to at least one of the output value and the first optimal solution, and initializes the reliability at a predetermined timing.

[0137] (Supplementary Note B3) The information processing method according to Supplementary Note B1 or B2, wherein in the derivation process, the at least one processor estimates a loss value corresponding to the first optimal solution, and updates parameters for deriving the first optimal solution by referring to the loss value.

[0138] (Supplementary Note B4) The information processing method according to Supplementary Note B3, wherein in the derivation process, the at least one processor calculates the loss value as an unbiased estimator.

[0139] (Supplementary Note B5) The information processing method according to Supplementary Note B3 or B4, wherein in the derivation process, the at least one processor further refers to a parameter for correction to update the first optimal solution.

[0140] (Supplementary Note B6) The information processing method according to any one of Supplementary Notes B1 to B5, wherein in the derivation process, the at least one processor sets a learning rate referred to for deriving the reliability to a value proportional to the −½ power of a total number of the first derivation processes.

[0141] (Supplementary Note B7) The information processing method according to any one of Supplementary Notes B1 to B5, wherein the Bregman divergence referred to in the derivation process to derive the first optimal solution is defined by a convex function including a logarithmic barrier term.

[0142] (Supplementary Note B8) The information processing method according to any one of Supplementary Notes B1 to B7, wherein the first derivation process and the second derivation process are online machine learning processes that refer to the output values ​​that are sequentially acquired.

[0143] [Appendix C] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0144] (Appendix C1) An information processing program that causes a computer to function as an information processing device, the information processing program causing the computer to function as: an acquisition means that acquires output values ​​obtained from each of one or more models; a plurality of first derivation processes that derive a first optimal solution by referring to the output values ​​acquired by the acquisition means; and a second derivation process that derives a second optimal solution depending on the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes.

[0145] (Supplementary Note C2) The information processing program according to Supplementary Note C1, wherein the derivation means derives the reliability by referring to at least one of the output value and the first optimal solution, and initializes the reliability at a predetermined timing.

[0146] (Supplementary Note C3) The information processing program according to Supplementary Note C1 or C2, wherein the derivation means estimates a loss value corresponding to the first optimal solution, and updates parameters for deriving the first optimal solution by referring to the loss value.

[0147] (Supplementary Note C4) The information processing program according to Supplementary Note C3, wherein the derivation means calculates the loss value as an unbiased estimator.

[0148] (Supplementary Note C5) The information processing program according to Supplementary Note C3 or C4, wherein the derivation means further refers to a parameter for correction to update the first optimal solution.

[0149] (Appendix C6) The information processing program according to any one of Appendices C1 to C5, wherein the derivation means sets a learning rate referred to for deriving the reliability to a value proportional to the −½ power of a total number of the first derivation processes.

[0150] (Appendix C7) The information processing program according to any one of Appendices C1 to C5, wherein the Bregman divergence referred to by the derivation means to derive the first optimal solution is defined by a convex function including a logarithmic barrier term.

[0151] (Supplementary Note C8) The information processing program according to any one of Supplementary Notes C1 to C7, wherein the first derivation process and the second derivation process are online machine learning processes that refer to the output values ​​that are sequentially acquired.

[0152] [Appendix D] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0153] (Appendix D1) An information processing device comprising at least one processor, the at least one processor executing a derivation process including: an acquisition process that acquires output values ​​obtained from each of one or more models; a plurality of first derivation processes that derive a first optimal solution by referring to the output values ​​acquired by the acquisition process; and a second derivation process that derives a second optimal solution depending on the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes.

[0154] The information processing device may further include a memory, and the memory may store a program for causing the at least one processor to execute each of the processes.

[0155] (Supplementary Note D2) The information processing device according to Supplementary Note D1, wherein in the derivation process, the at least one processor derives the reliability by referring to at least one of the output value and the first optimal solution, and initializes the reliability at a predetermined timing.

[0156] (Supplementary Note D3) The information processing device according to Supplementary Note D1 or D2, wherein in the derivation process, the at least one processor estimates a loss value corresponding to the first optimal solution, and updates parameters for deriving the first optimal solution by referring to the loss value.

[0157] (Supplementary Note D4) The information processing device according to Supplementary Note D3, wherein in the derivation process, the at least one processor calculates the loss value as an unbiased estimator.

[0158] (Supplementary Note D5) The information processing device according to Supplementary Note D3 or D4, wherein in the derivation process, the at least one processor further refers to a parameter for correction to update the first optimal solution.

[0159] (Appendix D6) The information processing device according to any one of appendices D1 to D5, wherein in the derivation process, the at least one processor sets a learning rate referred to for deriving the reliability to a value proportional to the −½ power of a total number of the first derivation processes.

[0160] (Appendix D7) The information processing device according to any one of appendices D1 to D5, wherein the Bregman divergence referred to in the derivation process to derive the first optimal solution is defined by a convex function including a logarithmic barrier term.

[0161] (Supplementary Note D8) The information processing device according to any one of Supplementary Notes D1 to D7, wherein the first derivation process and the second derivation process are online machine learning processes that refer to the output values ​​that are sequentially acquired.

[0162] [Appendix E] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0163] (Appendix E1) A non-transitory recording medium having recorded thereon an information processing program that causes a computer to function as an information processing device, the information processing program causing the computer to execute derivation processes including: an acquisition process that acquires output values ​​obtained from each of one or more models; a plurality of first derivation processes that derive a first optimal solution by referring to the output values ​​acquired by the acquisition process; and a second derivation process that derives a second optimal solution depending on the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes.

[0164] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.

[0165] REFERENCE SIGNS LIST 1, 1A... Information processing device 100, 100A... Information processing system 10A... Control unit 11... Acquisition unit 12... Derivation unit 121... First derivation unit 122... Second derivation unit S1, S100... Information processing method

Claims

1. An information processing apparatus comprising: acquisition means for acquiring output values obtained from each of one or more models; a plurality of first derivation processes for deriving a first optimal solution by referring to the output values acquired by the acquisition means; and second derivation means for deriving a second optimal solution according to the first optimal solutions derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes.

2. The information processing apparatus according to claim 1, wherein the derivation means derives the reliability by referring to at least one of the output value and the first optimal solution, and initializes the reliability at a predetermined timing.

3. The information processing apparatus according to claim 1 or 2, wherein the derivation means estimates a loss value corresponding to the first optimal solution, and updates a parameter for deriving the first optimal solution by referring to the loss value.

4. The information processing apparatus according to claim 3, wherein the derivation means calculates the loss value as an unbiased estimator.

5. The information processing apparatus according to claim 3 or 4, wherein the derivation means further refers to a parameter for correction and updates the first optimal solution.

6. The information processing apparatus according to any one of claims 1 to 5, wherein the learning rate referred to by the derivation means for deriving the reliability is set to a value proportional to the reciprocal of the square root of the total number of the first derivation processes.

7. The information processing apparatus according to any one of claims 1 to 5, wherein the Bregman information amount referred to by the derivation means for deriving the first optimal solution is defined by a convex function including a logarithmic barrier term.

8. The information processing apparatus according to any one of claims 1 to 7, wherein the first derivation process and the second derivation process are online machine learning processes that refer to the output values sequentially acquired.

9. An information processing method in which an information processing apparatus acquires output values obtained from each of one or more models, performs a plurality of first derivation processes for deriving a first optimal solution with reference to the acquired output values, and performs a second derivation process for deriving a second optimal solution according to the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes.

10. A program for causing a computer to function as an information processing apparatus, the program causing the computer to acquire output values obtained from each of one or more models, perform a plurality of first derivation processes for deriving a first optimal solution with reference to the acquired output values, and perform a second derivation process for deriving a second optimal solution according to the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes.

11. An information processing system including an information processing apparatus and a terminal device, the information processing apparatus including an acquisition unit that acquires output values obtained from each of one or more models, a derivation unit that performs a plurality of first derivation processes for deriving a first optimal solution with reference to the output values acquired by the acquisition unit, and a second derivation process for deriving a second optimal solution according to the first optimal solution derived by each of the plurality of first derivation processes and the reliability of each of the plurality of first derivation processes, and the terminal device including an execution unit that executes the second optimal solution derived by the information processing apparatus.

Citation Information

Patent Citations

  • Estimation apparatus and estimation method of prediction model of converter

    JP2014201770A

  • Index presentation system and index presentation method

    JP2018139050A

  • Systems and methods for AI-assisted surgery

    JP2023523560A

  • Measure determination system, measure determination method, and measure determination program

    WO2019220479A1