Learning system, learning method and program

The learning system addresses the challenge of unknown opponent utility in automated negotiations by using history information and noise addition to determine proposals, enabling effective negotiations without requiring opponent strategy knowledge.

JP7823334B2Active Publication Date: 2026-03-04NEC CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-08-31
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Existing automated negotiation systems struggle to perform effectively when the opponent's utility function or utility value is unknown.

Method used

A learning system that includes history information acquisition, evaluation value acquisition, learning, and noise addition to determine proposals based on past negotiations, allowing for automated negotiations without requiring knowledge of the opponent's utility function or strategy.

Benefits of technology

Enables automated negotiations even when the opponent's utility function or strategy is unknown, by learning from past interactions and adding noise to proposals, thus improving negotiation outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007823334000008
    Figure 0007823334000008
  • Figure 0007823334000009
    Figure 0007823334000009
  • Figure 0007823334000010
    Figure 0007823334000010
Patent Text Reader

Abstract

To enable automatic negotiation even if a negotiating partner's utility function, utility value, or strategy is unknown.SOLUTION: A learning apparatus comprises: history information acquisition means for acquiring history information of proposals made in negotiations; evaluation value acquisition means for acquiring an evaluation value for the proposal from a negotiating partner; and learning means for learning a method of determining his / her own proposal for the negotiating partner based on the history information and the evaluation value.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention , studies Learning System , studies Regarding learning methods and programs. [Background technology]

[0002] Several techniques for automated negotiation have been proposed. For example, the order-receiving side negotiation device described in Patent Document 1 determines a negotiation candidate for an order proposal from among multiple negotiation candidates based on the order-receiving side's utility value for the negotiation candidate and an estimate of the order source's utility value for the negotiation candidate. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] International Publication No. 2021 / 033302 Summary of the Invention [Problem to be solved by the invention]

[0004] It is desirable to be able to perform automated negotiations even when the opponent's utility function or utility value, or strategy, is unknown.

[0005] An example of the object of the present invention is to provide a method for solving the above-mentioned problems. Study Learning System , studies The purpose of this project is to provide learning methods and programs. [Means for solving the problem]

[0007] According to a second aspect of the present invention, a learning system includes history information acquisition means for acquiring history information of proposals made in negotiations, evaluation value acquisition means for acquiring an evaluation value for a proposal from a negotiation partner, learning means for learning a method for determining a proposal of one's own to the negotiation partner based on the history information and the evaluation value, and outputting one's own proposal to the negotiation partner based on the determination method, other party model execution means for receiving an input of one's own proposal to the negotiation partner and outputting a proposal from the negotiation partner, and evaluation means for receiving an input of a proposal from the negotiation partner and outputting the evaluation value. The counterparty model execution means includes a proposal determination means for determining a proposal that corresponds to one's own proposal to the negotiating partner, and a noise addition means for adding noise to the proposal from the negotiating partner determined by the proposal determination means.

[0009] According to a fourth aspect of the present invention, a learning method includes a computer acquiring history information of proposals made in negotiations, acquiring evaluation values ​​for proposals from a negotiating partner, and learning a method for determining its own proposal to the negotiating partner based on the history information and the evaluation values. outputting its own proposal to the negotiation partner based on the determination method, receiving an input of its own proposal to the negotiation partner and outputting a proposal from the negotiation partner, and receiving an input of a proposal from the negotiation partner and outputting the evaluation value. Including and outputting a proposal from the negotiating partner includes the computer determining a proposal corresponding to its own proposal to the negotiating partner and adding noise to the determined proposal from the negotiating partner.

[0010] According to a fifth aspect of the present invention, a program causes a computer to acquire history information of proposals made in negotiations, acquire evaluation values ​​for proposals from a negotiation partner, and learn a method for determining a proposal to the negotiation partner based on the history information and the evaluation values. and outputting its own proposal to the negotiating partner based on the determination method. To do, receiving an input of a proposal from the negotiation partner and outputting a proposal from the negotiation partner; receiving an input of a proposal from the negotiation partner and outputting the evaluation value; Run In the outputting of the proposal from the negotiation partner, the computer is caused to determine a proposal corresponding to the proposal of the negotiation partner and to add noise to the determined proposal from the negotiation partner. It is a program. [Effects of the Invention]

[0011] According to the present invention, automatic negotiation can be performed even when the other party's utility function or utility value, or strategy, is unknown. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of a learning system according to an embodiment. [Figure 2] FIG. 1 is a diagram illustrating an example of a configuration of a learning device according to an embodiment. [Figure 3]FIG. 1 is a diagram illustrating an example of a configuration of a data generating device according to an embodiment. [Figure 4] FIG. 2 is a diagram illustrating an example of input and output of data in the learning system according to the embodiment. [Figure 5] FIG. 10 is a diagram illustrating an example of a self-behavior model when a neural network is used in the embodiment. [Figure 6] FIG. 1 is a diagram illustrating an example of the configuration of a proposal determination device according to an embodiment. [Figure 7] FIG. 2 is a diagram illustrating an example of input and output of data in the proposal decision device according to the embodiment. [Figure 8] FIG. 1 is a diagram illustrating an example of the configuration of a negotiation device that predicts the behavior of a negotiation partner. [Figure 9] FIG. 10 is a diagram illustrating an example of a learning system according to a modified example of an embodiment. [Figure 10] FIG. 10 is a diagram showing the characteristics of domains used in an experiment according to an embodiment. [Figure 11] FIG. 10 is a diagram illustrating an example of a conflict degree. [Figure 12] FIG. 10 is a diagram showing evaluation index values ​​of experimental results according to the embodiment. [Figure 13] FIG. 10 is a diagram illustrating an example of a negotiation situation in the learning system according to the embodiment when the negotiation partner is a time-dependent agent. [Figure 14] FIG. 10 is a diagram illustrating an example of a change in a user's own utility value according to the embodiment when the negotiation partner is a time-dependent agent. [Figure 15] FIG. 10 is a diagram illustrating an example of a change in the counterpart utility value according to the embodiment when the negotiating counterpart is a time-dependent agent. [Figure 16] FIG. 10 is a diagram showing an example of a state of negotiation in the learning system according to the embodiment when the negotiation partner is a behavior-dependent agent. [Figure 17] FIG. 10 is a diagram illustrating an example of a change in a user's own utility value according to the embodiment when the negotiation partner is a time-dependent agent. [Figure 18] FIG. 10 is a diagram illustrating an example of a change in the counterpart utility value according to the embodiment when the negotiating counterpart is a time-dependent agent. [Figure 19] FIG. 1 is a diagram illustrating an example of a configuration of a learning device according to an embodiment. [Figure 20] FIG. 1 is a diagram illustrating an example of the configuration of a learning system according to an embodiment. [Figure 21] FIG. 1 is a diagram illustrating an example of the configuration of a proposal determination device according to an embodiment. [Figure 22] FIG. 10 is a diagram illustrating an example of a processing procedure in a learning method according to the embodiment. [Figure 23] FIG. 1 is a schematic block diagram illustrating an example configuration of a computer according to at least one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0013] The following describes embodiments of the present invention, but the following embodiments do not limit the scope of the invention as claimed. Furthermore, not all of the combinations of features described in the embodiments are necessarily essential to the solution of the invention. 1 is a diagram showing an example of the configuration of a learning system according to an embodiment. In the configuration shown in FIG. 1, learning system 1 includes learning device 100 and data generating device 200.

[0014] The learning system 1 is a system for learning a behavioral decision-making method in automated negotiation. The learning device 100 includes a model for determining an action in automated negotiation. The learning device 100 updates the parameter values ​​of this model to learn a method for determining an action. The learning device 100 is treated as a party in the automated negotiation and is also referred to as "itself" or "itself." The behavior determined by the learning device 100 corresponds to the behavior of the user of the learning device 100. The behavior determined by the learning device 100 is also referred to as the own action. Information indicating the own action is also referred to as the own action information. The model for determining the own action is also referred to as the own action model.

[0015] Learning device 100 is configured using, for example, a computer. Alternatively, learning device 100 may be configured using dedicated hardware for learning device 100, such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0016] The data generating device 200 generates data that the learning device 100 uses to learn the behavior decision-making method. In particular, the data generating device 200 has a model that simulates the behavior of the negotiating partner, determines the behavior of the negotiating partner based on the behavior determined by the learning device 100, and outputs information indicating the determined behavior of the negotiating partner to the learning device 100. The behavior of the negotiating partner is also referred to as opponent's action. Information indicating the opponent's action is also referred to as opponent's action information.

[0017] Furthermore, data generating device 200 calculates a utility value indicating an evaluation of the other party's action when learning device 100 determines an action. The utility value for the other party's action when learning device 100 determines an action corresponds to the utility value for the user of learning device 100. The utility value for the other party's action when learning device 100 determines an action is also referred to as the own utility value. Among the actions in negotiations, the "proposal" described below is the subject of evaluation based on utility value.

[0018] The data generating device 200 is configured using, for example, a computer. Alternatively, the data generating device 200 may be configured using dedicated hardware for the data generating device 200, such as an ASIC or FPGA. The learning device 100 and the data generating device 200 may be configured as a single device, for example, by implementing the learning device 100 and the data generating device 200 on a single computer.

[0019] Fig. 2 is a diagram showing an example of the configuration of learning device 100. In the configuration shown in Fig. 2, learning device 100 includes a first communication unit 110, a first display unit 120, a first operation input unit 130, a first memory unit 180, and a first control unit 190. First control unit 190 includes a history information generation unit 191, a behavior determination unit 192, and a learning control unit 193. The first communication unit 110 communicates with other devices. For example, the first communication unit 110 communicates with the data generating device 200 to transmit its own behavior information to the data generating device 200 and to receive other party behavior information and its own utility value from the data generating device 200.

[0020] First display unit 120 has a display screen such as a liquid crystal panel or an LED (Light Emitting Diode) panel, and displays various images. For example, first display unit 120 may display negotiation progress information, such as one's own behavior information and the other party's behavior information. The first operation input unit 130 includes input devices such as a keyboard and a mouse, and receives user operations. For example, the first operation input unit 130 may receive a user operation to instruct the start of learning a behavior determination method.

[0021] The first storage unit 180 stores various data. For example, the first storage unit 180 stores history information of proposals made in negotiations. This history information is used as input to the own behavior model. Furthermore, if the own behavior model is implemented as software, the first storage unit 180 stores software representing the own behavior model. The first storage unit 180 is configured using a storage device included in the learning device 100.

[0022] The first control unit 190 executes various processes by controlling each unit of the learning device 100. The functions of the first control unit 190 are executed, for example, by a CPU (Central Processing Unit) included in the learning device 100 reading and executing a program from the first storage unit 180.

[0023] The history information generation unit 191 generates history information of proposals made in negotiations. (The history information will be described later with reference to formula (1).) Here, time is expressed in time steps, such as time 0, time 1, ..., where one step is the time when learning device 100 determines its own action and data generating device 200 determines the other party's action in response to its own action. Here, time 0 is the time when learning system 1 starts operating.

[0024] Furthermore, actions in negotiations are considered to be either "Offer," "Accept," or "Reject." In an "Offer," specific details of the transaction are presented (proposed). Making one's own proposal after receiving a proposal from the other party is also called a "Counter Offer." In an "Accept" situation, the other party's proposal is accepted and the negotiation ends. In a "Reject" situation, the other party's proposal is rejected and the negotiation is terminated. In cases where negotiations must continue until the time limit, for example, it may be possible to decide on only one of the actions in negotiations: offer or accept. Information indicating the specific details of the transaction presented in the proposal is also referred to as proposal information, or simply as a proposal.

[0025] A proposal in which the own action is a proposal is also called an Own Offer. A proposal in which the other party's action is a proposal is also called an Opponent's Offer. Let ω be the self-proposal at time i. i and the opponent's proposal at time i is denoted as ω' i When the negotiation is ongoing at time t, all history information D from time 0 to time t is expressed as in equation (1).

[0026]

number

[0027] To facilitate the use of history information as input to the self-behavior model, the history information generation unit 191 may perform one-hot encoding of proposals for each proposal item. An item represents a type of proposal content. For example, in travel negotiations, the items may represent the destination and the travel cost. For example, destinations in travel negotiations may be represented by dummy variables such as Montreal as "100," Yokohama as "010," and Macau as "001." Furthermore, travel costs may be represented by dummy variables such as 100 dollars as "100," 500 dollars as "010," and 1,000 dollars as "001." In this case, the history information generation unit 191 may represent a proposal for a destination of Macau with a travel cost of 500 dollars as a vector of dummy variables for each item, such as [001, 010]. However, the conversion performed by the history information generation unit 191 is not limited to a specific one, and various conversions can be performed to convert the history information into data in a format that is easy to handle in the own behavior model. The data format into which the history information generation unit 191 converts the history information is not limited to the vector of dummy variables as described above, and various data formats expressed as vectors can be used. For example, the elements of the vector may be in one-to-one correspondence with the input nodes of the own behavior model configured using a neural network.

[0028] The behavior decision unit 192 inputs history information into the own behavior model and calculates own behavior information. The learning control unit 193 controls the learning of the own behavior model based on the own utility value obtained from the data generating device 200. Specifically, the learning control unit 193 updates the parameter values ​​of the own behavior model so that the evaluation indicated by the own utility value becomes as high as possible.

[0029] Fig. 3 is a diagram showing an example of the configuration of the data generating device 200. In the configuration shown in Fig. 3, the data generating device 200 includes a second communication unit 210, a second display unit 220, a second operation input unit 230, a second storage unit 280, and a second control unit 290. The second control unit 290 includes an opponent model executing unit 291 and an evaluation value calculating unit 292.

[0030] The second communication unit 210 communicates with other devices. For example, the second communication unit 210 communicates with the learning device 100 to receive its own behavior information from the learning device 100, and transmits the other party's behavior information and its own utility value for the other party's behavior information to the learning device 100. The second display unit 220 has a display screen such as a liquid crystal panel or an LED panel, and displays various images. For example, the second display unit 220 may display information indicating various settings related to the negotiation partner, such as the strategy settings of the negotiation partner. A strategy in this context may be a criterion, rule, or algorithm for determining an action or part of an action, such as a method for determining an offer based on a utility function, or a method for determining whether to accept a proposal from a negotiating partner. The second operation input unit 230 includes input devices such as a keyboard and a mouse, and receives user operations. For example, the second operation input unit 230 may receive a user operation for making settings related to the negotiation partner.

[0031] The second storage unit 280 stores various data. For example, the second storage unit 280 stores an opponent's utility function for calculating an opponent's utility value. The opponent's utility value is a value indicating an evaluation of the negotiating partner simulated by the data generating device 200 on the negotiating partner's own behavior. Among actions in negotiations, proposals that continue negotiations require evaluation by utility values. Therefore, when an action is a proposal, the utility function receives the proposal as input and outputs a utility value. The counterparty utility function receives the proposal as input and outputs a counterparty utility value that indicates the negotiating counterpart's evaluation of the proposal.

[0032] The second storage unit 280 also stores a self utility function for calculating a self utility value. The self utility function receives an input of a proposal and outputs a self utility value. Furthermore, when the model that simulates the negotiating partner is implemented in software, the second storage unit 280 stores the software that represents the model that simulates the negotiating partner.

[0033] The second control unit 290 executes various processes by controlling each unit of the data generating device 200. The functions of the second control unit 290 are executed, for example, by a CPU included in the data generating device 200 reading out a program from the second storage unit 280 and executing it.

[0034] The other party model execution unit 291 determines the other party's behavior using a model that simulates the negotiating party. The evaluation value calculation unit 292 inputs the other party's behavior information into the own utility function to calculate the own utility value. As described above with respect to the utility function, only when the other party's behavior is a proposal, the evaluation value calculation unit 292 may input the proposal into the own utility function to calculate the own utility value.

[0035] FIG. 4 is a diagram showing an example of data input and output in the learning system 1. In the configuration shown in FIG. 4, the behavior determining unit 192 inputs the history information input from the history information generating unit 191 into the own behavior model to calculate own behavior information. The other party model execution unit 291 receives input of the own behavior information from the behavior determination unit 192 and determines the behavior of the negotiation partner. Specifically, the other party model execution unit 291 inputs the own behavior information into the other party utility function and calculates the utility value of the own behavior for the negotiation partner. As described above, it is necessary to calculate the utility value when the behavior is a proposal, and when the own behavior is a proposal, the other party model execution unit 291 inputs the proposal into the other party utility function and calculates the utility value. The other party model execution unit 291 determines the other party's behavior using the calculated utility value and the other party's strategic model and outputs the other party's behavior information. On the other hand, if the own action is either acceptance or refusal, the negotiation ends. In this case, the other party model execution unit 291 does not determine a new other party action.

[0036] The opponent strategy model is a model that determines the behavior of the negotiating opponent based on input of one's own behavior information or negotiation history information. The opponent utility function and opponent strategy model are not limited to specific ones. For example, the opponent's strategy model may include a predetermined formula for setting a threshold value for the utility value, such as formula (6) or formula (7) described below, a rule for determining whether to accept the opponent's own proposal based on the threshold value, and a rule for selecting one of the proposal options based on the threshold value when making a proposal.The opponent's model execution unit 291 may then set the threshold value based on the predetermined formula, and determine to accept the opponent's own proposal when the opponent's utility value for the opponent's own proposal is equal to or greater than the threshold value.Also, when the opponent's utility value for the opponent's own proposal is less than the threshold value, the opponent's model execution unit 291 may select a proposal from the proposal options whose opponent's utility value is equal to or greater than the threshold value, and make the opponent's proposal as the opponent's behavior. However, the opponent strategy model is not limited to this, and may be configured using a neural network, for example, as in the case of the self-behavior model.

[0037] It is also possible to prepare multiple opponent strategy models and opponent utility functions, and to switch between the opponent strategy models and opponent utility functions to learn the user's own behavior model for various strategies and utilities (values). This is expected to enable negotiations to be conducted in a beneficial manner for the user with negotiation partners of various strategies and utilities during operation.

[0038] The evaluation value calculation unit 292 inputs the other party's action information calculated by the other party model execution unit 291 into the own utility function to calculate the own utility value. As described above, it is only necessary to calculate the utility value when the action is a proposal. Therefore, the evaluation value calculation unit 292 may be configured to calculate the own utility value only when the other party's action is a proposal. The own utility value calculated by the evaluation value calculation unit 292 is input to the learning control unit 193 of the learning device 100. The learning control unit 193 controls the learning of the own behavior model so that the evaluation indicated by the own utility value becomes as high as possible. As a learning method used by the learning control unit 193, a known method such as backpropagation can be used.

[0039] Furthermore, the other party's behavior information is input to the history information generation unit 191, which generates negotiation history information. The history information is used as input to the own behavior model. As described above, the history information generation unit 191 may perform one-hot encoding on the history information.

[0040] In the configuration shown in Fig. 4, the learning system 1 uses reinforcement learning to learn its own behavior model. Reinforcement learning here is machine learning that learns a policy, which is an agent's behavioral rule in a certain environment, based on the actions taken by the agent, the observed state of the environment or the agent, and a reward that represents an evaluation of the state or the agent's actions.

[0041] In learning the own behavior model, the behavior decision unit 192, which decides the own behavior using the own behavior model, is treated as an agent, and the own behavior is treated as an action in reinforcement learning. The own behavior model is treated as a policy. Furthermore, the history information output by the history information generation unit 191 to the behavior decision unit 192 is treated as information representing a state, and the own utility value calculated by the evaluation value calculation unit 292 or a value based on the own utility value is treated as a reward value.

[0042] The learning control unit 193 updates the parameter values ​​of the own behavior model so that the evaluation indicated by the own utility value becomes as high as possible. Here, the own behavior model corresponds to a model in machine learning, and updating the parameter values ​​of the own behavior model corresponds to learning the model. Here, it is expressed that the behavior decision unit 192 or the own behavior model learns the own behavior model, and the learning control unit 193 controls the learning of the own behavior model by the behavior decision unit 192. Alternatively, it may be expressed that the learning control unit 193 learns the own behavior model.

[0043] The reinforcement learning used in learning the self behavior model is not limited to a specific type, and various types of reinforcement learning can be used. Furthermore, the configuration of the learning system 1 may be modified depending on the type of reinforcement learning applied to learning the self behavior model. For example, when performing reinforcement learning that explicitly handles the reward function, the evaluation value calculation unit 292 and the learning control unit 193 may be integrally configured and provided in the learning device 100, so that the learning control unit 193 knows the self utility function, which is an example of a reward function.

[0044] The history information generating unit 191 may use a part of all the history information D from time 0 shown in formula (1) as the history information to be output to the behavior determining unit 192. For example, the history information generating unit 191 may generate and output, as the history information at time t, history information from L steps before in the time step as shown in formula (2).

[0045]

number

[0046] L is a positive integer constant. s t indicates historical information at time t. This historical information is also called state information (in reinforcement learning). T represents a predetermined end time of the negotiation. If the negotiation is not concluded by the end time T, the negotiation is deemed to have failed and is terminated.

[0047] In this way, the history information generating unit 191 uses the most recent partner proposal ω' t In addition, the past proposals of the other party and the most recent and past proposals of the user are stored as history information s t This allows the learning control unit 193 to control the learning of the own behavior model so that the behavior decision unit 192 decides on the own behavior based on the negotiation history.

[0048] As a reward function in reinforcement learning, a reward function that outputs a reward at the end of negotiation may be used. For example, if one accepts, the final proposal ω' is the adopted proposal. t The utility value U(ω') for t ) may be used as a reward. In this case, the reward is expressed as in equation (3).

[0049]

number

[0050] The function r represents the reward function. t} represents one episode in reinforcement learning. t+1 " represents acceptance as one's final action. If the other party accepts, the final self-proposal ω is the adopted proposal. t The utility value U(ω t ) may be used as a reward. In this case, the reward is expressed as in equation (4).

[0051]

number

[0052] As in equation (3), the function r represents the reward function. t+1" indicates the other party's acceptance. "{...,ω' t ,η' t+1}" represents one episode in reinforcement learning. t " represents a proposal as one's final action. If the negotiation ends without reaching an agreement, a penalty K may be given as a reward. K may be a predetermined constant. However, the reward function used in learning the self-behavior model is not limited to a specific one, and various reward functions can be used depending on the reinforcement learning employed. Depending on the type of reinforcement learning, the reward function may be unknown.

[0053] FIG. 5 is a diagram showing an example of a self-behavior model when a neural network is used. 5 includes an input layer, multiple hidden layers, and an output layer. History information output by history information generation unit 191 is input to the nodes of the input layer. As described above, history information that has been one-hot encoded may be input to the nodes of the input layer.

[0054] The output layer has one node indicating whether or not to accept, and multiple nodes indicating the content of the proposal when a proposal is made. If rejection can be selected as an action in the negotiation, a node indicating whether or not to reject may be provided. Alternatively, if the party itself does not accept, a node indicating whether or not to accept may not be provided. The method of expressing the content of the proposal is not limited to a specific method. For example, the neural network 300 may output data in the same one-hot encoding format as the input data, that is, the history information.

[0055] However, the configuration of the own behavior model is not limited to a specific one. The own behavior model may be configured using a machine learning model other than a neural network. For example, the own behavior model may be configured as a policy function-based model. Then, the behavior decision unit 192 may input history information into the policy function to calculate the own behavior information. Alternatively, the self behavior model may be configured to include both a value function such as Actor-Critic and a policy function, and the behavior decision unit 192 may update the policy function based on an evaluation of the policy function by the evaluation function.

[0056] In addition, when the self-behavior model is configured using a neural network, the configuration of the neural network is not limited to a specific one. For example, the self-behavior model may be configured using a convolutional neural network (CNN) or a recurrent neural network (RNN), but is not limited thereto.

[0057] Fig. 6 is a diagram showing an example of the configuration of a proposal determination device according to an embodiment. In the configuration shown in Fig. 6, the proposal determination device 400 includes a first communication unit 110, a first display unit 120, a first operation input unit 130, a first storage unit 180, and a first control unit 190. The first control unit 190 includes a history information generation unit 191 and an action determination unit 192. 6, parts having the same functions as those in FIG. 2 are given the same reference numerals (110, 120, 130, 180, 190, 191, 192), and detailed description thereof will be omitted here.

[0058] The proposal determination device 400 is a device that performs automatic negotiation using a self-behavior model that has been learned by the learning system 1. The proposal determination device 400 in Fig. 6 has a configuration in which the learning control unit 193 is removed from the learning device 100 in Fig. 2. In other respects, the proposal determination device 400 is similar to the learning device 100. The data format of the input / output data of the own behavior model in the proposal determination device 400 can be the same as that of the own behavior model in the learning control unit 193. In addition, since the own behavior model in the proposal determination device 400 has already been learned, the learning control unit 193 is not necessary.

[0059] The learning device 100 may be used as the proposal determination device 400 without using the learning control unit 193. Alternatively, when the learning device 100 is used as the proposal determination device 400 without using the learning control unit 193, the learning device 100 may be configured to periodically learn the own behavior model and update the own behavior model.

[0060] FIG. 7 is a diagram showing an example of input and output of data in the proposal determination device 400. As shown in FIG. Comparing the proposal determination device 400 shown in Fig. 7 with the learning device 100 shown in Fig. 4, the proposal determination device 400 does not include a learning control unit 193 and does not need to acquire its own utility value. In other respects, the proposal determination device 400 shown in Fig. 7 is similar to the learning device 100 shown in Fig. 4.

[0061] In learning of the own behavior model by the learning device 100, the learning control unit 193 controls the learning based on the own utility value so that the own behavior model outputs own behavior information that increases the evaluation indicated by the own utility value. In this way, learning is performed so that the own behavior model outputs own behavior information that takes into account the evaluation based on the own utility value. The proposal determination device 400 using the trained own behavior model can calculate and output own behavior information that reflects the evaluation indicated by the own utility value, without the need to explicitly state the own utility function or the own utility value.

[0062] One possible method for a negotiation device (proposal determination device) that performs automated negotiations to determine its actions in negotiations is to predict the evaluation or behavior of the negotiating partner regarding its own actions, and determine its own actions by taking into account the evaluation or behavior of the negotiating partner. The proposal determination device 400 will be further explained in comparison with this type of negotiation device.

[0063] FIG. 8 is a diagram showing an example of the configuration of a negotiation device that predicts the actions of a negotiation partner. The negotiation device 900 shown in FIG. 8 includes a partner model learning unit 901, an evaluation value calculation unit 902, a self utility value calculation unit 903, a proposal determination unit 904, and an acceptance determination unit 905. The other party model learning unit 901 uses other party behavior information or negotiation history information to learn an other party model that indicates the other party utility function. The other party model learning unit 901 outputs the utility function obtained by learning or a utility value calculated using the utility function.

[0064] The evaluation value calculation unit 902 outputs a self utility value that reflects the behavior history. More specifically, the evaluation value calculation unit 902 inputs the other party's behavior information into a self utility function to calculate a self utility value for the other party's behavior. As described above, it is only necessary to calculate a utility value when the behavior is a suggestion. Therefore, the evaluation value calculation unit 902 may be configured to calculate a self utility value only when the other party's behavior is a suggestion. The own utility calculation unit 903 inputs the own behavior information that the negotiation device 900 plans to propose to the negotiating partner into the own utility function to calculate the own utility.

[0065] The proposal determination unit 904 formulates a self proposal so that both the other party's utility value and the own utility value satisfy predetermined conditions, based on the self utility value from the own utility value calculation unit 903 and the other party's utility function or the other party's utility value from the other party model learning unit 901. For example, the proposal determination unit 904 solves an optimization problem to maximize the own utility function using the other party's utility function as one of the constraints, and calculates a self utility value such that the other party's utility value is large and the own utility value is also large. The proposal determination unit 904 also learns the strategy of the negotiating partner, and calculates a partner behavior prediction value that indicates the partner's predicted behavior.

[0066] The acceptance determination unit 905 determines its own behavior by inputting the own proposal determined by the own utility calculation unit 903 and the negotiation partner's behavior predicted by the proposal determination unit 904 into the acceptance strategy model. For example, the acceptance determination unit 905 determines its own behavior as either adopting the own proposal determined by the own utility calculation unit 903 or accepting the most recent proposal of the other party. The acceptance strategy model is a model for determining whether to accept a proposal from a partner. For example, similar to the above-described counterpart strategy model, the acceptance strategy model may include a predetermined formula for setting a threshold value for the utility value, rules for determining whether to accept the counterpart proposal based on the threshold value, and rules for selecting one of the proposal options based on the threshold value when making a proposal. The acceptance determination unit 905 may then set a threshold value based on the predetermined formula, and determine to accept the counterpart proposal when its own utility value for the counterpart proposal is equal to or greater than the threshold value. However, the acceptance strategy model is not limited to this, and may be configured using a neural network, for example, as in the case of the self-behavior model.

[0067] In operation, the negotiation device 900 of Figure 8 determines actions while evaluating the content of negotiations using the own utility function and the other party's utility function. However, utility is not always available in the form of a function. In the negotiation device 900 of Figure 8, if utility cannot be obtained in the form of a function, it may not be possible to evaluate actions. In this regard, the negotiation device 900 cannot be applied if utility cannot be obtained in the form of a function.

[0068] Furthermore, since the counterparty utility function varies depending on the negotiating partner, the negotiation device 900 needs to learn the utility function of the actual negotiating partner during operation. Until the learning of the counterparty utility function progresses, the negotiation device 900 may not be able to select an appropriate action.

[0069] In contrast to this, neither the own utility function nor the other party utility function is required during operation in the proposal determination device 400. Therefore, even when utility cannot be obtained in the form of a function, the proposal determination device 400 can be applied. Furthermore, there is no need to learn the utility function of the negotiating partner during operation in the proposal determination device 400. Therefore, the proposal determination device 400 does not encounter the problem of being unable to select an appropriate action until learning of the partner's utility function has progressed.

[0070] As described above, the history information generating unit 191 generates history information of proposals made in negotiations. History information generating unit 191 is an example of a history information acquiring means. Alternatively, when learning the own behavior model by reinforcement learning, the history information corresponds to information representing a state. For this reason, the data generating device 200 may be positioned as a device that simulates the environment in reinforcement learning, and the data generating device 200 may be equipped with a history information generating unit 191 to generate history information. The data generating device 200 may then transmit the history information to the learning device 100. In this case, the behavior determining unit 192, which acquires the history information in the learning device 100 and inputs it into the own behavior model, corresponds to an example of history information acquiring means.

[0071] Furthermore, the evaluation value calculation unit 292 calculates a self utility value corresponding to an evaluation value for a proposal from the negotiation partner. Then, the learning control unit 193 acquires the self utility value from the evaluation value calculation unit 292 and controls learning of the self behavior model. The learning control unit 193 corresponds to an example of an evaluation value acquisition means.

[0072] Furthermore, the behavior decision unit 192 learns the own behavior model based on the history information and the own utility value under the control of the learning control unit 193. The behavior decision unit 192 outputs own behavior information indicating the own proposal to the negotiation partner based on the proposal determination method using the own behavior model. The behavior decision unit 192 is an example of a learning means. The learning of the own behavior model by the behavior decision unit 192 is an example of learning a method for determining one's own proposal to a negotiating partner.

[0073] Furthermore, the counterpart model execution unit 291 receives an input of one's own proposal to the negotiation counterpart, and outputs counterpart behavior information indicating the proposal from the negotiation counterpart. The counterpart model execution unit 291 corresponds to an example of a counterpart model execution means. The evaluation value calculation unit 292 receives input of counterparty behavior information indicating a proposal from the negotiation counterpart, applies the counterparty behavior information to its own utility function, and outputs its own utility value corresponding to its own evaluation value for the proposal from the negotiation counterpart. The evaluation value calculation unit 292 corresponds to an example of evaluation means.

[0074] In the learning system 1, learning is performed using proposal history information, so that the reaction of the other party to one's own actions is reflected in the proposal determination method. Since the reaction of the other party to one's own actions is reflected in the proposal determination method, the proposal determination device 400 during operation does not need to set or learn the utility and strategy of the negotiation partner. In this way, with the learning system 1 during learning and the proposal determination device 400 during operation, there is no need to explicitly predict the behavior of the negotiating partner. Therefore, with the learning system 1 during learning and the proposal determination device 400 during operation, automatic negotiation can be performed even if the utility function or utility value, or strategy of the negotiating partner is unknown.

[0075] Furthermore, the history information generation unit 191 outputs the history information as vector-format data. The behavior decision unit 192 includes a neural network as its own behavior model, which receives input of history information in the form of vector-format data and outputs its own proposal to the negotiation partner. The learning device 100 can input history information to the neural network in the form of vector data suitable for input to the neural network. In this respect, the learning device 100 is expected to enable highly accurate learning of a self-behavior model constructed using a neural network.

[0076] Furthermore, in proposal determination device 400, history information generation unit 191 acquires history information of proposals made in negotiations. As described above, history information generation unit 191 corresponds to an example of history information acquisition means. The behavior decision unit 192 inputs the history information acquired by the history information generation unit 191 into a learned own behavior model based on the history information of proposals made in negotiations and the evaluation value of the proposal from the negotiation partner, and acquires and outputs the own proposal to the negotiation partner. The behavior decision unit 192 corresponds to an example of a proposal output means.

[0077] As described above, the self-behavior model can be learned by using a combination of the other party's utility function and the other party's strategy model during learning. As a result, the behavior decision method according to the utility and strategy of the negotiation partner is reflected in the self-behavior model, and the proposal determination device 400 does not need to set or learn the utility and strategy of the negotiation partner during operation. In this way, the proposal determination device 400 can perform automatic negotiations even when the utility function or utility value of the other party during negotiations, or the strategy, is unknown.

[0078] The learning system 1 and the proposal determination device 400 can be applied to various negotiations. The learning system 1 and the proposal determination device 400 can be used in various negotiations including, but not limited to, price negotiations. For example, the lending destination and lending period of the equipment may be adjusted using the learning system 1 or the proposal determination device 400. Furthermore, the learning system 1 or the proposal determination device 400 may carry out the procedure for lending the determined equipment. Furthermore, the learning system 1 or the proposal determination device 400 may be used for various adjustments related to time or place, or both, such as adjustment of delivery dates for manufactured parts, adjustment of logistics schedules, or adjustment of logistics routes.

[0079] Furthermore, the learning system 1 and the proposal determination device 400 can be applied to various decision-making processes that involve a combination of multiple strategies, not limited to negotiations. For example, the learning system 1 or the proposal determination device 400 may formulate a plant operation plan based on multiple evaluation criteria. Furthermore, the learning system 1 or the proposal determination device 400 may control the plant based on the formulated operation plan.

[0080] FIG. 9 is a diagram showing an example of learning system 1b according to a modified example of learning system 1. In learning system 1b shown in FIG. 9, opponent model execution unit 291b of data generation device 200b includes proposal determination unit 291c and noise addition unit 291d. In other respects, learning system 1b is similar to learning system 1. Learning system 1b is an example of learning system 1. Data generation device 200b is an example of data generation device 200.

[0081] The proposal determination unit 291c is similar to the other party model execution unit 291 of the learning system 1. Specifically, the proposal determination unit 291c inputs the own action information obtained from the learning device 100 into the other party utility function to calculate the utility value of the own action for the negotiation partner. As described above, it is necessary to calculate the utility value only when the action is a proposal. Therefore, the proposal determination unit 291c may calculate the utility value only when the own action is a proposal. The proposal determination unit 291c determines the other party's action using the calculated utility value and the other party's strategic model, and outputs the other party's action information. The proposal determination unit 291c corresponds to an example of a proposal determination means.

[0082] The noise adding unit 291d adds noise to the proposal from the negotiation partner determined by the proposal determining unit 291c. For example, the noise adding unit 291d may add noise to the other party utility value, which is the utility value used by the proposal determining unit 291c to determine the other party's action. In this case, the other party utility value after adding noise is expressed, for example, as in Equation (5).

[0083]

number

[0084] ω denotes the self-proposal obtained from the learning device 100. The function U opp denotes the opponent's utility function, and U opp (ω) indicates the opponent utility value before noise is added. ε indicates Gaussian noise. N(μ,σ 2) denotes a Gaussian distribution, and ε~N(μ,σ 2 ) where ε has mean μ and variance σ 2 It is shown that the distribution follows a Gaussian distribution. By adding noise to the opponent's utility value by the noise adding unit 291d, the proposal determining unit 291c calculates different opponent's behavior information for the same own utility value ω, and the variation of the learning data of the own behavior model increases. In particular, even if the opponent's strategy model is a deterministic model, by adding noise to the opponent's utility value by the noise adding unit 291d, a learning data set including various patterns of total history information D (see formula (1)) can be obtained. However, the method by which the noise adding unit 291d adds noise to the proposal is not limited to the method of adding noise to the other party utility value. For example, if quantitative information such as the number of items to be handed over is displayed in an editable manner in the other party proposal output by the proposal determining unit 291c as the other party behavior information, the noise adding unit 291d may add noise by rewriting this quantitative information.

[0085] As described above, the proposal determination unit 291c determines the other party's proposal in response to the negotiation partner's own proposal to the negotiation partner. The noise addition unit 291d adds noise to the other party's proposal determined by the proposal determination unit 291c. According to the learning system 1b, it is expected that the variation of the learning data of the own behavior model will increase, the learning of the own behavior model will progress, and overlearning in the learning of the own behavior model will be avoided.

[0086] An experiment using the configuration of learning system 1b will be described. In the following, the learning system will be referred to as learning system 1b both during learning and testing (operation). In the explanation of the experimental results, the notation learning system 1b and the notation proposal determination device 400 will not be used differently. Figure 10 shows the characteristics of the domains used in the experiment. The domains here refer to the negotiation settings. Specifically, the selectable proposals differ for each domain. In the experiment, learning and testing were carried out using five domains whose domain names are shown in Figure 10. Figure 10 shows the domain size and degree of conflict for each domain. Domain size is a type of proposal that can be selected. The degree of conflict is the Euclidean distance between the point (1,1) where both the own utility value and the opponent utility value are maximum and the KS solution (Kalai-Smorodinsky Solution).

[0087] FIG. 11 is a diagram showing an example of the degree of conflict. The horizontal axis of the graph in FIG. 11 indicates the self utility value, and the vertical axis indicates the other party utility value. Both the self utility value and the other party utility value take real values ​​in the range of [0, 1]. In the graph in FIG. 11, selectable proposals are plotted on the coordinates of the self utility value and the other party utility value of the proposal.

[0088] Point P11 indicates the KS solution. The KS solution indicates the intersection of the diagonal line connecting points (0,0) and (1,1) with the Pareto Frontier. In the example of FIG. 11, the proposal with the highest self-utility value among the proposals for each partner utility value corresponds to the Pareto Optimum Solution, and the boundary line obtained by connecting the Pareto solutions corresponds to the Pareto Frontier. Also, point (1,1) is represented as point P12.

[0089] The Euclidean distance D11 between the point P11 and the point P12 corresponds to the conflict degree. The conflict degree can be used as an index value showing the difficulty of negotiation. The greater the conflict degree, the more difficult the negotiation is considered to be.

[0090] In the experiment, three time-dependent agents and two action-dependent agents were used as negotiation partners. In explaining the experimental results, the term "agents" refers to parties in automated negotiations. The time-dependent agent changes the utility value setting according to the elapsed time of the negotiation. In the experiment, the utility value set by the time-dependent agent was determined as shown in Equation (6).

[0091]

number

[0092] t indicates time in time steps. Time T is the negotiation end time, and t takes an integer value of 0≦t≦T. U(ω t+1 ) indicates the utility value set at time t+1. An agent as a negotiating partner accepts a proposal that will make the other party's utility value equal to or greater than the set utility value. When making a proposal, the agent will make a proposal that will make the other party's utility value equal to or greater than the set utility value. U max is a constant that is determined as the maximum value of the utility value. min is a constant defined as the minimum value of the utility value.

[0093] e is a hyperparameter, and the larger the value of e, the smaller the utility value that is set at an earlier time. In the experiment, three different values ​​of e were set for the agents: "0.1," "1.0," and "5.0." According to equation (6), at time 0, the utility value is the maximum value U max As time passes, the utility value is set lower, and at time T, the utility value is the minimum value U min is.

[0094] The behavior-dependent agent changes the utility value setting depending on the other party's behavior. In the experiment, the utility value set by the behavior-dependent agent was determined as shown in Equation (7).

[0095]

number

[0096] δ is a hyperparameter, and the agent obtains the utility value U(ω' t ) and the utility value U(ω') at time t-δ t-δ In the experiment, two different agents were set with δ values ​​of "1" and "2". Here, the utility value U(ω' t ) and U(ω' t-δ ) is the agent's own utility value in relation to the proposal of the negotiating partner.

[0097] In the case of an agent as a negotiating partner, U min Above and U max The utility value setting is changed within the following range. The agent as the negotiating partner changes the other party's utility value U(ω' t-δ ), the other party's utility value U(ω' t ) is larger, the utility value set value U(ω t+1 ) is the utility value set at time t, U(ω t ) significantly reduced. The partner's utility value U(ω') at time t t ) is the partner utility value U(ω') at time t-δ t-δ ), the agent in the negotiation will set the utility value U(ω t+1 ) is the utility value set at time t, U(ω t ) larger than

[0098] In the experiment, the negotiation time limit was set to 40 rounds (40 time steps). Learning was performed using training data from 2000 episodes. The ε-greedy method was used as the strategy, with ε = 0.01. The penalty for unsuccessful negotiation was set to K = -1. The Gaussian noise during learning was (μ, σ) = (0.5 × 10 -3 ) was assumed to be noise that follows a Gaussian distribution. DDQN (Double Deep Q Network) was used as the deep reinforcement learning method.

[0099] The evaluation index used was the utility obtained when acting with a greedy policy after learning. The utility obtained here is the own utility value at the time of negotiation success. The greedy policy here corresponds to the case where ε = 0 in the ε-greedy method. Because the initial value has a large effect on performance, learning was performed using 100 different initial values, and the one with the best performance after learning was used as the evaluation target.

[0100] The evaluation index values ​​were compared with those obtained by two negotiation methods, Random Agent and RLBOA-agent. The Random Agent randomly selects a proposal from all selectable proposals. This proposal selection method corresponds to the proposal selection method used by the learning device 100 before learning. Since no learning is performed on the Random Agent, the average evaluation index value obtained from 100 negotiations was used. RLBOA-agent uses states and actions based on utility values, as proposed in previous research. DDQN was used as the learning method. As with the evaluation method for learning system 1b, learning was performed using 100 different initial values, and the system with the best performance after learning was used as the evaluation target.

[0101] Fig. 12 shows the evaluation index values ​​of the experimental results, which are compared with the Random Agent and the RLBOA-agent for each of five types of domains and five types of negotiation partners. Negotiating partner Boulware is a time-dependent agent with e=0.1. Linear is a time-dependent agent with e=1.0. Conceder is a time-dependent agent with e=5.0. TitForTat1 is an action-dependent agent with δ=1. TitForTa21 is an action-dependent agent with δ=2. Regarding the negotiation method, "Baseline" indicates the method using Random Agent. "DRBOA" indicates the case where RLBOA-agent and DDQN are used. "Ours" indicates the method using learning system 1b.

[0102] In the experimental results shown in Figure 12, the evaluation index value of learning system 1b was greater than that of Random Agent in all 25 combinations of domain and negotiation partner. Furthermore, in 23 of the 25 combinations, the evaluation index value of learning system 1b was greater than that of RLBOA-agent. Furthermore, in 15 of the 25 settings, the evaluation index value of the learning system 1b is the maximum value of "1." Thus, good experimental results were obtained.

[0103] Fig. 13 shows an example of the state of negotiation in learning system 1b when the negotiation partner is a time-dependent agent. Fig. 13 shows the state of negotiation when the negotiation partner is a time-dependent agent with e=1.0 and the domain is "thompson". The horizontal axis of the graph in Figure 13 represents the self-utility value. The vertical axis represents the other party's utility value. Triangles ('▲' or '▼') represent points where the utility values ​​for the proposals are plotted. Squares ('■') represent proposals that obtained the acquired utility that was adopted for evaluation. Circles ('●') represent Pareto solutions. Diamonds ('◆') represent KS solutions.

[0104] Fig. 14 shows an example of changes in the agent's own utility value when the negotiating partner is a time-dependent agent. The horizontal axis of the graph in Fig. 14 represents the elapsed time of the negotiation, normalized to a value of 1 when the negotiation ends. The vertical axis represents the agent's own utility value. Fig. 14 shows changes in the agent's own utility value during the negotiation process shown in Fig. 13.

[0105] Fig. 15 is a diagram showing an example of changes in the counterparty utility value when the negotiating counterparty is a time-dependent agent. The horizontal axis of the graph in Fig. 15 represents the elapsed time of the negotiation, normalized with the negotiation end time set to 1. The vertical axis represents the counterparty utility value. Fig. 15 shows changes in the counterparty utility value during the negotiation process shown in Fig. 13.

[0106] When the negotiating partner is a time-dependent agent, learning system 1b appears to explore the negotiating partner's utility in the early stages of the negotiation and take actions to obtain high utility in the late stages of the negotiation. In the examples of Figures 13 to 15, in the early stages of negotiation, the other party's utility value shown in Figure 15 is large, while the own utility value shown in Figure 14 is small. This can be interpreted as meaning that in the early stages of negotiation, learning system 1b prioritizes investigating the utility of the negotiating partner over increasing its own utility value.

[0107] On the other hand, in the final stages of the negotiation, the self-utility value shown in Figure 14 becomes large, while the other party's utility value shown in Figure 15 becomes small. This can be interpreted as learning system 1b deciding its actions by prioritizing the increase of its own utility value in the final stages of the negotiation.

[0108] Fig. 16 shows an example of the state of negotiation in learning system 1b when the negotiation partner is a behavior-dependent agent. Fig. 16 shows the state of negotiation when the negotiation partner is a behavior-dependent agent with δ=2 and the domain is "Grocery." The horizontal axis of the graph in Figure 16 represents the self-utility value. The vertical axis represents the other party's utility value. Triangles ('▲' or '▼') represent points where the utility values ​​for the proposals are plotted. Squares ('■') represent proposals that obtained the acquired utility that was adopted for evaluation. Circles ('●') represent Pareto solutions. Diamonds ('◆') represent KS solutions.

[0109] Fig. 17 shows an example of changes in the agent's own utility value when the negotiating partner is a time-dependent agent. The horizontal axis of the graph in Fig. 17 represents the elapsed time of the negotiation, normalized to a value of 1 when the negotiation ends. The vertical axis represents the agent's own utility value. Fig. 17 shows changes in the agent's own utility value during the negotiation process shown in Fig. 16.

[0110] Figure 18 is a diagram showing an example of changes in the counterparty utility value when the negotiating counterparty is a time-dependent agent. The horizontal axis of the graph in Figure 18 represents the elapsed time of the negotiation, normalized to a value of 1 when the negotiation ends. The vertical axis represents the counterparty utility value. Figure 18 shows changes in the counterparty utility value during the negotiation process shown in Figure 16.

[0111] When the negotiating partner is a behavior-dependent agent, learning system 1b appears to explore the negotiating partner's utility in the early stages of the negotiation and take actions to obtain high utility in the late stages of the negotiation. In the examples of Figures 16 to 18, learning system 1b often makes proposals that result in a low counterparty utility value. The counterparty utility value shown in Figure 18 repeatedly increases and decreases. When the counterparty utility value changes from a small value to a large value, the counterparty utility value is set to a relatively small value, making it easier for learning system 1b to conclude a negotiation. A negotiation is concluded when the counterparty utility value is larger than the value at the previous time step.

[0112] 19 is a diagram illustrating an example of the configuration of a learning device according to an embodiment. In the configuration illustrated in FIG. 19, a learning device 610 includes a history information acquisition unit 611, an evaluation value acquisition unit 612, and a learning unit 613. With this configuration, the history information acquisition unit 611 acquires history information of proposals made in negotiations. The evaluation value acquisition unit 612 acquires evaluation values ​​for proposals from the negotiation partner. The learning unit 613 learns a method for determining its own proposal to the negotiation partner based on the history information and the evaluation values. The history information acquisition unit 611 is an example of a history information acquisition means, the evaluation value acquisition unit 612 is an example of an evaluation value acquisition means, and the learning unit 613 is an example of a learning means.

[0113] The learning device 610 can learn using the proposal history information so that the other party's reaction to its own actions is reflected in the proposal decision-making method. Since the other party's reaction to its own actions is reflected in the proposal decision-making method, there is no need to set or learn the utility and strategy of the negotiating partner during operation. In this way, learning device 610 can perform automatic negotiations even when the utility function or utility value of the other party in the negotiation, or the strategy, is unknown.

[0114] The history information acquisition unit 611 can be realized, for example, by using the functions of the history information generation unit 191 shown in Fig. 2. The evaluation value acquisition unit 612 and the learning unit 613 can be realized, for example, by using the functions of the behavior determination unit 192 shown in Fig. 2.

[0115] 20 is a diagram illustrating an example of the configuration of a learning system according to an embodiment. In the configuration illustrated in FIG. 20, the learning system includes a history information acquisition unit 621, an evaluation value acquisition unit 622, a learning unit 623, an opponent model execution unit 624, and an evaluation unit 625. With this configuration, the history information acquisition unit 621 acquires history information of proposals made in negotiations. The evaluation value acquisition unit 622 acquires an evaluation value for a proposal from the negotiation partner. The learning unit 623 learns a method for determining one's own proposal to the negotiation partner based on the history information and the evaluation value, and outputs one's own proposal to the negotiation partner based on the learned determination method. The other party model execution unit 624 receives input of one's own proposal to the negotiation partner and outputs a proposal from the negotiation partner. The evaluation unit 625 receives input of a proposal from the negotiation partner and outputs an evaluation value. The history information acquisition unit 621 corresponds to an example of history information acquisition means. The evaluation value acquisition unit 622 corresponds to an example of evaluation value acquisition means. The learning unit 623 corresponds to an example of learning means. The opponent model execution unit 624 corresponds to an example of opponent model execution means. The evaluation unit 625 corresponds to an example of evaluation means.

[0116] The learning system 620 can learn using the proposal history information so that the other party's reaction to one's own actions is reflected in the proposal decision-making method. Because the other party's reaction to one's own actions is reflected in the proposal decision-making method, there is no need to set or learn the utility and strategy of the negotiation partner during operation. In this way, the learning system 620 allows automatic negotiation even when the utility function or utility value of the other party in the negotiation, or the strategy, is unknown.

[0117] The history information acquisition unit 621 can be realized, for example, by using the functions of the history information generation unit 191 shown in Fig. 2, etc. The evaluation value acquisition unit 622 and the learning unit 623 can be realized, for example, by using the functions of the behavior determination unit 192 shown in Fig. 2, etc. The opponent model execution unit 624 can be realized, for example, by using the functions of the opponent model execution unit 291 shown in Fig. 3, etc. The evaluation unit 625 can be realized, for example, by using the functions of the evaluation value calculation unit 292 shown in Fig. 3, etc.

[0118] 21 is a diagram illustrating an example of the configuration of a proposal determination device according to an embodiment. In the configuration illustrated in FIG. 21, a proposal determination device 630 includes a history information acquisition unit 631 and a proposal output unit 632. With this configuration, the history information acquisition unit 631 acquires history information of proposals made in negotiations. The proposal output unit 632 inputs the historical information acquired by the historical information acquisition unit 631 into a trained self-behavior model based on the historical information of proposals made in negotiations and the evaluation value of the proposal from the negotiating partner, and acquires and outputs the negotiating partner's own proposal to the negotiating partner. The history information acquisition unit 631 is an example of a history information acquisition means, and the proposal output unit 632 is an example of a proposal output means.

[0119] The proposal determination device 630 uses a self-behavior model that has been trained based on the history information of proposals made in negotiations and the evaluation values ​​of proposals from the negotiating counterpart, and is expected to reflect the other party's reaction to its own behavior in the proposal determination method. Since the other party's reaction to its own behavior is reflected in the proposal determination method, the proposal determination device 630 does not need to set or learn the utility and strategy of the negotiating counterpart. In this way, the proposal determination device 630 can perform automatic negotiations even when the utility function or utility value of the other party during negotiations, or the strategy, is unknown. The history information acquisition unit 631 can be realized, for example, by using the functions of the history information generation unit 191 shown in Fig. 6. The proposal output unit 632 can be realized, for example, by using the functions of the behavior determination unit 192 shown in Fig. 6.

[0120] 22 is a diagram illustrating an example of a processing procedure in a learning method according to an embodiment. The learning method illustrated in FIG. 22 includes acquiring history information (step S611), acquiring an evaluation value (step S612), and learning a proposal determination method (step S613). In acquiring history information (step S611), the computer acquires history information of proposals made in negotiation. In acquiring evaluation values ​​(step S612), the computer acquires evaluation values ​​for proposals from the negotiating counterparty. In learning a proposal determination method (step S613), the computer learns a method for determining its own proposal to the negotiating counterparty based on the history information and the evaluation values.

[0121] 22, learning is performed using proposal history information, so that the other party's reaction to one's own actions is reflected in the proposal decision-making method. Because the other party's reaction to one's own actions is reflected in the proposal decision-making method, there is no need to set or learn the utility and strategy of the negotiating partner during operation. In this way, according to the learning method shown in FIG. 22, automatic negotiation can be carried out even when the utility function or utility value of the other party during negotiation, or the strategy, is unknown.

[0122] FIG. 23 is a schematic block diagram illustrating an example configuration of a computer according to at least one embodiment. In the configuration shown in FIG. 23, a computer 700 includes a CPU 710, a main memory device 720, an auxiliary memory device 730, an interface 740, and a non-volatile recording medium 750.

[0123] One or more of the learning device 100, data generating device 200, proposal determining device 400, data generating device 200b, learning device 610, learning system 620, and proposal determining device 630, or a portion thereof, may be implemented in a computer 700. In this case, the operation of each of the above-described processing units is stored in the auxiliary storage device 730 in the form of a program. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program. The CPU 710 also allocates storage areas in the main storage device 720 corresponding to each of the above-described storage units in accordance with the program. Communication between each device and other devices is performed by an interface 740 having a communication function and performing communication under the control of the CPU 710. The interface 740 also has a port for a nonvolatile storage medium 750, and reads and writes information from and to the nonvolatile storage medium 750.

[0124] When learning device 100 is implemented in computer 700, the operations of first control unit 190 and each of its units are stored in the form of a program in auxiliary storage device 730. CPU 710 reads the program from auxiliary storage device 730, loads it into main storage device 720, and executes the above-described processing in accordance with the program.

[0125] Furthermore, the CPU 710 allocates a storage area corresponding to the first storage unit 180 in the main storage device 720 in accordance with the program. Communication with other devices by first communication unit 110 is performed by interface 740 having a communication function and operating under the control of CPU 710. Display of images by first display unit 120 is performed by interface 740 having a display device and displaying images under the control of CPU 710. Reception of user operations by first operation input unit 130 is performed by interface 740 having an input device and receiving user operations.

[0126] When the data generating device 200 is implemented in a computer 700, the operations of the second control unit 290 and each of its units are stored in the form of a program in an auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0127] Furthermore, the CPU 710 allocates a storage area corresponding to the second storage unit 280 in the main storage device 720 in accordance with the program. Communication with other devices by second communication unit 210 is performed by interface 740 having a communication function and operating under the control of CPU 710. Display of images by second display unit 220 is performed by interface 740 having a display device and displaying images under the control of CPU 710. Reception of user operations by second operation input unit 230 is performed by interface 740 having an input device and receiving user operations.

[0128] When the proposal determination device 400 is implemented in a computer 700, the operations of the first control unit 190 and each of its units are stored in the form of a program in an auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0129] Furthermore, the CPU 710 allocates a storage area corresponding to the first storage unit 180 in the main storage device 720 in accordance with the program. Communication with other devices by first communication unit 110 is performed by interface 740 having a communication function and operating under the control of CPU 710. Display of images by first display unit 120 is performed by interface 740 having a display device and displaying images under the control of CPU 710. Reception of user operations by first operation input unit 130 is performed by interface 740 having an input device and receiving user operations.

[0130] When the data generating device 200b is implemented in a computer 700, the operations of the opponent model executing unit 291b, the evaluation value calculating unit 292, and each of these units are stored in the form of a program in an auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0131] Furthermore, the CPU 710 allocates a storage area in the main storage device 720 for the processing performed by the data generating device 200b in accordance with the program. Communication between the data generating device 200b and other devices is performed by the interface 740 having a communication function and operating under the control of the CPU 710. Interaction between the data generating device 200b and the user is carried out by the interface 740 having an input device and an output device, presenting information to the user via the output device under the control of the CPU 710, and accepting user operations via the input device.

[0132] When the learning device 610 is implemented in the computer 700, the operations of the history information acquisition unit 611, the evaluation value acquisition unit 612, and the learning unit 613 are stored in the form of a program in the auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0133] Furthermore, the CPU 710 allocates a storage area in the main storage device 720 for the processing performed by the learning device 610 in accordance with the program. Communication between the learning device 610 and other devices is performed by the interface 740, which has a communication function and operates under the control of the CPU 710. Interaction between learning device 610 and a user is carried out by interface 740 having an input device and an output device, presenting information to the user via the output device under the control of CPU 710, and accepting user operations via the input device.

[0134] When the learning system 620 is implemented in a computer 700, the operations of the history information acquisition unit 621, evaluation value acquisition unit 622, learning unit 623, opponent model execution unit 624, and evaluation unit 625 are stored in the form of a program in an auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-mentioned processing in accordance with the program.

[0135] Furthermore, the CPU 710 allocates a storage area in the main storage device 720 for the processing performed by the learning system 620 in accordance with the program. Communication between the learning system 620 and other devices is performed by the interface 740, which has a communication function and operates under the control of the CPU 710. Interaction between the learning system 620 and the user is carried out by the interface 740 having an input device and an output device, which presents information to the user via the output device under the control of the CPU 710, and receives user operations via the input device.

[0136] When the proposal determination device 630 is implemented in the computer 700, the operations of the history information acquisition unit 631 and the proposal output unit 632 are stored in the form of a program in the auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-mentioned processing in accordance with the program.

[0137] Furthermore, the CPU 710 allocates a storage area in the main storage device 720 for the processing performed by the proposal determination device 630 in accordance with the program. Communication between the proposal determination device 630 and other devices is performed by the interface 740 having a communication function and operating under the control of the CPU 710 . Interaction between the proposal determination device 630 and the user is carried out by the interface 740 having an input device and an output device, presenting information to the user via the output device under the control of the CPU 710, and accepting user operations via the input device.

[0138] One or more of the above-described programs may be recorded on nonvolatile recording medium 750. In this case, interface 740 may read the programs from nonvolatile recording medium 750. CPU 710 may then directly execute the programs read by interface 740, or may temporarily store the programs in main storage device 720 or auxiliary storage device 730 and then execute them.

[0139] Note that the processing of each part may be performed by recording a program for executing all or part of the processing performed by learning device 100, data generating device 200, proposal determining device 400, data generating device 200b, learning device 610, learning system 620, and proposal determining device 630 on a computer-readable recording medium, and loading and executing the program recorded on this recording medium into a computer system. Note that the term "computer system" here includes hardware such as an OS (Operating System) and peripheral devices. Furthermore, "computer-readable recording media" refers to portable media such as flexible disks, optical magnetic disks, ROMs (Read Only Memory), and CD-ROMs (Compact Disc Read Only Memory), as well as storage devices such as hard disks built into computer systems. The program may be one that realizes part of the aforementioned functions, or may be one that can realize the aforementioned functions in combination with a program already stored in the computer system.

[0140] Although an embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention. [Explanation of symbols]

[0141] 1,620 Learning System 100, 610 Learning Device 110 First Communications Department 120 First display section 130 First operation input unit 180 First memory section 190 First Control Section 191 History Information Generation Unit 192 Action Decision-Making Department 193 Learning control unit 200, 200b Data generating device 210 Second Communications Department 220 Second display section 230 Second operation input unit 280 Second memory section 290 Second Control Section 291, 291b Opponent model execution part 291c Proposal decision section 291d Noise Adder 292 Evaluation value calculation unit 300 Neural Networks 400, 630 Proposal decision device 611, 621, 631 History information acquisition unit 612, 622 Evaluation value acquisition unit 613, 623 Learning Department 624 Opponent Model Execution Unit 625 Evaluation Department 632 Proposal Output Unit

Claims

1. a history information acquisition means for acquiring history information of proposals made in negotiations; evaluation value acquisition means for acquiring an evaluation value for a proposal from a negotiation partner; a learning means for learning a method for determining a proposal to be made to the negotiation partner based on the history information and the evaluation value, and for outputting a proposal to the negotiation partner based on the determination method; a counterpart model execution means for receiving an input of a proposal from the counterpart and outputting a proposal from the counterpart; evaluation means for receiving an input of a proposal from the negotiating partner and outputting the evaluation value; Equipped with The opponent model execution means a proposal determination means for determining a proposal corresponding to the negotiating party's proposal; a noise adding means for adding noise to the proposal from the negotiating partner determined by the proposal determining means; Equipped with Learning system.

2. The history information acquisition means outputs the history information as data in a vector format, the learning means includes a neural network that receives the history information in the form of vector data and outputs its own proposal to the negotiating partner; The learning system of claim 1 .

3. The computer Obtain historical information about proposals made in negotiations; Obtain an evaluation value for the proposal from the negotiating partner, learning a method for determining a proposal to be made to the negotiation partner based on the history information and the evaluation value, and outputting a proposal to the negotiation partner based on the determination method; receiving an input of one's own proposal to the negotiating partner and outputting a proposal from the negotiating partner; receiving an input of a proposal from the negotiating partner and outputting the evaluation value; This includes: The outputting of the proposal from the negotiation partner is performed by the computer. determining a proposal in response to its proposal to said negotiating partner; Adding noise to the proposal from the determined negotiating partner; A learning method that includes:

4. On the computer, Obtaining historical information about proposals made in negotiations; Obtaining an evaluation value for a proposal from a negotiation partner; learning a method for determining a proposal to the negotiation partner based on the history information and the evaluation value, and outputting a proposal to the negotiation partner based on the determination method; receiving an input of one's own proposal to the negotiation partner and outputting a proposal from the negotiation partner; receiving an input of a proposal from the negotiating partner and outputting the evaluation value; Execute In the outputting of the proposal from the negotiating partner, the computer determining a proposal in response to said proposal to said negotiating partner; adding noise to the proposal from the determined negotiation partner; Run program.

Citation Information

Patent Citations

  • Method and system for performing negotiation task using reinforcement learning agent

    JP2020013568A

  • Negotiation device, estimation method, program, and estimation device

    WO2019146044A1

  • Order-receiving-side negotiation device, order-receiving-side negotiation method, and order-receiving side negotiation program

    WO2021033302A1