Automatic negotiation agent adaptation
By detecting and adapting to changes in the counterparty's utility function through training samples, the negotiation agent enhances negotiation efficiency and profit outcomes in dynamic negotiation scenarios.
Patent Information
- Application Number
- JP2024536105
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-01-14
- Filing Date
- 2022-12-21
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-12-21
AI Technical Summary
Existing automated negotiation strategies become less effective when the utility function of the counterparty agent changes during frequent negotiations, leading to inefficiencies and reduced profit outcomes.
A system and method for detecting changes in the utility function of a counterparty agent, generating training samples from negotiations, and training a negotiation strategy model using these samples to adapt to the new utility function, enabling continuous learning and improvement.
Enables the negotiation agent to negotiate more efficiently and achieve higher profit results by adapting to changes in the counterparty's utility function, ensuring effective and dynamic negotiation strategies.
Smart Images

Figure 0007697598000037 
Figure 0007697598000038 
Figure 0007697598000039
Abstract
Description
Technical Field
[0001] The present disclosure relates to computer-readable media, computer-implemented methods, and apparatuses. This application claims the benefit of U.S. Provisional Application No. 63 / 292,383, filed Dec. 21, 2021, and U.S. Application No. 17 / 575,908, filed Jan. 14, 2022, the entire contents of which are incorporated herein by reference in their entirety.
Background Art
[0002] Negotiation is a decision-making process between two or more parties aiming to reach a mutually beneficial agreement. Automated negotiation involves negotiation between automated agents that act on behalf of real-world entities in order to reduce the time and effort associated with negotiation and achieve a mutually beneficial agreement.
Summary of the Invention
[0003] According to a first exemplary aspect of the present disclosure, a computer-readable medium includes instructions executable by a computer to cause the computer to detect a change in a utility function associated with an automated negotiation between a support agent and a counterparty agent while the support agent operates according to a first negotiation strategy model, generate a plurality of training samples from an automated negotiation between the support agent and the counterparty agent while the support agent operates according to a baseline negotiation strategy model, and train a negotiation strategy model initialized with the plurality of training samples to generate a second negotiation strategy model.
[0004] According to a second exemplary aspect of the present disclosure, a computer-implemented method includes detecting a change in a utility function associated with an automated negotiation between a support agent and an opponent agent while the support agent operates according to a first negotiation strategy model, generating a plurality of training samples from the automated negotiation between the support agent and the opponent agent while the support agent operates according to a baseline negotiation strategy model, and training a negotiation strategy model initialized with the plurality of training samples to generate a second negotiation strategy model.
[0005] According to a third exemplary aspect of the present disclosure, an apparatus includes a controller including a circuit configured to detect a change in a utility function associated with an automated negotiation between a support agent and an opponent agent while the support agent operates according to a first negotiation strategy model, generate a plurality of training samples from the automated negotiation between the support agent and the opponent agent while the support agent operates according to a baseline negotiation strategy model, and train a negotiation strategy model initialized with the plurality of training samples to generate a second negotiation strategy model.
Brief Description of the Drawings
[0006] Aspects of the present disclosure are best understood from the following detailed description when read in conjunction with the accompanying drawings. Note that various features are not drawn to scale in accordance with industry standard practice. In fact, the dimensions of various features may be arbitrarily enlarged or reduced for clarity of explanation.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
[0007] The following disclosure provides many different embodiments or examples for implementing different features of the provided subject matter. Hereinafter, to simplify the present disclosure, specific examples of components, values, operations, materials, arrangements, etc. are described. Of course, these are merely examples and are not intended to be limiting. Other components, values, operations, materials, arrangements, etc. are also conceivable. In addition, the present disclosure can repeat reference numerals and / or letters in various examples. This repetition is for the purpose of simplification and clarity and does not itself define the relationship between the various embodiments and / or configurations described.
[0008] Bilateral automated negotiation is a negotiation between two automated agents. A negotiation repeated between two fixed entities is a continuous negotiation. Exemplary negotiation scenarios consist of the utility functions of each agent and the negotiation domain. An exemplary negotiation protocol is the stacked alternating offer protocol. In at least some embodiments, each agent operates according to a negotiation strategy that is a combination of an acceptance strategy and a bidding strategy.
[0009] In at least some embodiments, the negotiation domain consists of one or more issues. The negotiation outcome space is the set of all possible negotiation outcomes and, in at least some embodiments, is defined as follows.
Number
Number
Number
Number
[0010] If historical data such as records of previous negotiation traces is available, machine learning techniques can be used to create domain-specific negotiation strategies. However, such strategies become less qualified when the utility function of the counter-agent in the negotiation changes during frequent negotiations.
[0011] According to at least some embodiments described herein, the negotiation agent negotiates from historical data and is trained to adapt to changes in the utility function of the counterparty agent. By doing so, at least some embodiments herein enable the negotiation agent to negotiate agreements more efficiently and obtain higher profit results when the utility function changes. At least some embodiments herein include detecting changes in the utility function of the counterparty agent. By doing so, at least some embodiments herein enable the negotiation agent to continue learning and improving over time.
[0012] In at least some embodiments herein, the parameter-based transfer learning approach involves transferring knowledge via shared parameters of the currently trained negotiation strategy model and the newly trained negotiation strategy model. In at least some embodiments, the currently trained negotiation strategy model trained in the source domain has already learned a well-defined structure, and since the task does not change due to any change to the utility function, the structure is transferable to the newly trained negotiation strategy model.
[0013] Figure 1 is a schematic diagram of a system for automatic negotiation agent adaptation according to at least some embodiments of the present disclosure. The system includes an apparatus 100, a support agent 112, and a counterparty agent 119. In at least some embodiments, the apparatus 100 and the support agent 112 are one or more computers such as a personal computer, a server, a mainframe, an instance of cloud computing, etc., including instructions executed by a controller to implement automatic negotiation agent adaptation.
[0014] Device 100 communicates with assistance agent 110 and includes a detection section 170, a generation section 172, and a training section 174. In at least some embodiments, detection section 170 is configured to detect changes in utility functions such as the assistance utility function of assistance agent 110 and the opponent utility function of opponent agent 119. In at least some embodiments, detection section 170 is configured to receive the most recent trace 124 and past traces 122 from negotiation trace storage 120 for opponent utility function change detection. In at least some embodiments, detection section 170 is configured to receive the current and updated assistance utility functions from assistance utility function storage 116 for assistance utility function change detection.
[0015] In at least some embodiments, generation section 172 is configured to generate training samples in response to the detection section detecting a change in the utility function. In at least some embodiments, generation section 172 is configured to receive a baseline trace 126 for training sample generation. In at least some embodiments, generation section 172 is configured to transmit a plurality of training samples to training section 174 for use in negotiating strategy model training.
[0016] In at least some embodiments, training section 174 is configured to train a negotiation strategy model. In at least some embodiments, training section 174 is configured to receive training samples 184 and context information 115 for negotiating strategy model training. In at least some embodiments, training section 174 is configured to transmit the trained negotiation strategy model to assistance agent 110 for use in live automated negotiation.
[0017] In at least some embodiments, the assistance agent 110 is configured to perform an automatic negotiation with the counterparty agent 119, whereby offers 118 are transmitted alternately until an offer is accepted. In at least some embodiments, the assistance agent 110 is configured to operate according to a negotiation strategy model such as a baseline negotiation strategy model 112 or a trained negotiation strategy model 114. In at least some embodiments, the assistance agent 110 is configured to select a counteroffer based on an unprocessed offer and context information 115 from the counterparty agent 119. In at least some embodiments, the context information 115 is dynamic information that affects the amount of profit.
[0018] In at least some embodiments, the baseline negotiation strategy model 112 is an algorithm for selecting offers and counteroffers based on an assistance utility function. In at least some embodiments, the baseline negotiation strategy model 112 is not trained specifically for negotiation with the counterparty agent 119. In at least some embodiments, the trained negotiation strategy model 114 is a machine learning model trained to select offers and counteroffers. In at least some embodiments, the trained negotiation strategy model 114 is trained specifically for negotiation with the counterparty agent 119.
[0019] FIG. 2 is an operation flow for automatic negotiation agent adaptation according to at least some embodiments of the present disclosure. The operation flow provides a method for automatic negotiation agent adaptation. In at least some embodiments, the method is performed by a controller of a device that includes sections for performing specific operations, such as the controller and device shown in FIG. 7 described below.
[0020] In S230, the detection section detects a change in the utility function. In at least some embodiments, the detection section detects a change in the utility function associated with an automated negotiation between the assisting agent and the counterparty agent while the assisting agent operates according to the first negotiation strategy model. In at least some embodiments, the detection section detects a change in the counterparty utility function of the counterparty agent and the first assisting utility function of the assisting agent. In at least some embodiments, the detection of the change in the counterparty utility function proceeds as shown in FIG. 3 described below. In at least some embodiments, the detection of the change in the assisting utility function proceeds as shown in FIG. 4 described below. In at least some embodiments, the detection section detects a change in only one utility function at a time.
[0021] In S240, the generation section generates training samples from the negotiation according to the baseline strategy model. In at least some embodiments, the generation section generates a plurality of training samples from an automated negotiation between the assisting agent and the counterparty agent while the assisting agent operates according to the baseline negotiation strategy model. In at least some embodiments, the generation section instructs the assisting agent to negotiate according to the baseline strategy model. In at least some embodiments, the generation section generates training samples from a plurality of complete negotiations. In at least some embodiments, the generation section generates a plurality of training samples from each complete negotiation.
[0022] In S250, the generation section determines whether a sufficient number of training samples have been generated. In at least some embodiments, the generation section determines whether a sufficient number of training samples have been generated based on the number of training samples. In at least some embodiments, the generation section determines whether a sufficient number of training samples have been generated based on one or more qualifications of the training samples. If the generation section determines that a sufficient number of training samples have been generated, the operation flow proceeds to the training of the new negotiation strategy model in S260. If the generation section determines that a sufficient number of training samples have not yet been generated, the operation flow returns to the training sample generation in S240.
[0023] In S260, the training section trains a new negotiation strategy model. In at least some embodiments, the training section trains a negotiation strategy model initialized with a plurality of training samples and generates a second negotiation strategy model. In at least some embodiments, the training section transmits the trained negotiation strategy model to the support agent and instructs the support agent to use the trained negotiation strategy model in the automated negotiation.
[0024] FIG. 3 is an operation flow for detecting a change in the opponent's utility function according to at least some embodiments of the present disclosure. The operation flow provides a method for detecting a change in the opponent's utility function. In at least some embodiments, the method is implemented by a detection section of a device such as the device shown in FIG. 7 described below.
[0025] In S332, the detection section or its sub-section obtains the most recent negotiation trace. In at least some embodiments, the detection section obtains the most recent negotiation trace from an automated negotiation between the assisting agent and the counterparty agent while the assisting agent is operating according to the first negotiation strategy model. The negotiation trace includes a plurality of time steps, and each time step in the plurality of time steps includes a counterparty agent offer and an assisting agent offer. In at least some embodiments, the most recent negotiation trace is the negotiation trace from the most recent complete negotiation. In at least some embodiments, the detection section obtains two or more most recent negotiation traces, such as the negotiation traces from the five most recent complete negotiations. In at least some embodiments, the detection section obtains the most recent negotiation trace directly from the assisting agent or from the negotiation trace storage unit.
[0026] In S333, the detection section or its sub-section obtains the previous negotiation trace. In at least some embodiments, the detection section obtains the previous negotiation trace from an automated negotiation between the assisting agent and the counterparty agent while the assisting agent is operating according to the first negotiation strategy model. In at least some embodiments, the previous negotiation trace is the negotiation trace used to generate the training samples that were then used to train the first negotiation strategy model. In at least some embodiments, the detection section obtains two or more previous negotiation traces, such as the negotiation traces from the first five complete negotiations, using the first negotiation strategy model. In at least some embodiments, the detection section obtains the previous negotiation trace directly from the assisting agent or from the negotiation trace storage unit.
[0027] In S335, the detection section or its sub-section compares the counterparty offers of the negotiation traces obtained in S332 and S333. In at least some embodiments, the detection section compares the counterparty agent offer of the most recent negotiation trace from an automated negotiation between the assisting agent and the counterparty agent while the assisting agent is operating according to the first negotiation strategy model, with the counterparty agent offer of a previous negotiation trace. In at least some embodiments, the detection section applies a classifier to a tuple of the average frequency distribution of counterparty offers in previous negotiation traces and the frequency distribution of counterparty offers in the most recent negotiation trace.
[0028] In at least some embodiments, the classifier is a binary classifier that classifies whether the negotiation strategy model on which the assisting agent is currently trained should be continued or adapted to a new negotiation strategy model for the negotiation strategy. In at least some embodiments, the classifier assumes that the counterparty utility function and the assisting utility function do not change significantly for a set of consecutive negotiations. In at least some embodiments,
Number
Number
Number
[0029] In at least some embodiments, the classifier is an XGBOOST-based classifier for training the classifier, and the training data is, as input
Number
Number
Number
Number
Number
[0030] In S336, the detection section or its sub-section estimates the change value of the opponent utility function. In at least some embodiments, the detection section estimates a change value representing the amount of change in the opponent utility function. In at least some embodiments, the change value is the output of an opponent offer comparison such as the operation in S335. In at least some embodiments where a classifier is used for opponent offer comparison, the output of the classifier is a boolean value, and the final classification between boolean values is based on a numerical value representing the difference between groups of opponent offers. In at least some embodiments where a distance-based algorithm is used for opponent offer comparison, the distance-based algorithm is applied to the tuple
Number
Number
Number
[0031] In S338, the detection section or its sub-section determines whether the change value estimated in S336 exceeds a threshold value. In at least some embodiments, the threshold value is an adjustable hyperparameter. In at least some embodiments, increasing the threshold value results in fewer instances of new negotiation strategy model training, which is more computationally efficient but yields less effective results, while decreasing the threshold value results in more instances of new negotiation strategy model training, which is less computationally efficient but yields more effective results. In at least some embodiments, the threshold value is adjusted to balance the trade-off between efficiency and effectiveness. In at least some embodiments where a classifier is used for opponent offer comparison, the output of the classifier is a boolean value, and the final classification between boolean values involves comparing a numerical value representing the difference between groups of opponent offers with the threshold value. In at least some embodiments where a distance-based algorithm is used for opponent offer comparison such as Equation 7, the estimated change value m is compared with the threshold value α M is compared with.
[0032] If the detection section determines that the change value exceeds the threshold value, the operation flow proceeds to an operation of waiting, for example, for a predetermined period in S339 before returning to the acquisition of the most recent negotiation trace in S332. In other words, in at least some embodiments, detection is performed periodically.
[0033] If the detection section determines that the change value exceeds the threshold value, the operation flow ends. In at least some embodiments, the end of the operation flow for opponent utility function change detection leads to training sample generation, and ultimately new negotiation strategy model training, such as the operations in S240 and S260 of FIG. 2. In at least some embodiments, generating and training are performed in response to determining that the change value exceeds the threshold value.
[0034] Figure 4 is an operation flow for detecting a change in the support utility function according to at least some embodiments of the present disclosure. The operation flow provides a method for detecting a change in the support utility function. In at least some embodiments, the method is implemented by a detection section of a device such as the device shown in FIG. 7 described below.
[0035] In S431, the detection section or its sub-section receives an updated support utility function. In at least some embodiments, the detection section receives a second support utility function. In at least some embodiments, receiving the updated support utility function triggers the detection of a change in the support utility function. In at least some embodiments, the detection of a change in the support utility function (S430) includes operations in S434, S437, and S438 and is performed in response to receiving the updated support utility function in S431. In other words, in at least some embodiments, the detection is performed in response to receiving the second support utility function.
[0036] In at least some embodiments, one simple method is to always train a new negotiation strategy model in response to a change in the support utility function, that is,
Number
[0037] In S434, the detection section or a sub-section thereof compares a previous utility function with an updated utility function. In at least some embodiments, the detection section compares a first utility function with a second utility function. In at least some embodiments where the utility function includes an ordered result set such as Equation 2, the detection section compares the order of results between the previous utility function and the updated utility function.
[0038] In at least some embodiments, the detection section is a metric based on the Levenshtein distance
Number
Number
Number
Number
Number
Number
Number
Number
[0039] For example, [Number] to [Number] In at least some embodiments, the detection section uses [Number] to [Number] to calculate the number of editing operations (insertion, deletion, or replacement) required to convert to, and compare the top n offers for both utility functions.
[0040] In S437, the detection section or a sub-section thereof estimates the change value of the utility function. In at least some embodiments, the detection section estimates a change value representing the amount of change between a first utility function and a second utility function. In at least some embodiments, the change value is the output of a utility function comparison such as the operation in S434. When the detection section uses a metric [Number] based on the Levenshtein distance to measure the change between utility functions, the detection section derives the change value l using the following relationship: [Number] [Number]
Number
Number
[0041] In S438, the detection section or its sub-section determines whether the change value estimated in S437 exceeds a threshold. In at least some embodiments, the threshold is an adjustable hyperparameter. In at least some embodiments, increasing the threshold results in fewer instances of new negotiation strategy model training, which is more computationally efficient but yields less effective results, while decreasing the threshold results in more instances of new negotiation strategy model training, which is less computationally efficient but yields more effective results. In at least some embodiments, the threshold is adjusted to balance the trade-off between efficiency and effectiveness. In at least some embodiments where the detection section uses a metric based on the Levenshtein distance, such as Equations 9 to 11, to measure the change between utility functions, the estimated change value l is compared with the threshold
Number
[0042] If the detection section determines that the change value exceeds the threshold, the operation flow returns to receiving the updated utility function in S431.
[0043] When the detection section determines that the change value exceeds the threshold, the operation flow ends. In at least some embodiments, the end of the operation flow for opponent utility function change detection leads to training sample generation, and ultimately new negotiation strategy model training, such as the operations at S240 and S260 in FIG. 2. In at least some embodiments, generating and training are performed in response to determining that the change value exceeds the threshold.
[0044] FIG. 5 is an operation flow for training sample generation according to at least some embodiments of the present disclosure. The operation flow provides a method for training sample generation. In at least some embodiments, the method is implemented by a generation section of a device such as the device shown in FIG. 7 described below.
[0045] In S541, the generation section or a sub-section thereof obtains a baseline negotiation trace. In at least some embodiments, the generation section obtains a negotiation trace from an automated negotiation between the assisting agent and the opponent agent while the assisting agent is operating according to the baseline negotiation strategy model, the negotiation trace including a plurality of time steps, and each time step in the plurality of time steps including an opponent agent offer and an assisting agent offer. In at least some embodiments, the generation section obtains a plurality of negotiation traces while the assisting agent is operating according to the baseline negotiation strategy model. In at least some embodiments, a legacy system is temporarily used as the baseline negotiation strategy model instead of the currently trained negotiation strategy model to negotiate with the opponent agent for a few complete negotiations, and H smallGenerate a set of a small number of negotiation traces with the ultimately agreed offer shown as. In at least some embodiments, the legacy system includes one or more instances of any type of compatible automated negotiation agent. In at least some embodiments, the generation section obtains the most recent negotiation trace directly from the assistance agent or from the negotiation trace storage.
[0046] In S543, the generation section or its sub-section starts generating a single training sample by including the first i time steps. In at least some embodiments, the generation section includes a complete set of offers in the training sample from the first i - 1 time steps. In at least some embodiments, the complete set of offers from the time step includes the counterparty offer and the assistance offer. In at least some embodiments, i is from 2 to n, increasing by 1 for each iteration, and n is the number of time steps in the negotiation trace.
[0047] In S544, the generation section or its sub-section continues to generate a single training sample by including the counterparty offer at time step i. In at least some embodiments, the counterparty offer at time step i is the counterparty offer after the complete set of offers in the training sample from the first i - 1 time steps included in S543.
[0048] In S545, the generation section or its sub - section continues to generate a single training sample by labeling the training samples with the assistance offers at time step i. In at least some embodiments, the assistance offer at time step i is included in the training sample such that the training section can identify the assistance offer as a label rather than as an input to the model. In at least some embodiments, in the operations of S543, S544, and S545, the generation section generates each training sample among the plurality of samples to include, as an input, a portion of consecutive time steps among the plurality of time steps from the first time step and the counter - agent offer of subsequent time steps among the plurality of time steps following that portion, and further includes the assisting - agent offer of the subsequent time steps as a label. In at least some embodiments, the generation section generates a set of input sequences and labels from all
Number
[0049] In S547, the generation section or its sub-section determines whether all time steps within the baseline negotiation trace have been processed. In at least some embodiments, the generation section determines that all time steps within the baseline negotiation trace have been processed after the operations in S543, S544, and S545 have been performed for time step n. If the generation section determines that there are unprocessed time steps remaining within the baseline negotiation trace, the operation flow returns to S543, increments i by 1 in S548, and then starts generating another single training sample for the next time step. If the generation section determines that all time steps within the baseline negotiation trace have been processed, the operation flow proceeds to the training sample weighting in S549.
[0050] In S549, the generation section or its sub-section weights the training samples. In at least some embodiments, the generation section weights each training sample based on the utility value of the finally agreed offer in the negotiation trace, and the utility value is obtained by applying the support utility function to the finally agreed offer. In at least some embodiments, the generation section uses the generated input-output set from the training samples generated from the negotiation trace T as the loss weight for the cross-entropy loss while (U s (ω * T )) k to overcome the poor model training caused by the negotiation traces within the negotiation history H D with a low utility value for the finally agreed offer, where ω * T is the finally agreed offer, U S (ω * T ) is the utility value for the finally agreed offer, and k is an adjustable hyperparameter.
[0051] Figure 6 is an operation flow for training a negotiation strategy model according to at least some embodiments of the present disclosure. The operation flow provides a method for training a negotiation strategy model. In at least some embodiments, the method is implemented by a training section of a device such as the device shown in FIG. 7 described below.
[0052] In S662, the training section or a sub-section thereof initializes at least a part of the negotiation strategy model being currently trained. In at least some embodiments, the training section initializes a part of the first negotiation strategy model with random values. In at least some embodiments, the negotiation strategy model uses a bidirectional long short-term memory (LSTM)-based architecture that first has an embedding layer and finally has a single dense layer with softmax activation. In at least some embodiments, the embedding layer captures partial information regarding the utility functions of both the assisting agent and the opponent agent.
[0053] In at least some embodiments, the change in the utility function degrades the performance of the negotiation strategy model being currently trained due to the use of the embedding layer that captures partial information of both utility functions. In at least some embodiments, the embedding layer is retrained when any of the utility functions changes. In at least some embodiments, both the embedding layer and the dense layer are retrained to recapture new information regarding the updated opponent or assisting utility function. In at least some embodiments where new strategy model training responds to the updated assisting utility function, the training section re-uses the weights and biases of the bidirectional LSTM layer from the negotiation strategy model being currently trained to reduce the total number of trainable parameters and to reduce the time and data utilized to train the model. In at least some embodiments, the training section initializes one or more of the embedding layer and the dense layer with random values while maintaining the bidirectional LSTM layer.
[0054] In S664, the training section or a sub-section thereof applies a negotiation strategy model to a training sample. In at least some embodiments, the training section inputs the training sample into the negotiation strategy model and reads an output from the negotiation strategy model. In at least some embodiments, the training section inputs a sequence of offers of the training sample and reads an output support offer from the negotiation strategy model. In at least some embodiments, the training section also inputs context data into the negotiation strategy model.
[0055] In S666, the training section or a sub-section thereof adjusts the negotiation strategy model based on the output from the negotiation strategy model. In at least some embodiments, the training section compares the output from the negotiation strategy model with the label of the training sample. In at least some embodiments, the training section compares the output support offer with the actual support offer from the baseline negotiation trace where the training sample is labeled. In at least some embodiments, the training section applies a loss function to the output and the label and adjusts the negotiation strategy model by deriving a loss value used to adjust the weights of the negotiation strategy model. In at least some embodiments, the strategy model adjustment in S666 is not performed for each iteration of the operations of S664, S666, S668, and S669, but is performed periodically with respect to the number of iterations, or in response to the loss value exceeding a threshold.
[0056] In S668, the training section or a sub-section thereof determines whether an end condition is met. In at least some embodiments, the end condition is the number of training iterations such as a number of epochs. In at least some embodiments, the end condition is met when the loss value falls below a threshold. If the training section determines that the end condition has not yet been met, the operation flow proceeds to select the next training sample (S669) before returning to the strategy model application in S664. If the training section determines that the end condition is met, the operation flow ends.
[0057] In at least some embodiments, the training section fine-tunes the entire negotiation strategy model with a very small learning rate and a small number of epochs. In at least some embodiments, the fine-tuning results in better model accuracy, but the utility value of the finally agreed offer is not significantly greater compared to the finally agreed offer achieved with the negotiation strategy model without fine-tuning.
[0058] FIG. 7 is a block diagram of a hardware configuration for automatic negotiation agent adaptation according to at least some embodiments of the present disclosure.
[0059] An exemplary hardware configuration includes a device 700 that interacts with a support agent 710 and communicates with a network 707. In at least some embodiments, the device 700 is integrated with the support agent 710. In at least some embodiments, the device 700 is a computer system that executes computer-readable instructions for performing operations for physical network function device access.
[0060] Device 700 includes a controller 702, a memory unit 704, a communication interface 706, and an input / output interface 708. In at least some embodiments, controller 702 includes a processor or programmable circuit that executes instructions that cause the processor or programmable circuit to perform operations in accordance with the instructions. In at least some embodiments, controller 702 includes an analog or digital programmable circuit, or any combination thereof. In at least some embodiments, controller 702 includes physically separated storage devices or circuits that interact through communication. In at least some embodiments, memory unit 704 includes a non-volatile computer-readable medium capable of storing executable data and non-executable data for access by controller 702 during execution of instructions. Communication interface 706 transmits and receives data to and from network 707. Input / output interface 708 connects to input device 708 via a parallel port, serial port, keyboard port, mouse port, monitor port, etc., for information exchange.
[0061] Controller 702 includes a detection section 770, a generation section 772, a training section 774, and a transmission section 886. Memory unit 704 includes a utility function 780, a negotiation trace 782, training samples 784, and training parameters 786.
[0062] The detection section 770 is a circuit or instruction of the controller 702 configured to detect changes to the utility function. In at least some embodiments, the detection section 770 is configured to detect changes in the utility function associated with an automated negotiation between the assisting agent and the counterparty agent while the assisting agent operates according to a first negotiation strategy model. In at least some embodiments, the detection section 770 utilizes information within the storage unit 704 such as the negotiation trace 782 and the utility function 780. In at least some embodiments, the detection section 770 includes subsections for implementing additional functionality as described in the aforementioned flowchart. In at least some embodiments, such subsections are referenced by names associated with the corresponding functionality.
[0063] The generation section 772 is a circuit or instruction of the controller 702 configured to generate training samples. In at least some embodiments, the generation section 772 is configured to generate a plurality of training samples from an automated negotiation between the assisting agent and the counterparty agent while the assisting agent operates according to a baseline negotiation strategy model. In at least some embodiments, the generation section 772 utilizes information within the storage unit 704 such as the negotiation trace 782 and records information such as the training samples 784 in the storage unit 704. In at least some embodiments, the generation section 772 includes subsections for implementing additional functionality as described in the aforementioned flowchart. In at least some embodiments, such subsections are referenced by names associated with the corresponding functionality.
[0064] The training section 774 is a circuit or instruction of the controller 702 configured to train the negotiation strategy model. In at least some embodiments, the training section 774 is configured to train a negotiation strategy model initialized using a plurality of training samples and generate a second negotiation strategy model. In at least some embodiments, the training section 774 utilizes information from the storage unit 704 such as the training samples 784 and the training parameters 786. In at least some embodiments, the training section 774 includes subsections for implementing additional functions as described in the aforementioned flowchart. In at least some embodiments, such subsections are referred to by names associated with the corresponding functions.
[0065] In at least some embodiments, the apparatus is another device capable of processing logical functions to perform the operations of this specification. In at least some embodiments, the controller and the storage unit need not be completely separate devices, and in some embodiments share a circuit or one or more computer-readable media. In at least some embodiments, the storage unit includes a hard drive that stores both computer-executable instructions and data accessed by the controller, the controller includes a combination of a central processing unit (CPU) and RAM, and the computer-executable instructions can be copied in whole or in part for execution by the CPU during the performance of the operations of this specification.
[0066] In at least some embodiments where the apparatus is a computer, a program installed on the computer can cause the computer to function as the apparatus of the embodiments described herein or perform operations associated with the apparatus. In at least some embodiments, such a program is executable by a processor to cause the computer to perform specific operations associated with some or all of the blocks of the flowcharts and block diagrams described herein.
[0067] At least some embodiments are described with reference to flowcharts and block diagrams, where the blocks represent (1) steps of a process in which an operation is performed, or (2) sections of a controller responsible for performing the operation. In at least some embodiments, certain steps and sections are implemented by dedicated circuits, programmable circuits supplied with computer-readable instructions stored on a computer-readable medium, and / or processors supplied with computer-readable instructions stored on a computer-readable medium. In at least some embodiments, the dedicated circuits include digital and / or analog hardware circuits, including integrated circuits (ICs) and / or discrete circuits. In at least some embodiments, the programmable circuits include reconfigurable hardware circuits having, for example, logical AND, OR, XOR, NAND, NOR, and other logical operations, flip-flops, registers, memory elements, etc., such as field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), and the like.
[0068] In at least some embodiments, a computer-readable storage medium includes a tangible device that can hold and store instructions for use by an instruction execution device. In some embodiments, the computer-readable storage medium includes, for example, but is not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVDs), memory sticks, floppy disks, punch cards, or mechanically encoded devices such as raised structures within grooves in which instructions are recorded, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.
[0069] In at least some embodiments, the computer-readable program instructions described herein are downloadable from a computer-readable storage medium to respective computing / processing devices or to an external computer or external storage device via a network, such as, for example, the Internet, a local area network, a wide area network, and / or a wireless network. In at least some embodiments, the network includes copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. In at least some embodiments, a network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium within each respective computing / processing device.
[0070] In at least some embodiments, the computer-readable program instructions for performing the operations described above are source code or object code written in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages. In at least some embodiments, the computer-readable program instructions are executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In at least some embodiments, in the latter scenario, the remote computer is connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or is connected to an external computer (e.g., through the Internet using an Internet service provider). In at least some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) executes the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to customize the electronic circuit in order to implement aspects of the present disclosure.
[0071] As described above, embodiments of the present disclosure have been described. However, the technical scope described in the claims is not limited to the above-described embodiments. Those skilled in the art will understand that various modifications and improvements to the above-described embodiments are possible. Those skilled in the art will also understand from the description of the claims that embodiments with such modifications or improvements are included in the technical scope of the present disclosure.
[0072] The operations, procedures, steps, and stages of each process implemented by the apparatus, system, program, and method shown in the claims, embodiments, or figures are not indicated by an order such as "before" or "preceding", and can be implemented in any order unless the output from the previous process is used in the subsequent process. Even if the process flow is described using terms such as "first" or "next" in the claims, embodiments, or figures, such description does not necessarily mean that the process must be implemented in the described order.
[0073] According to at least some embodiments of the present disclosure, the adaptation of the automated negotiation agent is implemented by detecting a change in the utility function associated with the automated negotiation between the assisting agent and the counterparty agent while the assisting agent operates according to a first negotiation strategy model, generating a plurality of training samples from the automated negotiation between the assisting agent and the counterparty agent while the assisting agent operates according to a baseline negotiation strategy model, and training a negotiation strategy model initialized using the plurality of training samples to generate a second negotiation strategy model.
[0074] Some embodiments include instructions in a computer program, a method implemented by a processor that executes the instructions of the computer program, and an apparatus that implements the method. In some embodiments, the apparatus includes a controller that includes a circuit configured to perform the operations within the instructions.
[0075] The above has outlined the features of several embodiments so that those skilled in the art can better understand the aspects of the present disclosure. Those skilled in the art should understand that they can readily use the present disclosure as a basis for designing or modifying other processes and structures that perform the same purposes and / or achieve the same advantages as the embodiments introduced herein. Those skilled in the art should also recognize that such equivalent configurations do not depart from the spirit and scope of the present disclosure, and that various changes, substitutions, and modifications can be made herein without departing from the spirit and scope of the present disclosure.
[0076] Some or all of the above-exemplified embodiments can be described as follows, but are not limited thereto.
[0077] (Appendix 1) A computer detecting a change in a utility function associated with an automated negotiation between the support agent and a counterparty agent while the support agent operates according to a first negotiation strategy model; generating a plurality of training samples from an automated negotiation between the support agent and the counterparty agent while the support agent operates according to a baseline negotiation strategy model; training a negotiation strategy model initialized with the plurality of training samples to generate a second negotiation strategy model A computer-readable medium including computer-executable instructions for causing the computer to perform operations including the above.
[0078] (Appendix 2) The computer-readable medium according to Appendix 1, wherein the utility function is a counterparty utility function of the counterparty agent.
[0079] (Appendix 3) The detecting of the change Obtaining a most recent negotiation trace from an automated negotiation between the support agent and the opponent agent while the support agent is operating according to the first negotiation strategy model, the negotiation trace including a plurality of time steps, each time step of the plurality of time steps including an opponent agent offer and a support agent offer; Comparing the opponent agent offer of the most recent negotiation trace by the support agent with opponent agent offers of previous negotiation traces from an automated negotiation between the support agent and the opponent agent while the support agent is operating according to the first negotiation strategy model; The computer-readable medium according to Appendix 2, comprising.
[0080] (Appendix 4) The detecting further includes estimating a change value representing a change amount of the opponent utility function; The generating and training are performed in response to determining that the change value exceeds a threshold value; The computer-readable medium according to Appendix 3.
[0081] (Appendix 5) The detecting is performed periodically, the computer-readable medium according to Appendix 2.
[0082] (Appendix 6) The utility function is a first support utility function of the support agent, the computer-readable medium according to Appendix 1.
[0083] (Appendix 7) The operation is Further including receiving a second support utility function And The detecting the change includes comparing the first support utility function with the second utility function; The computer-readable medium according to Appendix 6.
[0084] (Appendix 8) Said comparing includes estimating a change value representing a change amount between said first support utility function and said second support utility function, Said generating and training are performed in response to determining that the change value exceeds a threshold value, The computer-readable medium according to Appendix 7.
[0085] (Appendix 9) Said detecting is performed in response to receiving a second support utility function, the computer-readable medium according to Appendix 7.
[0086] (Appendix 10) Said generating the plurality of training samples includes obtaining a negotiation trace from an automated negotiation between the support agent and the opponent agent while the support agent operates according to a baseline negotiation strategy model, the negotiation trace includes a plurality of time steps, and each time step of the plurality of time steps includes an opponent agent offer and a support agent offer, the computer-readable medium according to Appendix 1.
[0087] (Appendix 11) Each training sample of the plurality of samples includes, as input, a part of consecutive time steps of the plurality of time steps from a first time step and an opponent agent offer of subsequent time steps of the plurality of time steps following the part, and further includes, as a label, a support agent offer of the subsequent time steps, the computer-readable medium according to Appendix 10.
[0088] (Appendix 12) Said generating the plurality of training samples is weighting each training sample based on the utility value of the finally agreed offer of the negotiation trace, the utility value being obtained by applying a support utility function to the finally agreed offer and includes, the computer-readable medium according to Appendix 11.
[0089] (Appendix 13) Said training comprises generating an initialized negotiation strategy model by initializing part of the first negotiation strategy model with random values The computer-readable medium according to Appendix 1, comprising the above.
[0090] (Appendix 14) detecting a change in the utility function associated with the automated negotiation between the assisting agent and the counterparty agent while the assisting agent is operating according to the first negotiation strategy model; generating a plurality of training samples from the automated negotiation between the assisting agent and the counterparty agent while the assisting agent is operating according to the baseline negotiation strategy model; training the negotiation strategy model initialized with the plurality of training samples to generate a second negotiation strategy model A computer-implemented method comprising the above.
[0091] (Appendix 15) The computer-implemented method according to Appendix 14, wherein the utility function is the counterparty utility function of the counterparty agent.
[0092] (Appendix 16) Said detecting the change comprises obtaining a recent negotiation trace from the automated negotiation between the assisting agent and the counterparty agent while the assisting agent is operating according to the first negotiation strategy model, the negotiation trace including a plurality of time steps, each of the plurality of time steps including a counterparty agent offer and an assisting agent offer, and comparing the counterparty agent offer of the recent negotiation trace with the counterparty agent offers of previous negotiation traces from the automated negotiation between the assisting agent and the counterparty agent while the assisting agent is operating according to the first negotiation strategy model The computer-implemented method according to Appendix 15, comprising the above.
[0093] (Appendix 17) Said detecting further includes estimating a change value representing a change amount of said opponent utility function, Said generating and training are performed in response to determining that said change value exceeds a threshold value. The computer-implemented method according to Appendix 16.
[0094] (Appendix 18) The utility function is the first support utility function of said support agent, the computer-implemented method according to Appendix 14.
[0095] (Appendix 19) Receiving a second support utility function further includes, Said detecting the change includes comparing said first support utility function with said second utility function. The computer-implemented method according to Appendix 18.
[0096] (Appendix 20) Detecting a change in the utility function associated with an automated negotiation between said support agent and an opponent agent while the support agent operates according to a first negotiation strategy model, Generating a plurality of training samples from an automated negotiation between said support agent and said opponent agent while the support agent operates according to a baseline negotiation strategy model, Training a negotiation strategy model initialized with said plurality of training samples to generate a second negotiation strategy model A controller including a circuit configured to comprise an apparatus.
Claims
1. causing a computer to detect a change in a utility function associated with an automated negotiation between the support agent and a counterparty agent while the support agent operates according to a first negotiation strategy model; generate a plurality of training samples from an automated negotiation between the support agent and the counterparty agent while the support agent operates according to a baseline negotiation strategy model; train a negotiation strategy model initialized with the plurality of training samples to generate a second negotiation strategy model; wherein the utility function is a counterparty utility function of the counterparty agent; wherein detecting the change includes: obtaining a most recent negotiation trace from an automated negotiation between the support agent and the counterparty agent while the support agent operates according to the first negotiation strategy model, the negotiation trace including a plurality of time steps, each of the plurality of time steps including a counterparty agent offer and a support agent offer; comparing the counterparty agent offer of the most recent negotiation trace with a counterparty agent offer of a previous negotiation trace from an automated negotiation between the support agent and the counterparty agent while the support agent operates according to the first negotiation strategy model; A program for causing the operations described above.
2. The detecting further includes estimating a change value representing an amount of change in the counterparty utility function, wherein generating and training are performed in response to determining that the change value exceeds a threshold value. The program according to claim 1.
3. The detecting is performed periodically. The program according to claim 1.
4. A method, comprising: detecting, by a computer, a change in a utility function associated with an automated negotiation between a support agent and a counterparty agent while the support agent operates according to a first negotiation strategy model; generating, by the computer, a plurality of training samples from an automated negotiation between the support agent and the counterparty agent while the support agent operates according to a baseline negotiation strategy model; training, by the computer, a negotiation strategy model initialized with the plurality of training samples to generate a second negotiation strategy model. The utility function is the opponent's utility function of the opponent agent, Detecting the change is obtaining a most recent negotiation trace from an automated negotiation between the assisting agent and the opponent agent while the assisting agent operates according to the first negotiation strategy model, the negotiation trace including a plurality of time steps, each of the plurality of time steps including an opponent agent offer and an assisting agent offer, and comparing the opponent agent offer of the most recent negotiation trace by the assisting agent with the opponent agent offers of previous negotiation traces from the automated negotiation between the assisting agent and the opponent agent while the assisting agent operates according to the first negotiation strategy model including a method. **Claim 5** A controller including a circuit configured to detect a change in a utility function associated with an automated negotiation between an assisting agent and an opponent agent while the assisting agent operates according to a first negotiation strategy model, generate a plurality of training samples from an automated negotiation between the assisting agent and the opponent agent while the assisting agent operates according to a baseline negotiation strategy model, and train a negotiation strategy model initialized with the plurality of training samples to generate a second negotiation strategy model is provided, an apparatus comprising: The utility function is the opponent's utility function of the opponent agent, Detecting the change is obtaining a most recent negotiation trace from an automated negotiation between the assisting agent and the opponent agent while the assisting agent operates according to the first negotiation strategy model, the negotiation trace including a plurality of time steps, each of the plurality of time steps including an opponent agent offer and an assisting agent offer, and comparing the opponent agent offer of the most recent negotiation trace by the assisting agent with the opponent agent offers of previous negotiation traces from the automated negotiation between the assisting agent and the opponent agent while the assisting agent operates according to the first negotiation strategy model including an apparatus.
Citation Information
Patent Citations
Automated negotiation
US20120290485A1
Method and system for performing negotiation task using reinforcement learning agents
US20200020061A1
Information processing device, information processing method, and program
WO2021070732A1