Negotiation information intelligent processing method, system, intelligent terminal and storage medium

By constructing a strategic multi-issue negotiation model through an extended game tree modeling framework, the inefficiency problem caused by relying on clear opponent modeling in existing technologies is solved, and efficient, stable and fair negotiations are achieved in an incomplete information environment, thereby improving negotiation efficiency and agreement quality.

CN120450830BActive Publication Date: 2025-09-16HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510955049.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-09-16
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

The existing technology for intelligent processing of negotiation information based on the BOA framework relies too much on explicit opponent modeling, resulting in low efficiency in negotiation information processing under incomplete information environments, high computational costs, and difficulty in achieving efficient and fair agreements in fast negotiation environments.

Method used

A strategic multi-issue negotiation model is constructed using the extended game tree modeling framework. By obtaining the negotiation agreement, negotiation domain, preference configuration of negotiation objects, negotiation strategy and outcome space, nodes and edges are constructed to represent negotiation status and actions. The strategy is dynamically adjusted to optimize the negotiation strategy and reduce the dependence on explicit opponent modeling.

Benefits of technology

It improves the efficiency of negotiation information processing in incomplete information environments, improves negotiation efficiency and agreement quality, and achieves efficient, stable and fair negotiation results in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120450830B_ABST
    Figure CN120450830B_ABST
Patent Text Reader

Abstract

The present application discloses a method, system, intelligent terminal, and storage medium for intelligent processing of negotiation information, relating to the field of data processing technology. The method comprises: obtaining a negotiation agreement, a negotiation domain, and the preference configuration, negotiation strategy, and result space corresponding to the negotiation partner; modeling based on the negotiation agreement, the negotiation domain, and the preference configuration, negotiation strategy, and result space corresponding to the negotiation partner, using an extended game tree modeling framework to construct a strategic multi-issue negotiation model, wherein the nodes in the strategic multi-issue negotiation model represent the negotiation state, and the edges in the strategic multi-issue negotiation model represent the negotiation actions of the negotiation partner, wherein the negotiation actions of the negotiation partner include making a new bid, accepting the current bid, rejecting the current bid, and leaving the negotiation; and determining and outputting the negotiation results between the negotiation partners based on the leaf nodes in the strategic multi-issue negotiation model. This is conducive to improving the efficiency of negotiation information processing in an incomplete information environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a negotiation information intelligent processing method, system, intelligent terminal and storage medium. Background Art

[0002] With the advancement of science and technology, especially artificial intelligence, the application of intelligent negotiation is becoming increasingly widespread. Existing technologies enable intelligent processing of negotiation information based on pre-defined frameworks, such as the BOA (Bidding, Opponent, Acceptance) framework. The BOA framework encompasses bidding strategies, opponent modeling, and acceptance conditions, aiming to modularize the key elements of the negotiation process.

[0003] The problem with existing technologies is that when performing intelligent processing of negotiation information based on existing frameworks, they rely too much on explicit opponent modeling, which is not conducive to improving the efficiency of negotiation information processing in an incomplete information environment.

[0004] Therefore, relevant technologies still need to be improved and developed. Summary of the Invention

[0005] The main purpose of this application is to provide a method, system, intelligent terminal and storage medium for intelligent processing of negotiation information, aiming to solve the technical problem in related technologies that when performing intelligent processing of negotiation information based on the existing framework, it relies too much on clear opponent modeling, which is not conducive to improving the efficiency of negotiation information processing in an incomplete information environment.

[0006] In order to achieve the above-mentioned objectives, the first aspect of the present application provides a method for intelligently processing negotiation information, wherein the method for intelligently processing negotiation information includes:

[0007] Obtain negotiation protocols, negotiation domains, and the corresponding preference configurations, negotiation strategies, and outcome spaces of negotiation partners;

[0008] Based on the negotiation protocol, the negotiation domain, and the preference configurations, negotiation strategies, and outcome spaces corresponding to the negotiation partners, a strategic multi-issue negotiation model is constructed based on an extended game tree modeling framework. Nodes in the strategic multi-issue negotiation model represent negotiation states, and edges in the strategic multi-issue negotiation model represent negotiation actions of the negotiation partners. The negotiation actions of the negotiation partners include making a new bid, accepting the current bid, rejecting the current bid, and leaving the negotiation.

[0009] The negotiation results between the negotiation parties are determined and output according to the leaf nodes in the strategic multi-issue negotiation model.

[0010] Optionally, the above negotiation agreement includes an alternating offer agreement.

[0011] Optionally, the preference configuration corresponding to the negotiation partner includes a linear additive utility function model for characterizing the preferences of the negotiation partner.

[0012] Optionally, the above-mentioned negotiation partner includes a first negotiation partner and a second negotiation partner;

[0013] Based on the negotiation agreement, the negotiation domain, and the preference configuration corresponding to the negotiation object, the modeling is performed based on the extended game tree modeling framework to construct a strategic multi-issue negotiation model, including:

[0014] Obtaining an initial state, and generating a root node of the expanded game tree according to the initial state;

[0015] Using the root node as the negotiation status node corresponding to the first negotiation object;

[0016] Determining a first negotiation action of the first negotiation partner based on the negotiation state node according to the negotiation state node corresponding to the first negotiation partner, as well as the preference configuration, negotiation strategy, and result space corresponding to the first negotiation partner, and generating a negotiation state node corresponding to the first negotiation action;

[0017] If the first negotiation action is to propose a new bid, then, based on the negotiation state node corresponding to the first negotiation action and the preference configuration, negotiation strategy, and result space corresponding to the second negotiation partner, determining a second negotiation action for the second negotiation partner based on the negotiation state node, and generating a negotiation state node corresponding to the second negotiation action;

[0018] If the second negotiation action is to reject the current bid, then determining a third negotiation action of the second negotiation partner based on the negotiation state node according to the negotiation state node corresponding to the second negotiation action, and the preference configuration, negotiation strategy, and result space corresponding to the second negotiation partner, and generating a negotiation state node corresponding to the third negotiation action;

[0019] If the third negotiation action is to make a new bid, the negotiation state node corresponding to the third negotiation action is used as the negotiation state node corresponding to the first negotiation partner, and the process returns to executing the negotiation state node corresponding to the first negotiation partner, as well as the preference configuration, negotiation strategy, and result space corresponding to the first negotiation partner, determining the first negotiation action of the first negotiation partner based on the negotiation state node, and generating the negotiation state node corresponding to the first negotiation action, until a preset generation termination condition is reached, thereby obtaining a strategic multi-issue negotiation model.

[0020] Optionally, the above-mentioned preset generation termination conditions include:

[0021] Any negotiation action obtained is accepting the current bid or leaving the negotiation, or the processing time reaches a preset time threshold.

[0022] Optionally, the negotiation strategy includes an acceptance strategy and a bidding strategy.

[0023] Optionally, if the second negotiation action is to reject the current bid, determining a third negotiation action of the second negotiation partner based on the negotiation state node according to the negotiation state node corresponding to the second negotiation action, and the preference configuration, negotiation strategy, and result space corresponding to the second negotiation partner, and generating a negotiation state node corresponding to the third negotiation action, including:

[0024] If the second negotiation action is to reject the current bid, determining, based on the negotiation state node corresponding to the second negotiation action and the preference configuration, bidding strategy, and result space corresponding to the second negotiation partner, that the third negotiation action of the second negotiation partner based on the negotiation state node is to propose a new bid;

[0025] Obtaining a regret value based on the negotiation state node corresponding to the second negotiation action, and the historical branches and historical nodes corresponding to the negotiation state node corresponding to the second negotiation action, wherein the regret value is used to represent the benefit gap between the actual choice of the negotiation partner and the optimal action;

[0026] Dividing the result space corresponding to the second negotiation partner into segments to obtain multiple bidding segments;

[0027] Determining the segment selection probability corresponding to each of the bidding segments according to the regret value corresponding to each of the bidding segments;

[0028] According to the segment selection probability, as well as the preference configuration, bidding strategy, and result space corresponding to the second negotiation partner, the bid corresponding to the second negotiation partner is obtained to generate a negotiation state node corresponding to the third negotiation action.

[0029] A second aspect of the present application provides a negotiation information intelligent processing system, wherein the negotiation information intelligent processing system includes:

[0030] The data acquisition module is used to obtain the negotiation agreement, negotiation domain, and the preference configuration, negotiation strategy and result space corresponding to the negotiation object;

[0031] a model building module for modeling based on the negotiation protocol, the negotiation domain, and the preference configurations, negotiation strategies, and outcome spaces corresponding to the negotiation partners, using an extended game tree modeling framework to construct a strategic multi-issue negotiation model, wherein nodes in the strategic multi-issue negotiation model represent negotiation states, and edges in the strategic multi-issue negotiation model represent negotiation actions of the negotiation partners, wherein the negotiation actions of the negotiation partners include making a new bid, accepting the current bid, rejecting the current bid, and leaving the negotiation;

[0032] The negotiation result acquisition module is used to determine the negotiation result between the above-mentioned negotiation parties based on the leaf nodes in the above-mentioned strategic multi-issue negotiation model and output it.

[0033] The third aspect of the present application provides an intelligent terminal, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, any one of the steps of the intelligent processing method for negotiation information is implemented.

[0034] A fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, any one of the steps of the above-mentioned method for intelligent processing of negotiation information is implemented.

[0035] As can be seen from the above, in the present application scheme, a negotiation agreement, a negotiation domain, and the preference configuration, negotiation strategy and result space corresponding to the negotiation object are obtained; based on the above negotiation agreement, the above negotiation domain, and the preference configuration, negotiation strategy and result space corresponding to the above negotiation object, modeling is performed based on the modeling framework of the extended game tree to construct a strategic multi-issue negotiation model, wherein the nodes in the above strategic multi-issue negotiation model represent the negotiation state, and the edges in the above strategic multi-issue negotiation model represent the negotiation actions of the above negotiation objects, and the negotiation actions of the above negotiation objects include making a new bid, accepting the current bid, rejecting the current bid and leaving the negotiation; the negotiation results between the above negotiation objects are determined and output based on the leaf nodes in the above strategic multi-issue negotiation model.

[0036] Compared with the existing technology, the intelligent processing method of negotiation information provided in this application is based on the modeling framework of the extended game tree, thereby constructing a strategic multi-issue negotiation model, instead of using the existing BOA framework. The strategic multi-issue negotiation model used in this application provides structured modeling for multi-issue negotiations through an extended game framework, enabling the intelligent agent to systematically reason and optimize negotiation strategies. Based on the idea of ​​game theory, a negotiation framework is constructed to support the intelligent agent to dynamically adjust its strategy in an incomplete information environment, thereby not relying too much on clear opponent modeling, which is conducive to improving the efficiency of negotiation information processing in an incomplete information environment, thereby improving negotiation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0038] Figure 1 This is a flowchart of a method for intelligently processing negotiation information provided by an embodiment of the present application;

[0039] Figure 2 This is a schematic diagram of a utility space provided in an embodiment of the present application;

[0040] Figure 3 This is a schematic diagram of the structure of a strategic multi-issue negotiation model provided in an embodiment of the present application;

[0041] Figure 4 This is a schematic diagram of solving the regret value in the negotiation result space provided by an embodiment of the present application;

[0042] Figure 5 This is a schematic diagram comparing the performance indicators of an intelligent agent corresponding to the method of the present application and other intelligent agents provided in an embodiment of the present application;

[0043] Figure 6 This is a schematic diagram comparing the performance indicators of another intelligent agent corresponding to the method of the present application and other intelligent agents provided in an embodiment of the present application;

[0044] Figure 7 This is a schematic diagram comparing the performance indicators of another intelligent agent corresponding to the method of the present application and other intelligent agents provided in an embodiment of the present application;

[0045] Figure 8 This is a schematic diagram of the components of a negotiation information intelligent processing system provided in an embodiment of the present application;

[0046] Figure 9 This is a block diagram of the internal structure principle of a smart terminal provided in an embodiment of the present application. DETAILED DESCRIPTION

[0047] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it should be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0048] It will be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0049] The following is a clear and complete description of the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0050] Intelligent negotiation is a key research area in artificial intelligence (AI), with increasing application across various fields, such as e-commerce. Currently, ensuring that intelligent agents can reach agreements that are both efficient and fair remains a key challenge. Traditional negotiation strategies often struggle to maintain this balance, resulting in suboptimal negotiation outcomes or significantly uneven distribution of benefits, with one party benefiting disproportionately while the other suffers, thus undermining negotiation effectiveness.

[0051] Existing technologies use the BOA framework for intelligent processing of negotiation information. While highly modular, this framework relies heavily on explicit opponent modeling, resulting in high computational costs in fast-paced negotiation environments. Furthermore, traditional time-dependent strategies arbitrarily make concessions based on timelines during negotiations, easily leading to suboptimal agreements. While behavior-dependent strategies can dynamically adjust offers based on the opponent's actions, they also encounter computational bottlenecks when the opponent's behavior is difficult to accurately predict.

[0052] In game theory, regret matching algorithms have been widely used in strategic decision-making due to their stability and efficiency. By continuously adjusting the regret value based on historical decisions, intelligent agents can gradually approach a Nash equilibrium. However, research on the application of regret matching algorithms in automated negotiation remains limited, especially in multi-issue negotiations, which require further exploration.

[0053] In summary, the current field of intelligent negotiation information processing faces numerous challenges, such as a vast outcome space, incomplete information, difficulty balancing strategies, and the pressure of real-time decision-making. These issues impact the effectiveness of intelligent agents in negotiations. For example, they hinder fairness in negotiations under incomplete information, incur high computational costs, hinder negotiation efficiency, and hinder the achievement of optimal negotiation outcomes.

[0054] In order to solve at least one of the above-mentioned technical problems, in the solution of the present application, a negotiation agreement, a negotiation domain, and the preference configuration, negotiation strategy and result space corresponding to the negotiation object are obtained; based on the above-mentioned negotiation agreement, the above-mentioned negotiation domain, and the preference configuration, negotiation strategy and result space corresponding to the above-mentioned negotiation object, modeling is performed based on the modeling framework of the extended game tree to construct a strategic multi-issue negotiation model, wherein the nodes in the above-mentioned strategic multi-issue negotiation model represent the negotiation state, and the edges in the above-mentioned strategic multi-issue negotiation model represent the negotiation actions of the above-mentioned negotiation objects, and the negotiation actions of the above-mentioned negotiation objects include making a new bid, accepting the current bid, rejecting the current bid and leaving the negotiation; the negotiation results between the above-mentioned negotiation objects are determined and output according to the leaf nodes in the above-mentioned strategic multi-issue negotiation model.

[0055] Compared with the existing technology, the intelligent processing method of negotiation information provided in this application is based on the modeling framework of the extended game tree, thereby constructing a strategic multi-issue negotiation model, instead of using the existing BOA framework. The strategic multi-issue negotiation model used in this application provides structured modeling for multi-issue negotiations through an extended game framework, enabling the intelligent agent to systematically reason and optimize negotiation strategies. Based on the idea of ​​game theory, a negotiation framework is constructed to support the intelligent agent to dynamically adjust its strategy in an incomplete information environment, thereby not relying too much on clear opponent modeling, which is conducive to improving the efficiency of negotiation information processing in an incomplete information environment, thereby improving negotiation efficiency.

[0056] Furthermore, the intelligent negotiation information processing method provided in this application can implement multi-issue bilateral intelligent negotiation based on dynamic bid regret matching, effectively improving negotiation efficiency and agreement quality. Intelligent negotiation involves multiple interrelated challenges, making the efficient performance of negotiating agents extremely difficult. First, the outcome space in multi-issue negotiations is enormous. As the number of negotiation topics increases, the number of optional bidding agreements increases exponentially, making it impossible for the agent to fully evaluate all potential bids within limited computing resources. Second, negotiations are often conducted with incomplete information, and the agent makes decisions under the condition of unknown opponent preferences, reservations, and strategies, further increasing the complexity of negotiations. In addition, the agent must strike a balance between maximizing its own utility and ensuring a cooperative agreement to avoid negotiation failure. Finally, in a real-time decision-making environment, the agent must quickly select the optimal bid within a vast action space, which places extremely high demands on computing resources and strategy optimization capabilities. To address these challenges, this application dynamically adjusts the bidding strategy to improve negotiation efficiency and agreement quality without relying on opponent modeling, ensuring that the negotiating agent can achieve efficient, stable, and fair negotiation results in complex environments.

[0057] like Figure 1As shown, the embodiment of the present application provides a method for intelligently processing negotiation information. Specifically, the method includes the following steps:

[0058] Step S100, obtaining the negotiation agreement, negotiation domain, and the preference configuration, negotiation strategy, and result space corresponding to the negotiation object;

[0059] Step S200: Based on the negotiation protocol, the negotiation domain, and the preference configurations, negotiation strategies, and outcome spaces corresponding to the negotiation partners, a model is constructed based on an extended game tree modeling framework to construct a strategic multi-issue negotiation model. Nodes in the strategic multi-issue negotiation model represent negotiation states, and edges in the strategic multi-issue negotiation model represent negotiation actions of the negotiation partners. The negotiation actions of the negotiation partners include making a new bid, accepting the current bid, rejecting the current bid, and leaving the negotiation.

[0060] Step S300 : determining the negotiation results between the negotiation partners based on the leaf nodes in the strategic multi-issue negotiation model and outputting the results.

[0061] It should be noted that the intelligent processing method for negotiation information provided in the embodiment of the present application can be applied to multi-topic multilateral negotiation scenarios. The embodiment of the present application is specifically explained by taking the application in a multi-topic bilateral negotiation scenario as an example, and is specifically explained by taking the application in a car purchase and sale negotiation scenario as an example, in which two negotiating parties conduct bilateral negotiations.

[0062] Specifically, a negotiation scenario consists of three parts: a negotiation protocol, a negotiation domain, and the preference profiles of the participants. The protocol characterizes the interaction method, the domain characterizes the topics and options available for negotiation, and the preference profile characterizes the order of preference or utility value of each negotiating agent (i.e., the negotiating partner) for different bids. In this embodiment, the proposed intelligent negotiation information processing method is used as an example in a real-world multi-issue negotiation, illustrating the application of the proposed method in this context.

[0063] In the intelligent processing method for negotiation information provided in this application, modeling is performed based on the modeling framework of the extended game tree, thereby constructing a strategic multi-issue negotiation model, instead of using the existing BOA framework. The strategic multi-issue negotiation model used in this application provides structured modeling for multi-issue negotiations through an extended game framework, enabling the intelligent agent to systematically reason and optimize negotiation strategies. Based on game theory, a negotiation framework is constructed to support the intelligent agent to dynamically adjust its strategy in an incomplete information environment, thereby not relying too much on clear opponent modeling, which is conducive to improving the efficiency of negotiation information processing in an incomplete information environment, thereby improving negotiation efficiency.

[0064] Specifically, the aforementioned negotiation protocol includes an alternating offer protocol. In the multi-issue bilateral negotiation scenario addressed by this application, two agents (i.e., negotiating parties) interact using the alternating offer protocol. Specifically, within a limited timeframe, two agents (i.e., negotiating parties) take turns submitting offers (i.e., a complete set of issue options). In each negotiation round, one agent presents a new offer, and the other agent can choose to accept it, thereby reaching an agreement, or reject it and submit a counteroffer in the next round. This process continues until both parties reach agreement on a single offer, or the negotiation fails due to time constraints. The entire negotiation process constitutes an ordered sequence of bids (bidding history), with response decisions in each round constrained by factors such as the current round, the historical bid sequence, known preferences, and a preset deadline. The protocol emphasizes turn-taking and bid completeness (each round presents bids on a full set of issues). A maximum number of rounds, a time limit, or a discount factor can be set as needed to constrain the negotiation process. In actual implementation, the agent dynamically adjusts its bidding and acceptance strategies based on its own utility function and its prediction model of the opponent, in order to reach an efficient, fair and mutually acceptable agreement within the rounds allowed by the protocol.

[0065] The negotiation domain is used to represent all possible topics involved in the negotiation and their discrete value sets. It is the basic element that constitutes the outcome space. The example negotiation domain in this application is based on the background of "automobile sales configuration negotiation" and involves six topics, namely: CD player, extra speakers, air conditioning, paint finish, tire and wheel configuration (Tires & Rims), and navigation and multimedia system (Navigation & Media). Each topic is discrete and has five configuration options, from high to low: good, fairly good, standard, meager, and none. The details are shown in Table 1 below:

[0066] Table 1

[0067]

[0068] In this setting, each issue has 5 possible options, so any complete offer can be represented as a six-tuple .

[0069] Among them, for a set of issues in the negotiation , Indicates in The option value on the topic. For example: , which means: select a car with a better CD player (fairly good), standard speakers, no air conditioning (none), better paint (fairly good), average tires (meager), and a good navigation system (good).

[0070] Each topic has five options, and the entire result space includes: possible bid combinations, forming a discrete multidimensional outcome space in the negotiation. Each bid will be evaluated by the respective utility functions of both parties.

[0071] This type of high-dimensional, discrete, and interactive negotiation domain modeling is representative and widely exists in real-world business scenarios (such as customized sales, resource allocation, etc.). It is also one of the standard test benchmarks for evaluating the performance of automatic negotiation algorithms.

[0072] Furthermore, the preference configuration corresponding to the negotiation partner includes a linear additive utility function model for characterizing the preference of the negotiation partner.

[0073] In automated negotiations, agents evaluate the value of each bid based on their own preference functions. This application uses a linear additive utility function model to represent the preference configurations of both parties. Specifically, each agent assigns a weight to each issue. , and for each option under the topic Specify a normalized utility value The utility value of the overall bid is the sum of the weighted utilities of each issue: To illustrate the calculation process, let's take Party A (the first negotiating party) as an example to show its utility function for a specific bid: The utility value calculation process is shown in Table 2 below:

[0074] Table 2

[0075]

[0076] The total utility value obtained by Party A under this bid is 0.49807. Similarly, Party B (the second negotiating party) can be used to evaluate this bid, determining Party B's satisfaction with this configuration and, therefore, whether the bid is likely to be accepted. This non-zero-sum preference structure demonstrates that the utility functions of both parties are not completely opposed, thus creating ample room for negotiation and making it possible to reach a win-win agreement through strategic proposals.

[0077] It should be noted that the result space can also be obtained (the corresponding result space can also be generated or determined according to the topic) to realize the preference configuration, negotiation strategy and result space corresponding to the above-mentioned negotiation agreement, the above-mentioned negotiation domain, and the above-mentioned negotiation object, and to combine the result space with the modeling framework based on the extended game tree to build a strategic multi-issue negotiation model.

[0078] Figure 2 This is a utility space diagram provided by an embodiment of the present application. Specifically, the outcome space (OutcomeSpace) consists of all possible bid combinations in the negotiation domain. In this example, the size of the outcome space is 15625, and each bid is mapped to two scalar values ​​under the utility functions of Party A and Party B respectively. and Therefore, the entire result space can be represented as a set of points in a two-dimensional coordinate system , forming a utility scatter plot. A key structure in this space is the Pareto frontier. A bid is called "Pareto optimal" if and only if there is no other bid that increases the utility of one party without decreasing the utility of the other. The set of all such bids is the Pareto frontier, denoted by: The Pareto front is an important reference line for measuring negotiation efficiency and fairness. Our Dynamic Offer-Regret Matching (DORM) method uses a regret-driven learning mechanism and importance sampling strategy during bid generation to guide the agent toward the Pareto front, thereby improving agreement quality and accelerating the speed of reaching agreements.

[0079] In the embodiment of the present application, a strategic multi-issue negotiation (SMN) model based on an extended-form game tree is constructed. Specifically, the negotiation objects include a first negotiation object and a second negotiation object.

[0080] Based on the negotiation agreement, the negotiation domain, and the preference configuration corresponding to the negotiation object, the modeling is performed based on the extended game tree modeling framework to construct a strategic multi-issue negotiation model, including:

[0081] Obtaining an initial state, and generating a root node of the expanded game tree according to the initial state;

[0082] Using the root node as the negotiation status node corresponding to the first negotiation object;

[0083] Determining a first negotiation action of the first negotiation partner based on the negotiation state node according to the negotiation state node corresponding to the first negotiation partner, as well as the preference configuration, negotiation strategy, and result space corresponding to the first negotiation partner, and generating a negotiation state node corresponding to the first negotiation action;

[0084] If the first negotiation action is to propose a new bid, then, based on the negotiation state node corresponding to the first negotiation action and the preference configuration, negotiation strategy, and result space corresponding to the second negotiation partner, determining a second negotiation action for the second negotiation partner based on the negotiation state node, and generating a negotiation state node corresponding to the second negotiation action;

[0085] If the second negotiation action is to reject the current bid, then determining a third negotiation action of the second negotiation partner based on the negotiation state node according to the negotiation state node corresponding to the second negotiation action, and the preference configuration, negotiation strategy, and result space corresponding to the second negotiation partner, and generating a negotiation state node corresponding to the third negotiation action;

[0086] If the third negotiation action is to make a new bid, the negotiation state node corresponding to the third negotiation action is used as the negotiation state node corresponding to the first negotiation partner, and the process returns to executing the negotiation state node corresponding to the first negotiation partner, as well as the preference configuration, negotiation strategy, and result space corresponding to the first negotiation partner, determining the first negotiation action of the first negotiation partner based on the negotiation state node, and generating the negotiation state node corresponding to the first negotiation action, until a preset generation termination condition is reached, thereby obtaining a strategic multi-issue negotiation model.

[0087] The preset termination conditions include: any negotiation action obtained is accepting the current bid or leaving the negotiation, or the processing time reaches a preset time threshold. The negotiation strategies include acceptance strategies and bid strategies. Specifically, the processing time is the time required for intelligent processing of negotiation information or the time required to construct a strategic multi-issue negotiation model.

[0088] Furthermore, if the second negotiation action is to reject the current bid, determining a third negotiation action of the second negotiation partner based on the negotiation state node according to the negotiation state node corresponding to the second negotiation action, and the preference configuration, negotiation strategy, and result space corresponding to the second negotiation partner, and generating a negotiation state node corresponding to the third negotiation action, including:

[0089] If the second negotiation action is to reject the current bid, determining, based on the negotiation state node corresponding to the second negotiation action and the preference configuration, bidding strategy, and result space corresponding to the second negotiation partner, that the third negotiation action of the second negotiation partner based on the negotiation state node is to propose a new bid;

[0090] Obtaining a regret value based on the negotiation state node corresponding to the second negotiation action, and the historical branches and historical nodes corresponding to the negotiation state node corresponding to the second negotiation action, wherein the regret value is used to represent the benefit gap between the actual choice of the negotiation partner and the optimal action;

[0091] Dividing the result space corresponding to the second negotiation partner into segments to obtain multiple bidding segments;

[0092] Determining the segment selection probability corresponding to each of the bidding segments according to the regret value corresponding to each of the bidding segments;

[0093] According to the segment selection probability, as well as the preference configuration, bidding strategy, and result space corresponding to the second negotiation partner, the bid corresponding to the second negotiation partner is obtained to generate a negotiation state node corresponding to the third negotiation action.

[0094] In bilateral, multi-issue negotiations, intelligent agents need to reach a mutually beneficial agreement among a combination of multiple alternative issues. To structure this modeling process, this application uses a modeling framework based on an extended game tree to construct a strategic multi-issue negotiation model. This tree-based structure is well suited to representing the decision sequence and possible outcomes throughout the negotiation process, highlighting the strategic choices of the agents at each stage. Unlike traditional models that focus on a single-step bidding strategy, this approach continuously adjusts the strategy during the negotiation process to form a sequential bidding strategy, better supporting the access of subsequent algorithmic components (such as regret value updates and strategy evolution) to structured information.

[0095] Figure 3 This is a schematic diagram of the structure of a strategic multi-issue negotiation model provided by the embodiment of this application. Specifically, the modeling in the embodiment of this application takes "auto sales configuration negotiation" as an example: the negotiation domain contains 6 issues, each issue has 5 options, and the negotiation is carried out between a car dealer and a car buyer. In actual implementation, the modeling and data processing process of the SMN model is as follows, and the construction method is as follows: Figure 3 shown.

[0096] (1) Negotiation domain and outcome space: Negotiators are a finite set Indicates that and There are two negotiation agents. In the scenario of "car sales configuration negotiation", A corresponds to a car dealer and B corresponds to a potential car buyer. A represents the negotiation partner and B represents the second negotiation partner. Conduct negotiations where: ={CD player, additional speakers, air conditioning system, paint quality, tires and wheels, navigation and multimedia system}. Each topic Have a set of options (e.g. {"good", "fairly good", "standard", "meager", "none"}) and for each option under the issue Specify a normalized utility value These utility values ​​constitute the utility function configuration of each agent. The distribution of these utility values ​​can be discrete or continuous, and the issues themselves can be numerical or non-numerical.

[0097] The preference of each agent is represented by its utility function, which is denoted as and , each bid is represented by ,in Indicates in The option values ​​on the topic, is the set of all possible bids. The preference of each agent is determined by its utility function It is defined as: ,in, , For negotiators Let's talk about the topic The importance weight of Is a negotiator On the topic Next, select all the options preference value.

[0098] For example, a specific bid It can be expressed as: = ("fairly good", "standard", "none", "fairly good", "meager", "good"), the bid is a combination of options on all issues. Each Each topic can find a unique correspondence in the result space. Have a finite set of options, denoted as , represents all possible option values ​​for the issue. Each issue and all possible bid combinations constitute the result space: , the result space is a six-dimensional discrete space containing 15625 negotiation solutions. For example, a bid The utility of Party A is 0.81, and the utility of Party B is 0.42, indicating that it is relatively favorable to Party A but not very attractive to Party B. Although agents tend to choose bids that have a higher utility for themselves, reaching an agreement often depends on both parties finding a mutually acceptable solution in the utility space. In the negotiation scenario of this application, the two agents know each other's utility function structure before the negotiation begins, but the reservation value (the minimum utility threshold they are willing to accept) is not disclosed.

[0099] (2) Generate game tree structure and nodes: To model the timing and strategy changes in the negotiation process, this application represents the negotiation process as an extensive-form game tree. A "game tree" is a tree-shaped decision diagram used to model the negotiation process, which is used to describe the bid and response paths of the intelligent agent in multiple rounds of negotiations. The tree is rooted at the initial state and expands hierarchically along the time sequence until the negotiation ends. Each node represents a negotiation state, and each edge represents an action or bid. The path from the root node to any node constitutes a historical path.

[0100] Historical Path Represents the action sequence from the root node of the game tree to a specific node, which reflects all the operations that have occurred during the negotiation process (such as bidding, acceptance, rejection, etc.). The set of all possible historical sequences is recorded as , each history path is a state path, which consists of several bidding nodes and response nodes. To model the negotiation evolution process, let Indicates the number of steps in the current negotiation, that is, the historical path length. Figure 1 The structure of the game tree is shown in the figure. Diamond nodes represent response (accept / leave) nodes, circular nodes represent nodes that reject the current bid (then propose a new bid instead of ending the negotiation), square nodes are corresponding nodes that propose a new bid, and terminal history collection The corresponding leaf node in the game tree indicates the final result of the negotiation. There are three conditions for the terminal state: reaching an agreement (such as Party A accepting Party B's bid); ending the negotiation (such as one party voluntarily withdrawing); timeout (when the number of steps is greater than the number of steps). The process ends automatically when the preset maximum number of rounds T is reached and no agreement is reached.

[0101] (3) In any non-terminal history path Afterwards, possible actions The agent represents the next action that the agent can take in this state, including accepting the current bid (accept), making a counter-proposal (reject and propose, i.e., rejecting the current bid), or terminating the negotiation (i.e., leaving the negotiation). The agent performs operations alternately, dynamically generating a set of actions that the current negotiator can perform based on the current game tree node and historical path. When the agent chooses to make a counter-proposal, the system will create a new bid generation node under the current node, and the agent will select the next action from the result space. Select a new bid ,For example = (standard, fairly good, none, fairly good, meager, good). The system will calculate the utility value of the bid under the utility function of both agents and , and use this value as the numerical identifier of the node on the game tree.

[0102] Assume that the current history path is: , then the operation options of party A are Definition. If Party A chooses to make a new bid , the system will perform the following data processing steps:

[0103] a) Constructing a new historical path ;

[0104] b) Store the path and corresponding bid into the game tree node set middle;

[0105] c) Calculate the bid utility value under the utility function of both parties and , and is used for acceptance evaluation in subsequent strategy modules.

[0106] (4) Bidding and acceptance strategies: In the SMN model, the negotiation strategy of each agent consists of two key parts: bidding strategy and acceptance strategy. The bidding strategy is used to determine the current historical path when the agent has the right to bid. , result space , opponent behavior and own utility function, and select a new bid The acceptance strategy is used to decide whether the agent accepts the bid, makes a counter-bid, or ends the negotiation.

[0107] The agent's strategy is denoted by ,in The two parties involved in the negotiation, Indicates the number of steps in the current negotiation. In each round of negotiation, the strategy Is a decision function that receives the current history path As input, output current action In the game tree, the bidding strategy controls the paths branching out from the bidding node ( Figure 3 The acceptance policy controls the path from the response node to the terminal or to continue negotiation ( Figure 3 Indicated by the solid line in the middle). Taking the actual data processing process as an example, assuming that the current historical path is: , in this state, the agent Its strategy function can be called Returns one of the following action results: Accept the current bid , move the negotiation to the successful termination point; reject and make a new bid , continue to generate the next round of bidding nodes; terminate the negotiation and generate the termination node of negotiation failure.

[0108] The effectiveness of a strategy can be evaluated by its expected utility. Representing an agent In the Step 1 Strategy Down to the historical path The probability of . According to the strategy The expected utility calculation formula is:

[0109] ;

[0110] in Indicates the terminal state, Indicates terminal status The utility obtained below.

[0111] These configurations provide a dynamic and comprehensive framework for modeling negotiation processes. By combining sequential strategies and utility-based decision making, the SMN model captures the complexity and interactivity of multi-issue negotiations to more effectively analyze the behavior of intelligent agents.

[0112] Under the above SMN model framework, in each round of negotiation, the agent needs to select from a large-scale result space. However, due to the high-dimensional discreteness of the result space and the nonlinear structure of the utility function, directly traversing or evaluating all bids is not only computationally expensive but may also lead to unstable agent strategies. To address this problem, this application innovatively constructs a bid abstraction technology based on utility segments during the strategy solving process. It clusters similar bids into different zones based on utility, thereby simplifying the decision-making process, effectively reducing computational complexity, improving the scalability of complex negotiation problems, and enabling the agent to focus on areas in the negotiation space that are negotiable.

[0113] This method consists of three core calculation processes: (1) Regret value calculation: measuring the profit gap between the agent's actual choice and the possible optimal action in the historical game tree branch; (2) Bid abstraction: dividing the high-dimensional result space into several segments based on utility similarity as the basic operation unit of the bidding strategy; (3) Segment selection probability calculation: based on historical regret feedback, dynamically adjust the sampling probability of each bidding segment to guide the strategy selection to tend to the potentially better area. These three processes together constitute the bidding strategy function The strategy scheduling and update basis is the key support for the "strategy evolution and bid generation" in the SMN framework of this application. In the overall data flow, this module receives input as the current historical path , result space , and completes the following steps internally: regret value calculation, utility segmentation, and segment selection probability generation, and finally outputs a probabilistic strategy preference result about the segment as the input basis for the subsequent bid generation module.

[0114] Figure 4 This is a schematic diagram of solving the regret value in the negotiation result space provided by the embodiment of the present application. Specifically, in a negotiation with multiple rounds of bidding, the utility obtained in each round of bidding may not be optimal. The difference between the strategy of the first step and the possible optimal strategy is introduced. The regret value reflects: The gap between the actual utility obtained under the premise of keeping the opponent's utility unchanged or improving, and the maximum possible utility that we can obtain.

[0115] In the result space, each bid Each corresponds to a pair of utility values In this application, it is defined that Agent A is responsible for any bid The regret value is: ,in, The bid corresponding to the best utility among all bids in the outcome space that are equivalent to the current bid in terms of opponent utility is the optimal bid. It represents the optimal utility point on the optimal frontier of the current target along the opponent's utility contour. This frontier gradually transitions from the Pareto frontier to the Nash negotiation solution region as the negotiation progresses.

[0116] In the "Car Sales Configuration Negotiation" scenario, assume that Party A's current bid is , its effect is , and the utility of opponent B is At this time, if a bid with higher utility is found on the Pareto frontier along the contour line where B’s utility level remains unchanged, , making and , then the regret value of the current bid is: .

[0117] This calculation process does not rely on the enumeration of the strategy space, but is directly based on the projection search in the result space. It has strong computability and immediate feedback, making it suitable as a basic indicator for strategy optimization after each bidding round.

[0118] In order to reduce the computational burden of bid selection, this application adopts bid abstraction technology to convert the entire result space According to the agent's own utility function and Divided into different sections Each segment represents a range of bids with similar utility values, and each segment contains a set of bid combinations with similar utility values. Such segments enable the agent to operate within the computationally feasible range while maintaining strategic flexibility and simplifying decision-making by processing a group of bids instead of a single bid.

[0119] For example, if 10 equal-width segments are constructed on the utility interval [0,1], then a bid If its utility is 0.74, it is classified into segment 8 ,in, Represents the segment obtained by division The set of segments formed.

[0120] Under this structure, the negotiation Segment-level regret value of the step is defined as the average of all bid regrets within the segment. , the calculation formula of regret value is:

[0121] ;

[0122] here, Is a bid The expected utility depends on and The expected utility is defined as:

[0123] ;

[0124] in, Indicates when bidding The maximum utility that agent A can achieve when adjusting horizontally in the outcome space while keeping the opponent's utility constant until it intersects the Pareto front. Corresponding agent Utility in Nash negotiation solutions. It is a time-dependent weighted function that changes as the negotiation progresses. In the early stages, it emphasizes Pareto optimality, while as stability becomes the goal, it gradually shifts the focus to the Nash negotiation solution.

[0125] By averaging the regret values ​​of all bids within each segment, the agent is able to capture the overall performance of the segment and identify segments where decision-making needs improvement.

[0126] The probability of selecting a segment is determined by the cumulative regret value of the segment. Segments with higher regret values ​​indicate that they have missed the opportunity to obtain higher utility, so they are considered more likely to be selected in subsequent steps. Based on the above regret value calculation, in the next step Select segment The probability of is defined as follows:

[0127] ;

[0128] in, Indicates the selected segment The probability of is the positive regret value of the segment (that is, the regret value is greater than zero), is the total number of segments. When the positive regret values ​​of all segments are zero, all segments are assigned the same selection probability to ensure that the agent continues to explore the entire negotiation space.

[0129] For example, if a segment Include bid , the corresponding regrets are 0.1, 0.2, and 0.15 respectively, then the regret of the segment is .

[0130] This technology allows segments with high regret values ​​to be assigned higher probabilities in subsequent rounds, thereby guiding the agent to "compensate" for historical decision-making losses in actual decision-making and achieve gradual optimization of the strategy.

[0131] Thus, by introducing bid abstraction and regret calculation techniques, this application effectively reduces the complexity of the high-dimensional negotiation space to a small number of utility segments, and continuously adjusts the strategic direction based on historical decision performance. Bid abstraction significantly reduces computational complexity by grouping similar bid utilities together, allowing the agent to focus on strategically important areas of the negotiation space. Combined with regret calculation, this method helps the agent learn from past decisions, identify areas for improvement, and effectively guide future segment selection, making the bidding strategy more optimal.

[0132] In the above process, through regret value calculation, bid abstraction, and segment selection probability mechanisms, strategic reduction and feedback optimization of the high-dimensional negotiation space are achieved, forming a probabilistic preference distribution for utility segments. However, this process only completes the construction of strategic preferences and does not yet generate actual executable bids. This application proposes a dynamic regret matching bidding strategy solution (DORM) method. This method provides a structured strategic decision-making approach in which the agent dynamically adjusts its behavior by combining regret updates sensitive to opponent behavior, regret matching, and a Monte Carlo sampling-based bidding strategy. Its goal is to improve the agent's decision-making ability over time, thereby achieving more effective negotiation outcomes. The overall data flow sequence is: historical bid trajectory analysis → segment regret value calculation → segment selection probability update → segment selection based on regret matching techniques → bid sampling within the segment → final bid output.

[0133] Specifically, the opponent's response data (bid changes) from the previous round is processed first, and based on the observed opponent concessions, the regret value of each utility segment is dynamically adjusted, and the probability of segment selection in the next round is updated. The DORM method continuously adjusts the agent's regret value by analyzing the opponent's behavior. The agent will evaluate the opponent's response (such as rejection without concessions, small concessions, or significant concessions) to update the regret value of different segments of the negotiation space. The opponent's concession degree is quantified by comparing the utility of its current bid with the previous round of bids, which is based on the opponent's utility function. In more complex negotiations, this utility can be estimated through opponent modeling or utility function prediction. Whenever the opponent makes a new bid , the agent is based on the opponent's current (in step) and its previous step The utility change at that time is used to determine whether the current concession is made. The calculation of the utility change is shown in the following formula:

[0134] ;

[0135] The types of concessions are categorized as follows:

[0136] ;

[0137] in, is a predefined threshold that distinguishes small concessions from significant concessions. This adaptive process allows the agent to optimize its strategy based on the observed opponent's behavior. For bids that fail to induce concessions or induce only small concessions, the agent increases the regret value of the corresponding segment, indicating that the negotiation progress is not ideal. On the contrary, significant concessions will reduce the regret value, indicating that the negotiation is moving towards an agreement. The regret value of the step and the actions of the negotiating parties are updated and calculated The regret value of the step, the regret update follows the following rules: ,in, Indicates that in step Time zone The updated regret value of is a time-varying factor that adjusts the magnitude of regret updates based on the opponent’s concessions. This mechanism ensures that the agent gradually focuses on the most promising areas in the negotiation space and optimizes its decision-making process as the negotiation progresses.

[0138] In each section After the regret value of is updated, its selection probability is calculated based on the accumulated regret value. The agent dynamically adjusts its strategy through a regret matching method, iteratively updating the strategy based on the regret value of each segment. The core idea is to select segments with higher regret values ​​to minimize future regret, thereby guiding the agent to make better decisions. Based on the regret calculation in the above formula, the selection probability of each segment when making the next negotiation decision is given by the following formula:

[0139] ;

[0140] This enables the strategy to dynamically adapt and optimize to the opponent's behavior. Represents a collection of segments A traversal of all bid segments in , used for normalization, Indicates in Step, section Regret value.

[0141] In the car sales configuration negotiation scenario, if the opponent rejects the agent's bid in the 6th round And become tougher, the system improves after analyzing the historical bidding trajectory Section If the opponent shows significant concessions, the regret of this segment will be reduced.

[0142] Furthermore, the segment sampling results obtained after the above processing are further processed based on the following steps. Weighted sampling is further performed within the selected segment according to the bid importance to generate the final actual bid , Representative bids, Represents the result space. It should be noted that Represents a bid that is not specified. Represents a specific bids. After the agent selects a segment based on regret, it generates bids from that segment using Importance Monte Carlo Sampling (IMCS). Within the selected segment, each bid is assigned an importance weight based on its potential utility to both parties, with preference given to bids closer to the Pareto frontier—the set of optimal agreements that balances the utilities of both parties.

[0143] bid Importance weight Determined by the following formula:

[0144] ;

[0145] in, Indicate a bid The horizontal distance to the Pareto frontier. This distance represents the distance when keeping the opponent's utility unchanged. The maximum achievable utility value corresponds to adjusting bids horizontally across the outcome space until they intersect the Pareto front. Bids closer to the Pareto front are considered more advantageous and therefore have higher weights.

[0146] Then, by normalizing the importance weights, The probability of selecting a bid is as follows:

[0147] ;

[0148] This approach ensures that bids closer to the Pareto front are more likely to be selected, thus promoting solutions that balance utility for both parties. However, bids with lower importance weights are still considered to maintain flexibility in bid selection.

[0149] Furthermore, to verify the effectiveness of the intelligent processing method for negotiation information provided in the embodiments of this application, experiments were conducted in five different negotiation fields and corresponding experimental results were provided. These experiments involved 10 agents from the 2024 Automated Negotiating Agents Competition (ANAC), aiming to verify the advantages of the dynamic regret matching agent (i.e., the agent corresponding to the DORM method) in terms of adaptability, utility performance, and convergence, and to compare it with existing negotiation agents. A total of 945 negotiations were conducted in the experiment, and indicators such as self-utility, opponent utility, social welfare, steps to agreement, and distance to Nash point were used to measure the performance of the agents.

[0150] This application conducts experiments across five main categories of multi-issue negotiation domains, each containing nine different negotiation domains, for a total of 45 domains. These experiments aim to evaluate the adaptability and utility of a Dynamic Regret Matching agent (DORM) under different distributional characteristics, as well as its ability to mitigate exploitation risk compared to the 10 agents used in ANAC 2024. Each agent competes against all other agents (including itself) in these 45 domains, completing a total of 5,445 negotiations.

[0151] Since the competition field in the current open source anl-agents package is relatively simple and it is difficult to fully evaluate the effectiveness of the agent, the embodiment of this application reproduces the competition environment of ANL 2024 and sets the negotiation deadline to 3000 steps to ensure the adequacy of the experiment. The experiment was conducted using the NegMAS and ANL platforms (both are Python-based automated negotiation research frameworks), and the operating environment is a server equipped with an Nvidia GeForce RTX 3090 GPU, an 80-core CPU, and 80 GiB of memory. The embodiment of this application evaluates the performance of the agent based on the following indicators: Self Utility, Opponent Utility, Social Welfare, Steps to Agreement, and Distance to Nash Point.

[0152] Negotiation Domain Setup: This experiment covers five main categories of multi-issue negotiation domains, each containing nine different negotiation scenarios, for a total of 45 negotiation domains. These categories include ADG, ItexVsCypress, Supermarket, Thompson, and Travel. These categories cover a large range of negotiation scenarios, and the domains within each category exhibit distinct result space distribution characteristics, ranging in size from approximately 180 to 188,160. These domains simulate realistic negotiation scenarios across multiple industries and environments, making them ideal for testing the adaptability and effectiveness of dynamic regret matching agents.

[0153] Opponents: The opponents in this experiment are the top 10 agents from the Automated Negotiation Agent Competition (ANL 2024). These agents include Shochan, UOAgent, AgentRenting2024, AntiAgent, HardChaosNegotiator, KosAgent, Nayesian2, CARCAgent, BidBot, and AgentNyan. These agents employ a variety of strategies. These agents are from the anl-agents Python package, which provides negotiation strategies from the ANL league.

[0154] In order to evaluate the performance of the dynamic regret matching agent, the key indicators in the ANAC competition are used in the embodiment of this application:

[0155] (1) Self Utility: The personal utility obtained by the dynamic regret matching agent in the negotiation, reflecting its ability to maximize its own benefits.

[0156] (2) Opponent Utility: The utility gained by the opponent in the negotiation, which measures whether the dynamic regret matching agent tends to negotiate in a mutually beneficial manner or adopt a highly exploitative strategy.

[0157] (3) Social Welfare: The sum of the utilities of both parties, reflecting the level of cooperation during the negotiation process.

[0158] (4) Steps to Agreement: The number of negotiation rounds required to reach an agreement, which measures negotiation efficiency.

[0159] (5) Distance to Nash Point: The degree of closeness between the final agreement and the Nash negotiation solution, reflecting the fairness and stability of the negotiation results.

[0160] The experimental results are shown in Table 3, and Figure 5 、 Figure 6 and Figure 7 As shown, Figure 5 This is a schematic diagram comparing the performance indicators of an intelligent agent corresponding to the method of the present application and other intelligent agents provided in an embodiment of the present application. Figure 6 This is another schematic diagram comparing the performance indicators of an intelligent agent corresponding to the method of the present application and other intelligent agents provided in an embodiment of the present application. Figure 7 This is another schematic diagram comparing the performance indicators of an agent corresponding to the method of the present application with other agents provided in an embodiment of the present application. Specifically, Figure 5 、 Figure 6 and Figure 7 The individual utility value, social welfare value and distance from the Nash negotiation solution of the dynamic regret matching agent based on the method of this application are shown respectively. Figure 5 、 Figure 6 and Figure 7 In , the corresponding indicator of any agent is the comprehensive value of multiple scenarios obtained after testing in multiple scenarios, for example, Figure 5 The individual utility value corresponding to the Zhongxiaojiang agent is the average of the individual utility values ​​in multiple scenarios. The same applies to the indicators corresponding to other agents, which will not be repeated here.

[0161] Table 3

[0162]

[0163] Specifically, Table 3 shows the performance of the dynamic regret matching agent on multiple indicators, including individual utility, opponent utility, social welfare, number of steps to reach an agreement, and distance to the Nash point. It should be noted that Figure 5 The upward arrow corresponding to the individual utility value in the middle represents that the higher the individual utility value, the better the corresponding effect; Figure 6 The upward arrow corresponding to the medium social welfare value means that the higher the social welfare value, the better the corresponding effect; Figure 7 The down arrow corresponding to the distance from the Nash negotiation solution indicates that the smaller the value, the better the corresponding effect. The up or down arrow corresponding to each indicator in Table 3 is similar and will not be repeated here.

[0164] This application example analyzes the results of each negotiation field in detail. The experimental results show that the dynamic regret matching agent performs well in multiple performance indicators:

[0165] (1) The Self Utility is the highest, at 0.890, indicating that the dynamic regret matching agent can effectively maximize its own utility in the negotiation.

[0166] (2) Social Welfare is the highest, at 1.494, indicating that the dynamic regret matching agent not only ensures its own benefits but also takes into account the utility of its opponent, thus achieving a high level of cooperation.

[0167] (3) It takes the fewest steps to reach an agreement, only 1853 steps. Compared with other high-performance agents, the dynamic regret matching agent can reach an agreement in a shorter number of negotiation rounds, thus improving negotiation efficiency.

[0168] (4) The distance to the Nash point is the smallest, only 0.276, indicating that the negotiation results of the dynamic regret matching agent are more stable and fair, reducing the risk of exploitation.

[0169] (5) The agreement ratio is as high as 0.999, indicating that the dynamic regret matching agent is able to reach an agreement in almost all negotiations, ensuring the consistency and effectiveness of the negotiations.

[0170] These experimental results fully demonstrate the effectiveness of the dynamic regret matching agent in maximizing individual utility, improving overall benefits, and optimizing negotiation efficiency, while maintaining stability and fairness, making it a powerful method in the field of automated negotiation. This demonstrates that this application solution is beneficial in improving the efficiency of negotiation information processing, thereby enhancing the efficiency of intelligent negotiation and improving the quality of negotiation outcomes.

[0171] Specifically, this application proposes an intelligent negotiation information processing method for dynamic bid regret matching in multi-issue bilateral negotiation scenarios. This method aims to address existing automated negotiation technologies' challenges, such as the massive size of the multi-issue outcome space, incomplete information, and the complexity of real-time decision-making. Based on the Strategic Multi-Issue Negotiation (SMN) model, this application combines bid abstraction technology with a dynamic bid regret matching method to optimize the negotiation process and achieve efficient, fair, and stable negotiation outcomes. The SMN model, using a generalized extended game framework, provides structured modeling for multi-issue negotiations, enabling intelligent agents to systematically reason about and optimize negotiation strategies. The bid abstraction technology divides the negotiation space into several important segments based on individual utility, reducing the complexity of the outcome space and improving strategy computation efficiency. The dynamic bid regret matching method dynamically updates the regret value of historical bids within each segment, enabling intelligent agents to adjust their bid strategies in real time, ensuring adaptability and stability in negotiations. Furthermore, to further optimize bid quality, this application employs Importance Monte Carlo sampling to generate bids that approach the Pareto frontier and Nash negotiation solutions, thereby improving negotiation efficiency and fairness.

[0172] This technology was extensively experimented with 10 adversary agents across various negotiation domains. The method corresponding to the dynamic regret matching agent not only achieved the highest individual utility (0.890), but also consistently reached agreements closest to the Nash bargaining solution (0.276). Experimental results demonstrate that this proposed method can achieve near-Nash bargaining solutions in a variety of negotiation scenarios, significantly outperforming existing techniques in terms of improving average individual utility and speed of agreement. This approach enables the agent to adaptively adjust its negotiation strategy, ensuring efficient, stable, and fair negotiation outcomes. This proposed method supports automated multi-issue negotiation.

[0173] It should be noted that in the embodiments of this application, regret value calculation is based on dynamic updates to optimize the bidding strategy, but this is not a specific limitation. In actual applications, other alternative regret calculation methods can be used, such as introducing a regret estimation mechanism based on adversarial learning or reinforcement learning, and optimizing the strategy through model training, rather than directly adjusting it dynamically based on historical regret values.

[0174] In one application scenario, a negotiation algorithm based on multi-agent reinforcement learning (MARL) can also be introduced to gradually optimize the agent strategy through collaborative learning during the negotiation process without relying on regret value.

[0175] like Figure 8 As shown in , corresponding to the above-mentioned negotiation information intelligent processing method, the embodiment of the present application further provides a negotiation information intelligent processing system, and the above-mentioned negotiation information intelligent processing system includes:

[0176] Data acquisition module 810, for acquiring negotiation protocols, negotiation domains, and preference configurations, negotiation strategies, and result spaces corresponding to negotiation partners;

[0177] A model building module 820 is configured to build a strategic multi-issue negotiation model based on the negotiation protocol, the negotiation domain, and the preference configurations, negotiation strategies, and outcome spaces corresponding to the negotiation partners, using an extended game tree modeling framework. The model comprises nodes representing negotiation states, edges representing negotiation actions of the negotiation partners, and negotiation actions of the negotiation partners including making a new bid, accepting a current bid, rejecting a current bid, and leaving the negotiation.

[0178] The negotiation result acquisition module 830 is configured to determine and output the negotiation result between the negotiation partners based on the leaf nodes in the strategic multi-issue negotiation model.

[0179] In this way, modeling is performed based on the extended game tree modeling framework to construct a strategic multi-issue negotiation model, rather than using the existing BOA framework. The strategic multi-issue negotiation model used in this application provides structured modeling for multi-issue negotiations through an extended game framework, enabling intelligent agents to systematically reason and optimize negotiation strategies. Based on game theory, a negotiation framework is constructed that supports intelligent agents to dynamically adjust strategies in an incomplete information environment, thereby not relying too much on clear opponent modeling, which is conducive to improving the efficiency of negotiation information processing in an incomplete information environment, thereby improving negotiation efficiency.

[0180] It should be noted that the specific structure and implementation of the above-mentioned negotiation information intelligent processing system and its various modules or units can refer to the corresponding description in the above-mentioned method embodiment, and will not be repeated here.

[0181] It should be noted that the division method of the various modules of the above-mentioned negotiation information intelligent processing system is not unique and is not used as a specific limitation here.

[0182] Based on the above embodiment, the present application also provides a smart terminal, whose principle block diagram can be as follows: Figure 9 As shown. The above-mentioned intelligent terminal includes a processor, a memory, a network interface and a display screen connected through a system bus. Among them, the processor of the intelligent terminal is used to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the intelligent terminal is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the steps of any one of the above-mentioned negotiation information intelligent processing methods are implemented. The display screen of the intelligent terminal can be a liquid crystal display or an electronic ink display.

[0183] Those skilled in the art will understand that Figure 9 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the solution of the present application, and does not constitute a limitation on the smart terminal to which the solution of the present application is applied. The specific smart terminal may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0184] In one embodiment, a smart terminal is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of any one of the negotiation information intelligent processing methods provided in the embodiments of the present application are implemented.

[0185] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the negotiation information intelligent processing methods provided in the embodiments of the present application are implemented.

[0186] It should be understood that the serial numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0187] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the above-mentioned device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0188] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0189] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0190] In the embodiments provided herein, it should be understood that the disclosed systems / terminal devices and methods can be implemented in other ways. For example, the system / terminal device embodiments described above are merely illustrative. For example, the division of the modules or units described above is merely a logical functional division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another system, or omitting or not implementing certain features.

[0191] If the above-mentioned integrated modules / units are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the present application can also implement all or part of the process steps in the above-mentioned method embodiments by using a computer program to instruct the relevant hardware. The above-mentioned computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The above-mentioned computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The above-mentioned computer-readable medium can include: any entity or device capable of carrying the above-mentioned computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, and software distribution medium. It should be noted that the content contained in the above-mentioned computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction.

[0192] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for intelligent processing of negotiation information, characterized in that: The method comprises: Obtain negotiation protocols, negotiation domains, and the corresponding preference configurations, negotiation strategies, and outcome spaces of negotiation partners; Based on the negotiation protocol, the negotiation domain, and the preference configurations, negotiation strategies, and outcome spaces corresponding to the negotiation partners, a model is constructed based on an extended game tree modeling framework to construct a strategic multi-issue negotiation model, wherein nodes in the strategic multi-issue negotiation model represent negotiation states, and edges in the strategic multi-issue negotiation model represent negotiation actions of the negotiation partners, wherein the negotiation actions of the negotiation partners include proposing a new bid, accepting a current bid, rejecting a current bid, and leaving the negotiation; Determine and output the negotiation results between the negotiation parties according to the leaf nodes in the strategic multi-issue negotiation model; Wherein, the negotiation partner includes a first negotiation partner and a second negotiation partner; The method of modeling based on the extended game tree modeling framework according to the negotiation agreement, the negotiation domain, and the preference configuration corresponding to the negotiation object to construct a strategic multi-issue negotiation model includes: Obtaining an initial state, and generating a root node of the expanded game tree according to the initial state; Using the root node as the negotiation status node corresponding to the first negotiation object; determining, based on the negotiation state node corresponding to the first negotiation partner, and the preference configuration, negotiation strategy, and result space corresponding to the first negotiation partner, a first negotiation action of the first negotiation partner based on the negotiation state node, and generating a negotiation state node corresponding to the first negotiation action; If the first negotiation action is to make a new bid, determining a second negotiation action of the second negotiation partner based on the negotiation state node according to the negotiation state node corresponding to the first negotiation action and the preference configuration, negotiation strategy, and result space corresponding to the second negotiation partner, and generating a negotiation state node corresponding to the second negotiation action; If the second negotiation action is to reject the current bid, determining a third negotiation action of the second negotiation partner based on the negotiation state node according to the negotiation state node corresponding to the second negotiation action, and the preference configuration, negotiation strategy, and result space corresponding to the second negotiation partner, and generating a negotiation state node corresponding to the third negotiation action; If the third negotiation action is to make a new bid, the negotiation state node corresponding to the third negotiation action is used as the negotiation state node corresponding to the first negotiation partner, and the steps of determining the first negotiation action of the first negotiation partner based on the negotiation state node and the preference configuration, negotiation strategy, and result space corresponding to the first negotiation partner and generating the negotiation state node corresponding to the first negotiation action are returned to execution until a preset generation termination condition is reached, thereby obtaining a strategic multi-issue negotiation model.

2. The intelligent processing method for negotiation information according to claim 1, characterized in that: The negotiation agreement includes an alternating offer agreement.

3. The intelligent processing method for negotiation information according to claim 2, characterized in that: The preference configuration corresponding to the negotiation partner includes a linear additive utility function model for characterizing the preference of the negotiation partner.

4. The intelligent processing method for negotiation information according to claim 1, characterized in that: The preset generation termination conditions include: Any negotiation action obtained is accepting the current bid or leaving the negotiation, or the processing time reaches a preset time threshold.

5. The intelligent processing method for negotiation information according to claim 1 or 4, characterized in that: The negotiation strategy includes an acceptance strategy and a bidding strategy.

6. The intelligent processing method for negotiation information according to claim 5, characterized in that: If the second negotiation action is to reject the current bid, determining a third negotiation action of the second negotiation partner based on the negotiation state node according to the negotiation state node corresponding to the second negotiation action and the preference configuration, negotiation strategy, and result space corresponding to the second negotiation partner, and generating a negotiation state node corresponding to the third negotiation action, including: If the second negotiation action is to reject the current bid, determining, based on the negotiation state node corresponding to the second negotiation action and the preference configuration, bidding strategy, and result space corresponding to the second negotiation partner, that a third negotiation action for the second negotiation partner based on the negotiation state node is to make a new bid; Obtaining a regret value based on the negotiation state node corresponding to the second negotiation action, and the historical branches and historical nodes corresponding to the negotiation state node corresponding to the second negotiation action, wherein the regret value is used to represent the benefit gap between the actual choice of the negotiation partner and the optimal action; Dividing the result space corresponding to the second negotiation partner into segments to obtain a plurality of bidding segments; Determining the segment selection probability corresponding to each bidding segment according to the regret value corresponding to each bidding segment; According to the segment selection probability, and the preference configuration, bidding strategy and result space corresponding to the second negotiation partner, a bid corresponding to the second negotiation partner is obtained to generate a negotiation state node corresponding to the third negotiation action.

7. A negotiation information intelligent processing system, characterized in that: The system comprises: The data acquisition module is used to obtain the negotiation agreement, negotiation domain, and the preference configuration, negotiation strategy and result space corresponding to the negotiation object; a model building module for modeling based on the extended game tree modeling framework according to the negotiation protocol, the negotiation domain, and the preference configurations, negotiation strategies, and outcome spaces corresponding to the negotiation partners, so as to construct a strategic multi-issue negotiation model, wherein nodes in the strategic multi-issue negotiation model represent negotiation states, and edges in the strategic multi-issue negotiation model represent negotiation actions of the negotiation partners, wherein the negotiation actions of the negotiation partners include making a new bid, accepting a current bid, rejecting a current bid, and leaving the negotiation; A negotiation result acquisition module, configured to determine and output the negotiation result between the negotiation partners based on the leaf nodes in the strategic multi-issue negotiation model; Wherein, the negotiation partner includes a first negotiation partner and a second negotiation partner; The model building module is specifically used to: obtain an initial state, and generate a root node of the extended game tree according to the initial state; Using the root node as the negotiation status node corresponding to the first negotiation object; determining, based on the negotiation state node corresponding to the first negotiation partner, and the preference configuration, negotiation strategy, and result space corresponding to the first negotiation partner, a first negotiation action of the first negotiation partner based on the negotiation state node, and generating a negotiation state node corresponding to the first negotiation action; If the first negotiation action is to make a new bid, determining a second negotiation action of the second negotiation partner based on the negotiation state node according to the negotiation state node corresponding to the first negotiation action and the preference configuration, negotiation strategy, and result space corresponding to the second negotiation partner, and generating a negotiation state node corresponding to the second negotiation action; If the second negotiation action is to reject the current bid, determining a third negotiation action of the second negotiation partner based on the negotiation state node according to the negotiation state node corresponding to the second negotiation action, and the preference configuration, negotiation strategy, and result space corresponding to the second negotiation partner, and generating a negotiation state node corresponding to the third negotiation action; If the third negotiation action is to make a new bid, the negotiation state node corresponding to the third negotiation action is used as the negotiation state node corresponding to the first negotiation partner, and the steps of determining the first negotiation action of the first negotiation partner based on the negotiation state node and the preference configuration, negotiation strategy, and result space corresponding to the first negotiation partner and generating the negotiation state node corresponding to the first negotiation action are returned to execution until a preset generation termination condition is reached, thereby obtaining a strategic multi-issue negotiation model.

8. An intelligent terminal, characterized in that: The intelligent terminal includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of the negotiation information intelligent processing method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the negotiation information intelligent processing method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Supply chain material data management method and data management system

    CN119721910A

  • Automated negotiation agent with opponent's behavior prediction

    US20220172264A1