Decision negotiation method and device based on depth deterministic policy gradient

By employing a decision negotiation method based on deep deterministic policy gradients, and utilizing Actor-Critic networks and utility models to optimize negotiation strategies, this approach addresses the issue of proposing reasonable negotiation values ​​in negotiation scenarios, thereby improving negotiation efficiency and the rationality of the results.

CN120875627APending Publication Date: 2025-10-31厦门工学院

Patent Information

Application Number
CN202511386873.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In negotiation scenarios, proposing a reasonable negotiation value is an urgent problem to be solved.

Method used

A decision negotiation method based on deep deterministic policy gradient is adopted. By receiving negotiation scenario information, a utility model is created, and a pre-trained Actor-Critic network is used to predict and update negotiation values. The negotiation strategy is optimized by combining trapezoidal fuzzy membership function and overall satisfaction function.

Benefits of technology

It enables the proposal of reasonable negotiation values ​​in negotiation scenarios, improves negotiation efficiency and the reasonableness of the results, and helps both parties reach a win-win solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875627A_ABST
    Figure CN120875627A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a decision negotiation method and device based on a depth deterministic policy gradient. The method comprises the following steps: receiving negotiation scene information input by a first negotiation party; wherein the negotiation scene information comprises a plurality of negotiation topics, a preference value interval for each negotiation topics and a preference weight value for each negotiation topics; for each negotiation topic, creating a invention for the negotiation topic according to the preference value interval and the preference weight value of the negotiation topic; wherein the utility model comprises corresponding trapezoidal fuzzy membership functions when negotiation values are in different preference value intervals, and an overall satisfaction function for performing weighted summation on the trapezoidal fuzzy membership functions of all negotiation dimensions; a first reference value of the first negotiation party is predicted through a pre-trained Actor-Critic network; and receiving a first negotiation value input by the first negotiation party according to the first reference value, and sending the first negotiation value to a second negotiation party, thereby facilitating determination of the negotiation value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information technology, and in particular to a decision negotiation method and apparatus based on deep deterministic strategy gradient. Background Technology

[0002] In negotiation scenarios, the two negotiating parties typically take turns sending negotiation values, proposing their own offers. After receiving the corresponding negotiation value, one party can update its own negotiation value by referring to the other party's. However, how to propose a reasonable negotiation value in negotiation scenarios becomes a pressing issue that needs to be addressed. Summary of the Invention

[0003] The purpose of this application is to provide a decision negotiation method and apparatus based on deep deterministic policy gradients, to solve the problem of how to propose a reasonable negotiation value in a negotiation scenario. The specific technical solution is as follows: A first aspect of this application provides a decision negotiation method based on deep deterministic policy gradient, applied to a first negotiating party in a negotiation scenario, the method comprising: The system receives negotiation scenario information input by the first negotiating party; wherein the negotiation scenario information includes: multiple negotiation topics, preference value ranges for each negotiation topic, and preference weight values ​​for each negotiation topic. For each negotiation topic, a utility model is created based on the preference value range and preference weight value of the negotiation topic; wherein, the utility model includes: trapezoidal fuzzy membership functions corresponding to the negotiation values ​​when they are in different preference value ranges, and an overall satisfaction function that is a weighted sum of the trapezoidal fuzzy membership functions of each negotiation dimension. A first reference value for the first negotiating party is predicted through a pre-trained Actor-Critic network. The first reference value is obtained by using the utility model to predict the predicted value through the Actor network in the pre-trained Actor-Critic network, calculating the reward value of the predicted value through the Critic network, and iteratively updating the value based on the reward value through the Actor network. The first negotiation value, input by the first negotiating party with reference to the first reference value, is received and sent to the second negotiating party.

[0004] In one possible implementation, after receiving the first negotiated value input by the first negotiating party with reference to the first reference value and sending it to the second negotiating party, the method further includes: Receive a second negotiation value fed back by the second negotiating party after receiving the first negotiation value; calculate a value estimate corresponding to the second negotiation value; wherein the value estimate is used to characterize the expected return corresponding to the second negotiation value; If the first negotiating party rejects the second negotiated value, a second reference value is calculated using the pre-trained Actor-Critic network.

[0005] In one possible implementation, the utility model includes: the trapezoidal fuzzy membership function and the overall satisfaction function; The trapezoidal fuzzy membership function is:

[0006] Where a, b, c, and d are the values ​​corresponding to the preference value ranges of each negotiation topic input by the first negotiating party; k represents the normalization constant. The overall satisfaction function is:

[0007] Where n represents the total number of negotiation topics; w i Representing the The weight of each issue; f i (o i ) is a given number The trapezoidal fuzzy membership function for the negotiated value o of each issue; o represents the issue value vector that constitutes a specific negotiated value.

[0008] In one possible implementation, the step of calculating the reward value of the predicted value through a Critic network and iteratively updating it based on the reward value through an Actor network includes: Using the Critic network and the formula:

[0009] Calculate the return value of the predicted value, where r t Represents the immediate reward at time step t; γ is the discount factor; The reward function of the Critic network; s t+1 and a t+1 These represent the next state and the next action, respectively. Based on the reward value, the Actor network uses the formula:

[0010] Iterative updates are performed, among which, This represents the gradient of the policy function with respect to the Actor network parameters θ; Represents the strategy function; This represents the estimated value of the Critic network. This indicates the parameters of the Actor network. Find the gradient. This indicates that the gradient of action a is calculated.

[0011] In one possible implementation, calculating the value estimate corresponding to the second negotiated value includes: The behavior value corresponding to the second negotiated value is calculated using the preset receiving strategy network, activation function, and output function; the behavior value is normalized using a normalization function to obtain a normalization result; wherein, the normalization result includes whether the second negotiated value is received; and the value estimate is calculated using a value function to calculate the value of the second negotiated value and the behavior value.

[0012] In one possible implementation, the step of calculating a second reference value through the pre-trained Actor-Critic network if the first negotiating party rejects the second negotiated value includes: If the first negotiating party rejects the second negotiated value, the mean and standard deviation of each issue are calculated based on the second negotiated value and the utility value of each issue through a preset pricing strategy network; the corresponding normal distribution is identified based on the mean and standard deviation of each issue; the normal distribution is sampled to obtain the utility value of each issue; and the second reference value is generated based on the utility value of each issue through the pre-trained Actor-Critic network.

[0013] In one possible implementation, after receiving the second negotiation value fed back by the second negotiating party after receiving the first negotiation value, the method further includes: If the first negotiating party accepts the second negotiated value, then a preset reward function is applied:

[0014] Calculate the reward for concluding the negotiation; where, A reward indicating the conclusion of negotiations; Indicates the final negotiation round t f The issue weight vector at the time; U(o) represents the utility value of the final agreed solution o; T represents the deadline round defined before the start of negotiations; The weights of the preset receiving strategy network and / or the preset bidding strategy network are adjusted based on the reward for the completion of the negotiation.

[0015] A second aspect of this application provides a decision negotiation apparatus based on a deep deterministic policy gradient, applied to a first negotiating party in a negotiation scenario, the apparatus comprising: The scenario receiving module is used to receive negotiation scenario information input by the first negotiating party; wherein, the negotiation scenario information includes: multiple negotiation topics, preference value ranges for each negotiation topic, and preference weight values ​​for each negotiation topic; The satisfaction calculation module is used to create a utility model for each negotiation topic based on the preference value range and preference weight value of the negotiation topic. The utility model includes: trapezoidal fuzzy membership functions corresponding to the negotiation values ​​when they are in different preference value ranges, and an overall satisfaction function that is a weighted sum of the trapezoidal fuzzy membership functions of each negotiation dimension. The negotiation value prediction module is used to predict the first negotiation value of the first negotiating party through a pre-trained Actor-Critic network. The first negotiation value is obtained by using the utility model to predict the predicted value through the Actor network in the pre-trained Actor-Critic network, calculating the reward value of the predicted value through the Critic network, and iteratively updating the Actor network based on the reward value. The negotiation value sending module is used to receive the second negotiation value input by the first negotiating party with reference to the first negotiation value, and send it to the second negotiating party.

[0016] In one possible implementation, the device further includes: The value estimation calculation module is used to receive a second negotiation value fed back by the second negotiating party after receiving the first negotiation value; and to calculate a value estimate corresponding to the second negotiation value; wherein the value estimate is used to characterize the expected return corresponding to the second negotiation value; The second reference calculation module is used to calculate a second reference value through the pre-trained Actor-Critic network if the first negotiating party rejects the second negotiated value.

[0017] In one possible implementation, the utility model includes: the trapezoidal fuzzy membership function and the overall satisfaction function; The trapezoidal fuzzy membership function is:

[0018] Where a, b, c, and d are the values ​​corresponding to the preference value ranges of each negotiation topic input by the first negotiating party; k represents the normalization constant. The overall satisfaction function is:

[0019] Where n represents the total number of negotiation topics; w i Representing the The weight of each issue; f i (oi ) is a given number The trapezoidal fuzzy membership function for the negotiated value o of each issue; o represents the issue value vector that constitutes a specific negotiated value.

[0020] In one possible implementation, the negotiated value prediction module is specifically used to predict the value using the formula via the Critic network:

[0021] Calculate the return value of the predicted value, where r t Represents the immediate reward at time step t; γ is the discount factor; The reward function of the Critic network; s t+1 and a t+1 These represent the next state and the next action, respectively. Based on the reward value, the Actor network uses the formula:

[0022] Iterative updates are performed, among which, This represents the gradient of the policy function with respect to the Actor network parameters θ; Represents the strategy function; This represents the estimated value of the Critic network. This indicates the parameters of the Actor network. Find the gradient. This indicates that the gradient of action a is calculated.

[0023] In one possible implementation, the value estimation calculation module is specifically used to calculate the behavior value corresponding to the second negotiated value through the preset reception policy network using an activation function and an output function; normalize the behavior value using a normalization function to obtain a normalization result; wherein the normalization result includes whether the second negotiated value is received; and calculate the value estimate corresponding to the second negotiated value and the behavior value using a value function.

[0024] In one possible implementation, the second reference calculation module is specifically used to: if the first negotiating party rejects the second negotiated value, calculate the mean and standard deviation of each issue based on the second negotiated value and the utility value of each issue through a preset bidding strategy network; identify the corresponding normal distribution based on the mean and standard deviation of each issue; sample the normal distribution to obtain the utility value of each issue; and generate the second reference value based on the utility value of each issue through the pre-trained Actor-Critic network.

[0025] In one possible implementation, the device further includes: The weight adjustment module is used to, if the first negotiating party accepts the second negotiated value, apply a preset reward function:

[0026] Calculate the reward for concluding the negotiation; where, A reward indicating the conclusion of negotiations; Indicates the final negotiation round t f The issue weight vector at the time; U(o) represents the utility value of the final agreed solution o; T represents the deadline round defined before the start of negotiations; The weights of the preset receiving strategy network and / or the preset bidding strategy network are adjusted based on the reward for the completion of the negotiation.

[0027] Another aspect of the embodiments of this application also provides an electronic device, including: Memory, used to store computer programs; When the processor executes the program stored in memory, it implements any of the above decision negotiation methods based on deep deterministic policy gradients.

[0028] In another aspect of the embodiments of this application, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, it implements any of the above-described decision negotiation methods based on deep deterministic policy gradients.

[0029] In another aspect of the embodiments of this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the above-described decision negotiation methods based on deep deterministic policy gradients.

[0030] Beneficial effects of the embodiments in this application: This application provides a decision negotiation method and apparatus based on deep deterministic policy gradient. The method includes: receiving negotiation scenario information input by a first negotiating party; wherein the negotiation scenario information includes: multiple negotiation topics, preference value ranges for each negotiation topic, and preference weight values ​​for each negotiation topic; for each negotiation topic, creating a utility model for the negotiation topic based on the preference value range and preference weight values; wherein the utility model includes: trapezoidal fuzzy membership functions corresponding to negotiation values ​​in different preference value ranges, and an overall satisfaction function that is a weighted sum of the trapezoidal fuzzy membership functions for each negotiation dimension; predicting a first reference value for the first negotiating party through a pre-trained Actor-Critic network, wherein the first reference value is obtained by predicting the predicted value using the utility model through the Actor network in the pre-trained Actor-Critic network, calculating the reward value of the predicted value through the Critic network, and iteratively updating the value through the Actor network based on the reward value; receiving a first negotiation value input by the first negotiating party with reference to the first reference value, and sending it to a second negotiating party. The solution proposed in this application allows for receiving negotiation scenario information at the outset of negotiations. Then, based on this information, a utility model is created for each negotiation topic according to its preference value range and preference weight. A Critic network calculates the predicted return value, and an Actor network iteratively updates this return value to obtain a reference value. This facilitates the first negotiating party in proposing a reasonable negotiation value within the negotiation scenario by referring to the first reference value and inputting a first negotiation value, thus simplifying the determination of the negotiation value.

[0031] Of course, implementing any product or method of this application does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other embodiments can be obtained based on these drawings.

[0033] Figure 1 A flowchart illustrating a decision negotiation method based on deep deterministic policy gradient provided in an embodiment of this application; Figure 2 A flowchart illustrating the calculation of a second reference value provided in an embodiment of this application; Figure 3A schematic diagram of a trapezoidal fuzzy membership function provided in an embodiment of this application; Figure 4 A schematic diagram of a decision negotiation device based on deep deterministic policy gradient provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0034] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of this application.

[0035] In a first aspect, this application provides a decision negotiation method based on deep deterministic policy gradients, applied to a first negotiating party in a negotiation scenario. (See also...) Figure 1 , Figure 1 A flowchart illustrating a decision negotiation method based on deep deterministic policy gradient provided in this application embodiment, the method comprising: Step S11: Receive negotiation scenario information input by the first negotiating party; wherein, the negotiation scenario information includes: multiple negotiation topics, preference value ranges for each negotiation topic, and preference weight values ​​for each negotiation topic. Step S12: For each negotiation topic, create a utility model for the negotiation topic based on the preference value range and preference weight value of the negotiation topic; wherein, the utility model includes: trapezoidal fuzzy membership functions corresponding to the negotiation values ​​when they are in different preference value ranges, and an overall satisfaction function that is a weighted sum of the trapezoidal fuzzy membership functions of each negotiation dimension. Step S13: Predict the first reference value of the first negotiating party using a pre-trained Actor-Critic network (an algorithmic framework combining policy optimization and value assessment). The first reference value is obtained by using the utility model to predict the predicted value through the Actor network in the pre-trained Actor-Critic network, calculating the reward value of the predicted value through the Critic network, and iteratively updating the value based on the reward value through the Actor network. Step S14: Receive the first negotiation value input by the first negotiating party with reference to the first reference value, and send it to the second negotiating party.

[0036] Corresponding to step S11 above, in this embodiment of the application, the negotiation scenario includes a first negotiating party and a second negotiating party. A specific negotiation environment is defined. This negotiation scenario includes two agents: a doctor agent (DA) and a patient agent (PA). When receiving negotiation scenario information input by the first negotiating party, the negotiation scenario information includes: multiple negotiation topics, preference value ranges for each negotiation topic, and preference weight values ​​for each negotiation topic. Specifically, before negotiation begins, both parties need to reach a consensus on a negotiation agreement, which clarifies the effective actions each party can take in any negotiation state. In one example, the widely used protocol framework in multi-topic automated negotiation—AOP (Alternating Offer) protocol—can be adopted. According to this protocol, the doctor and patient take turns making offers and concessions until a consensus is reached or time runs out.

[0037] Corresponding to step S12 above, in this embodiment of the application, for each negotiation topic, a utility model is created based on the preference value range and preference weight value input by the first negotiating party for that negotiation topic. Specifically, the utility model includes: trapezoidal fuzzy membership functions corresponding to negotiation values ​​falling within different preference value ranges, and an overall satisfaction function that is a weighted sum of the trapezoidal fuzzy membership functions for each negotiation dimension. Specifically, when setting up the negotiation environment, each Agent negotiates in an SDM (Science, Technology, Management, and Application) environment, which includes an issue domain. This issue domain includes... The negotiation topics are distinct and interdependent, each representing an important aspect of the decision-making process. For example, in doctor-patient negotiations, these could include treatment costs, efficacy, side effects, and risks. Each topic is considered a discrete, finite set containing several possible discrete or continuous values. In this scheme, the value of each topic is set as a discrete value.

[0038] Corresponding to step S13 above, in this embodiment, the negotiation value of the second negotiating party, i.e., the first reference value, is predicted through a pre-trained Actor-Critic network. This pre-trained Actor-Critic network can be a model trained based on historical data. Specifically, the first reference value is obtained by using the Actor network within the pre-trained Actor-Critic network to predict the value using the utility model, then calculating the reward value of the predicted value through the Critic network, and iteratively updating the value based on the reward value through the Actor network. Specifically, the AutoSDM-DDPG (Automatic Negotiation Framework Based on Deep Deterministic Policy Gradient Algorithm) in this application consists of multiple modules specifically optimized for key aspects of SDM negotiation. These modules cover multiple core elements such as negotiation environment setup, decision-making strategy formulation, bidding strategy planning, and acceptance strategy adjustment. In this framework, the SDM Agent, as the core component of the negotiation process, learns and optimizes strategies in the SDM scenario through DRL (Deep Reinforcement Learning). The Agent's main responsibility is to dynamically select the optimal strategy based on the state information of the negotiation environment, maximizing the utility of both parties and reaching a consensus. Specifically, the SDM (Sequential Deep Matching Model) agent consists of two main parts: an acceptance strategy and a bid strategy. Throughout the negotiation process, the agent uses the DDPG (Deep Deterministic Policy Gradient, a deep reinforcement learning algorithm based on the Actor-Critic framework) algorithm in DRL to continuously optimize its decision-making process.

[0039] Corresponding to step S14 above, when the first negotiating party inputs a first negotiation value based on the first reference value, this first negotiation value can be the first reference value or it can be different from the first reference value. In one example, the first reference value may include multiple specific values, and then the first negotiating party selects one of them as the first negotiation value and sends it to the second negotiating party. In one example, when applied to doctor-patient negotiation, the negotiation process in this embodiment may include: Initialization: The doctor agent and the patient agent define their own preferences according to their respective needs and assign weights to each issue. The weights can reflect the importance of each issue in the negotiation decision-making process. Price exchange: Both parties take turns proposing treatment plans (i.e., prices), and each plan includes specific values ​​for several issues (such as treatment costs, efficacy, side effects, etc.). Utility evaluation: Both parties evaluate the prices according to their own utility models, calculate the satisfaction with the price for each issue, and determine the total utility of all issues. Decision: If both parties accept the price, an agreement is reached. If one party rejects the price, the exchange continues until an agreement is reached or the time expires. Through this process, both parties can make reasonable trade-offs and concessions in a multi-issue negotiation environment, ultimately reaching a consensus and obtaining the best solution. In one example, negotiation values ​​are not sent indefinitely. If the sent negotiation value is accepted by the counterparty, the negotiation ends and no further negotiation values ​​need to be generated; if the counterparty's negotiation value is accepted by one's own side, the negotiation also ends. Whether or not acceptance is made is determined by the acceptance policy network.

[0040] As can be seen, the solution of this application can receive negotiation scenario information at the beginning of the negotiation, and then, based on the negotiation scenario information, create a utility model for each negotiation topic according to the preference value range and preference weight value of the negotiation topic. Then, the reward value of the predicted value is calculated by the Critic network, and the reference value is obtained by iteratively updating the reference value through the Actor network. This makes it easier for the first negotiating party to solve the problem of how to propose a reasonable negotiation value in the negotiation scenario by referring to the first negotiation value input by the first reference value, and facilitates the determination of the negotiation value.

[0041] In one possible implementation, see Figure 2 After receiving the first negotiated value input by the first negotiating party with reference to the first reference value and sending it to the second negotiating party, the method further includes: Step S21: Receive the second negotiation value fed back by the second negotiating party after receiving the first negotiation value; calculate the value estimate corresponding to the second negotiation value; wherein the value estimate is used to characterize the expected return corresponding to the second negotiation value; Step S22: If the first negotiating party rejects the second negotiated value, the second reference value is calculated through the pre-trained Actor-Critic network.

[0042] In this embodiment, each Agent first constructs negotiation domain knowledge based on its own objectives, including issue domain I and the possible value range of each issue. And the weight assigned to each issue. Building upon this, the Agent uses a fuzzy preference model to quantify and model the preferences and objectives on these issues, ensuring an accurate representation of the fuzzy and uncertain preferences of each negotiating party on different issues. Next, the Agent employs an Actor-Critic network architecture based on the DDPG algorithm to select and evaluate negotiation actions. In the first round of negotiation, the Agent selects a offer from the optimal offer set and sends it to the opponent. If the opponent accepts the offer, the negotiation ends; if rejected, the opponent Agent will make a counter-offer. At this point, the opponent's counter-offer... With time step Together, they constitute the environmental feedback state for the next round. Starting from the second round, the Agent, through the Actor network, determines the current negotiation state and its fuzzy preference utility model. Choosing the next action means accepting the competitor's offer. (Agreement reached), or a new offer B can be made and rejected. t+1 (Developing a new treatment plan). The Agent then feeds back the selected behavior to the environment. Next, the Critic network calculates the loss (or evaluation value) based on the behavior selected by the Actor network. The Critic's role is to evaluate the effectiveness of the Agent's behavior by calculating the difference between the utility of the current behavior and the expected utility, i.e., the reward value. The reward value loss is then used to guide the Actor network's adjustments, continuously improving its strategy to ensure the Agent makes more effective and beneficial decisions in subsequent actions. Through this iterative process, the Agent can continuously learn and adjust using the DDPG algorithm, gradually optimizing its bidding and acceptance strategies. Specifically, in each round of negotiation, the Agent calculates satisfaction based on the utility of the current offer and evaluates the satisfaction of each issue using fuzzy membership functions. These evaluations are used as the basis for the next round of decision-making, allowing each Agent to adjust its bidding or acceptance strategy based on the opponent's response. Ultimately, after multiple rounds of alternating bidding, the Agent is able to find the optimal balance among multiple issues, helping doctors and patients obtain a treatment plan acceptable to both parties, with a high sum of satisfaction and a low difference in satisfaction.

[0043] In one possible implementation, the utility model includes: the trapezoidal fuzzy membership function and the overall satisfaction function; The trapezoidal fuzzy membership function is: (1) Among them, the values of a, b, c, and d can be referred to Figure 3 the values corresponding to the preference value intervals of each negotiation issue input by the first negotiation party; k represents a normalization constant; The overall satisfaction function is: (2) Among them, n represents the total number of negotiation issues; w i represents the th issue; i (o i ) is the trapezoidal fuzzy membership function when the negotiation value o of the th issue is given; o represents the issue value vector constituting a specific negotiation value.

[0044] Specifically, in the embodiments of the present application, a utility model can be constructed. By constructing a utility model for the SDM multi-issue negotiation problem, the influence of each issue on the decision-making process of each negotiation party is quantified. To solve the ambiguity of the preferences of each negotiation party, this solution uses the fuzzy constraint theory to model the utility models of each negotiation party. By combining the fuzzy membership function, the uncertainty of negotiation preferences is effectively captured, enabling the utility model to more accurately reflect the actual needs of both negotiation parties. In this solution, the satisfaction of each issue value is represented by a trapezoidal fuzzy membership function. Finally, corresponding intervals are defined to describe the satisfaction of each negotiation party with factors such as treatment cost, efficacy, and side effects. The definition and shape of the trapezoidal membership function are shown in detail in formula (1) and Figure 3 . Here, the parameter defines the starting point, ending point, and core interval of the trapezoidal fuzzy membership function, which are used to determine the acceptance interval and transition range; represents the negotiation value of the specific issue under consideration, such as the value offer, time, etc. In one example, the interval represents the range with the highest Agent satisfaction; x d represents unacceptable values; or represents the gradual transition interval of satisfaction. The utility models of each negotiation party need to consider the influence of multiple factors. Therefore, a weighted sum is used to calculate the overall satisfaction of DA or PA with respect to the offer. Specifically, the overall satisfaction function is shown in formula (2). Among them, n represents the total number of negotiation issues; w i represents the th issue; i (o i ) is the trapezoidal fuzzy membership function when the negotiation value o of the th issue is given; o represents the issue value vector constituting a specific negotiation value.

[0045] In one possible implementation, the step of calculating the reward value of the predicted value through a Critic network and iteratively updating it based on the reward value through an Actor network includes: Using the Critic network and the formula: (3) Calculate the return value of the predicted value, where r t Represents the immediate reward at time step t; γ is the discount factor; The reward function of the Critic network; s t+1 and a t+1 These represent the next state and the next action, respectively. Based on the reward value, the Actor network uses the formula: (4) Iterative updates are performed, among which, This represents the gradient of the policy function with respect to the Actor network parameters θ; Represents the strategy function; This represents the estimated value of the Critic network. This indicates the parameters of the Actor network. Find the gradient. This indicates that the gradient of action a is calculated.

[0046] Specifically, in this application, pre-training is a key step in improving the learning efficiency and stability of the negotiation framework within the AutoSDM-DDPG framework, especially when dealing with complex multi-issue negotiation tasks. The main purpose of pre-training the Actor network is to provide a good initial strategy, reduce random exploration, and improve the learning efficiency and convergence stability of the AutoSDM-DDPG framework. By initializing the Actor network with pre-trained prior knowledge, the negotiation agent can adapt to the multi-issue negotiation environment more quickly, thereby achieving better performance during formal training and actual deployment. The data used for pre-training is generated through simulation based on the negotiation scenario and preference model established in this study. Specifically, doctor-patient negotiation instances are constructed and simulated by sampling preferences and issue weights within a clinically reasonable range to reflect the variability and diversity encountered in actual medical decision-making. This method can ensure that the real patterns are captured during the pre-training phase and prepare for subsequent model learning using real or more complex data. In this scheme, a pre-training method based on simulated data is adopted, aiming to provide an optimized starting point for the subsequent formal training phase. In this application, the first step of pre-training is to generate labeled data suitable for the Actor network, mainly by generating the training dataset through formula (5): (5) Here, Indicates the utility value of a specific issue; Indicates the utility value of the treatment plan; The preference value for each issue reflects the degree of preference of each negotiating party for different issues. In one example, the 20 sets of datasets generated by formula (5) will be used as training data in the pre-training stage of the Actor network. These data contain the preference values ​​and corresponding utility values ​​of different issues, which are used to train the network's decision-making ability in various scenarios. During the pre-training process, the Actor network learns through the backpropagation algorithm, gradually adjusts its parameters, and learns reasonable strategies from historical data. Specifically, the goal of the Actor network is to generate the optimal offer based on the current state and historical decisions, and optimize its strategy through the feedback provided by the Critic network. The Critic network is responsible for evaluating the value of the current offer, calculating the reward value, and providing feedback to the Actor network to guide it in adjusting its strategy, thereby maximizing long-term utility. The Critic network uses formula (3) to evaluate the reward value of the current state-behavior pair and update its parameters: In this model, γ t γ represents the immediate reward at time step t; γ is the discount factor. ; The reward function representing the target Critic network; S t+1 and a t+1 These represent the next state and the next action, respectively. The Critic network updates its parameters by minimizing the error between the actual reward value and the target reward value. After receiving feedback from the Critic network, the Actor network optimizes its strategy based on the gradient update in Equation (4). After receiving feedback from the Critic network, the Critic network uses Equation (4) to perform gradient update and optimize its strategy: In Equation (4): This represents the gradient of the policy objective function with respect to the Actor network parameters θ; For policy functions; This represents the Critic network's estimate of the state-action values. This update mechanism ensures that the Actor network can progressively optimize its strategy based on feedback from the Critic network, making more effective decisions in complex negotiation environments.

[0047] In one possible implementation, calculating the value estimate corresponding to the second negotiated value includes: calculating the behavior value corresponding to the second negotiated value using the preset reception policy network, an activation function, and an output function; normalizing the behavior value using a normalization function to obtain a normalization result; wherein the normalization result includes whether the second negotiated value is received; and calculating the value estimate corresponding to the second negotiated value and the behavior value using a value function.

[0048] Specifically, the acceptance policy network dynamically optimizes acceptance and rejection decisions in multi-issue negotiation through an Actor-Critic structure. The Actor network generates acceptance probabilities based on the current state and behavior, and selects the final behavior using Softmax. The Critic network provides a state-behavior value function to assist the Actor network in adjusting its strategy to improve the negotiation outcome. The Actor network is responsible for generating the acceptance or rejection decision. Its main function is to generate a probability distribution of the acceptance behavior based on the current state. The input to the Actor network is the current state S. t This includes a statement indicating that the opponent is Bid B at any time t (e.g., cost, risk, treatment time, efficacy, convenience) vectors. Then, through a series of fully connected layers, using the ReLU activation function and the Softplus output function, the Actor network finally outputs a non-negative action value 'a'. Action selection: After generating an action value 'a', the Actor network normalizes it using the Softmax function, calculating the probability distribution for each action. Then, a discrete sampling strategy is used to select a specific action 'a' from this distribution (0: reject, 1: accept). Critic network: Used to evaluate the value function Q(s,a) for a given state-action pair, i.e., the expected reward of the action. The input to the Critic network is the next state S. t+1 and the next action a t+1 The network consists of multiple fully connected layers and uses the ReLU activation function. The Critic network ultimately outputs a behavioral value estimate, Q(s,a).

[0049] In one possible implementation, the step of calculating the second reference value through the pre-trained Actor-Critic network if the first negotiating party rejects the second negotiated value includes: if the first negotiating party rejects the second negotiated value, calculating the mean and standard deviation corresponding to each issue through a preset bidding strategy network based on the second negotiated value and the utility value of each issue; identifying the corresponding normal distribution based on the mean and standard deviation corresponding to each issue; sampling the normal distribution to obtain the utility value corresponding to each issue; and generating the second reference value through the pre-trained Actor-Critic network based on the utility value corresponding to each issue.

[0050] In this application, the bidding strategy network is responsible for generating bid vectors for treatment options. In a multi-issue negotiation environment, the agent generates bids containing multiple issue values ​​based on the current negotiation state. Specifically, the bidding strategy network outputs a five-dimensional vector. Each value represents the utility of a specific issue, ranging from 0 to 1. These issues include treatment cost, efficacy, side effects, risks, and treatment convenience. If the agent rejects the current offer, it will make a counter-offer based on the current negotiation status. This process is achieved through continuous control, and the decision-making utilizes a sample drawn from a normal distribution. A normal distribution is used here. The utility of the quote is modeled, and its probability density function is shown in Equation (6): (6) Here, f(x) represents the probability density function of the normal distribution; x is a random variable (in this scheme, it refers to the utility value of the bid); It is the mean parameter; Here are the variance parameters. These two parameters are typically generated by a neural network to characterize the expected utility of each issue and the uncertainty of utility assessment. Based on this distribution, the utility of each issue is generated, and then the final offer is determined. If the agent evaluates and accepts the current offer, the negotiation ends using the acceptance strategy network. If the agent rejects the current offer, a counter-offer is generated through the offer strategy network. The input to the offer strategy network is a six-dimensional vector containing the opponent's offer (cost, risk, treatment duration, effectiveness, convenience) and the time step t. Then, the Actor network generates the mean and standard deviation of the utility value for each issue based on the current state. These utility values ​​are generated by sampling from a normal distribution. During the offer generation process, the average value for each issue is first output by multiple Actor networks. and standard deviation Then, sampling is performed using formula (7): (7) Here, z i These are random variables sampled from a standard normal distribution N(0,1), used to generate the utility value for each issue. Through this mechanism, the bidding strategy network can balance the expected utility and uncertainty of each issue and generate a bid B that meets the negotiation requirements. i The Critic network evaluates the utility of a bid by calculating the payoff value of the current state and the bid-ask pair. The input to the Critic network is the next state S. t+1 and the next action a t+1 The Critic network connects to the Actor network and outputs a corresponding reward value, representing the utility of the offer in the current state. The output of the Critic network provides feedback to the Actor network to optimize the offer strategy. The Critic network has a similar structure to the Actor network, consisting of multiple fully connected layers, and outputs a scalar reward value representing the action value.

[0051] In a possible implementation, after receiving the second negotiation value fed back by the second negotiation party after receiving the first negotiation value, the method further includes: if the first negotiation party accepts the second negotiation value, then through a preset reward function: (7) calculate the reward for the end of negotiation; where represents the reward for the end of negotiation; represents the issue weight vector at the final negotiation round t f When; U(o) represents the utility value of the finally agreed-upon plan o; T represents the defined deadline before the negotiation starts; the weights of the preset receiving policy network and / or the preset offering policy network are corrected through the reward for the end of negotiation. Specifically, in the embodiments of the present application, for the preset reward function, the acceptance policy network and the bidding policy network share the same reward function. Given the defined deadline T before the negotiation starts, the issue weights and the utility U(o) of the final offer o. If two Agents reach an agreement before the negotiation deadline (t f <T and t f ≠0), the reward is set to the weighted utility of reaching an agreement, which directly reflects the overall satisfaction of the negotiation result. If no agreement is reached at the deadline T (conflicting transaction), the reward is set to zero, which is a penalty for the negotiation failure and discourages the Agent from extending the negotiation without reaching an agreement. This reward structure encourages the two Agents to effectively reach a mutually beneficial agreement and maximize the weighted sum of their own preferences, enabling the negotiation to dynamically achieve immediate goals and satisfactory joint decisions.

[0052] In the present application, the data structures involved in each step include negotiation issue values, negotiation issue weight values, negotiation issue utility values, the state vector of the Agent, the offer vector, etc. Specifically, it includes: In the setting of the negotiation scenario, the issue domain ( negotiation issues, the value range of each negotiation issue) defines the specific scenario of the Agent's negotiation.

[0053] The negotiation preference data defined in the construction process of the utility model provides a basis for the subsequent automatic negotiation of the Agent, that is, the preference interval of the Agent for the values of each negotiation issue , each issue weight w i , and the overall issue satisfaction function U.

[0054] The negotiation process data defined in the construction process of the negotiation framework is the data involved when the Agent performs automatic negotiation and is a specific manifestation of the negotiation. The specific data includes the current offer B tThe opponent's counter-offer The time step t, and the next round of environmental feedback state composed of the counter-offer and the time step; the fuzzy preference utility model U, and the generated next action a. t+1 Next state S t+1, Current state S t, Current behavior a t And the calculated return value.

[0055] During the pre-training phase, the acceptance policy network, the offer policy network, and the reward function work together to specifically transform the data defined in the process of constructing the negotiation framework.

[0056] In this application, the relationship between data structures and actual processing objects is explained. When processing negotiation data, the process essentially involves performing various mathematical operations on the Agent's values ​​for the negotiation agenda, read from the runtime memory, to obtain a new set of data. The new output is then stored on the computer in numerical format.

[0057] In this application, the hardware components and execution process are as follows: This solution is primarily based on a computer software system. The hardware components include computing devices such as servers and workstations, as well as storage devices for storing negotiated data and model parameters. The execution process is as follows: The computing devices run the AutoSDM-DDPG framework software, which interacts with other devices (such as doctor's and patient's terminals) via a network to complete the negotiation process. This application presents AutoSDM-DDPG, a multi-issue SDM automated negotiation framework based on DDPG, designed to address SDM problems in the healthcare decision-making field. Its beneficial effects are manifested in the following aspects: First, under different negotiation deadlines and issue number constraints, AutoSDM-DDPG consistently achieves higher social welfare (total individual satisfaction) and better fairness (difference in individual satisfaction) compared to other agent-based SDM models. Furthermore, even with increasing complexity of negotiation issues, this scheme maintains high individual satisfaction among negotiation participants, demonstrating strong adaptability and scalability. Second, AutoSDM-DDPG significantly reduces the average number of negotiation rounds required to reach consensus, indicating improved efficiency in solving complex multi-issue SDM problems. Finally, AutoSDM-DDPG better captures the nonlinearity and fuzziness of preferences among negotiating parties in the real world, thereby achieving more balanced and equitable outcomes. Overall, this scheme combines deep reinforcement learning with fuzzy preference modeling, effectively addressing the inherent uncertainty and diversity of preferences in SDM. Moreover, AutoSDM-DDPG can provide a promising and scalable solution for intelligent, fair, and efficient automated negotiation in healthcare and other fields.

[0058] A second aspect of this application provides a decision negotiation device based on deep deterministic policy gradient, applied to a first negotiating party in a negotiation scenario, see [link to previous document]. Figure 4 , Figure 4 A schematic diagram of a decision negotiation device based on deep deterministic policy gradient provided in this application embodiment, the device comprising: The scenario receiving module 401 is used to receive negotiation scenario information input by the first negotiating party; wherein, the negotiation scenario information includes: multiple negotiation topics, preference value ranges for each negotiation topic, and preference weight values ​​for each negotiation topic. The satisfaction calculation module 402 is used to create a utility model for each negotiation topic based on the preference value range and preference weight value of the negotiation topic; wherein, the utility model includes: trapezoidal fuzzy membership functions corresponding to the negotiation values ​​when they are in different preference value ranges, and an overall satisfaction function that is a weighted sum of the trapezoidal fuzzy membership functions of each negotiation dimension. The negotiation value prediction module 403 is used to predict the first negotiation value of the first negotiating party through a pre-trained Actor-Critic network, wherein the first negotiation value is obtained by using the utility model to predict the predicted value through the Actor network in the pre-trained Actor-Critic network, calculating the reward value of the predicted value through the Critic network, and iteratively updating the value through the Actor network based on the reward value. The negotiation value sending module 404 is used to receive the second negotiation value input by the first negotiating party with reference to the first negotiation value, and send it to the second negotiating party.

[0059] In one possible implementation, the device further includes: The value estimation calculation module is used to receive a second negotiation value fed back by the second negotiating party after receiving the first negotiation value; and to calculate a value estimate corresponding to the second negotiation value; wherein the value estimate is used to characterize the expected return corresponding to the second negotiation value; The second reference calculation module is used to calculate a second reference value through the pre-trained Actor-Critic network if the first negotiating party rejects the second negotiated value.

[0060] In one possible implementation, the utility model includes: the trapezoidal fuzzy membership function and the overall satisfaction function; The trapezoidal fuzzy membership function is:

[0061] Where a, b, c, and d are the values ​​corresponding to the preference value ranges of each negotiation topic input by the first negotiating party; k represents the normalization constant. The overall satisfaction function is:

[0062] Where n represents the total number of negotiation topics; w i Representing the The weight of each issue; f i (o i ) is a given number The trapezoidal fuzzy membership function for the negotiated value o of each issue; o represents the issue value vector that constitutes a specific negotiated value.

[0063] In one possible implementation, the negotiated value prediction module is specifically used to predict the value using the formula via the Critic network:

[0064] Calculate the return value of the predicted value, where r t Represents the immediate reward at time step t; γ is the discount factor; The reward function of the Critic network; s t+1 and a t+1 These represent the next state and the next action, respectively. Based on the reward value, the Actor network uses the formula:

[0065] Iterative updates are performed, among which, This represents the gradient of the policy function with respect to the Actor network parameters θ; Represents the strategy function; This represents the estimated value of the Critic network. This indicates the parameters of the Actor network. Find the gradient. This indicates that the gradient of action a is calculated.

[0066] In one possible implementation, the value estimation calculation module is specifically used to calculate the behavior value corresponding to the second negotiated value through the preset reception policy network using an activation function and an output function; normalize the behavior value using a normalization function to obtain a normalization result; wherein the normalization result includes whether the second negotiated value is received; and calculate the value estimate corresponding to the second negotiated value and the behavior value using a value function.

[0067] In one possible implementation, the second reference calculation module is specifically used to: if the first negotiating party rejects the second negotiated value, calculate the mean and standard deviation of each issue based on the second negotiated value and the utility value of each issue through a preset bidding strategy network; identify the corresponding normal distribution based on the mean and standard deviation of each issue; sample the normal distribution to obtain the utility value of each issue; and generate the second reference value based on the utility value of each issue through the pre-trained Actor-Critic network.

[0068] In one possible implementation, the device further includes: The weight adjustment module is used to, if the first negotiating party accepts the second negotiated value, apply a preset reward function:

[0069] Calculate the reward for concluding the negotiation; where, A reward indicating the conclusion of negotiations; Indicates the final negotiation round t f The issue weight vector at the time; U(o) represents the utility value of the final agreed solution o; T represents the deadline round defined before the start of negotiations; The weights of the preset receiving strategy network and / or the preset bidding strategy network are adjusted based on the reward for the completion of the negotiation.

[0070] As can be seen, the device of this application can receive negotiation scenario information at the beginning of the negotiation, and then, based on the negotiation scenario information, create a utility model for each negotiation topic according to the preference value range and preference weight value of the negotiation topic. The reward value of the predicted value is calculated by the Critic network, and the reference value is obtained by iteratively updating the reference value through the Actor network. This makes it easier for the first negotiating party to solve the problem of how to propose a reasonable negotiation value in the negotiation scenario by referring to the first negotiation value input by the first reference value, and facilitates the determination of the negotiation value.

[0071] This application also provides an electronic device, such as... Figure 5 As shown, it includes: Memory 501 is used to store computer programs; When processor 502 executes the program stored in memory 501, it performs the following steps: The system receives negotiation scenario information input by the first negotiating party; wherein the negotiation scenario information includes: multiple negotiation topics, preference value ranges for each negotiation topic, and preference weight values ​​for each negotiation topic. For each negotiation topic, a utility model is created based on the preference value range and preference weight value of the negotiation topic; wherein, the utility model includes: trapezoidal fuzzy membership functions corresponding to the negotiation values ​​when they are in different preference value ranges, and an overall satisfaction function that is a weighted sum of the trapezoidal fuzzy membership functions of each negotiation dimension. A first reference value for the first negotiating party is predicted through a pre-trained Actor-Critic network. The first reference value is obtained by using the utility model to predict the predicted value through the Actor network in the pre-trained Actor-Critic network, calculating the reward value of the predicted value through the Critic network, and iteratively updating the value based on the reward value through the Actor network. The first negotiation value, input by the first negotiating party with reference to the first reference value, is received and sent to the second negotiating party.

[0072] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0073] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0074] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0075] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0076] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of any of the above-described decision negotiation methods based on deep deterministic policy gradients.

[0077] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the decision negotiation methods based on deep deterministic policy gradients in the above embodiments.

[0078] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a solid-state drive (SSD), etc.

[0079] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0080] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, the embodiments for apparatus, electronic devices, and storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0081] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.

Claims

1. A decision negotiation method based on deep deterministic policy gradient, characterized in that, The method includes: Receive negotiation scenario information input by the first negotiating party; wherein, the negotiation scenario information includes: multiple negotiation topics, preference value ranges for each negotiation topic, and preference weight values ​​for each negotiation topic; For each negotiation topic, a utility model is created based on the preference value range and preference weight value of the negotiation topic; wherein, the utility model includes: trapezoidal fuzzy membership functions corresponding to the negotiation values ​​when they are in different preference value ranges, and an overall satisfaction function that is a weighted sum of the trapezoidal fuzzy membership functions of each negotiation dimension. A first reference value for the first negotiating party is predicted through a pre-trained Actor-Critic network. The first reference value is obtained by using the utility model to predict the predicted value through the Actor network in the pre-trained Actor-Critic network, calculating the reward value of the predicted value through the Critic network, and iteratively updating the value based on the reward value through the Actor network. The system receives the first negotiation value input by the first negotiating party with reference to the first reference value, and sends the first negotiation value to the second negotiating party.

2. The method according to claim 1, characterized in that, After receiving the first negotiated value input by the first negotiating party with reference to the first reference value and sending it to the second negotiating party, the method further includes: Receive a second negotiated value from the second negotiating party; calculate a value estimate corresponding to the second negotiated value; wherein the value estimate is used to characterize the expected return corresponding to the second negotiated value; If the first negotiating party rejects the second negotiated value, a second reference value is calculated using the pre-trained Actor-Critic network.

3. The method according to claim 1, characterized in that, The utility model includes: the trapezoidal fuzzy membership function and the overall satisfaction function; The trapezoidal fuzzy membership function is: Where a, b, c, and d are the values ​​corresponding to the preference value ranges of each negotiation topic input by the first negotiating party; k represents the normalization constant. The overall satisfaction function is: Where n represents the total number of negotiation topics; w i Representing the The weight of each issue; f i (o i ) is a given first The trapezoidal fuzzy membership function for the negotiated value o of each issue; o represents the issue value vector that constitutes a specific negotiated value.

4. The method according to claim 1, characterized in that, The step of calculating the reward value of the predicted value through the Critic network and iteratively updating it based on the reward value through the Actor network includes: Using the Critic network and the formula: Calculate the return value of the predicted value, where r t Represents the immediate reward at time step t; γ is the discount factor; The reward function of the Critic network; s t+1 and a t+1 These represent the next state and the next action, respectively. Based on the reward value, the Actor network uses the formula: Iterative updates are performed, among which, This represents the gradient of the policy function with respect to the Actor network parameters θ; Represents the policy function; This represents the estimated value of the Critic network. This indicates the parameters of the Actor network. Find the gradient. This indicates that the gradient of action a is calculated.

5. The method according to claim 2, characterized in that, The calculation of the value estimate corresponding to the second negotiated value includes: By using a preset receiving strategy network and an activation function and an output function, the behavior value corresponding to the second negotiated value is calculated; the behavior value is normalized by a normalization function to obtain a normalization result; wherein, the normalization result includes whether the second negotiated value is received; the value estimate is obtained by calculating the second negotiated value and the behavior value using a value function.

6. The method according to claim 5, characterized in that, If the first negotiating party rejects the second negotiated value, the calculation of the second reference value through the pre-trained Actor-Critic network includes: If the first negotiating party rejects the second negotiated value, the mean and standard deviation of each issue are calculated based on the second negotiated value and the utility value of each issue through a preset pricing strategy network; the corresponding normal distribution is identified based on the mean and standard deviation of each issue; the normal distribution is sampled to obtain the utility value of each issue; and the second reference value is generated based on the utility value of each issue through the pre-trained Actor-Critic network.

7. The method according to claim 6, characterized in that, After receiving the second negotiation value fed back by the second negotiating party after receiving the first negotiation value, the method further includes: If the first negotiating party accepts the second negotiated value, then a preset reward function is applied: Calculate the reward for concluding the negotiation; where, A reward indicating the conclusion of negotiations; Indicates the final negotiation round t f The issue weight vector at the time; U(o) represents the utility value of the final agreed solution o; T represents the deadline round defined before the start of negotiations; The weights of the preset receiving strategy network and / or the preset bidding strategy network are adjusted based on the reward for the completion of the negotiation.

8. A decision negotiation device based on deep deterministic policy gradient, characterized in that, The device, used as a first negotiating party in a negotiation scenario, includes: The scenario receiving module is used to receive negotiation scenario information input by the first negotiating party; wherein, the negotiation scenario information includes: multiple negotiation topics, preference value ranges for each negotiation topic, and preference weight values ​​for each negotiation topic; The satisfaction calculation module is used to create a utility model for each negotiation topic based on the preference value range and preference weight value of the negotiation topic. The utility model includes: trapezoidal fuzzy membership functions corresponding to the negotiation values ​​when they are in different preference value ranges, and an overall satisfaction function that is a weighted sum of the trapezoidal fuzzy membership functions of each negotiation dimension. The negotiation value prediction module is used to predict the first negotiation value of the first negotiating party through a pre-trained Actor-Critic network. The first negotiation value is obtained by using the utility model to predict the predicted value through the Actor network in the pre-trained Actor-Critic network, calculating the reward value of the predicted value through the Critic network, and iteratively updating the Actor network based on the reward value. The negotiation value sending module is used to receive the second negotiation value input by the first negotiating party with reference to the first negotiation value, and send it to the second negotiating party.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains a computer program that, when executed by a processor, implements the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Negotiation agent method based on deep reinforcement learning

    CN119250104A

  • Multi-modal feature extraction network enhanced negotiation decision-making method based on space-time directed graph

    CN119357627A

  • Doctor-patient negotiation method and device fusing LLM and multi-agent architecture, equipment and medium

    CN120544955A

  • Convergent actor critic-based fuzzy reinforcement learning apparatus and method

    US20020198854A1

  • Quantum Thermal System

    US20240354625A1

Cited By

  • User role analysis method and system in network community

    CN121256595A