Digital human based on large model and method and system for data processing of procurement negotiation
By employing a large-scale model-based digital human procurement negotiation method, utilizing multi-flow parallel expansion and cross-flow grafting mechanisms, combined with multi-dimensional state vectors and multi-round adversarial deduction, the problems of flexibility and decision-making black box in procurement negotiations are solved, enabling prudent decision-making in complex environments.
Patent Information
- Application Number
- CN202511285343.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2045-09-10
AI Technical Summary
Existing technologies lack flexibility and make the decision-making process opaque in procurement negotiations. Traditional models are unable to cope with dynamic games and irrational behavior of counterparties, and single-threaded reasoning is prone to getting trapped in local optima.
A digital human procurement negotiation method based on a large model is adopted. Multiple thought streams are initialized in parallel and expanded to generate candidate thought trees. When the confidence level is low, cross-stream grafting is triggered. Combined with multi-dimensional negotiation state vectors and multi-round adversarial inference, the minimum-maximum regret criterion is used to determine the optimal strategy.
It enables flexible responses in complex negotiation environments, avoids local optima, ensures that decisions are balanced among multi-dimensional objectives, anticipates long-term impacts and potential risks, and avoids significant losses caused by misjudgments.
Smart Images

Figure CN120822618B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and system for processing procurement negotiation data based on a large-scale digital human model. Background Technology
[0002] As a specific application of artificial intelligence in the field of business interaction, digital humans can autonomously analyze competitor quotes, formulate response strategies, and generate communication scripts, playing a crucial role in complex procurement negotiations. Early procurement negotiation support systems mainly relied on expert systems with preset rules or statistical models based on historical data.
[0003] However, while expert systems have clear logic, their rules are rigid and they struggle to cope with the dynamic game dynamics and irrational behavior of opponents in real negotiations, resulting in poor adaptability. Statistical models based on traditional machine learning, such as regression analysis or classification algorithms, can learn price patterns or opponent behavior patterns from historical data, but they often treat negotiation as a single-point decision problem, lacking a holistic understanding of the multi-round, long-term negotiation process. Furthermore, the decision-making process is often a black box, making it difficult to explain the logic behind the strategies—a fatal flaw in high-risk business negotiations.
[0004] With the rise of large language models, techniques such as chain reasoning have provided new directions for simulating complex human reasoning processes. By constructing reasoning chains step by step, chain reasoning enhances the model's logical analysis capabilities and the interpretability of decisions in complex problems, solving the black-box problem of traditional models. However, most chain reasoning techniques adopt a single-threaded linear reasoning mode, that is, reasoning along a single thought path. If the initial direction is chosen improperly or there is a deviation in the intermediate links, the entire reasoning chain is prone to getting stuck in local optima or even going astray. When facing negotiating opponents with ever-changing strategies, unidirectional reasoning cannot predict the opponent's reactions and the future multi-round adversarial evolution, which has significant limitations. Summary of the Invention
[0005] The purpose of this invention is to propose a procurement negotiation data processing method and system based on a large-scale digital human model, in order to solve the problems of lack of flexibility in negotiation and "black box" decision-making process in the prior art; to this end, this invention provides solutions in the following two aspects.
[0006] In a first aspect, the present invention provides a method for processing procurement negotiation data based on a large-scale digital human model, comprising the following steps:
[0007] The process involves: acquiring the counterparty's negotiation data for the current negotiation round and the team's pre-defined multi-dimensional negotiation objectives; initializing multiple thought streams in parallel within a pre-constructed negotiation knowledge graph based on the counterparty's negotiation data and multi-dimensional negotiation objectives, with each thought stream generating candidate thought trees through hierarchical expansion; triggering cross-stream grafting when the confidence level of any thought tree's branch to be expanded falls below a preset threshold, merging nodes from other thought streams with confidence levels above the preset threshold into the current thought tree as new branches; calculating a multi-dimensional negotiation state vector for the leaf nodes of each candidate thought tree, including price advantage, long-term cooperation value, and performance reliability, and selecting non-dominated leaf nodes as candidate strategy sets; calling multiple pre-defined counterparty profile game models to perform adversarial deductions for each candidate strategy over the next N rounds, generating a probability distribution matrix containing multiple deduction outcomes; determining the optimal strategy for the current round using the minimum-maximum regret criterion based on the probability distribution matrix, and generating structured negotiation response data based on the thought path corresponding to the optimal strategy.
[0008] Preferably, based on the counterparty's negotiation data and multi-dimensional negotiation objectives, multiple thought streams are initialized in parallel within a pre-constructed negotiation knowledge graph. Each thought stream generates a candidate thought tree through hierarchical expansion, including: initializing three independent thought streams with the three dimensions of optimal price, shortest delivery cycle, and most stable cooperative relationship as the main objectives, and setting each main objective as the root node of the corresponding thought stream; for any node in each thought stream, taking the negotiation point represented by the node as the center, retrieving related arguments, potential risks, and solutions as child nodes in the negotiation knowledge graph, and expanding layer by layer with breadth priority until the expansion depth reaches a preset number of layers to form a candidate thought tree.
[0009] Preferably, when the confidence level of any thought tree's branch to be expanded is lower than a preset threshold, triggering cross-flow grafting, and merging nodes from other thought flows with confidence levels higher than the preset threshold into the current thought tree as new branches, includes: calculating the confidence level of the branch to be expanded as the cumulative product of the confidence levels of all nodes on the path from the root node to the current branch to be expanded; pausing the expansion of the branch when the confidence level of the branch to be expanded in thought tree A is lower than the preset threshold; retrieving all nodes in other thought trees whose path confidence levels from the root node to the current node are higher than the preset threshold, and selecting the node with the highest path confidence level as the grafting node; copying the grafting node and its complete subtree and connecting them to the branch to be expanded in thought tree A that triggered the pause, forming a new expanded branch.
[0010] Preferably, the step of calculating a multi-dimensional negotiation state vector for each candidate thought tree leaf node, including price advantage, long-term cooperation value, and performance reliability, includes: calculating price advantage. The calculation formula is: ,in, To offer competitive pricing, This is the historical average bid from the opposing party. Offer a strategy quote for the current leaf node; calculate the long-term cooperation value. The historical cooperation period, historical average order amount, and supplier rating score are normalized, and then weighted and summed. The calculation formula is as follows: ,in, For the sake of long-term cooperation value, , , These are the normalized values of historical cooperation duration, historical average order amount, and supplier rating score, respectively. , , For the corresponding preset weight coefficients, and ; Calculate the reliability of contract performance The calculation formula is: ,in, For the sake of reliable performance, This refers to the number of orders delayed in the past year. This represents the total number of orders placed within the past year.
[0011] Preferably, the step of calling multiple preset opponent profile game models to conduct adversarial simulations for each candidate strategy over the next N rounds, generating a probability distribution matrix containing multiple simulation outcomes, includes: calling an aggressive price suppression model, a robust value-oriented model, and a compromise relationship maintenance model as multiple opponent profile game models; taking any candidate strategy as our first round input, simulating the opponent's response and our subsequent two rounds of responses in the three models respectively, until three rounds of simulation are completed; and counting the frequency of occurrence of the three outcomes of reaching a transaction, negotiation breakdown, and negotiation stalemate at the end of each model, and converting the occurrence frequency into probability to obtain the probability distribution matrix of the simulation outcomes.
[0012] Preferably, the step of determining the optimal strategy for this round based on the probability distribution matrix and using the minimum-maximum regret criterion in the set of candidate strategies includes: pre-setting payoff values for three hypothetical outcomes—a deal being reached, negotiations breaking down, and negotiations reaching a stalemate—and, in conjunction with the probability distribution matrix, calculating the expected payoff value of each candidate strategy under different opponent profile game models to form a payoff matrix; identifying the highest expected payoff value that all candidate strategies can achieve under each opponent profile game model; calculating the regret matrix, wherein each element is calculated by subtracting the expected payoff value of a specific strategy under the opponent profile game model from the highest expected payoff value under the opponent profile game model; determining the maximum regret value of each candidate strategy across all opponent profile game models; and selecting the candidate strategy with the smallest maximum regret value as the optimal strategy for this round.
[0013] In the second aspect, a procurement negotiation data processing system based on a large-scale model of digital humans includes the following modules: a mind expansion module, used to acquire the opponent's negotiation data for the current negotiation round and the multi-dimensional negotiation objectives preset by the user; based on the opponent's negotiation data and the multi-dimensional negotiation objectives, multiple mind flows are initialized in parallel in a pre-constructed negotiation knowledge graph, and each mind flow generates candidate mind trees through hierarchical expansion; during the expansion process, when the confidence of any branch to be expanded in any mind tree is lower than a preset threshold, cross-flow grafting is triggered, and nodes in other mind flows with confidence higher than the preset threshold are merged into the current mind tree as new branches. The deduction module is used to calculate a multi-dimensional negotiation state vector for each leaf node of the candidate thought tree, including the bid advantage, long-term cooperation value, and performance reliability, and to select non-dominated leaf nodes as the candidate strategy set. It calls multiple preset opponent profile game models to perform adversarial deduction for each candidate strategy in the next N rounds, generating a probability distribution matrix containing multiple deduction outcomes. The negotiation data generation module is used to determine the optimal strategy for this round in the candidate strategy set according to the probability distribution matrix and the minimum maximum regret criterion, and to generate structured negotiation response data based on the thought path corresponding to the optimal strategy.
[0014] Preferably, based on the counterparty's negotiation data and multi-dimensional negotiation objectives, multiple thought streams are initialized in parallel within a pre-constructed negotiation knowledge graph. Each thought stream generates a candidate thought tree through hierarchical expansion, including: initializing three independent thought streams with the three dimensions of optimal price, shortest delivery cycle, and most stable cooperative relationship as the main objectives, and setting each main objective as the root node of the corresponding thought stream; for any node in each thought stream, taking the negotiation point represented by the node as the center, retrieving related arguments, potential risks, and solutions as child nodes in the negotiation knowledge graph, and expanding layer by layer with breadth priority until the expansion depth reaches a preset number of layers to form a candidate thought tree.
[0015] Preferably, when the confidence level of any thought tree's branch to be expanded is lower than a preset threshold, triggering cross-flow grafting, and merging nodes from other thought flows with confidence levels higher than the preset threshold into the current thought tree as new branches, includes: calculating the confidence level of the branch to be expanded as the cumulative product of the confidence levels of all nodes on the path from the root node to the current branch to be expanded; pausing the expansion of the branch when the confidence level of the branch to be expanded in thought tree A is lower than the preset threshold; searching other thought trees for all nodes on the path from the root node to the current node with confidence levels higher than the preset threshold, and selecting the node with the highest path confidence level as the grafting node; copying the grafting node and its complete subtree and connecting it to the branch to be expanded in thought tree A that triggered the pause, forming a new expanded branch.
[0016] Preferably, the step of calculating a multi-dimensional negotiation state vector for each candidate thought tree leaf node, including price advantage, long-term cooperation value, and performance reliability, includes: calculating price advantage. The calculation formula is: ,in, To offer competitive pricing, This is the historical average bid from the opposing party. Offer a strategy quote for the current leaf node; calculate the long-term cooperation value. The historical cooperation period, historical average order amount, and supplier rating score are normalized, and then weighted and summed. The calculation formula is as follows: Among these, the value lies in long-term cooperation. , , These are the normalized values of historical cooperation duration, historical average order amount, and supplier rating score, respectively. , , For the corresponding preset weight coefficients, and ; Calculate the reliability of contract performance The calculation formula is: Among these, reliability of performance is crucial. This refers to the number of orders delayed in the past year. This represents the total number of orders placed within the past year.
[0017] The beneficial effects of this invention are as follows: By initializing multiple thought streams in parallel and utilizing a cross-stream grafting mechanism, this invention overcomes the problem of traditional single thought chains easily getting trapped in local optima. Furthermore, by constructing a multi-dimensional state vector that includes pricing advantage, cooperation value, and contract fulfillment reliability, it ensures that the selected candidate strategies can achieve effective equilibrium among multiple conflicting objectives, avoiding one-sided decisions that sacrifice long-term interests by focusing only on short-term prices. In addition, the introduction of multiple rounds of adversarial simulations allows for the anticipation of the long-term impact and potential risks of strategies. The use of the minimum-maximum regret criterion for final decision-making enables the selection of the most prudent solution in an uncertain negotiation environment, avoiding significant losses that may result from misjudging the opponent's intentions. Attached Figure Description
[0018] Figure 1 The flowchart illustrating the steps of the procurement negotiation data processing method based on a large-model digital human in this embodiment is shown in the illustration. Detailed Implementation
[0019] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0020] like Figure 1As shown, the procurement negotiation data processing method based on a large-scale model of digital humans in this embodiment includes the following steps:
[0021] Step S1: Obtain the counterparty's negotiation data for the current negotiation round and the multi-dimensional negotiation objectives preset by the party; based on the counterparty's negotiation data and the multi-dimensional negotiation objectives, initialize multiple thought streams in parallel in the pre-constructed negotiation knowledge graph, and generate candidate thought trees through hierarchical expansion of each thought stream; during the expansion process, when the confidence of any thought tree's branch to be expanded is lower than a preset threshold, trigger cross-stream grafting, and merge nodes in other thought streams with confidence higher than the preset threshold into the current thought tree as new branches.
[0022] The system receives structured data from the counterparty via an API interface, such as a JSON object containing a price quote of 105 yuan, a delivery period of 60 days, and payment terms of cash on delivery. Simultaneously, it loads its own multi-dimensional negotiation objectives from an internal database. These objectives are quantified as a price target range of 95 to 102 yuan, an optimal delivery period of 30 days, acceptable payment terms including a 30% deposit, and a desired long-term cooperation score of 4.5 or higher (out of 5). A pre-built negotiation knowledge graph stores entities and relationships such as supplier historical performance records, industry cost benchmarks, and market supply and demand. Based on this, three thought flows are initialized in parallel: Thought Flow A is price-optimal oriented, Thought Flow B is long-term cooperation oriented, and Thought Flow C is rapid performance oriented. Specifically, when constructing the knowledge graph, the spaCy library is used for basic entity recognition, such as company name and date, and pre-trained models such as BERT are used for fine-tuning to identify negotiation-specific entities, such as negotiation topics and concessions. Relationship extraction preferably uses the OpenNRE library to identify relationships between entities. When generating the mind tree, a stateful Actor model is used, where each thought flow is a Ray Actor.
[0023] Based on the counterparty's negotiation data and multi-dimensional negotiation objectives, multiple thought streams are initialized in parallel within a pre-constructed negotiation knowledge graph. Each thought stream generates a candidate thought tree through hierarchical expansion. This includes: initializing three independent thought streams based on three main objectives: optimal price, shortest delivery cycle, and most stable cooperative relationship, with each main objective set as the root node of its corresponding thought stream; for any node in each thought stream, taking the negotiation point represented by the node as the center, retrieving related arguments, potential risks, and solutions from the negotiation knowledge graph as child nodes, and expanding layer by layer with a breadth-first approach until the expansion depth reaches the preset number of layers, thus forming a candidate thought tree.
[0024] The negotiation process establishes three core objectives: optimal price, shortest delivery cycle, and most stable partnership. For each objective, a separate thought process is initiated, using the objective itself as the starting point of a thought tree. Each thought tree is then constructed layer by layer. Starting from any key negotiation point, the knowledge graph is searched for supporting arguments, potential risks, and corresponding countermeasures, which are then added as new nodes to the next layer. This process continues in a breadth-first manner until a preset limit for the number of expansion layers is reached, ultimately forming multiple structured candidate thought trees.
[0025] When the confidence level of any branch to be expanded in any thought tree is lower than a preset threshold, cross-flow grafting is triggered, merging nodes from other thought flows with confidence levels higher than the preset threshold into the current thought tree as new branches. This includes: calculating the confidence level of the branch to be expanded as the cumulative product of the confidence levels of all nodes on the path from the root node to the current branch to be expanded; pausing the expansion of the branch when the confidence level of the branch to be expanded in thought tree A is lower than the preset threshold; retrieving all nodes in other thought trees whose path confidence levels from the root node to the current node are higher than the preset threshold, and selecting the node with the highest path confidence level as the grafting node; copying the grafting node and its complete subtree and connecting them to the branch to be expanded in thought tree A that triggered the pause, forming a new expanded branch.
[0026] When constructing a mind tree, a confidence level is calculated for each branch. The confidence level is preferably obtained by multiplying the confidence levels of all nodes along the path from the root node to the current node. Once the cumulative confidence level of a branch, for example in mind tree A, drops below a preset threshold of 0.3, the growth of that branch is paused. Candidate nodes with a path confidence level higher than 0.3 are searched in other mind trees B and C, and the node with the highest confidence level is selected. This high-confidence node and all its subordinate substructures are then completely copied and transplanted to the paused branch in mind tree A, effectively replacing the original weak branch with a more reliable argument path.
[0027] Step S2: Calculate a multi-dimensional negotiation state vector for each leaf node of the candidate thought tree, including the bid advantage, long-term cooperation value, and performance reliability, and select non-dominated leaf nodes as the candidate strategy set; call multiple preset opponent profile game models to perform adversarial simulations for each candidate strategy in the next N rounds, and generate a probability distribution matrix containing multiple simulation outcomes.
[0028] For each leaf node of a candidate thought tree, calculate a multi-dimensional negotiation state vector containing the bid advantage degree, long-term cooperation value, and performance reliability, including: calculating the bid advantage degree. The calculation formula is: ,in, To offer competitive pricing, This is the historical average bid from the opposing party. Offer a strategy quote for the current leaf node; calculate the long-term cooperation value. The historical cooperation period, historical average order amount, and supplier rating score are normalized, and then weighted and summed. The calculation formula is as follows: ,in, For the sake of long-term cooperation value, , , These are the normalized values of historical cooperation duration, historical average order amount, and supplier rating score, respectively. , , For the corresponding preset weight coefficients, and ; Calculate the reliability of contract performance The calculation formula is: ,in, For the sake of reliable performance, This refers to the number of orders delayed in the past year. This represents the total number of orders placed within the past year.
[0029] The pricing advantage is calculated using the formula described above: subtract the current strategy's price from the competitor's historical average price, and then divide by that historical average price, thus quantifying price competitiveness. The assessment of long-term cooperation value involves normalizing indicators such as historical cooperation duration, average order amount, and supplier rating, and then weighting and summing them according to preset weights. Delivery reliability is calculated by statistically analyzing the proportion of orders delivered without delays in the past year out of the total number of orders. These three indicators together constitute a comprehensive evaluation system for each strategy, providing comprehensive and objective data support for selecting the optimal strategy and ensuring that decisions achieve the best balance between cost, value, and risk.
[0030] Multiple pre-defined opponent profile game models are invoked to conduct adversarial simulations for each candidate strategy over the next N rounds, generating a probability distribution matrix containing multiple simulation outcomes. This includes: invoking an aggressive price suppression model, a robust value-oriented model, and a compromise relationship maintenance model as the multiple opponent profile game models; using any candidate strategy as the first-round input, simulating the opponent's response and the player's subsequent two rounds of responses in the three models, until all three rounds of simulation are completed; and counting the frequency of occurrence of the three outcomes—a deal being reached, negotiations breaking down, and negotiations reaching a stalemate—at the end of the simulation under each model, converting the frequency into probability to obtain the probability distribution matrix of the simulation outcomes.
[0031] To predict the actual effects of each candidate strategy, they were placed in a simulated adversarial environment for simulation. Three typical adversary behavior patterns were preset: aggressive price suppression, conservative value orientation, and compromising relationship maintenance, and the simulation was set to last for three rounds. For any candidate strategy, its interaction process under the above three adversary patterns was simulated for three rounds to observe the dynamic evolution results. After the simulation, the number of times a deal was reached, negotiations broke down, and a stalemate occurred under each adversary pattern was counted, and the probability of each outcome was calculated and summarized to form an outcome probability distribution matrix.
[0032] Step S3: Based on the probability distribution matrix, the minimum-maximum regret criterion is used to determine the optimal strategy for this round from the set of candidate strategies, and structured negotiation response data is generated based on the thought path corresponding to the optimal strategy.
[0033] Construct a regret matrix based on the probability distribution matrix. Specifically, for each deduced outcome (each column), find the highest probability value achievable across all strategies. Then, subtract the probability value of each strategy in that column from this highest probability value to obtain the regret value for each strategy under that outcome. For each row of the regret matrix (each strategy), find its maximum regret value across all outcomes. Select the strategy with the smallest maximum regret value as the optimal strategy for this round. After determining the optimal strategy, trace its complete path in the mind tree from the root node to the leaf node. Transform the logic of each node along the path (e.g., the expected price adjustment due to our commitment to long-term procurement) into natural language and encapsulate it along with the strategy's final quote, delivery date, and other parameters into a structured JSON object as a formal negotiation response.
[0034] In an optional embodiment, based on the probability distribution matrix, the optimal strategy for the current round is determined using the minimum-maximum regret criterion from the set of candidate strategies. This includes: pre-setting payoff values for three hypothetical outcomes—a deal being reached, negotiations breaking down, and negotiations reaching a stalemate—and, in conjunction with the probability distribution matrix, calculating the expected payoff value of each candidate strategy under different opponent profile game models to form a payoff matrix; identifying the highest expected payoff value achievable by all candidate strategies under each opponent profile game model; calculating the regret matrix, where each element is calculated by subtracting the expected payoff value of a specific strategy under the opponent profile game model from the highest expected payoff value under that model; determining the maximum regret value for each candidate strategy across all opponent profile game models; and selecting the candidate strategy with the smallest maximum regret value as the optimal strategy for the current round.
[0035] Assign a pre-defined payoff score to each of the three possible outcomes: reaching a deal, negotiation breakdown, and negotiation stalemate. Combined with a derived probability distribution, calculate the expected payoff for each strategy when dealing with different types of opponents, forming a payoff matrix. Calculate the regret value: for each opponent type, subtract the expected payoff of the current strategy from the highest possible payoff in that scenario. The resulting value represents the opportunity missed by choosing that strategy instead of the optimal strategy. Find the maximum possible regret value for each strategy across all opponent types, and select the strategy with the smallest maximum regret value as the optimal, lowest-risk choice for this round of negotiations.
[0036] This invention also provides a procurement negotiation data processing system based on a large-scale digital human model, comprising the following modules:
[0037] The thinking expansion module is used to acquire the opponent's negotiation data for the current negotiation round and the multi-dimensional negotiation objectives preset by the party. Based on the opponent's negotiation data and the multi-dimensional negotiation objectives, multiple thinking streams are initialized in parallel in a pre-constructed negotiation knowledge graph. Each thinking stream generates candidate thinking trees through hierarchical expansion. During the expansion process, when the confidence of any branch to be expanded in any thinking tree is lower than a preset threshold, cross-stream grafting is triggered, and nodes in other thinking streams with confidence higher than the preset threshold are merged into the current thinking tree as new branches.
[0038] The deduction module is used to calculate a multi-dimensional negotiation state vector for each leaf node of the candidate thought tree, which includes the bid advantage, long-term cooperation value and performance reliability, and to select non-dominated leaf nodes as the set of candidate strategies. It calls multiple preset opponent profile game models to conduct adversarial deductions for each candidate strategy in the next N rounds, generating a probability distribution matrix containing multiple deduction outcomes.
[0039] The negotiation data generation module is used to determine the optimal strategy for this round from the set of candidate strategies based on the probability distribution matrix and the minimum maximum regret criterion, and to generate structured negotiation response data based on the thought path corresponding to the optimal strategy.
[0040] In a preferred embodiment, based on the counterparty's negotiation data and multi-dimensional negotiation objectives, multiple thought streams are initialized in parallel within a pre-constructed negotiation knowledge graph. Each thought stream generates a candidate thought tree through hierarchical expansion, including: initializing three independent thought streams with the three dimensions of optimal price, shortest delivery cycle, and most stable cooperative relationship as the main objectives, and setting each main objective as the root node of the corresponding thought stream; for any node in each thought stream, taking the negotiation point represented by the node as the center, retrieving related arguments, potential risks, and solutions as child nodes in the negotiation knowledge graph, and expanding layer by layer with breadth priority until the expansion depth reaches a preset number of layers to form a candidate thought tree.
[0041] In a preferred embodiment, when the confidence level of any thought tree's branch to be expanded is lower than a preset threshold, cross-flow grafting is triggered, merging nodes from other thought flows with confidence levels higher than the preset threshold into the current thought tree as new branches. This includes: calculating the confidence level of the branch to be expanded as the cumulative product of the confidence levels of all nodes on the path from the root node to the current branch to be expanded; pausing the expansion of the branch when the confidence level of the branch to be expanded in thought tree A is lower than the preset threshold; retrieving all nodes in other thought trees whose path confidence levels from the root node to the current node are higher than the preset threshold, and selecting the node with the highest path confidence level as the grafting node; copying the grafting node and its complete subtree and connecting them to the branch to be expanded in thought tree A that triggered the pause, forming a new expanded branch.
[0042] In a preferred embodiment, a multi-dimensional negotiation state vector, comprising price advantage degree, long-term cooperation value, and performance reliability, is calculated for the leaf nodes of each candidate thought tree, including: calculating the price advantage degree. The calculation formula is: ,in, To offer competitive pricing, This is the historical average bid from the opposing party. Offer a strategy quote for the current leaf node; calculate the long-term cooperation value. The historical cooperation period, historical average order amount, and supplier rating score are normalized, and then weighted and summed. The calculation formula is as follows: Among these, the value lies in long-term cooperation. , , These are the normalized values of historical cooperation duration, historical average order amount, and supplier rating score, respectively. , , For the corresponding preset weight coefficients, and ; Calculate the reliability of contract performance The calculation formula is: Among these, reliability of performance is crucial. This refers to the number of orders delayed in the past year. This represents the total number of orders placed within the past year.
[0043] In a preferred embodiment, multiple preset opponent profile game models are invoked to perform adversarial simulations for each candidate strategy over the next N rounds, generating a probability distribution matrix containing multiple simulation outcomes. This includes: invoking an aggressive price suppression model, a robust value-oriented model, and a compromise relationship maintenance model as the multiple opponent profile game models; using any candidate strategy as our first-round input, simulating the opponent's response and our subsequent two-round responses in the three models, until three rounds of simulation are completed; and counting the frequency of occurrence of the three outcomes—a deal being reached, negotiations breaking down, and negotiations reaching a stalemate—at the end of the simulation under each model, and converting the frequency of occurrence into probability to obtain the probability distribution matrix of the simulation outcomes.
[0044] In a preferred embodiment, based on the probability distribution matrix, the optimal strategy for the current round is determined using the minimum-maximum regret criterion from the set of candidate strategies. This includes: pre-setting payoff values for three hypothetical outcomes—a deal being reached, negotiations breaking down, and negotiations reaching a stalemate; calculating the expected payoff value of each candidate strategy under different opponent profile game models, based on the probability distribution matrix, to form a payoff matrix; identifying the highest expected payoff value achievable by all candidate strategies under each opponent profile game model; calculating the regret matrix, where each element is calculated by subtracting the expected payoff value of a specific strategy under the opponent profile game model from the highest expected payoff value under that model; determining the maximum regret value for each candidate strategy across all opponent profile game models; and selecting the candidate strategy with the smallest maximum regret value as the optimal strategy for the current round.
[0045] In the description of this specification, "multiple" means at least two, such as two, three or more, etc., unless otherwise expressly and specifically defined.
[0046] While various embodiments of the invention have been shown and described in this specification, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and essence of the invention.
Claims
1. A method for processing procurement negotiation data based on a large-scale digital human model, characterized in that, Includes the following steps: Obtain the counterparty's negotiation data for the current negotiation round and the multi-dimensional negotiation objectives preset by our side; Based on the counterparty's negotiation data and multi-dimensional negotiation objectives, multiple thought streams are initialized in parallel within a pre-constructed negotiation knowledge graph, and each thought stream generates a candidate thought tree through hierarchical expansion. During the expansion process, when the confidence of any branch of the thought tree to be expanded is lower than a preset threshold, cross-flow grafting is triggered, and nodes in other thought flows with a confidence higher than the preset threshold are merged into the current thought tree as new branches. For each candidate thought tree, calculate a multi-dimensional negotiation state vector containing the bid advantage, long-term cooperation value, and performance reliability for the leaf nodes, and select non-dominated leaf nodes as the set of candidate strategies; call multiple preset opponent profile game models to perform adversarial simulations for each candidate strategy in the next N rounds, and generate a probability distribution matrix containing multiple simulation outcomes. Based on the probability distribution matrix, the minimum maximum regret criterion is used to determine the optimal strategy for this round from the set of candidate strategies, and structured negotiation response data is generated based on the thought path corresponding to the optimal strategy.
2. The procurement negotiation data processing method based on a large-scale digital human according to claim 1, characterized in that, Based on the counterparty's negotiation data and multi-dimensional negotiation objectives, multiple thought processes are initialized in parallel within a pre-constructed negotiation knowledge graph. Each thought process generates a candidate thought tree through hierarchical expansion, including: With the three main objectives of optimal price, shortest delivery cycle, and most stable cooperative relationship as the primary objectives, three independent thought processes are initialized, and each primary objective is set as the root node of the corresponding thought process. For any node in each thought process, taking the negotiation point represented by the node as the center, retrieve related arguments, potential risks, and solutions as child nodes from the negotiation knowledge graph, and expand the breadth-first approach layer by layer until the expansion depth reaches the preset number of layers to form a candidate thought tree.
3. The procurement negotiation data processing method based on a large-scale model of digital humans according to claim 1, characterized in that, When the confidence level of any branch of a thought tree to be expanded is lower than a preset threshold, cross-flow grafting is triggered, merging nodes from other thought flows with a confidence level higher than the preset threshold into the current thought tree as new branches, including: The confidence score of the branch to be expanded is calculated as the cumulative product of the confidence scores of all nodes on the path from the root node to the current branch to be expanded. When the confidence level of the branch to be expanded in mind tree A is lower than a preset threshold, the expansion of that branch is paused. Retrieve all nodes in other mind trees whose path confidence from the root node to the current node is higher than the preset threshold, and select the node with the highest path confidence as the grafting node. The grafting node and its complete subtree are copied and connected to the branch to be expanded in mind tree A that triggers a pause, forming a new branch.
4. The procurement negotiation data processing method based on a large-scale digital human according to claim 1, characterized in that, The calculation of a multi-dimensional negotiation state vector for each candidate thought tree leaf node, including price advantage, long-term cooperation value, and performance reliability, includes: Calculate the price advantage The calculation formula is: ,in, To offer competitive pricing, This is the historical average bid from the opposing party. Offer a quote for the current leaf node strategy; Calculating the value of long-term cooperation The historical cooperation period, historical average order amount, and supplier rating score are normalized, and then weighted and summed. The calculation formula is as follows: ,in, For the sake of long-term cooperation value, , , These are the normalized values of historical cooperation duration, historical average order amount, and supplier rating score, respectively. , , For the corresponding preset weight coefficients, and ; Calculate the reliability of contract performance The calculation formula is: ,in, For the sake of reliable performance, This refers to the number of orders delayed in the past year. This represents the total number of orders placed within the past year.
5. The procurement negotiation data processing method based on a large-scale digital human according to claim 1, characterized in that, The process involves calling multiple pre-defined opponent profile game models to perform adversarial simulations for each candidate strategy over the next N rounds, generating a probability distribution matrix containing multiple simulation outcomes, including: The aggressive price suppression model, the robust value-oriented model, and the compromise relationship maintenance model are used as game models for profiling multiple counterparties. Using any candidate strategy as our first input, we simulate the opponent's response and our subsequent two responses in three models, until we complete three rounds of simulation. The frequency of occurrence of the three outcomes—a deal being reached, negotiations breaking down, and negotiations reaching a stalemate—at the end of the simulation under each model was statistically analyzed, and the frequency of occurrence was converted into probability to obtain the probability distribution matrix of the simulation outcome.
6. The procurement negotiation data processing method based on a large-scale model of digital humans according to claim 1, characterized in that, The step of determining the optimal strategy for this round based on the probability distribution matrix and using the minimum-maximum regret criterion from the set of candidate strategies includes: To predict the payout values for three possible outcomes—a deal being reached, negotiations breaking down, and negotiations reaching a stalemate—and in conjunction with the probability distribution matrix, the expected payout value for each candidate strategy under different opponent profile game models is calculated, thus forming a payout matrix. Find the highest expected payoff value that all candidate strategies can achieve under each opponent profile game model; The regret matrix is calculated by subtracting the expected payoff of a specific strategy under the opponent profile game model from the highest expected payoff under the opponent profile game model. Determine the maximum regret value of each candidate strategy in all opponent profile game models; The candidate strategy with the smallest maximum regret value is selected as the optimal strategy for this round.
7. A procurement negotiation data processing system based on a large-scale digital human model, characterized in that, Includes the following modules: The thinking expansion module is used to obtain the opponent's negotiation data and the multi-dimensional negotiation objectives preset by the party in the current negotiation round; based on the opponent's negotiation data and multi-dimensional negotiation objectives, multiple thinking streams are initialized in parallel in the pre-constructed negotiation knowledge graph, and each thinking stream generates candidate thinking trees through hierarchical expansion; During the expansion process, when the confidence of any branch of the thought tree to be expanded is lower than a preset threshold, cross-flow grafting is triggered, and nodes in other thought flows with a confidence higher than the preset threshold are merged into the current thought tree as new branches. The deduction module is used to calculate a multi-dimensional negotiation state vector for the leaf nodes of each candidate thought tree, including the bid advantage, long-term cooperation value and performance reliability, and to select non-dominated leaf nodes as the set of candidate strategies; it calls multiple preset opponent profile game models to conduct adversarial deduction for each candidate strategy in the next N rounds, and generates a probability distribution matrix containing multiple deduction outcomes. The negotiation data generation module is used to determine the optimal strategy for this round from the set of candidate strategies based on the probability distribution matrix and the minimum maximum regret criterion, and to generate structured negotiation response data based on the thought path corresponding to the optimal strategy.
8. The procurement negotiation data processing system based on a large-scale digital human according to claim 7, characterized in that, Based on the counterparty's negotiation data and multi-dimensional negotiation objectives, multiple thought processes are initialized in parallel within a pre-constructed negotiation knowledge graph. Each thought process generates a candidate thought tree through hierarchical expansion, including: With the three main objectives of optimal price, shortest delivery cycle, and most stable cooperative relationship as the primary objectives, three independent thought processes are initialized, and each primary objective is set as the root node of the corresponding thought process. For any node in each thought process, taking the negotiation point represented by the node as the center, retrieve related arguments, potential risks, and solutions as child nodes from the negotiation knowledge graph, and expand the breadth-first approach layer by layer until the expansion depth reaches the preset number of layers to form a candidate thought tree.
9. The procurement negotiation data processing system based on a large-scale digital human according to claim 7, characterized in that, When the confidence level of any branch of a thought tree to be expanded is lower than a preset threshold, cross-flow grafting is triggered, merging nodes from other thought flows with a confidence level higher than the preset threshold into the current thought tree as new branches, including: The confidence score of the branch to be expanded is calculated as the cumulative product of the confidence scores of all nodes on the path from the root node to the current branch to be expanded. When the confidence level of the branch to be expanded in mind tree A is lower than a preset threshold, the expansion of that branch is paused. Retrieve all nodes from the root node to the current node whose path confidence is higher than the preset threshold in other mind trees, and select the node with the highest path confidence as the grafting node. The grafting node and its complete subtree are copied and connected to the branch to be expanded in mind tree A that triggers a pause, forming a new branch.
10. The procurement negotiation data processing system based on a large-scale digital human according to claim 7, characterized in that, The calculation of a multi-dimensional negotiation state vector for each candidate thought tree leaf node, including price advantage, long-term cooperation value, and performance reliability, includes: Calculate the price advantage The calculation formula is: ,in, To offer competitive pricing, This is the historical average bid from the opposing party. Offer a quote for the current leaf node strategy; Calculating the value of long-term cooperation The historical cooperation period, historical average order amount, and supplier rating score are normalized, and then weighted and summed. The calculation formula is as follows: Among these, the value lies in long-term cooperation. , , These are the normalized values of historical cooperation duration, historical average order amount, and supplier rating score, respectively. , , For the corresponding preset weight coefficients, and ; Calculate the reliability of contract performance The calculation formula is: Among them, for the reliability of contract performance, This refers to the number of orders delayed in the past year. This represents the total number of orders placed within the past year.
Citation Information
Patent Citations
Game behavior dynamic evolution and strategy deduction optimization method and system in space field
CN119539090A
Method for AI language self-improvement agent using language modeling and tree search techniques
US12210849B1