Medical decision negotiation system based on Monte Carlo tree search
Through the medical decision-making consultation system based on Monte Carlo tree search, the application barriers of the SDM model in the health care field are solved, efficient consultation and satisfaction improvement between doctors and patients in the decision-making process, and promoting social equity.
Patent Information
- Application Number
- CN202510352161.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-08
AI Technical Summary
The application of existing common decision-making (SDM) models in the healthcare field is limited by obstacles such as insufficient clinical time, information asymmetry and personal bias. Traditional patient decision-making aid tools and clinical decision-making support systems have failed to effectively balance the values and preferences of doctors and patients, resulting in insufficient decision-making process.
The medical decision-making consultation system based on Monte Carlo tree search is adopted, and the user interface module, co-decision module, Monte Carlo tree search negotiation framework, utility evaluation module and negotiation control module are used to combine the belief-wish-intention (BDI) architecture and trapezoidal fuzzy membership function to generate treatment preference protocols to guide doctors and patients to make reasonable decisions during the negotiation process.
It improves overall satisfaction of doctors and patients in decision-making, reduces satisfaction differences, promotes social equity, and maintains efficient and stable negotiation performance in complex scenarios through the Monte Carlo tree search algorithm.
Smart Images

Figure CN120280182A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of healthcare, and particularly relates to a medical decision-making negotiation system based on Monte Carlo tree search. Background Art
[0002] The shared decision-making (SDM) model has become an important framework in the field of healthcare, facilitating collaborative decision-making between healthcare providers and patients. Different from the traditional paternalistic and informed decision-making processes, SDM prioritizes patients' autonomy and the right to manage their personal health, making it an ideal model for medical decision-making. Many inventions have demonstrated the benefits of SDM. For example, Kioko et al. emphasized that SDM aligns with the ethical goal of providing technically accurate and morally reasonable healthcare decisions as elaborated by Pellegrino and Thomasma. Additionally, Drake et al. emphasized that implementing SDM can improve the quality of care, increase patient satisfaction, promote better self-management, enhance treatment compliance, and lead to more meaningful health outcomes.
[0003] With the development of SDM, various conceptual models and frameworks have been proposed to simplify the SDM process. The Makoul model identified nine basic elements of SDM, including defining the problem, presenting options, discussing pros and cons, etc. Elwyn et al. transformed these elements into a practical three-step model (choice talk, option talk, decision talk) for convenient practical application. Similarly, the Stiggelbout model simplified SDM into an easy-to-remember four-step process to further assist implementation.
[0004] Despite the great potential of SDM, its widespread application in clinical practice is still limited due to various obstacles. These challenges include the need for new educational structures, communication strategies, and decision-making formats for personalized medicine. Insufficient clinical time, a severe information asymmetry between patients and healthcare providers, and personal biases further hinder the effective implementation of SDM. Current efforts to alleviate these challenges mainly focus on the development of SDM programs (SDPs), patient decision aids (PtDAs), and clinical decision support systems (CDSS). SDPs usually involve interactive video programs, which are costly to produce and thus not widely used. PtDAs are designed to inform patients of available treatment options and related risks, and they are only successful when the information is accurate and accessible to patients with different literacy levels and cultural backgrounds. CDSS uses computerized clinical knowledge bases to generate patient-specific recommendations to help doctors and patients make informed decisions. However, the above methods mainly guide patients to make decisions on specific issues through traditional media, without considering or balancing the values and preferences of doctors and patients, and proposing reference treatment options to help doctors and patients jointly make informed treatment decisions, so there are limitations.
[0005] In response, the inventor proposes a medical decision negotiation system based on Monte Carlo tree search to solve the above problems. Summary of the Invention
[0006] The purpose of the present invention is to provide a medical decision negotiation system based on Monte Carlo tree search to solve the problems raised in the above background technology.
[0007] To achieve the above purpose, the present invention provides the following technical solutions:
[0008] A medical decision negotiation system based on Monte Carlo tree search, comprising:
[0009] User interface module: used to receive and display information in the medical decision negotiation process, the information including the treatment preferences of both doctors and patients, negotiation progress and results;
[0010] Joint decision-making module: used to construct the behavior models of doctor agents and patient agents based on the Belief-Desire-Intention (BDI) architecture and using trapezoidal fuzzy membership functions, for simulating the communication and joint decision-making process between doctors and patients to obtain negotiation results;
[0011] Monte Carlo Tree Search Negotiation Framework: Based on the information and negotiation results, it uses the Monte Carlo Tree Search algorithm to generate a treatment preference protocol, which includes a bidding strategy, opponent preferences, and acceptance strategies, and is used to guide the doctor agent and the patient agent to make reasonable decisions during the negotiation process;
[0012] Utility Evaluation Module: It is used to calculate the satisfaction degrees of both the doctor and the patient with the treatment preference protocol, and the indicators of the satisfaction degree include the difference between the combined ASV and ASV;
[0013] Negotiation Control Module: It is used to control the negotiation process and obtain the final medical decision based on the treatment preference protocol corresponding to the highest value of the satisfaction degree indicator. The negotiation process includes negotiation start, negotiation rounds, negotiation end, and result output.
[0014] Preferably, the expression of the trapezoidal fuzzy membership function is:
[0015]
[0016] where U(S) represents the overall satisfaction value, M i (S) represents the i-th membership degree of the solution S, w i represents the weight of the i-th issue, and n represents the total number of issues that the doctor and the patient need to negotiate.
[0017] Preferably, the Monte Carlo Tree Search algorithm expands new nodes based on the Upper Confidence Tree, and the conditions for expanding new nodes are:
[0018]
[0019] where n p is the number of times the parent node is simulated, n c is the number of its child nodes, α is a parameter of the model, and its value range is 0 < α < 1; if the result is not to expand a new node, the selected node is the maximizing node i:
[0020]
[0021] where n is the total number of simulations of the tree, n i is the number of times node i in the tree is simulated, s i is the score of node i, and C is a parameter of the model.
[0022] Preferably, the goal of the bidding strategy is to obtain the predicted value of x * rounds; the steps for obtaining the predicted value are: generating a Gaussian distribution and using Gaussian process regression in the Gaussian distribution for prediction;
[0023] The first step of the Gaussian process regression is to calculate the covariance matrix K, which represents the proximity between the turning angles (x i ) i∈[1,n] of the sequence based on the covariance function; Let
[0024]
[0025] Calculate the distance between the turning of the predicted proposal x * and the previous turning in the vector K * :
[0026] K * =(k(x * ,x1),…,k(x * ,x n ))
[0027] The predicted values relied on by the Gaussian process regression are all assumed to be of the dimension of a multivariate Gaussian; Using the results of the multivariate Gaussian, calculate
[0028]
[0029] where K ** =k(x * ,x * ) and the result corresponds to a Gaussian random variable with a mean of and a standard deviation of σ * .
[0030] Preferably, in the communication common decision of the opponent preference, the overall satisfaction value U(S) is used to measure the satisfaction of the agent with the negotiation content;
[0031] Obtain the overall satisfaction function U(S t ) of the opponent in the t-th round, and adopt the setting that the first hypothesis is the problem weight w i of the opponent and the second hypothesis is the fuzzy membership function M i (S t );
[0032] The problem weight w i is obtained by setting a set of all possible weight matrices H w , calculating real numbers and associating them with the weight hypothesis using the following linear function:
[0033]
[0034] where, is the sorting of the weight w j in the hypothesis h i , and n is the number of negotiation issues.
[0035] Preferably, the fuzzy membership function M i (S t ), the preferences of the physician agent and the patient agent are considered as membership functions;
[0036] Assign a membership degree to each hypothesis in the hypothesis space and model the fuzzy membership function as a probability distribution;
[0037] By assuming various fuzzy membership functions and their corresponding probabilities By association, we can approximate the shape of the true fuzzy membership function of the i-th negotiation issue;
[0038] The probability P(h) of the opponent's bid is calculated using the following formula j ∣S t ), as follows:
[0039]
[0040] U'(S t )=U'(S t-1 )-c(t)
[0041] Among them, P(h j ∣S t ) indicates that it is assumed that j Under this condition, the other party will offer S t The probability of U(S t ∣h j ) is based on the assumption that h j In the case of t The satisfaction value, U'(S t ) is the satisfaction value of the other party's expected next bid, and function c(t) is the negotiation concession strategy assumed to be used by the other party;
[0042] Fuzzy membership function The expected value of the shape is calculated as follows:
[0043]
[0044] in, is the j-th hypothesis of problem i, and the fuzzy membership function generates the satisfaction value of the bid, express probability;
[0045] use to represent the assumptions about the weight values of problem i according to hypothesis j, and the values of the associated weights, i.e., The expected value of the weight is calculated as follows:
[0046]
[0047] Use the following formula to calculate the total satisfaction S expected by the opponent based on the bid t :
[0048]
[0049] For each question, use the following formula to standardize the weights and probability distributions of the evaluation function:
[0050]
[0051] where represents that the sum of all possibilities related to the satisfaction value hypothesis is 1; represents that the sum of all possibilities related to the weight value hypothesis is 1.
[0052] Preferably, the acceptance strategy is based on utility, time, or a combination of utility and time, where the acceptance conditions of the agent are as follows:
[0053]
[0054] where represents the current acceptance condition, that is, based on the parameter r and the opponent's proposal judge the next action;
[0055] Accept: Accept the opponent's proposal;
[0056] Propose a new counteroffer, proposal;
[0057] r: The current negotiation round, an integer variable used to identify the current negotiation round;
[0058] t: Represents the time pressure threshold, which changes dynamically as the negotiation round r increases and is used to determine whether to accept the current offer, defined as:
[0059]
[0060] where:
[0061] α: The time weight parameter, usually a constant, and 0 ≤ α ≤ 1, controlling the influence of the initial time threshold on t;
[0062] β: The growth rate parameter, controlling the non-linear growth or decay of t with the negotiation round r;
[0063] N max : The maximum number of negotiation rounds, representing the total number of negotiation rounds;
[0064] The proposal made by the opponent in the r-th round, usually representing the value of the issue in the plan proposed by the opponent in this round;
[0065] Denote the counteroffer made by us in the previous round (r - 1);
[0066] The opponent's proposal in the current round r The utility value for oneself;
[0067] The proposal made by oneself in negotiation round r - 1 The utility value of.
[0068] Compared with the prior art, the beneficial effects of the present invention are:
[0069] In various scenarios with different numbers of problems, in terms of the average CASV and DASV metrics, the proposed MCTSAN model is always superior to the existing models (such as FCAN and ANFGA); specifically, even when the complexity of the negotiation scenario increases, MCTSAN can maintain a high level of social welfare and reduce the satisfaction gap between patients and healthcare providers; the model can maintain excellent performance and stability even in a large solution space, which highlights its potential as a valuable tool for implementing SDM in real medical environments; by leveraging agent-based negotiation techniques, the framework can not only improve the overall satisfaction of patients and healthcare providers in decision-making but also promote social fairness by minimizing satisfaction differences. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 It is a block diagram of a medical decision-making negotiation system based on Monte Carlo tree search according to the present invention;
[0071] Figure 2 It is a graph of experimental comparison data of the MCTSAN framework of the present invention with the FCAN and ANFGA models using different negotiation strategies in terms of the average CASV metric;
[0072] Figure 3 It is a graph of experimental comparison data of the MCTSAN model of the present invention with different strategies and the FCAN and ANFGA models for the average DASV metric;
[0073] Figure 4 It is a graph of experimental comparison data of the MCTSAN model of the present invention with the FCAN and ANFGA models using different strategies in terms of the Avg.R metric. DETAILED DESCRIPTION OF THE INVENTION
[0074] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0075] Embodiment 1:
[0076] Please refer to Figure 1 As shown in the figure, a medical decision-making negotiation system based on Monte Carlo tree search includes:
[0077] User interface module: used to receive and display information in the medical decision-making negotiation process, and the information includes the treatment preferences, negotiation progress and results of both doctors and patients;
[0078] Joint decision-making module: used to construct the behavior models of doctor agents (DAs) and patient agents (PAs) based on the Belief-Desire-Intention (BDI) architecture and using trapezoidal fuzzy membership functions, for simulating the process of shared decision-making (SDM) communication between doctors and patients to obtain negotiation results;
[0079] Monte Carlo tree search negotiation framework (MCTSAN): used to generate treatment preference protocols based on the information and negotiation results using the Monte Carlo tree search algorithm (MCTS), and the treatment preference protocols include bidding strategies, opponent preferences and acceptance strategies, for guiding DAs and PAs to make reasonable decisions during the negotiation process;
[0080] Utility evaluation module: used to calculate the satisfaction degrees of both doctors and patients with the treatment preference protocols, and the indicators of the satisfaction degrees include the combination of ASV (The combination of ASV, abbreviated as CASV) and the difference of ASV (The disparate of ASV, abbreviated as DASV);
[0081] Negotiation control module: used to control the negotiation process and obtain the final medical decision based on the treatment preference protocol corresponding to the highest value of the satisfaction degree indicator. The negotiation process includes negotiation start, negotiation rounds, negotiation end and result output, ensuring the efficient and orderly progress of the negotiation process.
[0082] As can be seen from the above, the Monte Carlo tree search negotiation framework (MCTSAN) uses the Monte Carlo tree search algorithm (MCTS) to generate favorable treatment preference protocols. This innovative method can quickly find the optimal or near-optimal solution among a large number of possible negotiation strategies, significantly improving the negotiation efficiency;
[0083] Guided by the bidding strategy, opponent preference, and acceptance strategy, the doctor agent (DA) and patient agent (PA) can make more reasonable and efficient decisions during the negotiation process, further improving the overall quality of the decisions.
[0084] By leveraging agent-based negotiation technology, this framework can not only enhance the overall satisfaction of patients and healthcare providers in decision-making but also promote social fairness by minimizing satisfaction differences.
[0085] Specifically, the expression of the trapezoidal fuzzy membership function is as follows:
[0086]
[0087] where U(S) represents the overall satisfaction value, and M i (S) represents the i-th membership degree of the solution S, which can be directly obtained from a group of doctors and patients and can flexibly and efficiently represent their preferences for certain issues; w i represents the weight of the i-th issue; n represents the number of issues that doctors need to negotiate and the i-th question.
[0088] Specifically, the Monte Carlo tree search algorithm expands new nodes based on the Upper Confidence Bound applied to Trees (UCT). The conditions for expanding new nodes are as follows:
[0089]
[0090] where n p is the number of times the parent node has been simulated, n c is the number of its child nodes, and α is a parameter of the model (usually in the range 0 < α < 1); if the result is not to expand a new node, the selected node is the maximizing node i:
[0091]
[0092] where n is the total number of simulations of the tree, n i is the number of simulations of node i in the tree, s i is the score of node i, and C is also a parameter of the model.
[0093] Specifically, the goal of the bidding strategy is to predict what proposal the opponent will make in the x * th round. Using Gaussian process regression, this method generates a Gaussian distribution, whose mean corresponds to the value predicted by the algorithm, and the standard deviation corresponds to the uncertainty caused by the model;
[0094] The first step of the Gaussian process regression is to calculate the covariance matrix K, which represents the proximity between the turning points (x i ) i∈[1,n] of the sequence based on the covariance function (also known as the kernel k); let
[0095]
[0096] Then, calculate the predicted proposal x * for the turn and the distance between the previous turn in the vector K * :
[0097] K * =(k(x * , x1), …, k(x * , x n ))
[0098] Gaussian process regression relies on the assumption that all these values are in dimensions of multivariate Gaussian; using the results of multivariate Gaussian, calculate
[0099]
[0100] where K ** =k(x * , x * ) corresponds to a Gaussian random variable with a mean of and a standard deviation of σ * ;
[0101] A major aspect of Gaussian process regression is the choice of kernel. The most common ones are the radial basis function (RBF), rational quadratic function (RQF), Matern kernel, and exponential squared sine (ESS). These kernels are used to define the distance between bargaining rounds, and the agent uses the rational quadratic function (RQF).
[0102] Specifically, in the SDM, the opponent preference uses the overall satisfaction value U(S) to measure the agent's satisfaction with the negotiation content;
[0103] Obtain the overall satisfaction function U(S t ) of the opponent in the t-th round, and adopt the first assumption as the opponent's problem weight w i and the second assumption as the fuzzy membership function M i (S t );
[0104] Therefore, assume w i and M i (S t ) respectively. The first assumption involves the problem weight w i , where a set of all possible weight matrices H w can be set first, and then real numbers are calculated and associated with the weight assumption using the following linear function:
[0105]
[0106] Among them, is the sorting of the weights w j in the hypothesis h i , and n is the number of negotiation issues.
[0107] Specifically, the second hypothesis involves the opponent's fuzzy membership function M i (S t ), and the preferences of DA and PA are regarded as membership functions;
[0108] Assign a membership degree to each hypothesis in the hypothesis space and model the fuzzy membership function as a probability distribution;
[0109] By associating various fuzzy membership function hypotheses with their corresponding probabilities , approximate the shape of the true fuzzy membership function of the i-th negotiation issue;
[0110] Generally, during the negotiation process, both parties must make more or less concessions to reach an agreement. Therefore, assume that the agent adopts a time-related strategy based on concessions; according to this strategy, the agent starts with the bid with the highest total satisfaction and turns to its reservation value when approaching the negotiation deadline;
[0111] To predict the possibility of the opponent's bid, use the following formula to calculate the conditional probability P(h j ∣S t ), as follows:
[0112]
[0113] U'(S t ) = U'(S t-1 ) - c(t)
[0114] Among them, P(h j ∣S t ) represents the probability that the opponent will make a bid S j under the condition of h t , U(S t ∣h j ) is the satisfaction value when the opponent also makes a bid S j under the hypothesis h t , U'(S t ) is the satisfaction value of the opponent's expected next bid; the function c(t) is the negotiation concession strategy assumed to be used by the opponent;
[0115] The expected value of the shape of the fuzzy membership function is calculated as follows:
[0116]
[0117] Among them, is the j-th hypothesis of problem i, and the fuzzy membership function generates the satisfaction value of the bid, indicating the probability of.
[0118] Use to represent the hypothesis of the weight value of problem i according to hypothesis j, and the related weight value, that is, The expected value of the weight is calculated as follows:
[0119]
[0120] Finally, based on the bid, calculate the total satisfaction S expected by the opponent t , as follows:
[0121]
[0122] For each problem, it is necessary to standardize the weights and probability distributions of the evaluation function:
[0123]
[0124] These two formulas indicate that for each problem i (where i = 1, 2,..., n), the possibilities or of all different hypotheses j sum to 1. Among them, indicates that the sum of all possibilities related to the satisfaction value hypothesis is 1; indicates that the sum of all possibilities related to the weight value hypothesis is 1.
[0125] Specifically, the acceptance strategy is based on utility, time, or a combination of both, and the acceptance conditions of the agent are as follows:
[0126]
[0127] Among them represents the current acceptance condition, that is, based on the parameter r and the opponent's proposal to judge the next action;
[0128] The output action can be: End: End the negotiation;
[0129] Accept: Accept the opponent's proposal;
[0130] Propose a new counteroffer (proposal);
[0131] r: The current negotiation round, which is an integer variable used to identify the current negotiation round.
[0132] t represents the time pressure threshold, which changes dynamically as the negotiation round r increases and is used to determine whether to accept the current offer. It is defined as:
[0133]
[0134] Where:
[0135] α: The time weight parameter, usually a constant, and 0 ≤ α ≤ 1, which controls the influence of the initial time threshold on t;
[0136] β: The growth rate parameter, which controls the non-linear growth or decay of t with the negotiation round r;
[0137] N max : The maximum number of negotiation rounds, representing the total number of negotiation rounds;
[0138] The proposal made by the opponent in the r-th round, usually representing the value of the issue in the proposed solution by the opponent in this round;
[0139] Represents the counteroffer made by us in the previous round (r - 1);
[0140] The opponent's proposal in the current round r The utility value for ourselves;
[0141] Our proposal in the negotiation round r - 1 The utility value.
[0142] As can be seen from the above, through the user interface module, the system can clearly display the information in the medical decision-making negotiation process, including the treatment preferences, negotiation progress and results of both doctors and patients, which greatly promotes the transparent communication and mutual understanding between doctors and patients;
[0143] The co-decision module adopts the Belief-Desire-Intention (BDI) architecture and the trapezoidal fuzzy membership function, which can simulate the complex communication process between doctors and patients, realize more accurate and user-friendly shared decision-making (SDM), thereby improving the scientific nature of decision-making and patient satisfaction;
[0144] The utility evaluation module can calculate the satisfaction degrees of both doctors and patients with the negotiation results. This quantitative index helps to objectively evaluate the effect of the negotiation process and provides a basis for subsequent improvement;
[0145] At the same time, this module also calculates social welfare indicators such as the combined ASV (CASV) and the difference in ASV (DASV). These indicators can reflect the impact of the negotiation results on the overall social welfare and help to achieve more fair and sustainable medical decisions.
[0146] Example 2:
[0147] Reference Figures 2 to 4 As shown, a comparative experiment will be conducted to effectively verify the performance of the proposed SDM negotiation framework. It will be compared with two existing frameworks: the Fuzzy Constraint-based Agent Negotiation Framework (FCAN) and the Agent-based Negotiation Framework using Fuzzy Constraints and Genetic Algorithms (ANFGA). Both of these frameworks are used for bilateral and multi-issue preference negotiation in SDM. It should be noted that there are relatively few frameworks that use agent-based methods to solve the preference negotiation problem in SDM. Therefore, the comparison scope of the present invention is limited.
[0148] Experimental Setup
[0149] Environment: The program is written in Java and runs on IntelliJ IDEA on the Windows 10 operating system. In addition, the present invention also proposes an SDM negotiation framework implemented in the open-source negotiation software Genius.
[0150] Among them, r represents the current negotiation round, N max represents the negotiation deadline, and t represents the negotiation time loss. α and β are constants; 1 > β > 0 and 0 ≤ α ≤ 1. Each agent executes 200 times in different experiments.
[0151] Dataset: The experiment uses the preference data on the treatment plans for childhood asthma collected by Lin et al. The preference negotiation issues involved in their treatment plans mainly include cost, effectiveness, side effects, risk, convenience, etc. The preference data includes the value preference and weight preference of the issues. The DA and PA preference data in the experiment are shown in Table 2. Table 2 gives an example of preference data containing 5 issues, which is used as input to more specifically demonstrate the model process. As shown in Table 2, the preference consists of the satisfaction function and weight of each issue, and the problem domain is below the issue name. The satisfaction of the participants with the value of each issue is determined by the trapezoidal membership function, which is represented by a quadruple. For example, the satisfaction of PA with the issue "Cost" is "(2, 3.5, 4, 6) F , that is, the acceptable treatment cost range for PA is 2000 - 6000 yuan, and the most willing to accept the treatment is 3500 - 4000 yuan. In addition, the weight is represented by a decimal number from 0 to 1, reflecting the importance of the issue to the participants. For this case, PA values the risk degree of the treatment plan the most, while DA values the side effects more.
[0152] Table 1: Preference Data of DA and PA
[0153]
[0154] Table 1
[0155] Performance Metrics
[0156] The proposed SDM negotiation framework is evaluated according to the following performance metrics.
[0157] Combined ASV (The combination of ASV, CASV): This metric represents the sum of Avg.ASV doctor and Avg.ASV patient and represents the social welfare of doctors and patients. It can be calculated as follows.
[0158]
[0159] CASV = Avg.ASV doctor + Avg.ASV patient
[0160] where Agr total represents the number of negotiation protocols. represents the solution reached in the i-th negotiation. U DA and U PA represent the doctor's ASV function and the patient's ASV function respectively, and are calculated by formula 1.
[0161] Disparity of ASV (The disparate of ASV, DASV): This metric represents the difference between Avg.ASV doctor and Avg.ASV patient The agreed solution may be satisfactory to both parties rather than extreme. It can be calculated as follows.
[0162] DASV = |Avg.ASV doctor + Avg.ASV patient |
[0163] Average number of negotiation rounds (Avg.R): This metric represents the average number of negotiation rounds.
[0164]
[0165] where R i represents the number of rounds required for the doctor-patient pair to reach an agreement on negotiation process i
[0166] Experimental results
[0167] To simulate the negotiation scenarios in different solution spaces, each negotiation was limited to 20 rounds, and comparative experiments were conducted under the conditions where the number of problems was set to 1, 3, 5, 7, 9. The results are shown in Table 2 and Figure 2 , Figure 3 and Figure 4 as follows:
[0168]
[0169]
[0170] Table 2
[0171] Figure 2 Shows the comparative experiments of the MCTSAN framework proposed by the present invention with the FCAN and ANFGA models adopting different negotiation strategies in terms of the average CASV index. It can be observed that although when the number of problems is set to 3, the average CASV of MCTSAN is lower than that of ANFGA adopting the competition strategy, the average CASV of MCTSAN is generally superior to that of FCAN and ANFGA. In addition, as the number of problems increases, the advantage of MCTSAN becomes more obvious. These results emphasize that the MCTSAN model can maintain a high level of social welfare in the case of the increasing complexity of the negotiation scenario, thereby improving the overall satisfaction of patients and healthcare providers with the decision-making. In addition, MCTSAN also shows adaptability to different strategies and scenarios.
[0172] Figure 3 Is the comparative experiment of the MCTSAN model adopting different strategies with the FCAN and ANFGA models on the average DASV index. In the case of a small number of problems, MCTSAN can maintain a low average DASV. As the number of problems increases, especially when the number of problems is 7 and 9, the average DASV of FCAN is lower than that of MCTSAN. However, the difference between MCTSAN and FCAN is within 0.01. Therefore, overall, as the complexity of the negotiation scenario increases, MCTSAN can effectively reduce the satisfaction gap between patients and healthcare providers, thereby promoting social fairness.
[0173] Figure 4 Shows the comparative experiments of the MCTSAN model with the FCAN and ANFGA models adopting different strategies in terms of the Avg.R index. When the negotiation deadline is 20 rounds, MCTSAN can always reach an agreement within this period. As the number of issues increases, the negotiation time required by MCTSAN becomes longer, which reflects that both parties need additional time to reach a consensus on all issues. However, the total number of negotiation rounds is still between 6 and 15 rounds, significantly lower than that of ANFGA.
[0174] As can be seen from the above results, as the number of problems increases, the negotiation time required prolongs, and the overall satisfaction between patients and healthcare providers decreases. The increase in the number of problems expands the solution space, making the search process more challenging. In addition, both parties need more time to reach a consensus on all issues. However, compared with FCAN and ANFGA, MCTSAN maintains better and more stable performance in a large solution space.
[0175] As can be seen from the above, in various scenarios with different numbers of problems, in terms of the average CASV and DASV metrics, the proposed MCTSAN model is always superior to the existing models (such as FCAN and ANFGA). Specifically, even when the complexity of the negotiation scenario increases, MCTSAN can maintain a high level of social welfare and reduce the satisfaction gap between patients and healthcare providers. The model can maintain excellent performance and stability even in a large solution space, which highlights its potential as a valuable tool for implementing SDM in a real medical environment. By leveraging agent-based negotiation techniques, the framework can not only improve the overall satisfaction of patients and healthcare providers in decision-making but also promote social fairness by minimizing satisfaction differences.
[0176] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A medical decision-making negotiation system based on Monte Carlo tree search, characterized in that Including: User interface module: used to receive and display information in the medical decision-making negotiation process, where the information includes the treatment preferences, negotiation progress and results of both doctors and patients; Shared decision-making module: used to construct the behavior models of doctor agents and patient agents based on the Belief-Desire-Intention (BDI) architecture and using trapezoidal fuzzy membership functions, for simulating the communication and shared decision-making process between doctors and patients to obtain negotiation results; Monte Carlo tree search negotiation framework: used to generate a treatment preference protocol based on the information and negotiation results using the Monte Carlo tree search algorithm, where the treatment preference protocol includes a bidding strategy, opponent preference and acceptance strategy, for guiding the doctor agent and the patient agent to make reasonable decisions in the negotiation process; Utility evaluation module: used to calculate the satisfaction degrees of both doctors and patients with the treatment preference protocol, where the indicators of the satisfaction degree include the difference between the combined ASV and ASV; Negotiation control module: used to control the negotiation process and obtain the final medical decision based on the treatment preference protocol corresponding to the highest value of the satisfaction degree indicator, where the negotiation process includes negotiation start, negotiation rounds, negotiation end and result output.
2. The medical decision-making negotiation system based on Monte Carlo tree search according to claim 1, wherein The expression of the trapezoidal fuzzy membership function is: Among them, U(S) represents the overall satisfaction value, and M i (S) represents the i-th membership degree of solution S, and w i represents the weight of the i-th issue, and n represents the total number of issues that doctors and patients need to negotiate.
3. The medical decision-making negotiation system based on Monte Carlo tree search according to claim 1, wherein The Monte Carlo tree search algorithm expands new nodes based on the upper confidence tree, where the conditions for expanding new nodes are: where n p is the number of times the parent node is simulated, n c is the number of its child nodes, and α is a parameter of the model with a value range of 0 < α < 1; if the result is not to expand new nodes, the selected node is the node i that maximizes: where n is the total number of simulations of the tree, n i is the number of simulations of node i in the tree, s i is the score of node i, and C is a parameter of the model.
4. The medical decision-making negotiation system based on Monte Carlo tree search according to claim 1, characterized in that, The goal of the bidding strategy is to obtain the predicted value of x * for the round; The step for obtaining the predicted value is: generating a Gaussian distribution and using Gaussian process regression in the Gaussian distribution for prediction; The calculation process of the Gaussian process regression includes: Calculate the covariance matrix K, which represents the proximity between the turning angles (x i ) i∈[1,n] of the sequence based on the covariance function; Let Calculate the predicted proposal x * The turn of and the vector K * The distance between the previous turn in K * = (k(x * , x1), …, k(x * , x n )) The predicted values on which the Gaussian process regression depends are all assumptions of the dimensions of a multivariate Gaussian; using the results of the multivariate Gaussian to calculate where K ** = k(x * , x * ) the result corresponds to a Gaussian random variable with a mean of and a standard deviation of σ * .
5. The medical decision-making negotiation system based on Monte Carlo tree search according to claim 2, wherein In the communication and shared decision-making of the opponent preference, the overall satisfaction value U(S) is used to measure the satisfaction degree of the agent with the negotiation content; Obtain the overall satisfaction function U(S of the opponent in the t-th round t ), and adopt the setting that the first hypothesis is the problem weight w of the opponent i and the second hypothesis is the fuzzy membership function M i (S t ); The problem weight w i By setting a set of all possible weight matrices H w , calculating a real number and using the following linear function to associate it with the weight hypothesis as follows: Among them, is the sorting of the weights w j in the hypothesis h i , and n is the number of negotiation topics.
6. The medical decision-making negotiation system based on Monte Carlo tree search according to claim 5, characterized in that, The fuzzy membership function M i (S t ) in which the preferences of the doctor agent and the patient agent are regarded as membership functions; Assigning a membership degree to each hypothesis in the hypothesis space and modeling the fuzzy membership function as a probability distribution; By associating various fuzzy membership function assumptions with their corresponding probabilities relate to approximate the shape of the true fuzzy membership function for the i-th negotiation issue; Calculate the probability \(P(h\) of the opponent's bid using the following formula j |S t ), as follows: U'(S t ) = U'(S t-1 ) - c(t) Among them, P(h j ∣S t ) represents the probability that the other party will make an offer S j under the condition of h t . U(S t ∣h j ) is the satisfaction value when the other party also makes an offer S j under the assumption of h t . U'(S t ) is the satisfaction value of the next offer expected by the other party, and the function c(t) is the negotiation concession strategy assumed to be used by the other party; Fuzzy membership function The expected value of the shape is calculated as follows: Among them, is the j-th hypothesis of problem i, and the fuzzy membership function generates the satisfaction value of the bid, represents the probability of; Use to represent the hypothesis regarding the weight value of problem i according to hypothesis j, and the value of the relevant weight, that is, The expected value of the weight is calculated as follows: Using the following formula, calculate the opponent's expected total satisfaction S based on the bid t : For each problem, the following formula is used to standardize the weights and probability distributions of the evaluation function: Among them, indicates that the sum of all possibilities related to the satisfaction value hypothesis is 1; indicates that the sum of all possibilities related to the weight value hypothesis is 1.
7. The medical decision-making negotiation system based on Monte Carlo tree search according to claim 1, characterized in that The acceptance strategy is based on utility, time or a combination of utility and time, where the conditions for the acceptance strategy of the agent are as follows: Among them represents the current acceptance condition, that is, based on the parameter r and the opponent's proposal judge the next action; Accept: Accept the opponent's proposal; Make a new counter-offer, proposal; r: the current negotiation round, an integer variable used to identify the current negotiation round; t: represents the time pressure threshold, which changes dynamically as the negotiation round r increases and is used to decide whether to accept the current offer, defined as: Where: α: the time weight parameter, usually a constant and 0 ≤ α ≤ 1, controlling the influence of the initial time threshold on t; β: the growth rate parameter, controlling the non-linear growth or decay of t with the negotiation round r; N max : The maximum number of negotiation rounds, representing the total number of negotiation rounds; The proposal put forward by the opponent in the r-th round usually represents the value of the issue in the plan proposed by the opponent in this round; Denote the counteroffer made by us in the previous round (r - 1); The opponent's proposal in the current round r The utility value for oneself; The utility value of the proposal made by oneself in negotiation round r-1 .