A traditional Chinese medicine personalized prescription decision-making method and system based on multi-agent game

CN122842840APending Publication Date: 2026-09-29BEIJING NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610868684.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-16
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0008]无法显式处理群体经验与个体需求的冲突,现有方法要么完全依赖群体统计规律,如主题模型、GNN,要么完全依赖个体输入,如 LLM 微调,缺乏一个能够在“普遍证候-治法-方剂体系”与“患者特异性禁忌、兼症、体质”之间进行显式博弈与权衡裁决的框架;

Benefits of technology

[0057]1.本发明所述基于多智能体博弈的中医个性化处方决策方法及系统,通过群体智能体生成群体适配方剂 、个体智能体提出个体化修改 、协调智能体裁决的博弈机制,能够在恪守群体经验规律的前提下,精准识别并移除对特定患者存在风险或不适宜的药物;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122842840A_ABST
    Figure CN122842840A_ABST
Patent Text Reader

Abstract

The application discloses a traditional Chinese medicine personalized prescription decision-making method and system based on multi-agent game, and multiple rounds of alternating games are carried out in the order of individual agent proposal, coordination agent decision, group agent proposal, coordination agent decision and environment evaluation, the individual agent is responsible for proposing a prescription modification proposal based on the individualized symptoms of a patient with single-ingredient medicine as the granularity, the group agent is responsible for maintaining that the prescription does not deviate from the syndrome type and treatment method system, the coordination agent makes a proposal decision based on a differentiated multi-dimensional scoring function, and an environment evaluation model outputs a comprehensive score and a reward signal from four dimensions of treatment method layer similarity, symptom coverage, monarch-minister-attendant-ingredient structure similarity and life quality improvement degree, the application adopts double-layer reinforcement learning under a centralized training and distributed execution framework to carry out training, and outputs a personalized prescription suitable for a patient and a complete interpretable game decision trajectory, and realizes individual adaptive game based on group experience rules.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and traditional Chinese medicine information technology, specifically relating to a method and system for personalized prescription decision-making in traditional Chinese medicine based on multi-agent game theory. Background Technology

[0002] Traditional Chinese medicine prescription decision-making is a core part of clinical diagnosis and treatment. Its essence is to determine the syndrome and treatment method based on the patient's symptoms, signs, tongue and pulse information under the guidance of the principle of "differentiation of syndromes and treatment", and then combine them to form a personalized herbal prescription. With the rapid development of artificial intelligence technology, researchers have tried to use data mining, machine learning and even large language models to realize intelligent TCM prescription recommendation in order to assist clinical decision-making.

[0003] Yao et al. proposed a topic modeling approach for traditional Chinese medicine prescriptions (Yao L, Zhang Y, Wei B, et al. A topic modeling approach for traditional Chinese medicine prescriptions. IEEE TKDE, 2018, 30(6): 1007-1021.). This model treats prescriptions as "documents" and symptoms and herbs as "words". By introducing the roles of "principal, assistant, adjuvant, and guide" and the relationship of "herb compatibility", it simulates the generation process of "theory, method, prescription, and medicine" in traditional Chinese medicine. The PTM (prescription topic model) can mine the implicit topics between syndrome, symptoms, and herbs from a large-scale prescription database and realize symptom → herb recommendation and herb → symptom prediction. However, this type of method is still essentially a bag-of-words model, which is difficult to capture the nonlinear and complex relationship between symptoms and herbs. It is also sensitive to the sparsity of prescription data and cannot explicitly handle the individual differences between different patients under the same syndrome.

[0004] To overcome the limitations of topic models, researchers have begun to utilize graph neural networks (GNNs) to model symptom-herb bipartite graphs, symptom co-occurrence graphs, and herb co-occurrence graphs. For example, Jin et al. proposed a syndrome-aware herbrecommendation with multi-graph convolution network (Jin Y, Zhang W, He X, et al. Syndrome-aware herbrecommendation with multi-graph convolution network. ICDE, 2020: 145-156.), which uses a multi-graph convolution network to learn the embedding representations of symptoms and herbs separately, and then uses a multilayer perceptron to fuse the symptom set into a latent "syndrome" vector before recommending herbs. Yang et al. proposed a multi-layer information fusion-based knowledge-driven herbrecommendation recommendation model based on graph convolutional networks (Yang Y, Rao Y, Yu M, et al. Multi-layer information fusion based on graph convolutional network for knowledge-driven herbrecommendation. Neural Networks, 2022, 146: 1-10.), further introducing a herbal knowledge graph (including properties, functions, targets, etc.), and using a graph neural network with multi-layer information fusion to improve the quality of herbal embedding, resulting in improved recommendation accuracy; Zhou et al. proposed a deep neural network integrating phenotype and molecule (Zhou W, Yang K, Zeng J, et al. FordNet: Recommending traditional Chinese medicine formula via deep neural network integrating phenotype and molecule. Pharmacological Research, 2021, 173: 105752.), which for the first time combined macroscopic symptom information with microscopic molecular networks (herbs-compounds-targets), used convolutional neural networks to extract diagnostic descriptive features, fused molecular information through network embedding, and used hierarchical sampling for data augmentation;

[0005] In recent years, large language models (LMs) have also been applied in the field of traditional Chinese medicine. For example, Lingdan (Hua R, Dong X, Wei Y, et al. Lingdan: enhancing encoding of traditional Chinese medicine knowledge for clinical reasoning tasks with large language models. JAMIA, 2024, 31(9): 2019-2029.) conducted large-scale continuous pre-training on ancient Chinese medicine books, textbooks, medical records and pharmacopoeias based on Baichuan2-13B, and made fine adjustments for the tasks of recommending traditional Chinese medicine preparations and prescriptions. Zhongjing (Yang S, Zhao H, Zhu S, et al. Zhongjing: Enhancing the Chinese medical capabilities of large language model through expert feedback and real-world multi-turn dialogue. AAAI, 2024, 38(17): 19368-19376.) first introduced reinforcement learning from human feedback (RLHF) into the large language model of traditional Chinese medicine and constructed 7 The CMtMedQA dataset, containing tens of thousands of real multi-turn doctor-patient dialogues, has achieved good results in terms of multi-turn consultations and security. The TCMLLM-PR (Tian H, Yang K, Dong X, et al. TCMLLM-PR: evaluation of large language models for prescription recommendation in traditional Chinese medicine. 2024.) implements prescription recommendation based on ChatGLM-6B through instruction fine-tuning and has been validated in cross-dataset transfer.ChatGLM-FGIDs-TCM (Zheng W, Tong Y, Gao T, et al. ChatGLM-FGIDs-TCM: A knowledge-integrated large language model for clinical decision support in traditional Chinese medicine. BIBM, 2025.) combines Retrieval Enhanced Generation (RAG) with a traditional Chinese medicine knowledge graph to achieve knowledge-enhanced question answering and reasoning for functional gastrointestinal disorders.

[0006] PresRecST (Dong X, Zhao C, Song X, et al. PresRecST: a novel herbalprescription recommendation algorithm for real-world patients withintegration of syndrome differentiation and treatment planning. JAMIA, 2024,31(6): 1268-1279.) proposes a progressive prescription recommendation model that decomposes the TCM clinical decision-making process into four sequential steps: "symptom collection → syndrome differentiation → treatment method determination → herbal recommendation". It adopts a residual neural network to fuse knowledge graph embedding step by step and realizes multi-task joint prediction of syndrome, treatment method, and herbal medicine. This method is superior to the traditional end-to-end model in terms of interpretability and task structure consistency. However, PresRecST is still essentially a deterministic forward reasoning model and lacks explicit modeling and game adjudication mechanism for conflicts between different decision-making subjects.

[0007] In summary, existing TCM prescription decision-making technologies, including topic models, graph neural networks, large language models, and progressive reasoning models, generally suffer from the following technical shortcomings:

[0008] The existing methods either rely entirely on group statistical patterns, such as topic models and GNNs, or rely entirely on individual inputs, such as LLM fine-tuning. They lack a framework that can make explicit decisions and trade-offs between the "general syndrome-treatment-prescription system" and "patient-specific contraindications, comorbidities, and constitution".

[0009] The decision-making process is opaque and lacks interpretability. Most models are black box outputs, failing to record the reasons, scoring criteria, and adjudication process for each drug modification, thus failing to meet the stringent requirements of traceable and auditable decision-making in medical settings.

[0010] The action granularity is too coarse, making fine adjustment difficult. Existing methods usually output the entire prescription at once or recommend multiple drugs at once, lacking a fine-grained modification mechanism based on "single drugs", which does not conform to the practice of "adding and subtracting according to symptoms" in TCM clinical practice.

[0011] The existing evaluation indicators are too simplistic and lack multi-dimensional evaluation of prescription quality. Most existing works use recommendation system indicators such as Precision@K, Recall@K, and F1@K, but fail to comprehensively evaluate prescription quality from multiple dimensions such as consistency of treatment methods, symptom coverage, the structure of principal, assistant, and adjuvant herbs, and patient quality of life. They also fail to effectively integrate evaluation signals into model training. Based on the above-mentioned technical problems of the existing technology, this invention provides a TCM personalized prescription decision-making method and system based on multi-agent game theory. Summary of the Invention

[0012] To address the aforementioned technical problems in the existing technology, this invention provides a method and system for personalized prescription decision-making in traditional Chinese medicine based on multi-agent game theory.

[0013] The present invention adopts the following technical solution:

[0014] This invention provides a method for personalized prescription decision-making in Traditional Chinese Medicine based on multi-agent game theory, comprising:

[0015] Step 1: Obtain the patient's clinical information, including chief complaint, present illness, four diagnostic methods, past medical history, and treatment stage information;

[0016] Step 2: Based on the patient's clinical information, the patient's syndrome type and treatment method are determined by querying the TCM knowledge graph and classic prescription database, and a traditional typical prescription is selected as the initial prescription for the game.

[0017] Step 3: The individual intelligent agent proposes prescription modification suggestions based on the patient's individual clinical information at the granular level of single herbs. The individual intelligent agent cannot see the patient's syndrome type and treatment information. The four diagnostic methods include observation, auscultation and olfaction, inquiry and palpation, which are the core basis for TCM syndrome differentiation and treatment.

[0018] Step 4: The coordinating agent receives the proposal from the individual agents, calls the multi-dimensional scoring function to score, and accepts the proposal when the score reaches the corresponding preset threshold; otherwise, it rejects the proposal. The coordinating agent has complete visibility of all clinical information.

[0019] The scoring functions for individual intelligent agents' proposals include: symptom demand matching degree, indirect coverage detection, prescription structure stability, and treatment deviation degree;

[0020] Step 5: The swarm intelligence agent proposes prescription modification suggestions at the granular level of single drugs based on the syndrome type and treatment method information. The swarm intelligence agent cannot see the four diagnostic methods information and past medical history information.

[0021] Step 6: The coordinating agent receives the proposal from the group of agents, calls the multi-dimensional scoring function to score it, and accepts the proposal when the score reaches the corresponding preset threshold; otherwise, it rejects the proposal. The coordinating agent has complete visibility of all clinical information.

[0022] The scoring function for the proposals of the swarm intelligence agent includes: evidence-based rationality, consistency of treatment methods, stability of prescription structure, and necessity;

[0023] Step 7: After each round of the game, the environmental assessment model quantitatively evaluates the current prescription from four dimensions: similarity of treatment methods, symptom coverage, similarity of the principal, assistant, and adjuvant structures, and degree of improvement in quality of life, and outputs a comprehensive score and reward signal.

[0024] Step 8: When the preset termination condition is met, terminate the game and output the generated personalized prescription and the complete game decision trajectory.

[0025] Furthermore, the sequential game rules described in the above steps are as follows: using the traditional typical prescription as the initial prescription for the game, and following the alternating order of individual agent proposal, coordinating agent decision, group agent proposal, coordinating agent decision, and environmental assessment, each agent can propose a maximum of a preset number of modification proposals per round.

[0026] Furthermore, in step 4, the comprehensive score for the individual agent's proposal is shown in equation (1):

[0027] (1),

[0028] in, For symptom-needs matching degree, For indirect coverage detection, The stability of the formula structure The degree of deviation of the treatment method, , , , For the preset weighting coefficients, when Accept the offer if it is made, otherwise refuse it;

[0029] The comprehensive score for the group's proposal is shown in Equation (2):

[0030] (2),

[0031] in, For the sake of evidence-based rationality, E2 represents consistency of treatment methods. For the structural stability of the prescription, E4 is necessary. , , , For the preset weighting coefficients, when Accept the offer if it is made, otherwise refuse.

[0032] Furthermore, in step 4, the coordinating agent generates a decision reason vector after each decision and sends it to the individual agent and the group agent respectively. The individual agent and the group agent update their respective belief state representations based on the decision reason vectors, and the belief state representations serve as the conditional inputs for generating the next proposal.

[0033] Furthermore, in step 7, the four-dimensional assessment of the environmental assessment model includes:

[0034] Treatment method similarity score: Based on the treatment method tree of TCM clinical diagnosis and treatment terminology, the semantic similarity between the current prescription treatment method and the labeled treatment method is calculated using the Wu-Palmer similarity algorithm;

[0035] Symptom coverage score: The large language model is used to identify the coverage of the patient's symptoms by the current prescription. A two-way verification mechanism of positive inquiry and negative verification is used to prevent hallucinations. The ratio of the number of covered symptoms to the total length of the symptom list is calculated.

[0036] Similarity scoring of the monarch-minister-assistant-envoy structure: The large language model is called to divide the current prescription and the tagged prescription into monarch-minister-assistant-envoy categories respectively, and the Jaccard similarity of the four levels of monarch, minister, assistant and envoy is calculated respectively, and the sum is calculated according to the preset weights;

[0037] Quality of life improvement rating: Five-level rating based on four dimensions: fatigue, sleep, pain, and digestion, with the arithmetic mean taken.

[0038] Furthermore, in step 7, the reward signal includes: a global reward as shown in equation (3):

[0039] (3) Used to train the coordinating agent, wherein This represents the reward value for the coordinating agent in this round of the game. This represents the environmental score for this round of prescriptions. This represents the environmental score for the previous prescription.

[0040] The local rewards are shown in equations (4) and (5):

[0041] (4) If the proposal is accepted, the reward value is positive; otherwise, the reward is negative, where β represents the decision threshold for the individual agent's proposal. I If the value is ≥β, the proposal is accepted; otherwise, it is rejected. (5) If the proposal is accepted, otherwise a negative penalty is imposed, which is used to train individual agents and the swarm agent respectively. γ represents the decision threshold for the swarm agent's proposal. If the value is ≥γ, the proposal is accepted; otherwise, it is rejected.

[0042] The overall score of the environmental model is shown in equation (6):

[0043] (6),

[0044] in, To score the similarity of the treatment methods, Assess symptom coverage score. Score the similarity of the monarch-minister-assistant-envoy structure. The degree of symptom relief was scored.

[0045] Furthermore, the policy networks of the swarm intelligence agents, individual intelligence agents, and coordinating intelligence agents are trained using a centralized training and distributed execution framework. The Critic network accesses the global state during the training phase, while the Actor network relies only on its local observations during the execution phase.

[0046] Furthermore, in step 8, the termination conditions include: if neither the individual agent nor the group agent makes any new proposals for two consecutive rounds, termination is triggered; or if the preset maximum number of game rounds is reached, termination is triggered; or if the comprehensive score of the environmental assessment model exceeds a preset threshold, all termination signals should be triggered by the environment.

[0047] The game decision trajectory includes: initializing a typical prescription, modification proposals and reasons from individual agents in each round, modification proposals and reasons from the collective agents in each round, multi-dimensional scoring and adjudication results of the coordinating agent for each proposal, comprehensive score and multi-dimensional score of the environmental assessment model after each round of the game, and the final personalized prescription.

[0048] This invention also provides a personalized TCM prescription decision-making system based on multi-agent game theory, comprising:

[0049] Based on the patient's clinical information, we consulted the TCM knowledge graph and classic formula database to determine the patient's syndrome type and treatment method, and selected a traditional typical formula as the initial formula for the game.

[0050] The individual intelligent agent module proposes prescription modification suggestions at the granular level of a single drug, based on the patient's individual clinical information.

[0051] The swarm intelligence module proposes prescription modification suggestions at the granular level of single herbs, based on patient syndrome and treatment information.

[0052] The coordinating agent module is used to receive proposals from individual agent modules and / or group agent modules, call the corresponding multi-dimensional scoring function to score them, and make a decision to accept or reject them.

[0053] The environmental assessment module is used to quantitatively evaluate prescriptions from four dimensions: similarity of treatment methods, symptom coverage, similarity of the principal, assistant, adjuvant, and guide structures, and degree of improvement in quality of life, and outputs a comprehensive score and reward signal.

[0054] The game control module is used to control the individual intelligent agent module and the group intelligent agent module to play multiple rounds of alternating games according to the preset order game rules, and to terminate the game when the termination condition is met.

[0055] The output module is used to output the generated personalized prescription and the complete game decision trajectory.

[0056] Compared with the prior art, the superior effects of the present invention are as follows:

[0057] 1. The TCM personalized prescription decision-making method and system based on multi-agent game theory described in this invention, through a game mechanism in which a group of intelligent agents generate group-appropriate prescriptions, individual intelligent agents propose individualized modifications, and coordinated intelligent agents make decisions, can accurately identify and remove drugs that are risky or unsuitable for specific patients while adhering to the rules of group experience.

[0058] 2. The TCM personalized prescription decision-making method and system based on multi-agent game theory described in this invention uses an environmental model to quantitatively evaluate prescriptions from four dimensions: treatment method, function, drug, and quality of life. The comprehensive score is improved from 5.5 to 8.72, which fully demonstrates that this invention can significantly improve the individual suitability and expected clinical efficacy of prescriptions while maintaining consistency in treatment methods and structural stability.

[0059] 3. The TCM personalized prescription decision-making method and system based on multi-agent game theory described in this invention fully records the proposals, scores, decision reasons, and environmental model evaluation results of each round of the game. All decisions are verifiable and form a traceable audit trail, which can be reviewed by doctors, used for teaching analysis, and for tracing the source of medical disputes. This solves the fundamental problem that existing "black box" models cannot explain the source of prescriptions.

[0060] 4. The TCM personalized prescription decision-making method and system based on multi-agent game theory described in this invention does not depend on specific syndrome types or diseases. It can be transferred to any TCM internal medicine, surgery, gynecology, pediatrics, and other departments for prescription decision-making tasks. Only adjustments need to be made to the knowledge graph, typical prescription library, contraindication rule library, and environmental assessment indicators to make it applicable to different specialty scenarios. Experiments have shown good results in various complex syndrome types such as metastatic pancreatic cancer and cardiovascular comorbidities, demonstrating the system's versatility and robustness. Attached Figure Description

[0061] Figure 1 This is a schematic diagram of the TCM personalized prescription decision-making method based on multi-agent game theory as described in this embodiment of the invention;

[0062] Figure 2 This is a schematic diagram of the multi-dimensional evaluation of the environmental model in the TCM personalized prescription decision-making method based on multi-agent game theory described in this embodiment of the invention. Detailed Implementation

[0063] To better understand the above-mentioned objectives, features and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.

[0064] Example

[0065] like Figure 1 As shown, the TCM personalized prescription decision-making method based on multi-agent game theory includes:

[0066] Step 1: Obtain and break down the patient's clinical information, including chief complaint, present illness, four diagnostic methods, past medical history, and treatment stage information;

[0067] Step 2: Based on the patient's clinical information, the patient's syndrome type and treatment method are determined by querying the TCM knowledge graph and classic prescription database, and a traditional typical prescription is selected as the initial prescription for the game.

[0068] Step 3: The individual intelligent agent proposes prescription modification suggestions based on the patient's individual clinical information at the granular level of single herbs. The individual intelligent agent cannot see the patient's syndrome type and treatment information. The four diagnostic methods include observation, auscultation and olfaction, inquiry and palpation, which are the core basis for TCM syndrome differentiation and treatment.

[0069] Step 4: The coordinating agent receives the proposal from the individual agents, calls the multi-dimensional scoring function to score, and accepts the proposal when the score reaches the corresponding preset threshold; otherwise, it rejects the proposal. The coordinating agent has complete visibility of all clinical information.

[0070] The scoring functions for individual intelligent agents' proposals include: symptom demand matching degree, indirect coverage detection, prescription structure stability, and treatment deviation degree;

[0071] Step 5: The swarm intelligence agent proposes prescription modification suggestions at the granular level of single drugs based on the syndrome type and treatment method information. The swarm intelligence agent cannot see the four diagnostic methods information and past medical history information.

[0072] Step 6: The coordinating agent receives the proposal from the group of agents, calls the multi-dimensional scoring function to score it, and accepts the proposal when the score reaches the corresponding preset threshold; otherwise, it rejects the proposal. The coordinating agent has complete visibility of all clinical information.

[0073] The scoring function for the proposals of the swarm intelligence agent includes: evidence-based rationality, consistency of treatment methods, stability of prescription structure, and necessity;

[0074] Step 7: After each round of the game, the environmental assessment model quantitatively evaluates the current prescription from four dimensions: similarity of treatment methods, symptom coverage, similarity of the principal, assistant, and adjuvant structures, and degree of improvement in quality of life, and outputs a comprehensive score and reward signal.

[0075] Step 8: When the preset termination condition is met, terminate the game and output the generated personalized prescription and the complete game decision trajectory.

[0076] Furthermore, the sequential game rules described in the above steps are as follows: using a traditional typical prescription as the initial prescription for the game, and following an alternating sequence of individual agent proposal, coordinating agent decision, group agent proposal, coordinating agent decision, and environmental evaluation, each agent can propose a maximum of a preset number of modification proposals per round; wherein the individual agent is responsible for adjusting the prescription according to the patient's personalized symptoms, the group agent is responsible for maintaining the prescription within the main syndrome type and treatment system, and the two agents propose to the coordinating agent in turn, and the coordinating agent makes a proposal decision based on comprehensive evaluation and environmental feedback, wherein the environmental model is responsible for evaluating the prescription during the game process and providing feedback signals to the coordinating agent to support its further decision-making;

[0077] The collective intelligence agent is positioned as a representative of collective experience. Its visible information (observation space) includes TCM syndrome type, TCM treatment method, chief complaint, present medical history, and treatment stage. Its action space consists of single-step operations: add herb, remove herb, ensuring that the prescription conforms to the syndrome type-treatment method system and maintaining the collective pattern.

[0078] The individual intelligent agent is positioned as an individual adaptation representative. The visible information (observation space) includes the chief complaint, present medical history, four diagnostic methods, past medical history, and treatment stage. The action space is a single-step operation: add herb, remove herb. Based on the group's rules, the agent improves the coverage and safety of prescriptions for individual symptoms.

[0079] The coordinating agent is positioned as an evidence-based decision-maker (coordinator). The visible information (observation space) is all clinical information (including the sum of the visible information of the group agent and individual agents). The action space is the output of accept or reject for each proposal of the group agent or individual agent. It balances the group rules and individual needs to maximize the quality of the overall prescription.

[0080] The environmental model is positioned as an environmental assessment. The visible information (observation space) includes all patient information, syndrome differentiation and treatment methods, real prescription labels, and current prescription. The action space is no action. It only outputs assessment / reward signals, calculates multi-dimensional comprehensive scores, and provides global rewards.

[0081] Information visibility constraints: Individual agents cannot see syndrome differentiation and treatment information, ensuring they focus on individual suitability; group agents cannot see individual diagnostic information and past medical history information, simulating the limitations of the group knowledge base; coordinating agents have complete information to help them make decisions after comprehensive consideration; the environment model has complete information and real prescription labels. This design simulates the division of roles in real clinical practice, ensuring the rationality of the game.

[0082] The coordinating agent uses different multi-dimensional scoring functions for proposals from individual agents and group agents, and sets decision thresholds for each.

[0083] Furthermore, in step 4, the comprehensive score for the individual agent's proposal is shown in equation (1):

[0084] (1),

[0085] in, To assess the symptom-needs match, whether the newly added drugs address the patient's actual symptoms; To indirectly cover the detection, whether the target symptoms are already covered by existing medications; The stability of the formula structure, whether it disrupts the principal, assistant, adjuvant, and guide structure (relative to the basic formula). To determine the degree of deviation of the treatment method, modify whether to weaken the main treatment method (i.e., whether it deviates from the group's pattern). , , , Recommended weights are those with preset weighting coefficients: Decision threshold ;when Accept the offer if it is made, otherwise refuse it;

[0086] The comprehensive score for the group's proposal is shown in Equation (2):

[0087] (2),

[0088] in, To ensure evidence-based rationality, the knowledge graph evidence chain of whether a drug has a syndrome type → treatment method → ​​drug is modified; E2 represents the consistency of treatment method, which is the semantic similarity between the drug efficacy vector and the target treatment method. E4 represents the stability of the prescription structure, specifically whether to restore or maintain the principal, assistant, adjuvant, and guide structure of the basic prescription; E5 represents necessity, specifically the degree to which the deviation between the modified prescription and the standard group prescription is reduced. , , , Recommended weights are those with preset weighting coefficients: Decision threshold ,when Accept the offer if it is made, otherwise refuse.

[0089] Furthermore, in step 4, the coordinating agent generates a decision reason vector after each decision and sends it to the individual agent and the group agent respectively. The individual agent and the group agent update their respective belief state representations based on the decision reason vectors, and the belief state representations serve as the conditional inputs for generating the next proposal.

[0090] Furthermore, such as Figure 2 As shown, in step 7, after each round of the game, the environmental model receives the current prescription sent by the coordinating agent, scores it from four dimensions, with each dimension having a maximum score of 10 points, and finally outputs a comprehensive score (0-10 points) and a corresponding reward signal. The four-dimensional evaluation of the environmental assessment model includes:

[0091] The similarity score at the treatment method level, with a weight of 0.3, is based on the treatment method tree of TCM clinical diagnosis and treatment terminology. LLM is called to match the current prescription to the most relevant node in the treatment method tree and find the node with the labeled treatment method. The semantic similarity between the current prescription treatment method and the labeled treatment method is calculated using the Wu-Palmer similarity algorithm. The score = similarity × 10.

[0092] Symptom coverage score, weighted at 0.2, uses a large language model to identify the coverage of the patient's symptoms by the current prescription. It prevents hallucinations through a two-way verification mechanism of positive inquiry and negative verification. The score is calculated as the ratio of the number of covered symptoms to the total length of the symptom list. The score is calculated as symptom coverage rate × 10.

[0093] The similarity score of the principal, assistant, adjuvant, and guide structures is scored with a weight of 0.2. The large language model is used to classify the current prescription and the tagged prescription into principal, assistant, adjuvant, and guide structures, and the Jaccard similarity of the principal, assistant, adjuvant, and guide structures is calculated separately. The similarity is then weighted and summed according to the preset weights. The total similarity is calculated as follows: total similarity = 0.5 × principal drug similarity + 0.3 × assistant drug similarity + 0.1 × adjuvant drug similarity + 0.1 × guide drug similarity. The score is calculated as total similarity × 10.

[0094] The assessment of the degree of improvement in quality of life, with a weight of 0.3, is based on a five-level scale across four dimensions: fatigue, sleep, pain, and digestion. Each dimension is rated on a scale of 0, 2.5, 5, 7.5, and 10. The scoring criteria are: 0 (no relief), 2.5 (indirect effect), 5 (mild relief), 7.5 (basic relief), and 10 (complete relief). The overall score is the arithmetic mean of the four dimensions (each with a weight of 0.25).

[0095] Furthermore, in step 7, the reward signal includes: a global reward as shown in equation (3):

[0096] (3) Used to train the coordinating agent, wherein This represents the reward value for the coordinating agent in this round of the game. This represents the environmental score for this round of prescriptions. This represents the environmental score for the previous prescription.

[0097] The local rewards are shown in equations (4) and (5):

[0098] (4) If the proposal is accepted, the reward value is positive; otherwise, the reward is negative, where β represents the decision threshold for the individual agent's proposal. I If the value is ≥β, the proposal is accepted; otherwise, it is rejected. (5) If the proposal is accepted, otherwise a negative penalty is imposed, which is used to train individual agents and the swarm agent respectively. γ represents the decision threshold for the swarm agent's proposal. If the value is ≥γ, the proposal is accepted; otherwise, it is rejected.

[0099] The overall score of the environmental model is shown in equation (6):

[0100] (6),

[0101] in, To score the similarity of the treatment methods, Assess symptom coverage score. Score the similarity of the monarch-minister-assistant-envoy structure. The degree of symptom relief was scored.

[0102] Furthermore, the policy networks of the swarm intelligence agents, individual intelligence agents, and coordinating intelligence agents are trained using a centralized training and distributed execution framework. The Critic network accesses the global state during the training phase, while the Actor network relies only on its local observations during the execution phase.

[0103] Furthermore, in step 8, the termination conditions include: the individual agent and the group agent have no new proposals for two consecutive rounds, triggering termination; or the preset maximum number of game rounds is reached, triggering termination; or the comprehensive score of the environmental assessment model exceeds a preset threshold and is triggered by the coordinating agent, and all termination signals should be triggered by the environment.

[0104] The game decision trajectory includes: initializing a typical prescription, modification proposals and reasons from individual agents in each round, modification proposals and reasons from the collective agents in each round, multi-dimensional scoring and adjudication results of the coordinating agent for each proposal, comprehensive score and multi-dimensional score of the environmental assessment model after each round of the game, and the final personalized prescription.

[0105] This invention also provides a personalized TCM prescription decision-making system based on multi-agent game theory, comprising:

[0106] Based on the patient's clinical information, we consulted the TCM knowledge graph and classic formula database to determine the patient's syndrome type and treatment method, and selected a traditional typical formula as the initial formula for the game.

[0107] The swarm intelligence module is used to query the TCM knowledge graph and classic prescription database based on the patient's syndrome type and treatment information, and generate a group-appropriate prescription.

[0108] The individual intelligent agent module is used to propose prescription modification suggestions based on individual patient clinical information, using single drugs as the granularity.

[0109] The swarm intelligence module proposes prescription modification suggestions at the granular level of single herbs, based on patient syndrome and treatment information.

[0110] The coordinating agent module is used to receive proposals from individual agent modules and / or group agent modules, call the corresponding multi-dimensional scoring function to score them, and make a decision to accept or reject them;

[0111] The environmental assessment module is used to quantitatively evaluate prescriptions from four dimensions: similarity of treatment methods, symptom coverage, similarity of the principal, assistant, adjuvant, and guide structures, and degree of improvement in quality of life, and outputs a comprehensive score and reward signal.

[0112] The game control module is used to control the individual intelligent agent module and the group intelligent agent module to play multiple rounds of alternating games according to the preset order game rules, and to terminate the game when the termination condition is met.

[0113] The output module is used to output the generated personalized prescription and the complete game decision trajectory.

[0114] Furthermore, the knowledge graph construction module uses a graph attention network to embed the TCM knowledge graph, and the policy network input of the swarm intelligence module includes dynamically updated drug node embedding vectors. The embedding vectors and policy network parameters are optimized synchronously during joint training.

[0115] The following is a detailed explanation of the implementation process of this invention, using a real and complex clinical case: a 78-year-old postoperative pancreatic cancer patient. This embodiment focuses on demonstrating the first round of the game, reflecting the core idea of ​​"individual adaptation game based on group rules".

[0116] Case input (JSON format):

[0117] {

[0118] "patient": {

[0119] "gender": "female",

[0120] "age": 78,

[0121] "chief_complaint": "Malignant pancreatic tumor was discovered more than 3 months ago, and surgery was performed more than 3 months ago."

[0122] "current_symptoms": {

[0123] "appetite": "poor appetite, low food intake",

[0124] "nausea_vomiting": "Significant nausea and vomiting due to food nausea".

[0125] "abdomen": "Persistent abdominal discomfort",

[0126] "energy": "significant fatigue"

[0127] "extremities": "numbness in hands and feet (significant in both feet)",

[0128] "sleep": "sleep is possible",

[0129] "stool": "Stool is on the dry side"

[0130] },

[0131] "Inspection": "Slightly dark complexion, listless"

[0132] "tongue": "Pale and dark tongue with a white and slippery coating",

[0133] "pulse": a wiry and hesitant pulse.

[0134] "listening_smelling": "short of breath and reluctant to speak",

[0135] "past_history": "Decades-long history of hypertension; underwent a hysterectomy ten years ago."

[0136] },

[0137] "syndrome": "Liver and spleen qi stagnation syndrome, qi stagnation and blood stasis syndrome, and toxin and blood stasis syndrome",

[0138] Treatment principle: Harmonize the Shao Yang meridian, clear heat from the bowels, invigorate blood circulation, and detoxify.

[0139] "base_formula": "Da Chai Hu Tang, Pi Ji Wan",

[0140] "initial_formula": {

[0141] "herbs": ["Bupleurum", "Scutellaria", "Pinellia", "Fresh Ginger", "Immature Bitter Orange", "White Peony Root", "Rhubarb", "Jujube", "Magnolia Bark", "Coptis", "Evodia Fruit", "Poria", "Alisma", "Amomum Fruit", "Ginseng", "Atractylodes Rhizome", "Dried Ginger", "Aconitum Carmichaelii", "Croton Seed", "Glycyrrhiza"]

[0142] }

[0143] }

[0144] Initialization: The swarm intelligence generates a swarm-fitting agent.

[0145] Based on the syndrome types "liver and spleen qi stagnation syndrome, qi stagnation and blood stasis syndrome, and toxic blood stasis syndrome" and the treatment principles "harmonizing the lesser yang, clearing the bowels and purging heat, and promoting blood circulation and detoxifying", the swarm intelligence agent queries the TCM knowledge graph and classic prescription database to generate a group-appropriate prescription, which is a modified version of Da Chai Hu Tang combined with Pi Ji Wan, containing a total of 20 herbs. This prescription represents the group's empirical rules and serves as the anchor point for subsequent individual adaptation games.

[0146] First round of game – Individual adaptation phase:

[0147] The individual intelligent agent receives all the patient's individual information, including advanced age of 78, history of hypertension, postoperative deficiency of vital energy, nausea and vomiting, fatigue, numbness in the hands and feet, dry stools but no actual heat, pale and dark tongue with white and slippery coating, etc. It evaluates the initial prescription and makes modification suggestions in order of priority. Each suggestion only operates on one drug.

[0148] Individual agent suggestion 1: The patient has a history of hypertension for decades. Aconitum carmichaelii is pungent, hot and dispersing, which may raise blood pressure and deplete yin and blood. Suggestion: Remove Aconitum carmichaelii.

[0149] Coordinating agent evaluation (Score_I calculation):

[0150] Symptom-needs match: The patient has no corresponding indication, and removal does not affect efficacy → Ssymptom=1.0.

[0151] Indirect coverage detection: No alternative drugs → Sredundancy=1.0.

[0152] Formula structure stability: Aconitum carmichaelii is not the principal ingredient; its removal does not affect the core structure → E3=1.0.

[0153] Deviance of treatment method: Aconitum carmichaelii is not necessary for harmonizing Shaoyang and promoting blood circulation and detoxification. Remove the treatment method that does not deviate from the treatment method → ​​Stherapy=1.0.

[0154] ScoreI = 0.4 × 1.0 + 0.25 × 1.0 + 0.2 × 1.0 + 0.15 × 1.0 = 1.0 ≥ 0.60 → Accept.

[0155] Updated prescription: Remove Aconitum carmichaelii.

[0156] Individual agent suggestion 2: The patient's tongue coating is white and slippery, without Yangming bowel obstruction and heat accumulation. Rhubarb, being bitter and cold, is not suitable for purging and may further damage the middle Yang. Suggestion: Remove rhubarb.

[0157] Coordinating agent evaluation: Reject.

[0158] Updated prescription: None.

[0159] Individual agent suggestion 3: The patient has obvious nausea and vomiting, but ginseng can cause stagnation of qi, which may hinder the stomach and worsen the fullness. Suggestion: Remove ginseng.

[0160] Coordinating agent evaluation: Accepted.

[0161] Update prescription: Remove ginseng.

[0162] After the first round of Phase I, the prescription has removed Aconitum carmichaelii and ginseng, leaving 18 herbs: Bupleurum chinense, Scutellaria baicalensis, Pinellia ternata, Zingiber officinale, Citrus aurantium, Paeonia lactiflora, Ziziphus jujuba, Rheum palmatum, Magnolia officinalis, Coptis chinensis, Evodia rutaecarpa, Poria cocos, Alisma plantago-aquatica, Amomum villosum, Atractylodes macrocephala, Zingiber officinale, Croton tiglium, and Glycyrrhiza uralensis.

[0163] First round of game – the group rule maintenance phase:

[0164] The swarm intelligence examines the current prescription and finds that although the individual matching is reasonable, it deviates significantly from the treatment principles of "promoting blood circulation and detoxifying" and "eliminating masses and dispersing nodules," and lacks key drugs for the syndrome of mutual accumulation of toxins and blood stasis. In accordance with the requirements of the treatment principles, the swarm intelligence proposes to add drugs in turn, with each proposal involving only one drug.

[0165] Swarm intelligence suggestion 1: The patient has pancreatic cancer, and the core pathogenesis is the mutual accumulation of toxins and blood stasis. It is necessary to add drugs that clear heat and detoxify, disperse nodules, and eliminate cancer. Suggestion: Add *Hedyotis diffusa* (white flower snake tongue grass).

[0166] Coordinating agent evaluation: Acceptance

[0167] Updated prescription: Add Hedyotis diffusa.

[0168] Swarm intelligence suggestion 2: Medications are needed to address qi stagnation and blood stasis, including those that promote blood circulation, relieve stagnation, and alleviate pain. Suggestion: Add Curcuma zedoaria.

[0169] Coordinating agent evaluation: Accepted.

[0170] Updated prescription: Add Curcuma zedoaria.

[0171] Swarm intelligence suggestion 3: The patient experiences significant shortness of breath, reluctance to speak, and fatigue, requiring tonification of Qi and support without causing stagnation. Suggestion: Add Astragalus membranaceus.

[0172] Coordinating agent evaluation: Reject.

[0173] Updated prescription: None.

[0174] After the first round of the collective intelligence phase, the prescription was updated to: Bupleurum, Scutellaria baicalensis, Pinellia ternata, fresh ginger, Citrus aurantium, Paeonia lactiflora, jujube, rhubarb, Magnolia officinalis, Coptis chinensis, Evodia rutaecarpa, Poria cocos, Alisma plantago-aquatica, Amomum villosum, Atractylodes macrocephala, dried ginger, Croton tiglium, Glycyrrhiza uralensis, Hedyotis diffusa, Curcuma zedoaria (20 ingredients in total).

[0175] Environmental model assessment after the first round:

[0176] The coordinating agent sends the current prescription to the environmental model for multi-dimensional scoring. The total score is 7.56, while the initial basic prescription has a score of about 5.5 due to the presence of contraindicated drugs. The ΔScore is +2.06, which provides a positive reward signal to the coordinating agent.

[0177] Subsequent rounds and the end of the game:

[0178] After the first round, individual agents still had unproposed individualized suggestions, such as adding Albizia bark, Tribulus terrestris, Medicated Leaven, and Costus root, as well as replacing Citrus aurantium with Citrus aurantium peel. The swarm agents also had unproposed rule additions, such as Lonicera japonica, Pear root, Turtle shell, Costus root, and Citrus reticulata peel. The system then entered the second and third rounds of the game, continuing to iterate at a rate of 3 suggestions from each individual agent and the swarm agent per round.

[0179] Game termination condition: After the fourth round, if neither the individual agents nor the group agents make any new proposals (the prescription has stabilized), and the overall score of the environment model does not improve significantly for two consecutive rounds, the system terminates the game.

[0180] Final personalized prescription output:

[0181] The final personalized prescription (19 ingredients in total): Bupleurum, Scutellaria baicalensis, Pinellia ternata, fresh ginger, Citrus aurantium (replacing Citrus aurantium), Paeonia lactiflora, jujube, Albizia julibrissin bark, Tribulus terrestris, Medicated leaven, Aucklandia lappa, Citrus reticulata peel, Curcuma zedoaria, Hedyotis diffusa, Lonicera japonica, Astragalus membranaceus, Pyrus pyrifolia root, Trionyx sinensis shell, Cannabis sativa seed.

[0182] To quantify the effect, let's take a 78-year-old pancreatic cancer patient after surgery as an example:

[0183] The initial group formula contained 20 herbs, of which 13 herbs, including Aconitum carmichaelii, Croton tiglium, Rheum palmatum, Coptis chinensis, Evodia rutaecarpa, Zingiber officinale, Ginseng, Atractylodes macrocephala, Magnolia officinalis, Poria cocos, Alisma plantago-aquatica, Amomum villosum, and Glycyrrhiza uralensis, were deemed unsuitable by the individual intelligent agent (risk of hypertension, stomach damage due to bitter and cold nature, purgative effects on the body's vital energy, and stagnation in the stomach). During the game, this invention successfully removed all 8 high-risk / clearly contraindicated herbs (Aconitum carmichaelii, Croton tiglium, Rheum palmatum, Coptis chinensis, Evodia rutaecarpa, Zingiber officinale, Ginseng, and Atractylodes macrocephala) and replaced or deleted the remaining unsuitable herbs (such as replacing Citrus aurantium with Citrus aurantium peel). The final prescription completely avoided herbs that conflicted with the patient's individual characteristics such as hypertension, advanced age, deficiency of vital energy, and vomiting, with a contraindication violation rate of 0%.

[0184] Meanwhile, based on individual symptoms, this invention adds 12 targeted drugs, including Hedyotis diffusa, honeysuckle, Curcuma zedoaria, pear root, turtle shell, Astragalus membranaceus, Costus root, green tangerine peel, medicated leaven, Albizia bark, Tribulus terrestris, and hemp seed, which significantly improves the coverage of the pathogenesis of "toxicity and blood stasis" and symptoms such as "fatigue, nausea, and numbness in the hands and feet".

[0185] The environmental model quantifies the prescription from four dimensions: treatment method, function, medication, and quality of life (out of 10). In the above case:

[0186] The initial base score for the similarity of the treatment layer was 6.5, and the final game-theoretic prescription score was 8.8, representing an improvement of +35.4%.

[0187] The initial baseline score for the functional layer (symptom coverage) was 3.3, and the final game-theoretic prescription score was 8.0, representing an improvement of +142%.

[0188] The initial base prescription score for the drug layer (principal, assistant, adjuvant, and guide structure) was 5.2, and the final game prescription score was 8.4, representing an improvement of +61.5%.

[0189] The initial baseline score for the quality of life layer was 4.5, and the final game theory prescription score was 8.5, representing an improvement of +88.9%.

[0190] The initial base score was 5.5, and the final game theory score was 8.72, representing an improvement of +58.5%.

[0191] The overall score improved from 5.5 to 8.72, demonstrating that the invention can significantly improve the individual suitability of prescriptions and expected clinical efficacy while maintaining consistency in treatment methods and structural stability.

[0192] The final prescription generated by this invention was compared with the final prescription issued by a senior TCM doctor in the real world for the same patient: the real prescription contained 18 drugs, while the prescription of this invention contained 19 drugs.

[0193] Common ingredients: Bupleurum, Scutellaria baicalensis, Pinellia ternata, fresh ginger, Citrus aurantium, Paeonia lactiflora, jujube, Albizia julibrissin bark, Tribulus terrestris, Medicated leaven, Aucklandia lappa, Citrus reticulata peel, Curcuma zedoaria, Hedyotis diffusa, Lonicera japonica, Astragalus membranaceus, Pyrus pyrifolia root, and Trionyx sinensis (17 of the 18 ingredients overlap; only one ingredient in the actual prescription is not included in the system, while this invention adds Hemp seed for lubricating the intestines and relieving constipation).

[0194] The overlap of core drugs was 94.4% (17 / 18). The only difference was the selection of individualized laxatives. The addition of hemp seed met the needs of patients with dry stools and was clinically reasonable.

[0195] The results demonstrate that the present invention can effectively reproduce the clinical decision-making logic of experienced TCM doctors and has high clinical practical value.

[0196] This invention fully records the proposals, scores, decision reasons, and environmental model evaluation results for each round of the game. For example, regarding the decision to "remove Aconitum carmichaelii," the coordinating agent's reasoning is: "The patient is elderly and has hypertension. Aconitum carmichaelii is pungent, hot, and dispersing, which may raise blood pressure and deplete yin and blood. Safety takes priority, so this decision is accepted." All decisions are documented and traceable, forming a traceable audit trail that can be reviewed by doctors, used for teaching analysis, and for tracing the source of medical disputes, thus solving the fundamental problem that existing "black box" models cannot explain the source of prescriptions.

[0197] This invention employs a two-layer reinforcement learning reward structure within the CTDE framework: local rewards provide dense, immediate feedback for both individual and swarm agents. Where β represents the decision threshold for an individual agent's proposal; if ScoreI ≥ β, the proposal is accepted; otherwise, it is rejected. γ represents the decision threshold for a group of agents' proposals. If the value is ≥γ, the proposal is accepted; otherwise, it is rejected. Reward I and reward G are the average scores given by the coordinating agent to each proposal of the individual agent and the group agent, respectively. The experiment shows that the average score of the individual agent's proposal is 0.93 and the average score of the group agent's proposal is 0.95, indicating that both can learn efficiently and propose reasonable modifications.

[0198] Global rewards guide the coordinating agent to learn the optimal decision-making strategy. Reward_H represents the reward given to the coordinating agent by the environment model in each round. Specifically, it is the difference between the current round and the previous round's rating of the current prescription. After 1000 training episodes, the coordinating agent's decision accuracy (consistency with expert judgment) increased from the initial 68% to 92%, and the global decision quality converged steadily.

[0199] Training process overview: The quality of life layer uses a dataset containing tens of thousands of real cases, prescription labels, and doctor modification trajectories. It employs the CTDE framework and the MAPPO algorithm for two-stage training.

[0200] CTDE (Centralized Training with Decentralized Execution) is a paradigm designed specifically for multi-agent reinforcement learning (MARL) to address the challenge of learning by a single agent in a dynamic, partially observable environment.

[0201] MAPPO (Multi-Agent Proximal Policy Optimization) is a successful upgrade and extension of the single-agent reinforcement learning algorithm PPO (Proximal Policy Optimization).

[0202] PPO is a very popular algorithm in the field of single-agent reinforcement learning. Its core is to ensure that there are no overly aggressive "mutations" when updating decision policies, thereby ensuring the stability of learning.

[0203] MAPPO extends the robust and efficient characteristics of PPO from a single AI to multi-AI collaborative scenarios. Its core idea is to retain the performance advantages of PPO and expand its application scope, enabling multiple intelligent agents to work collaboratively like a well-trained team.

[0204] 1. Behavioral cloning pre-training: Directly supervise the policy networks of individual agents and swarm agents by modifying trajectories using doctors' prescriptions.

[0205] 2. Reinforcement Learning Fine-Tuning: Using the environment model to calculate changes in the overall score as... , The reward value for agent H in each round of the environment model is the difference between the coordinating agent's decision and the threshold. I and reward G reward I The reward is the value given by the coordinating agent to the individual agents in each round. G The policy networks of the coordinating agent, individual agents, and swarm agents are updated alternately to provide the reward value for the coordinating agent to the swarm agent in each round.

[0206] This invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the invention as claimed. The scope of protection of this invention is defined by the appended claims.

Claims

1. A method for personalized prescription decision-making in Traditional Chinese Medicine based on multi-agent game theory, characterized in that, include: Step 1: Obtain and break down the patient's clinical information, including chief complaint, present illness, four diagnostic methods, past medical history, and treatment stage information; Step 2: Based on the patient's clinical information, the patient's syndrome type and treatment method are determined by querying the TCM knowledge graph and classic prescription database, and a traditional typical prescription is selected as the initial prescription for the game. Step 3: The individual intelligent agent proposes prescription modification suggestions based on the patient's individual clinical information at the granular level of single drugs. The individual intelligent agent cannot see the patient's syndrome type and treatment information. The four diagnostic methods include inspection, auscultation and olfaction, inquiry and palpation. Step 4: The coordinating agent receives the proposal from the individual agents, calls the multi-dimensional scoring function to score, and accepts the proposal when the score reaches the corresponding preset threshold; otherwise, it rejects the proposal. The coordinating agent has complete visibility of all clinical information. The scoring functions for individual intelligent agents' proposals include: symptom demand matching degree, indirect coverage detection, prescription structure stability, and treatment deviation degree; Step 5: The swarm intelligence agent proposes prescription modification suggestions at the granular level of single drugs based on the syndrome type and treatment method information. The swarm intelligence agent cannot see the four diagnostic methods information and past medical history information. Step 6: The coordinating agent receives the proposal from the group of agents, calls the multi-dimensional scoring function to score it, and accepts the proposal when the score reaches the corresponding preset threshold; otherwise, it rejects the proposal. The coordinating agent has complete visibility of all clinical information. The scoring function for the proposals of the swarm intelligence agent includes: evidence-based rationality, consistency of treatment methods, stability of prescription structure, and necessity; Step 7: After each round of the game, the environmental assessment model quantitatively evaluates the current prescription from four dimensions: similarity of treatment methods, symptom coverage, similarity of the principal, assistant, and adjuvant structures, and degree of improvement in quality of life, and outputs a comprehensive score and reward signal. Step 8: When the preset termination condition is met, terminate the game and output the generated personalized prescription and the complete game decision trajectory.

2. The method for personalized TCM prescription decision-making based on multi-agent game theory according to claim 1, characterized in that, The sequential game rules are as follows: using traditional typical prescriptions as the initial prescriptions for the game, and following the alternating order of individual agent proposals, coordinating agent adjudication, group agent proposals, coordinating agent adjudication, and environmental assessment, each agent can propose a maximum of a preset number of modification proposals per round.

3. The method for personalized TCM prescription decision-making based on multi-agent game theory according to claim 1, characterized in that, In step 4, the comprehensive score for the individual agent's proposal is shown in equation (1): (1), in, For symptom-needs matching degree, For indirect coverage detection, The stability of the formula structure For the degree of deviation of the treatment method, , , , For the preset weighting coefficients, when Accept the offer if it is made, otherwise refuse it; The comprehensive score for the group's proposal is shown in Equation (2): (2), in, For the sake of evidence-based rationality, E2 represents consistency of treatment methods. E4 is necessary for the structural stability of the prescription. , , , For the preset weighting coefficients, when Accept the offer if it is made, otherwise refuse.

4. The TCM personalized prescription decision-making method based on multi-agent game theory according to claim 1, characterized in that, In step 4, the coordinating agent generates a decision reason vector after each decision and sends it to the individual agent and the group agent respectively. The individual agent and the group agent update their respective belief state representations based on the decision reason vectors, and the belief state representations serve as the conditional inputs for generating the next proposal.

5. The method for personalized TCM prescription decision-making based on multi-agent game theory according to claim 1, characterized in that, In step 7, the four-dimensional assessment of the environmental assessment model includes: Treatment method similarity score: Based on the treatment method tree of TCM clinical diagnosis and treatment terminology, the semantic similarity between the current prescription treatment method and the labeled treatment method is calculated using the Wu-Palmer similarity algorithm; Symptom coverage score: The large language model is used to identify the coverage of the patient's symptoms by the current prescription. A two-way verification mechanism of positive inquiry and negative verification is used to prevent hallucinations. The ratio of the number of covered symptoms to the total length of the symptom list is calculated. Similarity scoring of the monarch-minister-assistant-envoy structure: The large language model is called to divide the current prescription and the tagged prescription into monarch-minister-assistant-envoy categories respectively, and the Jaccard similarity of the four levels of monarch, minister, assistant and envoy is calculated respectively, and the sum is calculated according to the preset weights; Quality of life improvement rating: Five-level rating based on four dimensions: fatigue, sleep, pain, and digestion, with the arithmetic mean taken.

6. The method for personalized TCM prescription decision-making based on multi-agent game theory according to claim 1, characterized in that, In step 7, the reward signal includes: the global reward as shown in equation (3): (3) Used to train the coordinating agent, wherein This represents the reward value for the coordinating agent in this round of the game. This represents the environmental score for this round of prescriptions. This represents the environmental score for the previous prescription. The local rewards are shown in equations (4) and (5): (4) If the proposal is accepted, the reward value is positive; otherwise, the reward is negative, where β represents the decision threshold for the individual agent's proposal. I If the value is ≥β, the proposal is accepted; otherwise, it is rejected. (5) If the proposal is accepted, otherwise a negative penalty is imposed, which is used to train individual agents and the swarm agent respectively. γ represents the decision threshold for the swarm agent's proposal. If the value is ≥γ, the proposal is accepted; otherwise, it is rejected. The overall score of the environmental model is shown in equation (6): (6), in, To score the similarity of the treatment methods, Assess symptom coverage score. Score the similarity of the monarch-minister-assistant-envoy structure. The degree of symptom relief is scored.

7. The method for personalized TCM prescription decision-making based on multi-agent game theory according to claim 6, characterized in that, The policy networks of the swarm intelligence agents, individual intelligence agents, and coordinating intelligence agents are trained using a centralized training and distributed execution framework. The Critic network accesses the global state during the training phase, while the Actor network relies only on its local observations during the execution phase.

8. The method for personalized prescription decision-making in traditional Chinese medicine based on multi-agent game theory according to claim 1, characterized in that, In step 8, the termination conditions include: no new proposals from individual agents and group agents for two consecutive rounds, triggering termination; or reaching the preset maximum number of game rounds, triggering termination; or the comprehensive score of the environmental assessment model exceeding the preset threshold, triggering termination. All termination signals should be triggered by the environment. The game decision trajectory includes: initializing a typical prescription, modification proposals and reasons from individual agents in each round, modification proposals and reasons from the collective agents in each round, multi-dimensional scoring and adjudication results of the coordinating agent for each proposal, comprehensive score and multi-dimensional score of the environmental assessment model after each round of the game, and the final personalized prescription.

9. A TCM personalized prescription decision-making system based on multi-agent game theory, applied to the TCM personalized prescription decision-making method based on multi-agent game theory as described in any one of claims 1 to 8, characterized in that, include: Based on the patient's clinical information, we consulted the TCM knowledge graph and classic formula database to determine the patient's syndrome type and treatment method, and selected a traditional typical formula as the initial formula for the game. The individual intelligent agent module proposes prescription modification suggestions at the granular level of a single drug, based on the patient's individual clinical information. The swarm intelligence module proposes prescription modification suggestions at the granular level of single herbs, based on patient syndrome and treatment information. The coordinating agent module is used to receive proposals from individual agent modules and / or group agent modules, call the corresponding multi-dimensional scoring function to score them, and make a decision to accept or reject them. The environmental assessment module is used to quantitatively evaluate prescriptions from four dimensions: similarity of treatment methods, symptom coverage, similarity of the principal, assistant, adjuvant, and guide structures, and degree of improvement in quality of life, and outputs a comprehensive score and reward signal. The game control module is used to control the individual intelligent agent module and the group intelligent agent module to play multiple rounds of alternating games according to the preset order game rules, and to terminate the game when the termination condition is met. The output module is used to output the generated personalized prescription and the complete game decision trajectory.