Large model inquiry system based on multi-dimensional reinforcement learning and Markov probability optimization
Through a large-scale consultation system with multi-dimensional reinforcement learning and Markov probability optimization, the problem of insufficient personalization and intelligence of the existing medical consultation system is solved, personalized diagnosis and cost-benefit analysis are realized, and the timeliness and accuracy of the knowledge base is ensured.
Patent Information
- Application Number
- CN202510004025.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-07-11
AI Technical Summary
The existing medical consultation systems lack personalization and intelligence, making it difficult to effectively utilize multi-dimensional information for dynamic optimization and cost-benefit analysis, and the knowledge base is not updated in time.
A large-model consultation system based on multi-dimensional reinforcement learning and Markov probability optimization is adopted, including feature information extraction module, dynamic optimization module, optimal cost-effective medical decision-making module and knowledge graph update module. Personalized diagnosis and cost prediction are carried out through multiple state prediction models and medical knowledge graphs, and the knowledge base is continuously optimized through feedback mechanisms.
A personalized and intelligent consultation system is realized, which can dynamically optimize diagnosis and treatment decisions, provide high-quality consultation prompt words and cost predictions, and ensure the timeliness and accuracy of the knowledge base.
Smart Images

Figure CN120299669A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of large medical models, and particularly relates to a large model consultation system based on multi-dimensional reinforcement learning and Markov probability optimization. Background Art
[0002] With the booming development of artificial intelligence, knowledge graphs have become a new hot spot in the field of knowledge services and have received extensive attention from scholars and the industrial community at home and abroad. A knowledge graph consists of a schema graph, a data graph, and the relationships between them: the schema graph describes the conceptual level of the human knowledge domain, emphasizing the formal expression of concepts and concept relationships. The nodes in the schema graph are conceptual entities, and the edges are semantic relationships between concepts; the data graph describes the physical world level, emphasizing a series of objective facts. Well-known general knowledge graphs include Google "Knowledge Graph", Sogou "Zhilitang", YAGO, DBpedia, etc., which are characterized by large scale, wide domain, and containing a large amount of common sense.
[0003] Currently, medicine is one of the vertical fields with the widest application of knowledge graphs. For example, the traditional Chinese medicine knowledge graph constructed by Shanghai Shuguang Hospital, the ontology medical knowledge base SNOMED-CT2, and applications such as IBM Watson Health have also come into people's sight in the past two years. The early research on medical question-and-answer systems mainly focused on information retrieval, extraction, and summarization technologies. Terol et al. used two knowledge bases, UMLS and WordNet, set 10 types of medical question types, and used the application of natural language processing technology to generate and process the logical form of questions, and extract answers from the knowledge bases. Abacha et al. compared medical question-and-answer systems based on medical ontologies, combined medical ontologies, domain knowledge, NLP-related technologies, and semantic relationships to implement an automatic medical system.
[0004] Obviously, the introduction of medical knowledge graphs can bring great advantages to medical diagnosis systems, and it can construct rich relationships between diseases and diseases, diseases and symptoms, and symptoms and symptoms at the information level.
[0005] In recent years, with the development of large model technology, there is an urgent need in this field to develop a more intelligent and personalized consultation system based on large models. Summary of the Invention
[0006] In order to implement a more intelligent and personalized consultation system based on large models, the present invention discloses a large model consultation system based on multi-dimensional reinforcement learning and Markov probability optimization, including:
[0007] A feature information extraction module that obtains the feature information distribution of the patient's consultation text information by using multiple state prediction models;
[0008] A dynamic optimization module that obtains patient disease information and patient personalized characteristics based on the distribution of feature information for personalized diagnosis of patients; at the same time, obtains symptom query information based on a medical knowledge graph and uses a multi-dimensional reinforcement learning model to dynamically optimize the patient consultation status to generate high-quality consultation instruction information;
[0009] An optimal cost-benefit medical decision-making module that deeply mines and analyzes the remaining available numbers of the patient's medical insurance, the current symptom manifestations, the available hospital drug inventory, and the effectiveness and feedback of past treatment plans according to the patient's personal basic information (such as age, gender, past medical history, etc.) to achieve multi-dimensional and all-round cost prediction and effect evaluation;
[0010] A knowledge graph update module that continuously updates based on the real-time feedback of patient diagnosis information and the knowledge of current medical research results to ensure that the medical knowledge base always remains up-to-date.
[0011] Preferably, the dynamic optimization module further includes:
[0012] A single prompt word unit that, based on a multi-dimensional reinforcement learning calculation paradigm, selectively traverses relevant disease knowledge graphs and associated massive medical literatures, and conducts heuristic and efficient searches to guide the large model to carry out the best consultation prompt words for high-performance single diagnosis and treatment decisions.
[0013] Preferably, the dynamic optimization module further includes:
[0014] A dynamic consultation prompt word unit that includes a time series optimization model based on a Markov model constructed according to the patient's individualized disease process for dynamic diagnosis and treatment auxiliary decision-making to generate the best consultation prompt words that can guide the large model to carry out high-performance dynamic diagnosis and treatment decisions.
[0015] Preferably, the dynamic optimization module further includes:
[0016] A fusion unit that fuses the single prompt word unit and the dynamic consultation prompt word unit to provide individualized best consultation prompt words and diagnosis and treatment decisions for patients with different disease histories.
[0017] Preferably,
[0018] A feature information extraction module that, for a large amount of text feature information obtained from the consultation record, uses various state prediction models such as feature correlation calculation, emotion library comparison, and statistical information conversion to obtain the feature information distribution of the patient's consultation text information.
[0019] Preferably, the optimal cost-benefit medical decision-making module further includes:
[0020] A data conversion unit, which is used to convert the scattered treatment records of each patient into an ordered and structured action sequence to retain all the details of the patient's treatment process and provide rich data support for subsequent cluster analysis, pattern recognition, and cost prediction.
[0021] Preferably, the optimal cost-benefit medical decision-making module further includes:
[0022] A clustering patient action sequence unit, which discretizes the entire cost range into multiple segments and quantifies the actual cost as the average value of each segment to achieve efficient management and utilization of cost information.
[0023] Preferably, the optimal cost-benefit medical decision-making module further includes:
[0024] An estimation unit, which is used to predict the treatment cost of a patient from the action sequence and provide a forward-looking estimate of future medical expenses by matching the patient action sequence with known treatment patterns.
[0025] Preferably, the knowledge graph update module further includes:
[0026] A feedback and optimization unit, which, through a feedback loop and a self-optimization mechanism, becomes the core driving force for the development of the knowledge graph (KG) and its trust model to build a closed-loop system, aiming to strengthen the accuracy of the knowledge graph and the reliability of the trust model through continuous cyclic iteration.
[0027] The present invention has the following characteristics:
[0028] The system described in the present invention belongs to a knowledge-driven medical interview large model system, which simultaneously understands the patient's disease information and the patient's personalized characteristics according to the distribution of feature information to serve the patient's personalized diagnosis; at the same time, it generates symptom query information based on the existing medical knowledge graph and uses a multi-dimensional reinforcement learning model to dynamically optimize the patient interview state to generate high-quality interview instruction information.
[0029] The core advantage of the present invention lies in its evolutionary design - it has a modular architecture, supports flexible replacement and upgrade of components, and ensures that the technology stack keeps pace with the times; at the same time, it also embeds a sensitive learning mechanism that can learn new knowledge from each medical interaction, realizes self-iteration and intelligent optimization, and thus always remains at the forefront of medicine to provide the most accurate and efficient health solutions for patients. Brief Description of the Drawings
[0030] By reading the detailed description in the following preferred specific embodiments, various other advantages and benefits of the present invention will become clear to those of ordinary skill in the art. The drawings in the specification are only for the purpose of showing the preferred embodiments and are not considered as limiting the present invention. Obviously, the drawings described below are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts. Moreover, throughout the drawings, the same reference numerals are used to represent the same components.
[0031] Figure 1 FIG. is a schematic structural diagram of a large model consultation system based on multi-dimensional reinforcement learning and Markov probability optimization in an embodiment of the present invention;
[0032] Figure 2 FIG. is a schematic workflow diagram of a large model consultation system based on multi-dimensional reinforcement learning and Markov probability optimization in another embodiment of the present invention;
[0033] Figure 3 FIG. is a schematic workflow diagram of cost prediction for patient disease treatment based on a large model in another embodiment of the present invention. Specific Embodiments
[0034] The following will refer to Figures 1 to 3 Specific embodiments of the present invention will be described in detail. Although specific embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0035] It should be noted that certain terms are used in the specification and claims to refer to specific components. Those skilled in the art should understand that technicians may use different names to refer to the same component. The specification and claims do not use the difference in names as a way to distinguish components, but rather use the difference in functions of components as the criterion for distinction. As mentioned throughout the specification and claims, "comprising" or "including" is an open-ended term and should be interpreted as "including but not limited to". The following description of the specification is for the purpose of implementing the preferred embodiments of the present invention, but the description is for the general principle of the specification and is not intended to limit the scope of the present invention. The scope of protection of the present invention shall be defined by the appended claims.
[0036] For ease of understanding of the embodiments of the present invention, the following will further explain with specific embodiments as examples in conjunction with the drawings, and each drawing does not constitute a limitation on the embodiments of the present invention.
[0037] See Figure 1, in one embodiment, the present invention discloses a large model consultation system based on multi-dimensional reinforcement learning and Markov probability optimization, including:
[0038] A feature information distribution module that obtains the feature information distribution of the patient's consultation text information using multiple state prediction models;
[0039] A dynamic optimization module that obtains the patient's disease information and personalized characteristics based on the feature information distribution for personalized diagnosis of the patient; at the same time, obtains symptom query information based on the medical knowledge graph, and uses a multi-dimensional reinforcement learning model to dynamically optimize the patient's consultation status to generate high-quality consultation instruction information;
[0040] An optimal cost-effective medical decision-making module that deeply mines and analyzes the remaining available digital medical insurance of the patient, the current symptom manifestations, the available hospital drug inventory, and the effectiveness and feedback of past treatment plans according to the patient's personal basic information (such as age, gender, past medical history, etc.) to achieve multi-dimensional and all-round cost prediction and effect evaluation; it should be noted that the optimal cost-effective medical decision-making module is a highly intelligent system that integrates the concepts of big data analysis, artificial intelligence algorithms, and personalized medicine, aiming to create an economically feasible, efficient, and accurate treatment plan for each patient;
[0041] A knowledge graph update module that continuously updates based on the real-time feedback of patient diagnosis information and the knowledge of current medical research results to ensure that the medical knowledge base always remains up-to-date; it can be understood that the knowledge graph update module constructs an efficient and dynamic update mechanism, aiming to inject continuous vitality into the medical knowledge base in the large model to ensure the scientificity, timeliness, and accuracy of medical decisions. It is a key component in the fields of medical big data and artificial intelligence, and it plays an important role in ensuring that the medical knowledge base always remains in the most accurate state.
[0042] In another embodiment, the dynamic optimization module further includes:
[0043] A single prompt word unit that, based on the multi-dimensional reinforcement learning calculation paradigm, selectively traverses relevant disease knowledge graphs and associated massive medical literature, and conducts heuristic and efficient searches to guide the large model to generate the best consultation prompt words for high-performance single diagnosis and treatment decisions.
[0044] Specific heuristic examples are as follows:
[0045] Under the calculation paradigm of multi-dimensional reinforcement learning, the design of the single prompt word unit aims to guide the large model to make high-performance single diagnosis and treatment decisions in a heuristic manner by efficiently traversing the disease knowledge graph and associated massive medical literature. The heuristic search mechanism mainly includes the following aspects:
[0046] 1. Context-based Semantic Understanding
[0047] Heuristic search first relies on a deep understanding of the current medical record context, including patient symptoms, past medical history, laboratory test results, etc. Through semantic analysis, the model can identify the most relevant keywords and concepts to the current situation, thus constructing a preliminary search scope and direction.
[0048] 2. Dynamic Exploration of Disease Knowledge Graph
[0049] The disease knowledge graph is a structured knowledge base that contains multi-dimensional information such as diseases, symptoms, treatment plans, drug action mechanisms, etc., as well as the complex relationships between them. Heuristic search will conduct dynamic exploration in the knowledge graph according to the preliminary keywords and concepts, and preferentially access those nodes that are highly relevant to the current situation to obtain more potential diagnostic clues and treatment suggestions.
[0050] 3. Literature Retrieval and Abstract Extraction
[0051] In addition to the knowledge graph, heuristic search will also automatically retrieve relevant medical literature, especially the latest research papers and clinical guidelines. Through natural language processing technology, the model can quickly abstract the key information in the literature, such as newly discovered disease markers, effective treatment plans, side effect warnings, etc., to provide a scientific basis for decision-making.
[0052] 4. Decision Optimization of Multi-dimensional Reinforcement Learning
[0053] Throughout the search process, the multi-dimensional reinforcement learning algorithm continuously evaluates the effectiveness and efficiency of the search path, and optimizes the search strategy through continuous trial and error and feedback. The model will automatically adjust the search depth, breadth, and keyword weights according to past successful experiences to achieve the best search effect.
[0054] 5. Generation of Personalized Prompt Words
[0055] Based on the above search results, the model can generate a set of personalized prompt word lists. These prompt words not only cover the key information of the current medical record but also include new knowledge extracted from the knowledge graph and literature. The selection and sorting of prompt words follow heuristic principles, aiming to guide doctors to quickly locate the core of the problem, reduce ineffective inquiries, and improve the diagnosis and treatment efficiency.
[0056] 6. Real-time Feedback and Learning
[0057] In practical applications, the model will continuously receive feedback from doctors and patients, including inquiry results, diagnosis confirmation, treatment effects, etc. This information will be used to update the knowledge graph and literature index, and at the same time as training data for the reinforcement learning algorithm to continuously optimize the search algorithm and decision-making model.
[0058] Through the above heuristic search mechanism, a single prompt word unit can quickly locate key knowledge in a vast amount of information, assist doctors in making accurate diagnosis and treatment decisions, while reducing medical errors and improving the quality and efficiency of medical services. This mechanism not only reflects the huge potential of artificial intelligence in the medical field but also provides important technical support for the construction of future intelligent medical systems.
[0059] In another embodiment, the dynamic optimization module further includes:
[0060] A dynamic consultation prompt word unit, which includes a time-series optimization model based on a Markov model constructed according to the individual disease progression of the patient for dynamic diagnosis and treatment assistance decision-making, to generate the best consultation prompt words that can guide the large model to carry out high-performance dynamic diagnosis and treatment decisions.
[0061] In another embodiment, the dynamic optimization module further includes:
[0062] A fusion unit, which fuses the single prompt word unit and the dynamic consultation prompt word unit to provide individualized best consultation prompt words and diagnosis and treatment decisions for patients with different disease histories.
[0063] In another embodiment, in the described system,
[0064] A feature information distribution module, which, for a large amount of text feature information obtained from the consultation records, uses various state prediction models such as feature correlation calculation, emotion library comparison, and statistical information conversion to obtain the feature information distribution of the patient's consultation text information.
[0065] In another embodiment, the optimal cost-effective medical decision-making module further includes:
[0066] A data conversion unit, which is used to convert the scattered treatment records of each patient into an ordered and structured action sequence to retain all the details of the patient's treatment process and provide rich data support for subsequent clustering analysis, pattern recognition, and cost prediction.
[0067] In another embodiment, the optimal cost-effective medical decision-making module further includes:
[0068] A clustering patient action sequence unit, which discretizes the entire cost range into multiple segments and quantifies the actual cost as the average value of each segment to achieve efficient management and utilization of cost information.
[0069] In another embodiment, the optimal cost-effective medical decision-making module further includes:
[0070] An estimation unit, which is used to predict the treatment cost of the patient from the action sequence, and by matching the patient action sequence with known treatment patterns, provides a forward-looking estimate of future medical expenses.
[0071] In another embodiment, the knowledge graph updating module further includes:
[0072] The feedback and optimization unit, through feedback loops and self-optimization mechanisms, becomes the core driving force for the development of the knowledge graph (KG) and its trust model, in order to build a closed-loop system, aiming to enhance the accuracy of the knowledge graph and the reliability of the trust model through continuous iteration. This process not only promotes the dynamic update and evolution of medical knowledge, but also ensures that the decision-making system based on these knowledge graphs can continue to adapt to the complex and changing medical environment.
[0073] See also Figure 2 , in another embodiment,
[0074] The large-model medical consultation system based on multi-dimensional reinforcement learning and Markov probability optimization disclosed in the present invention can extract a variety of text feature information based on the patient's current medical consultation record.
[0075] It should be noted that text feature information is a characteristic description of the patient's symptoms and can be used to extract the patient's symptom information block. Even if the present invention does not make any innovations to the medical knowledge graph, the large-model consultation system based on multi-dimensional reinforcement learning and Markov probability optimization disclosed by the present invention can also generate relevant information about diseases and symptoms in combination with the existing medical knowledge graph.
[0076] The large-model medical consultation system based on multidimensional reinforcement learning and Markov probability optimization disclosed in the present invention realizes feature correlation calculation, emotion library comparison, statistical information conversion and symptom completeness prediction on the basis of dynamic programming using multidimensional reinforcement learning, Markov chain and Monte Carlo method, so as to provide individualized optimal medical consultation prompts and treatment decisions for patients with different medical histories, and obtain the patient's diagnosis-related probability.
[0077] The large-model consultation system based on multi-dimensional reinforcement learning and Markov probability optimization disclosed in the present invention is exemplarily used to implement the following method, which includes:
[0078] (1) Modeling the diagnosis-driven interaction problem as a multidimensional Markov decision process:
[0079] The information inclusion rate of various interactive questions answered by patients: confidence level of misdiagnosis, patient's trust in the big model diagnosis, patient's current friendliness, patient's confidence in the diagnosis result, and patient's emotional state.
[0080] Define the 5-tuple (A, P, S, R, γ), and the meaning and function of each of its components, for example,
[0081] The set of S states is regarded as the information inclusion rate of the patient's answers. For example, for the misdiagnosis confidence level, S is considered as different intervals of the misdiagnosis probability, S = {high probability of misdiagnosis, relatively high probability of misdiagnosis, general probability of misdiagnosis, low probability of misdiagnosis, no probability of misdiagnosis}, and the corresponding accuracy intervals are 85%-100%, 80%-85%, 70%-80%, 60%-70%, and 0%-60%;
[0082] A represents the set of information for different aspects extracted from the patient's answers and electronic medical records, depending on demographics, longitudinal and physiological measurements, clinical characteristics, relevant emotional tone words, and the diagnostic relevance vocabulary relationship in the patient's answers. The corresponding action a is the treatment measure taken by the clinician for the patient (for example, prescribing certain drugs, arranging a surgical operation, etc.).
[0083] P represents the transition function. Exemplarily,
[0084] U t |S t = s can be regarded as the patient's own biological system. Given the current health and intervention, the patient will enter the next time step t + 1 from the current time step t, that is, after taking the action a in the current state s, it reaches the state U t |S t = s is the probability set, where U t |S t = s represents the probability distribution on S.
[0085] R represents the reward function. Exemplarily,
[0086] If the information status of the patient in the next answer regarding this aspect is improved, then a reward is assigned through the reward function U t |S t = s, where U t |S t = s represents the probability distribution on R.
[0087] γ represents the discount.
[0088] Furthermore, the large model interrogation system based on multi-dimensional reinforcement learning and Markov probability optimization disclosed by the present invention is committed to: after modeling the diagnosis and treatment-driven interaction problem as a Markov decision, using multi-dimensional reinforcement learning to maximize the reward function, so as to find the optimal diagnosis strategy, that is, to find the expectation of the strategy with the maximum cumulative reward:
[0089] U t |S t = s(1)
[0090] where, U t |S t = s is the strategy π of taking the action a in the s state.
[0091] (2) The probability distribution on S can be learned offline by fusing Monte Carlo and dynamic programming:
[0092] Define the state value function U t |S t = s, which is used to evaluate the quality of a piece of information in a certain aspect during the interaction of the patient. Among them, the value of each state is determined not only by the state of the patient's current information but also by the state of the information feedback made by the subsequent patients.
[0093] Therefore, taking the expectation of the cumulative reward of the state can obtain the state value function U of the current s t |S t = s:
[0094] U t |S t = s(2)
[0095] Among them,
[0096] U t |S t = s means: under the policy π, the state value function of state s, starting from state s and making decisions according to the policy π, the expected cumulative reward that can be obtained.
[0097] Among them,
[0098] U represents the utility function, which is the direct feedback given by the environment at a certain time step t, measuring the immediate benefits or costs obtained after taking a certain action, and reflecting the direct consequences of taking an action in a specific state.
[0099] U t |S t = s represents the immediate reward obtained at time step t, which depends on the state and action at that time. Specifically in your description, it is the patient's feedback on the information at time step t, such as the change in the confidence level of diagnosis, treatment satisfaction, etc., which are all feedback indicators directly affecting the quality of the current decision.
[0100] U t |S t = s, indicating the immediate reward obtained when the current state is given as s;
[0101] Performing a temporal difference decomposition on U t The above formula can be further expressed as:
[0102] V π (s) = E[R t+1 + γ[R t+2 + γ[…]]|S t= s] (3)
[0103] Introduce the current information state value function V of the t+2 patient π (s′):
[0104] V π (s) = E[R t+1 + γV π (s′)|S t = s] (4)
[0105] Therefore, the optimal cumulative expected utility V * (s) represents:
[0106]
[0107] Therefore, define the Q(s,a) state-action value function:
[0108] q π (s,a) = E π [r t+1 + γr t+2 + γ 2 r t+3 +…|A t = a,S t = s] = E π [G t |A t = a,S t = s] (6)
[0109] where G t is the total discounted reward starting from time t, and it can be seen from here that the decay value of r will affect the change of the Q(s,a) function. Construct the iterative form from time t to time t+1:
[0110] q π (s,a) = E π [R t+1 + γq π (S t+1 ,A t+1 )A t = a,S t = s] (7)
[0111] Optimal value action function The Q-table is updated as follows:
[0112]
[0113] (3) Construct the effective interaction information of the patient into a Markov process and dynamically adjust the confidence of various information:
[0114] By constructing various information that helps with interactive questioning from the patient's answers into a Markov decision process, the confidence levels of various information are dynamically adjusted. Specifically,
[0115] In the patient's t-th answer, by extracting relevant features of various information, taking actions, and updating the state (confidence level) according to the probability transition function;
[0116] For multi-dimensional reinforcement learning problems, relevant probabilities at time t in multiple aspects can be obtained: the misdiagnosis confidence level d1 t , the degree of trust d2 of the patient in the diagnosis of the large model t , the patient friendliness d3 t , the diagnosability d4 relying on the current information t , the degree of the patient's emotional state d5 t , the degree of the patient's understanding of the large model d6 t , the tolerance d7 of the diagnostic questioning process t , the credibility of the patient's information is d8 t , where,
[0117] The tolerance d7 of the diagnostic questioning process t is related to the number of questions and answers, so d7 t follows a binomial distribution, that is:
[0118]
[0119] where p represents the probability that the patient leaves the conversation process after each conversation.
[0120] It should be noted that during the medical interview process, the information ambiguity caused by some misunderstandings of the patient and the unclear description of the patient's symptoms will both lead to misdiagnosis of the model. Therefore, we need to model the misdiagnosis confidence level d1 t , specifically,
[0121] Assume that during the t-th medical interview process, the patient provides some relevant medical information, and through language information extraction, the feature F of the patient after this answer can be obtained N t , where, F N t = [f1 t , f2 t ,..., f N t , f i t represents the set of all feature information of the i-th type in the previous t iterations, that is By obtaining the features, a series of mapping relationships are constructed according to the knowledge graph:
[0122]
[0123] Among them, F k t represents the top N corresponding to the k-th disease obtained from the knowledge graph for the extracted significant symptom features. Then, it is necessary to obtain the average similarity level between the i-th feature described by the patient during the t-th consultation and the significant features of all K diseases:
[0124]
[0125] Among them, ψ s represents the cosine similarity, represents the scale factor.
[0126] Meanwhile, the misdiagnosis situation of the patient is also affected by the demographic information (age, ethnicity, gender) in the electronic medical record. For example, it is generally believed that patients aged 20 - 50 provide more accurate information, while patients aged 10 - 20 may provide false positive feature information due to their lower life experience. It is necessary to add the demographic information in the electronic medical record. Therefore, a statistical information factor is introduced
[0127]
[0128] Among them, H(·) represents the self-information entropy function of a certain basic statistical feature.
[0129] Therefore, the average similarity level can be corrected to the following equation:
[0130]
[0131] When the similarity level obtained at the t-th time has a large difference from the previous similarity level, it indicates that the i-th feature information provided by the patient at the t-th time has a great negative impact on the diagnosis of all possible K diseases, that is:
[0132]
[0133] Among them, represents the probability distribution of all features in the t-th Q&A, represents the probability distribution of all features in the (t + 1)-th Q&A. δ is the misdiagnosis threshold parameter.
[0134] Therefore, there is:
[0135]
[0136] It should be noted that the degree of trust d2 of the patient in the diagnosis of the large model t, on the one hand, it depends on the amount of information provided by the patient at the t-th time, and on the other hand, it is affected by the tolerable degree d7 of the diagnostic inquiry process t influence
[0137] Therefore, for the patient information obtained at the t-th time, the following probability formula can be given:
[0138]
[0139] wherein, λ represents the weight parameter
[0140] Furthermore, relying on the diagnosable degree d4 of the current information t , that is, after the t-th medical consultation, according to the known patient characteristic information, the diagnostic prediction probability of the patient's disease
[0141] (4) Use the features as the model input and construct possible disease classification labels:
[0142] For Find possible diseases on the knowledge graph module and use them as the disease label y
[0143] It should be noted that using the features as the model input and the possible diseases as the classification labels, training multiple disease information data sets in the reinforcement model to obtain the final result
[0144] Exemplarily, through the fully connected transformation of the reinforcement model, the prediction vector d4 can be obtained t .
[0145] Preferably, the present invention uses cross-entropy to calculate the loss between y and d4 t and uses it to train the reinforcement model:
[0146] L ce =-ylog(d4 t ) (17)
[0147] In this way, through the obtained multiple relevant probabilities, the large model medical consultation system based on multi-dimensional reinforcement learning and Markov probability optimization disclosed by the present invention drives and guides the language large model to ask questions in the optimal direction, so as to complete the patient information collection process and the complete diagnosis and treatment process
[0148] See Figure 3 , in another embodiment
[0149] The prediction model for the cost of treating a patient's disease based on a large model disclosed by the present invention can generate an optimal diagnosis and treatment decision-making plan suitable for the patient's own state according to the results generated in the previous process and the patient's own information, such as medical insurance, family status, past medical history, etc
[0150] The disease treatment cost prediction model for large models disclosed by the present invention, for example, is used to implement the following method, which includes:
[0151] Generate clustered patient action sequences: Explore similarities and cost patterns
[0152] In the process of constructing the prediction model, clustering analysis plays a crucial role. It helps us group a large number of patients according to the similarity of their treatment courses, laying the foundation for subsequent cost prediction. This process not only involves an in-depth understanding of the patient action sequences but also requires the introduction of an innovative cost concept. By discretizing the entire cost range into multiple segments and quantifying the actual cost as the average value of each segment, we can achieve efficient management and utilization of cost information. The core of this strategy is that it simplifies the complex and variable cost data into a set of easy-to-process categories, maintaining the relative integrity of the cost information while avoiding over-complication caused by minor differences in cost values.
[0153] Design distance metric: Treatment Pattern Difference (TPD)
[0154] Among many methods for measuring sequence similarity, we choose Treatment Pattern Difference (TPD) as a tool to evaluate the differences between patient action sequences. TPD not only inherits the basic idea of edit distance, that is, converting one sequence to another through insertion, deletion, and substitution operations, but also innovatively incorporates cost considerations, making it a distance metric suitable for medical scenarios. Specifically, the calculation formula of TPD is as follows:
[0155]
[0156] where C1 = y i , C2 = y j , C3 = ∣y i ―y j +∈∣ and C4 = ∣y i ―y j ∣ correspond to insertion, deletion, substitution, and matching costs respectively, while w1 and w2 represent the weights of different activities and similar activities. By reasonably setting these parameters, TPD can accurately reflect the comprehensive differences between two action sequences at the clinical activity and cost levels, providing strong support for subsequent clustering analysis.
[0157] Select clustering algorithm: Hierarchical DBSCAN
[0158] Among numerous clustering algorithms, we selected hierarchical DBSCAN, which is a density-based spatial clustering algorithm, especially suitable for dealing with clustering datasets with irregular shapes and unknown numbers. Compared with traditional clustering methods, hierarchical DBSCAN does not require pre-specifying the number of clusters but automatically identifies high-density regions in the dataset and divides them into independent subgroups. This feature makes the algorithm more flexible and robust when dealing with patient action sequences and can effectively capture the characteristics of patient groups under different treatment patterns.
[0159] Markov Chains Representing Groups: Revealing the Dynamic Changes in Treatment Patterns
[0160] To further analyze and predict the development trend of patient action sequences, we adopted a first-order Markov chain model to represent each clustering group. Markov chains are powerful statistical models that assume the next state of the system depends only on the current state and is not affected by previous states, making them very suitable for describing the sequential nature of clinical activities. By calculating the state transition probability distribution, that is
[0161]
[0162] where q i,j,l represents the number of transitions from action a i to action a j We can gain insights into the transition rules between various clinical activities in the patient's treatment process and provide data-driven insights for cost prediction and treatment plan optimization. Predicting the Treatment Cost of Patients from Action Sequences: Refining the Prediction Mechanism
[0163] Cost prediction is the core link of the entire medical expense prediction system. It provides a forward-looking estimate of future medical expenditures by matching the patient action sequence with known treatment patterns. This prediction process is divided into three key stages, aiming to ensure the accuracy and practicality of the prediction results through multi-dimensional analysis.
[0164] Cost Estimation Based on Specific Groups
[0165] In the initial stage of cost prediction, we focused on finding the treatment pattern most similar to the patient action sequence to be predicted. By using the innovative metric of Treatment Pattern Difference (TPD), we were able to screen out k nearest neighbor action sequences in the large patient database, and the groups G l to which these sequences belong will be used as the reference objects for cost estimation. Subsequently, we used the weighted average method to aggregate the actual costs c j,l of these k nearest neighbor sequences and calculated the total cost estimate l of the new action sequence P for group G The weight w j,lThe setting follows the inverse principle of the TPD metric, ensuring that sequences with higher similarity contribute more to the total cost estimation.
[0166]
[0167] Evaluating the likelihood of patient groups
[0168] To more comprehensively evaluate the fit between patient action sequences and various treatment modalities, we introduced the Markov chain model, a powerful tool that can reveal the inherent laws of clinical activity sequences. By calculating the likelihood L n (P) of an ordered action sequence P = (a1, a2, …, a l ) in a specific group G l , we can not only measure the matching degree between the sequence and the group but also gain insights into the transition probabilities between clinical activities in the sequence. This likelihood calculation is based on the state transition probabilities of the Markov chain, and the specific formula is as follows:
[0169]
[0170] where p l (a1) represents the initial probability of the starting clinical activity a1 in group G l , and p l (a i+1 │a i ) reflects the probability of transitioning from activity a i to activity a i+1 .
[0171] Final cost prediction
[0172] After integrating the cost estimation of a specific group and the likelihood of the patient action sequence, we enter the final stage of cost prediction. The goal of this stage is to integrate the estimated costs of all groups with the corresponding likelihood L l (P) through weighted summation to obtain the final predicted cost This process ensures that the prediction result fully considers the possibilities and cost distributions of different treatment modalities, improving the accuracy and reliability of the prediction.
[0173]
[0174] Through the careful design and implementation of the above three stages, our cost prediction method can not only provide personalized and refined medical expense predictions based on patient treatment sequences but also gain insights into the internal logic of treatment modalities, providing a scientific basis for medical resource planning, patient financial preparation, and medical policy formulation. The successful application of this framework marks an important step in the field of medical big data analysis and intelligent decision support.
[0175] In addition, since medical knowledge is not static but updated regularly, for the knowledge graph used in the large model, we also design an update scheme here to reduce the conflict problems in the graph. A knowledge graph (KG) is a structured knowledge base that represents entities, relations, and their interactions in the real world in the form of triples. Formally, a knowledge graph can be defined as a binary tuple KG = (E, R, T), where:
[0176] · E is the entity set, containing all the entities described in the graph, such as people, places, events, etc.
[0177] · R is the relation set, defining various types of relations that may exist between entities, such as "born in", "located in", etc.
[0178] · is the triple set, and each triple (e1, r, e2) indicates that there is a relation r between entities e1 and e2.
[0179] At any point in time t, the state of the knowledge graph can be represented as KG t = (E t , R t , T t ), where E t , R t and T t are the entity set, relation set, and triple set at that time point respectively. Over time, due to the addition of new data (such as newly discovered entities, newly confirmed relations) or the correction of existing data (such as the update of entity attributes, the re-evaluation of relations), the knowledge graph will continue to evolve.
[0180] Achieve feedback and self-optimization via a feedback loop and self-optimization mechanism
[0181] Building a closed-loop feedback system is an efficient self-optimization mechanism, and its core lies in enhancing the accuracy of the knowledge graph (KG) and the reliability of the trust model through continuous iteration. The working principle of this mechanism can be summarized by the following mathematical expression:
[0182] KG t+1 , TrustModel t+1 = Optimizer(KG t , TrustModel t , Data t,Feedback t ) Here, the Optimizer is a core optimization function that acts as the intelligent engine of the system. This function synthesizes four key input elements to drive the evolution of the system:
[0183] 1. The current state of the knowledge graph (KG t ): Represents the knowledge structure and relationship network accumulated within the system at time step t.
[0184] 2. The current trust model (TrustModel t ): Evaluates the credibility or accuracy of various elements (such as entities, relationships, etc.) in the knowledge graph and is an important basis for system decision-making.
[0185] 3. New data (Data t ): Includes newly collected triple (entity - relationship - entity) information and possible correction information, which are sourced from external data sources or user inputs and are used to enrich and correct the existing knowledge graph.
[0186] 4. Feedback (Feedback t ): This is a multi-dimensional information set that combines user feedback and internal system evaluations. User feedback may directly point out errors or suggest improvements, while internal system evaluations may be based on the algorithm's checks for data consistency and logical rationality.
[0187] The Optimizer function performs a series of complex operations, such as data fusion, error detection and correction, trustworthiness re-evaluation, etc., through comprehensive analysis of these inputs, and finally outputs the updated knowledge graph KG t+1 and the trust model TrustModel t+1 . This process ensures that the system can continuously learn, adapt, and evolve, so that when facing new data or environmental changes, it can maintain the high quality and reliability of its knowledge base.
[0188] Through such a closed-loop feedback mechanism, the system can not only optimize itself but also gradually improve its ability to process complex information, identify errors, and automatically correct them, providing more accurate, trustworthy, and useful knowledge services to end-users.
[0189] Next, we will introduce the knowledge graph update and the new model update respectively:
[0190] Dynamic iterative optimization
[0191] The evolution of the knowledge graph can be formalized as a transformation process from KG t to KG t+1 where KG t+1 =(E t+1 ,Rt+1 , T t+1 )。This transformation process can be described by the following formula or steps:
[0192] 1. Entity update: E t+1 = E t ∪ ΔE t , where ΔE t represents the set of entities newly added between time points t and t + 1.
[0193] 2. Relationship update: R t+1 = R t ∪ ΔR t , where ΔR t represents the set of newly defined or introduced relationships. It should be noted that in practice, the relationship set R is usually relatively stable, but theoretically it may still change over time.
[0194] 3. Triple update: T t+1 = T t ∪ ΔT t \ Γ t , where ΔT t represents the set of newly added triples, and Γ t represents the set of triples deleted or corrected during t to t + 1. The union and difference operations of sets are used here to represent the addition and deletion of triples.
[0195] Dynamic update formula for trust
[0196] The dynamic update of trust is not only the construction of a mathematical expression, but also the simulation and response to the complexity of the real world. In the optimized description, we consider the following key factors:
[0197] Continuation of the current trust state: The continuity of trust is reflected in the time series, that is, the current trust level Trust t (e) directly affects the trust score at future moments. This is adjusted by the coefficient α to ensure a smooth transition of trust.
[0198] Integration and evaluation of new evidence: Each new piece of evidence ne is assigned a different weight Weight(ne) according to the reliability, relevance, and novelty of its source, and is further screened by the validity index Validity(ne) to ensure that high-quality information has a positive impact on trust, while low-quality or irrelevant information has limited impact.
[0199] Immediate feedback on user feedback: The user feedback fb serves as a direct feedback loop for trust adjustment and is transformed into a quantified trust impact through sentiment analysis Sentiment(fb). Positive feedback directly enhances trust, negative feedback prompts a decrease, and neutral feedback maintains the status quo, ensuring the system's self-adjustment ability and user orientation.
[0200] Taking all these factors into account, the optimized trust update equation is directly reflected as:
[0201]
[0202] Decay(t,t + 1) is a time decay function used to reflect the change of trust over time. Here, it not only directly embodies the interaction mechanism of various factors but also flexibly adjusts the contribution of each part through the parameters α, β, and γ to ensure the adaptability and accuracy of the model. In this way, the trust update process not only considers time decay and the weighted evaluation of new evidence but also effectively integrates the emotional orientation of user feedback, forming a dynamically balanced and self-optimizing trust scoring system.
[0203] In another embodiment, in driving the large model consultation system based on multi-dimensional reinforcement learning and Markov probability optimization, the present invention can also record large fluctuations in the sensitivity of the model caused by factors at a certain moment.
[0204] Although the embodiments of the present invention have been described above in conjunction with the figures, the present invention is not limited to the above specific embodiments and application fields. The above specific embodiments are merely illustrative and guiding, rather than restrictive. Those of ordinary skill in the art can also make many forms under the inspiration of this specification and without departing from the scope protected by the claims of the present invention, and all of these fall within the scope of protection of the present invention.
Claims
1. A large model consultation system based on multi-dimensional reinforcement learning and Markov probability optimization, characterized in that, Including: A feature information extraction module that obtains the feature information distribution of the patient's interview text information using multiple state prediction models; An interview instruction dynamic optimization module that obtains the patient's disease information and personalized characteristics based on the feature information distribution for personalized diagnosis of the patient; meanwhile, obtains symptom query information based on the medical knowledge graph and uses a multi-dimensional reinforcement learning model to dynamically optimize the patient's interview state to generate high-quality interview instruction information; An optimal cost-benefit medical decision-making module that deeply mines and analyzes the remaining available digital medical insurance of the patient, the current symptom manifestations, the available hospital drug inventory, and the effectiveness and feedback of past treatment plans based on the patient's personal basic information, including age, gender, and past medical history, to achieve multi-dimensional and all-round cost prediction and effect evaluation; A knowledge graph update module that continuously updates based on the real-time feedback of patient diagnosis information and the knowledge of current medical research results to ensure that the medical knowledge base always remains up-to-date.
2. The system according to claim 1, wherein Preferably, the dynamic optimization module further includes: A single prompt word unit that, based on the multi-dimensional reinforcement learning calculation paradigm, selectively traverses relevant disease knowledge graphs and associated massive medical literature, and conducts heuristic and efficient searches to guide the large model to carry out the best interview prompt words for high-performance single diagnosis and treatment decisions.
3. The system according to claim 1, wherein The dynamic optimization module further includes: A dynamic interview prompt word unit that includes a time series optimization model based on the Markov model constructed according to the patient's individualized disease progression for dynamic diagnosis and treatment auxiliary decision-making to generate the best interview prompt words that can guide the large model to carry out high-performance dynamic diagnosis and treatment decisions.
4. The system according to claim 1, wherein The dynamic optimization module further includes: A fusion unit that fuses the single prompt word unit and the dynamic interview prompt word unit to provide individualized best interview prompt words and diagnosis and treatment decisions for patients with different disease histories.
5. The system according to claim 1, wherein A feature information extraction module that, for a large amount of text feature information obtained from the interview record, adopts multiple state prediction models such as feature correlation calculation, emotion library comparison, and statistical information conversion to obtain the feature information distribution of the patient's interview text information.
6. The system according to claim 1, wherein The optimal cost-benefit medical decision-making module further includes: A data conversion unit that is used to convert the scattered treatment records of each patient into an ordered and structured action sequence to retain all the details of the patient's treatment process and provide rich data support for subsequent clustering analysis, pattern recognition, and cost prediction.
7. The system according to claim 1, wherein The optimal cost-benefit medical decision-making module further includes: A clustering patient action sequence unit that discretizes the entire cost range into multiple paragraphs and quantifies the actual cost as the average value of each paragraph to achieve efficient management and utilization of cost information.
8. The system according to claim 6, wherein The optimal cost-benefit medical decision-making module further includes: An estimation unit that is used to predict the treatment cost of the patient from the action sequence and provides a forward-looking estimate of future medical expenses by matching the patient action sequence with known treatment patterns.
9. The system according to claim 1, wherein The knowledge graph update module further includes: The feedback and optimization unit, through a feedback loop and a self-optimization mechanism, becomes the core driving force for the development of the Knowledge Graph (KG) and its trust model, to build a closed-loop system aimed at strengthening the accuracy of the Knowledge Graph and the reliability of the trust model through continuous cyclic iterations.
Citation Information
Cited By
Rehabilitation evaluation and auxiliary tool adaptation method, system and equipment based on reinforcement learning
CN121071085A