Patient value and loss risk assessment system and method based on combination of ML and LLM
By using a patient value and churn risk assessment system based on ML combined with LLM, the problems of single assessment dimensions, poor coordination, and insufficient interpretability have been solved. This system enables accurate identification of high-value patients and dynamic management of churn risk, thereby improving the efficiency of refined operation of medical services.
Patent Information
- Application Number
- CN202610179477.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-09
- Publication Date
- 2026-03-20
AI Technical Summary
In existing technologies, medical institutions face problems such as a single assessment dimension, poor model synergy, lack of interpretability, and insufficient dynamic adaptability in patient value stratification and churn risk assessment. This leads to a disconnect between assessment results and clinical needs, failing to meet the requirements of refined operations.
A patient value and churn risk assessment system based on ML combined with LLM is adopted, including modules for data collection, feature construction, model training, churn risk assessment, interpretability and outcome linkage. Through multimodal data processing and specialty-adaptive design, an assessment-intervention closed loop is formed to support iterative optimization of the model.
It improves the accuracy and clinical interpretability of identifying high-value patients and predicting churn risk, enables dynamic adaptability of assessment results and precise allocation of resources, reduces the probability of patient churn, and enhances the targeting and effectiveness of medical services.
Smart Images

Figure CN121707286A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence medical technology, and in particular to a patient value and churn risk assessment system and method based on ML combined with LLM. Background Technology
[0002] With the deepening of medical informatization, hospital systems such as HIS, LIS, PACS, and EMR have accumulated massive amounts of multi-source, heterogeneous patient diagnosis and treatment data, covering laboratory test values, imaging report texts, electronic medical records, expense details, and registration and cancellation behavior. However, current medical institutions still have significant limitations in assessing patient value segmentation and churn risk: Firstly, the assessment dimensions are too narrow. Traditional solutions rely heavily on human experience or simple rules based on consumption amount and frequency of visits, which cannot capture potential high-value patients such as those with early-stage severe illnesses and high adherence to chronic diseases. Moreover, the risk warning of attrition often lags behind the actual behavior of patients who have not had follow-up visits for a long time. There is a lack of quantitative analysis on the rationality of diagnosis and treatment and treatment adherence, which is out of touch with clinical needs. Secondly, the models have poor synergy. Although machine learning (ML) models are good at numerical prediction, they are difficult to gain doctors' trust due to their black-box nature. Although large language models (LLM) have semantic understanding capabilities, they have problems such as high computational cost and low numerical prediction efficiency. Moreover, the two types of models are mostly used independently and have not formed a collaborative architecture of ML quantitative prediction + LLM semantic enhancement, which cannot adapt to the needs of multimodal medical data processing. Third, there is a lack of interpretability and applicability. The output of ML models lacks clinically understandable evidence, and LLM interpretations often deviate from medical logic, resulting in low adoption rates by doctors. Furthermore, the evaluation results are not linked to intervention strategies, and a prediction-intervention closed loop cannot be formed. Fourth, it lacks dynamic adaptability. It fails to design adaptive features and parameters for core indicators and treatment cycle differences across different specialties such as cirrhosis and end-stage renal disease, and lacks a continuous iteration mechanism, leading to a decline in assessment accuracy after long-term use. In summary, there is an urgent need for an assessment solution that integrates the advantages of ML and LLM, possesses specialty adaptability, strong interpretability, and can form a closed-loop management system to meet the needs of medical institutions for refined operations and high-quality diagnosis and treatment. Summary of the Invention
[0003] The purpose of this invention is to address the shortcomings of existing technologies by proposing a patient value and churn risk assessment system and method based on ML combined with LLM.
[0004] To achieve the above objectives, the present invention adopts the following technical solution: a patient value and churn risk assessment system based on ML combined with LLM, including a data acquisition module, a feature construction module, a model training module, a high-value prediction module, a churn risk assessment module, an interpretability module, a result linkage module, and a model iteration module; The data acquisition module is used to collect patients' historical and real-time multi-source medical data; the feature construction module is used to integrate the data collected by the data acquisition module to generate multimodal feature vectors. The model training module is used to build an evaluation model using ML and LLM. The input is the multimodal feature vector generated by the feature construction module, and the output is the trained joint evaluation model. The high-value prediction module is used to output patient value judgment results based on the joint assessment model. The input is the joint assessment model generated by the model training module and the real-time multimodal feature vector of the patient. The output is the patient's high-value probability value and binary judgment result. The churn risk assessment module is used to comprehensively analyze the possibility of patient churn. The input is the patient's real-time multimodal feature vector, and the output is the churn risk score and risk level. The interpretability module is used to generate clinically understandable assessment criteria. The inputs are the judgment results of the high-value prediction module, the risk score of the churn risk assessment module and the corresponding multimodal feature vector, and the output is a structured interpretation report. The result linkage module is used to match targeted intervention strategies. The inputs are the binary judgment result of the high-value prediction module and the risk level of the churn risk assessment module, and the output is a personalized intervention plan. The model iteration module is used to dynamically optimize the evaluation model. The inputs are doctors' feedback data on the evaluation results, actual patient churn and value conversion results data, and newly added multimodal feature data. The output is the updated joint evaluation model parameters.
[0005] Preferably, the specific working steps of the data acquisition module are as follows: The subsequent feature construction module divides the data into three levels of priority. For unstructured text data, field-level dynamic data masking is used; for structured data, range-based data masking is used. Data collection must pass integrity, consistency, and timeliness checks. Data that fails the checks will trigger a second data collection.
[0006] Preferably, the specific working steps of the feature construction module are as follows: A specialty feature priority database is constructed based on specialty clinical diagnosis and treatment guidelines, and the feature selection logic is dynamically adjusted through feature importance scoring; A time decay coefficient is introduced into the historical medical visit behavior characteristics to calculate the weighted features; Anomaly traceability labels are used to mark any abnormal features discovered during the cleaning process.
[0007] Preferably, the specific working steps of the model training module are as follows: The total loss of the joint evaluation model is calculated by weighting the ML model loss and the LLM loss; a specialized error correction term is added during the hyperparameter optimization process; After each round of training, the results are verified through the real-time inference interface of the high-value prediction module.
[0008] Preferably, the specific working steps of the high-value prediction module are as follows: The judgment threshold is derived from the model's real-time error factor and the intervention resource factor; Stratify the confidence levels of the high-value probability values in the output; Predict feature subsets based on data marked by the data acquisition module.
[0009] Preferably, the specific working steps of the churn risk assessment module are as follows: The LLM diagnosis and treatment trajectory reasoning module outputs diagnosis and treatment deviation labels, triggering the adjustment of indicator weights in the ML quantitative scoring card; when the ML scoring card outputs a high-risk warning, it triggers the deep semantic review of the LLM module; and the risk level threshold of the ML scoring card is corrected based on the specialty feature distribution of the feature construction module.
[0010] Preferably, the specific working steps of the interpretability module are as follows: The system generates explanatory content based on a list of core features of specialties built on the feature construction module; it supplements the explanatory logic with feedback results from the dual-engine system of the churn risk assessment module; and it automatically traces the cause of the error if the doctor's feedback explanation does not match the clinical reality.
[0011] Preferably, the specific working steps of the result linkage module are as follows: establishing intervention resource priorities based on the confidence level stratification of the high-value prediction module and the risk change rate of the churn risk assessment module; adjusting the intervention frequency according to the risk score changes tracked by the churn risk assessment module; and collecting post-intervention effect data to calculate the strategy effectiveness.
[0012] Preferably, the specific working steps of the model iteration module are as follows: hierarchical processing of multi-source feedback data, divided into three levels according to the iteration frequency; fluctuation upper limit control of the weight of the core features of the specialty during the iteration process; and ensuring system stability through full-link verification after each iteration.
[0013] This invention also proposes a method for assessing patient value and churn risk based on ML combined with LLM, comprising the following steps: S1: Dynamic priority data collection. Based on the list of core indicators for specialties in the subsequent feature construction steps, the data from the hospital's HIS, LIS, PACS, and EMR systems are subject to three-level priority scheduling. At the same time, layered desensitization processing is adopted. After the collected data undergoes integrity, consistency, and timeliness verification, it is synchronized to the data quality logs of the feature construction and model iteration steps. S2: Specialty Adaptive Feature Construction. Based on specialty diagnosis and treatment guidelines, a specialty feature priority library is built. Core features are selected through feature importance scoring. Missing value imputation, outlier removal, and normalization are performed on structured data. Semantic vectors are generated for text data. Time decay coefficients are introduced to calculate weighted features for time-series behavioral data. The selected multimodal feature vectors are synchronized to the model training step. Feature format standardization is synchronized to the data consistency verification step in S1. Abnormal feature annotation and source tracing labels are pushed to the interpretability step and the model iteration step. S3: Dual-model collaborative training, constructing a joint evaluation model of ML and medical-specific LLM, using a dynamic loss function to calculate the total loss, optimizing hyperparameters through a hyperparameter optimization framework, and adding a specialty error correction term. During the training process, each round is verified through the real-time inference interface of the high-value prediction step, and the finally trained joint model is pushed to the high-value prediction step. S4: Dynamic threshold high-value prediction. It receives the real-time multimodal feature vector of patients from S2 and the joint model of S3. It uses a dual-driving factor dynamic threshold to derive and determine the threshold, outputs the high-value probability value and stratifies it according to confidence level. The prediction results are synchronized to the prediction error log of the churn risk assessment step, the interpretability step and the model iteration step. S5: Dual-engine feedback loss risk assessment. Based on the high-value prediction results of S4, the dual engines of LLM diagnosis and treatment reasoning and ML quantitative scoring are launched. The LLM engine outputs deviation labels to trigger the weight adjustment of the ML scoring card. The ML engine calculates the risk score to trigger the deep semantic review of the LLM engine. Finally, the risk level and cause analysis are synchronized to the interpretability step and the result linkage step. The dual-engine feedback data is pushed to the model iteration step. S6: Specialty-adapted interpretability report generation. It receives the high-value judgment results from S4 and the risk assessment results from S5, generates a report according to clinical decision-making guidance, and synchronizes the report to the results linkage step after privacy compliance verification. The interpretation error reported by the doctor is pushed to the error source library of the model iteration step. S7: Dynamic resource scheduling intervention. Based on the confidence level of S4 and the risk change rate of S5, intervention resource priority is established, and a value-risk four-quadrant strategy is matched. The intervention resource occupancy rate is fed back to the intervention resource factor of S4 in real time, and the post-intervention effect data is pushed to the model iteration steps. S8: Multi-source feedback hierarchical iteration. Based on the feedback data of S1-S7, a three-level iteration is implemented. During the iteration process, the fluctuation upper limit of the weight of the core features of the specialty is controlled. After the iteration, the whole link is verified. If it fails, it is rolled back to the previous version.
[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows: Based on different specialty treatment guidelines, a core feature system is constructed. For example, the specialty of cirrhosis with upper gastrointestinal bleeding prioritizes key indicators such as hemoglobin and alanine aminotransferase, while the specialty of end-stage renal disease with stage 4 CKD focuses on core parameters such as serum creatinine and calcium-phosphorus product, ensuring that the feature dimensions are highly matched with the clinical needs of the specialty. At the same time, a time-decay weighted logic is introduced to strengthen the impact of recent medical behavior and indicator changes on the assessment results, which is more in line with the dynamic pattern of disease progression and patient behavior. In addition, a dual-model collaborative training architecture is adopted. The ML model accurately captures the nonlinear relationship of numerical features, while the LLM model deeply analyzes the semantic information of text data. The two complement each other and effectively improve the utilization efficiency of multimodal medical data, making the results of high-value patient identification and loss risk prediction more in line with clinical practice. Through a specialist-adaptive interpretation generation mechanism, the model's decision-making logic is transformed into professional terminology familiar to clinicians. For example, core logic such as the association between a persistent decrease in hemoglobin indicating an increased risk of bleeding and abnormal serum creatinine and deteriorating renal function is clearly marked. Key features and feature-conclusion causal chains that significantly affect the assessment results are clearly labeled. Simultaneously, diagnostic bias analysis (such as diagnostic omissions and departmental recommendation biases) output by LLM is used as supplementary evidence, making the judgment logic of the assessment results traceable and easy to understand. Furthermore, the module supports the function of tracing the source of errors reported by doctors, allowing problems to be located at both the feature level (such as missed key indicators or unreasonable time-series windows) and the model level (such as semantic parsing bias and numerical prediction errors), further improving the clinical acceptance and practicality of the interpretation report. For high-value, high-risk patients, dedicated health management and multidisciplinary consultation resources are provided; for high-value, low-risk patients, routine follow-up and health education services are offered; and for low-value, high-risk patients, basic follow-up and reminders for the rational allocation of medical resources are conducted, ensuring that intervention measures are accurately tailored to patient needs. Simultaneously, the real-time occupancy status of intervention resources can be fed back to the high-value prediction module, dynamically optimizing the judgment threshold to balance resource supply and assessment accuracy. Post-intervention data, such as patient follow-up visits and doctor adoption rates, can also serve as a basis for strategy optimization, forming a virtuous cycle of assessment-intervention-feedback-optimization. This effectively reduces the probability of high-value patient attrition and improves the targeting and effectiveness of medical services. The model iteration module employs a hierarchical iteration mechanism. High-frequency iterations fine-tune the LLM prompt word templates and ML model feature weights to adapt to short-term changes in clinical behavior. Mid-frequency iterations optimize the assessment threshold formula and specialty risk level standards to respond to adjustments in clinical practice. Low-frequency iterations retrain the joint model and update the specialty feature library to adapt to long-term clinical guideline updates and changes in data distribution. Simultaneously, fluctuation control is implemented on the weights of core specialty features during iteration to prevent a decrease in the model's sensitivity to key clinical indicators. This ensures that the assessment accuracy remains stable throughout the long-term use of the system, continuously meeting the dynamic needs of refined operations and high-quality medical care in healthcare institutions. Ultimately, this helps hospitals improve the retention rate of high-value patients, optimize the efficiency of medical resource allocation, reduce the risk of misdiagnosis and missed diagnosis, and improve treatment adherence and health management effectiveness. Attached Figure Description
[0015] Figure 1 This is a flowchart of the overall algorithm of the present invention; Figure 2 This is a high-value patient prediction chart for the present invention; Figure 3 This is a risk assessment diagram for the loss of this invention. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0017] In the description of this invention, it should be understood that the terms length, width, up, down, front, back, left, right, vertical, horizontal, top, bottom, inside, outside, etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "multiple" means two or more, unless otherwise explicitly specified.
[0018] Please see Figure 1-3 A patient value and churn risk assessment system based on ML combined with LLM includes a data acquisition module, a feature construction module, a model training module, a high-value prediction module, a churn risk assessment module, an interpretability module, a result linkage module, and a model iteration module. The data acquisition module is used to collect patients' historical and real-time multi-source medical data. It should be noted that the multi-source medical data collected by the data acquisition module includes users' test and examination data, image report text, microbial test reports, doctor's diagnosis conclusions, five histories (past medical history, family history, personal history, marital and reproductive history, current medical history), and patient basic information such as age, gender, medical insurance type, and contact information. It also connects to the hospital's four core systems (HIS, LIS, PACS, and EMR) through the HL7 / FHIR standard interface. The real-time data transmission latency is ≤5 seconds, and batch data is automatically synchronized every morning at midnight and the data integrity is verified by MD5 hash value. The feature construction module is used to integrate the data collected by the data acquisition module to generate multimodal feature vectors. It should be noted that the specific process of the feature construction module generating multimodal feature vectors is to perform feature extraction, format conversion, category encoding, outlier cleaning and data normalization on the collected data. The feature dimensions are determined by the range of core clinical indicators of the specialty. Among them, structured data uses the mean or median to impute missing values, IQR method to detect outliers, and Min-Max normalization. Text-type unstructured data generates semantic vectors based on the BioBERT pre-trained model through a medical NLP engine. Time-series behavioral data calculates statistical features by determining the sliding window length according to the specialty diagnosis and treatment cycle. Min-Max normalization is a method that standardizes the original data to the [0,1] interval. First, identify the feature set to which the data to be processed belongs, find the minimum and maximum values of the feature in this set, calculate the difference between the maximum and minimum values to obtain the numerical range of the feature, use each piece of raw data to be processed, subtract the minimum value of the feature found earlier, calculate the deviation of the raw data relative to the minimum value, divide the calculated deviation by the numerical range obtained earlier, and the final result is the data after standardization. The model training module is used to build an evaluation model using ML and LLM. The input is the multimodal feature vector generated by the feature construction module, and the output is the trained joint evaluation model. It should be noted that in the model training module, the ML model uses a gradient boosting tree model to model the nonlinear relationship of quantified features, while the LLM uses a medical-specific pre-trained model to handle the semantic parsing and knowledge enhancement of unstructured text features. The model hyperparameters are derived based on the three-dimensional balance principle of prediction accuracy, computational efficiency, and clinical interpretability. The objective function weights W1 (high-value identification accuracy weight) and W2 (loss risk prediction recall weight) satisfy (W2+W2=1), and their specific values are determined according to the clinical priority of the specialty. The high-value prediction module is used to output patient value judgment results based on the joint assessment model. The input is the joint assessment model generated by the model training module and the real-time multimodal feature vector of the patient. The output is the patient's high-value probability value and binary judgment result. It should be noted that the threshold for the high-value prediction module is derived through a clinical resource optimization and allocation model. Specifically, the threshold is calculated as follows: the product of the opportunity cost of missing a high-value patient and the prior probability of the high-value patient, divided by the sum of the product of the opportunity cost of missing a high-value patient and the prior probability of the high-value patient and the product of the ineffective service cost of misjudging a high-value patient and the prior probability of the non-high-value patient. At the same time, this threshold must meet two conditions: first, the intervention coverage rate of high-value patients is not lower than the pre-set threshold; second, the efficiency of medical resource input does not exceed the pre-set threshold. The churn risk assessment module is used to comprehensively analyze the possibility of patient churn. The input is the patient's real-time multimodal feature vector, and the output is the churn risk score and risk level. It should be noted that an LLM treatment trajectory reasoning + ML risk probability modeling mechanism is adopted. In this mechanism, the large language model infers the rationality of the treatment process based on the knowledge of specialty clinical guidelines, while the machine learning model calculates the risk probability based on the correlation between features. The final risk score consists of two parts: one is the risk score derived from the large language model inference, and the other is the risk score calculated by the machine learning model. These two scores are added together with certain weights. The weighting coefficients are determined by weighting the predictive contributions of the two engines in historical data using precision and recall. The risk level is divided into high risk, medium risk, and low risk based on the range of the final risk score. The interpretability module is used to generate clinically understandable assessment criteria. The inputs are the judgment results of the high-value prediction module, the risk score of the churn risk assessment module and the corresponding multimodal feature vector, and the output is a structured interpretation report. It should be noted that the interpretability module uses LLM to transform the model decision logic into clinical terminology, highlighting the top 5 core features with the highest weight influencing value judgment and risk assessment, such as high-frequency visit intervals, trends in key test indicators, and cancellation rates. The report must include a description of the feature-conclusion causal chain, that is, clearly define the relationship logic between one or more core features and the assessment conclusion. The report format adopts Markdown format and is compatible with the requirements of embedding electronic medical records in the doctor's workstation. The result linkage module is used to match targeted intervention strategies. The inputs are the binary judgment result of the high-value prediction module and the risk level of the churn risk assessment module, and the output is a personalized intervention plan. It should be noted that the result linkage module formulates different strategies based on a value-risk four-quadrant matrix, where the four quadrants are high-value-high-risk, high-value-low-risk, low-value-high-risk, and low-value-low-risk, respectively. The determination of the intervention timing parameters combines the patient's treatment cycle characteristics extracted from the large language model and the golden window period for clinical intervention. Specifically, the method is as follows: The smaller of the following two values is used: the product of the treatment cycle characteristics and the first adjustment factor, and the product of the golden window of clinical intervention and the second adjustment factor. Both adjustment factors are determined based on specialist clinical practice data. The model iteration module is used to dynamically optimize the evaluation model. The inputs are doctors' feedback data on the evaluation results, actual patient churn and value conversion results data, and newly added multimodal feature data. The output is the updated joint evaluation model parameters.
[0019] It should be noted that a batch update is triggered when the number of new samples reaches 20% of the historical sample size or when the model's prediction accuracy drops by more than 5% for three consecutive months. New data is absorbed in real time through incremental learning on a daily basis. During the update process, it is necessary to maintain the consistency of clinical knowledge of the LLM and the predictive stability of the ML model. As an optional embodiment, the specific working steps of the data acquisition module are as follows: The subsequent feature construction module divides the data into three levels of priority. For unstructured text data, field-level dynamic data masking is used; for structured data, range-based data masking is used. Data collection must pass integrity, consistency, and timeliness checks. Data that fails the checks will trigger a second data collection.
[0020] It should be noted that the subsequent feature construction module has a list of core indicators for each specialty, which are determined by the clinical diagnosis and treatment guidelines of each specialty. For example, the specialty of liver cirrhosis with upper gastrointestinal bleeding prioritizes the collection of hemoglobin and alanine aminotransferase values from the LIS, and the specialty of end-stage renal disease CKD stage 4 prioritizes the collection of serum creatinine and hemoglobin values from the LIS. The three priority levels are specifically divided into emergency / first aid data, key indicators of the current visit, and historical archived data. The latency requirements for the three priority data are: emergency / first aid data latency of no more than 1 second, key indicators of the current visit latency of no more than 5 seconds, and historical archived data are synchronized in batches every morning. The scheduling logic is linked in real time with the indicator timeliness requirement interface of the feature construction module to ensure that the real-time feature data required by the high-value prediction module is supplied first. The field-level dynamic desensitization rules are as follows: core clinical terms that need to be referenced in the interpretability module retain their original expressions, and patient privacy information is desensitized using the SHA-256 hash algorithm. Specifically, the hash value is equal to the result of the SHA-256 algorithm calculating the combination of the plaintext of the privacy information and the random salt value. The range anonymization rule is as follows: age is converted into an interval form composed of adjacent age segment nodes, and cost data is converted into a range form. The anonymization rule is linked with the privacy compliance verification interface of the interpretability module. Integrity verification is determined by the non-empty rate of fields, and the non-empty rate of core fields must be no less than 99%. Consistency verification refers to the feature format specification library of the feature construction module, requiring LIS values to conform to the number + unit format and dates to conform to the YYYY-MM-DD format. Timeliness verification is determined by the difference between the data generation time and the collection time, and the difference of real-time data must not exceed the preset delay threshold. After data that fails verification triggers a second collection, the collection results are synchronized to the data quality log of the model training module in real time. This log records the collection time, data source, verification result, and reason for the anomaly, providing a basis for data deviation correction in subsequent model iteration modules. As an optional embodiment, the specific working steps of the feature construction module are as follows: A specialty feature priority database is constructed based on specialty clinical diagnosis and treatment guidelines, and the feature selection logic is dynamically adjusted through feature importance scoring; A time decay coefficient is introduced into the historical medical visit behavior characteristics to calculate the weighted features; Anomaly traceability labels are used to mark any abnormal features discovered during the cleaning process.
[0021] It should be noted that the specialty feature priority database is constructed according to specialty classification. For example, for the specialty of liver cirrhosis with upper gastrointestinal bleeding, core features such as Hb value fluctuation range, semantic vector of hematemesis-related chief complaint, and frequency of hospitalization within the recent time window are prioritized in the LIS. The time window is determined according to the specialty treatment cycle and is usually 30 days. For the specialty of end-stage renal disease CKD stage 4, core features such as Cr value, calcium-phosphorus product, and completion rate of dialysis preparation-related examinations within the recent time window are prioritized in the LIS. The time window is usually 90 days. Feature importance scores are calculated by weighting the SHAP value of the ML model in the model training module with the semantic relevance of the LLM. Specifically, the feature importance score equals the product of the first weight coefficient and the feature's SHAP value, plus the product of the second weight coefficient and the feature's semantic relevance. The feature's SHAP value reflects its contribution to the ML model's prediction results, while the feature's semantic relevance is obtained by calculating the cosine similarity between the feature text and the description of the specialty disease using the LLM. The sum of the two weight coefficients is 1, and their specific values are determined based on the specialty clinical focus. The screening results are synchronized to the priority update interface of the data acquisition module to optimize the focus of subsequent data acquisition. The formula for calculating the time decay coefficient is: the time decay coefficient is equal to the power of the product of the negative decay coefficient of the natural index and the difference between the current time and the historical consultation time; where the decay coefficient is determined based on the follow-up period of the specialty disease, it is usually taken as 0.01 to 0.05, and the time decay coefficient of recent features such as the last 7 days is closer to 1 and has a higher weight; this coefficient is linked to the time series risk analysis interface of the churn risk assessment module. For example, the weight of the churn rate after decay in the last 1 day of a specific time window is equal to the product of the time decay coefficient and the original churn rate, where the specific time window 1 is usually taken as 14 days. This weight directly affects the score of the consultation behavior risk dimension of the ML scoring card; Abnormal features refer to feature values that exceed the clinically reasonable range, such as Hb values less than 30 g / L or greater than 200 g / L in LIS. Abnormal source tracing labels include equipment error, specimen problems, and data entry errors, which are automatically labeled based on the association rules between historical abnormal data and causes through the abnormal cause reasoning interface of the feature construction module. This label is synchronized to the report notes interface of the interpretability module to avoid abnormal features misleading the evaluation conclusions, and is also pushed to the badcase library of the model iteration module for subsequent model parameter correction, such as adjusting the weight of abnormal features or introducing abnormal handling branches.
[0022] As an optional embodiment, the specific working steps of the model training module are as follows: The total loss of the joint evaluation model is calculated by weighting the ML model loss and the LLM loss; a specialized error correction term is added during the hyperparameter optimization process; After each round of training, the results are verified through the real-time inference interface of the high-value prediction module.
[0023] It should be noted that the total loss of the joint evaluation model is calculated as follows: the total loss equals the product of the ML model loss weight coefficient and the ML model loss, plus 1 minus the product of the ML model loss weight coefficient and the LLM loss. Here, the ML model is a model whose classification loss is calculated based on the cross-entropy loss of high-value patient annotations. Specifically, the classification loss is equal to the inverse of the negative number of samples multiplied by the logarithmic product of the true labels of all samples and the probability that the model predicts the sample to be high-value, plus the sum of the logarithmic products of '1 minus the true label' and '1 minus the probability that the model predicts the sample to be high-value'. In the true label, 1 represents high value and 0 represents non-high value. The LLM is a medical-specific LLM model, whose semantic loss is calculated based on the consistency error of diagnostic and treatment semantics annotated by clinical experts. Specifically, the semantic loss is equal to 1 minus the cosine similarity between the semantic vector generated by the LLM parsing of the diagnostic and treatment text and the standard semantic vector annotated by clinical experts. The loss weight coefficients of the ML model are dynamically adjusted by the specialty feature quality score of the feature construction module. If the specialty feature quality score is not lower than the feature quality threshold, the first weight coefficient is used; if the specialty feature quality score is lower than the feature quality threshold, the second weight coefficient is used. The feature quality threshold is usually 0.95, determined according to the specialty clinical requirements for feature accuracy. The first weight coefficient is usually 0.6, the second weight coefficient is usually 0.6, and the first weight coefficient is greater than the second weight coefficient. The semantic loss of LLM is synchronized to the feature construction module through the semantic error feedback interface to correct the semantic vector generation weights of text features. For example, text features with high semantic errors have their weights increased in the feature vector. The specialty error correction term is used to optimize the hyperparameter search range. If the prediction error of the ML model for samples corresponding to specialty core features exceeds the error threshold, the hyperparameter search range is adjusted. For example, the search range for the number of leaf nodes of the model is adjusted from the interval formed by the upper and lower limits of the initial range to the interval formed by the lower limit of the initial range multiplied by '1 plus the adjustment coefficient' and the upper limit of the initial range multiplied by '1 plus the adjustment coefficient'. Among them, the specialty core features include hemoglobin values of less than 70 g / L for liver cirrhosis specialty and serum creatinine values of greater than 600 μmol / L for CKD stage 4 specialty. The sample prediction error is the prediction error of the specialty core feature samples. The error threshold is usually taken as 0.1, and the adjustment coefficient is usually taken as 0.5. The correction term parameters are derived from the specialty core feature dimension of the feature construction module. For example, the adjustment coefficient is positively correlated with the specialty core feature dimension. The specific method for real-time inference verification is as follows: input the validation set feature vector through the interface of the high-value prediction module and calculate the predicted AUC value; if the absolute value of the difference between the current round AUC value and the previous round AUC value exceeds the fluctuation threshold, feature re-selection or LLMpoppt fine-tuning is triggered. Feature re-selection means calling the feature construction module interface to recalculate the feature importance score, and LLMpoppt fine-tuning means adjusting the input prompt word template of LLM to ensure that the trained joint model is adapted to the inference requirements of the downstream prediction module; the fluctuation threshold is usually set to 0.05. As an optional embodiment, the specific working steps of the high-value prediction module are as follows: The judgment threshold is derived from the model's real-time error factor and the intervention resource factor; Stratify the confidence levels of the high-value probability values in the output; Predict feature subsets based on data marked by the data acquisition module.
[0024] It should be noted that the model's real-time error factor is calculated using linear interpolation based on the real-time AUC value pushed by the model training module. Specifically, the model's real-time error factor equals the high threshold minus the real-time AUC value minus the high AUC threshold divided by the product of the high AUC threshold minus the low AUC threshold and the high threshold minus the low threshold. The high AUC threshold is typically 0.92, and the threshold corresponding to when the real-time AUC value is not lower than the high AUC threshold is the high threshold, typically 0.6. The low AUC threshold is typically 0.88, and the threshold corresponding to when the real-time AUC value is not higher than the low AUC threshold is the low threshold, typically 0.55. The intervention resource factor is dynamically adjusted based on the resource occupancy rate of the result linkage module. Specifically, the intervention resource factor equals the base threshold plus the product of an adjustment coefficient and the resource occupancy rate minus the resource occupancy rate threshold. The base threshold is typically 0.58, the adjustment coefficient is typically 0.1, and the resource occupancy rate threshold is typically 0.5. When the resource occupancy rate is higher than the resource occupancy rate threshold, the intervention resource factor is increased; when the resource occupancy rate is lower than the resource occupancy rate threshold, the intervention resource factor is decreased. The resource occupancy rate is obtained in real-time through the resource status interface. The final threshold is calculated as follows: The final threshold is equal to the product of the model real-time error factor weight coefficient and the model real-time error factor itself, plus the product of the intervention resource factor weight coefficient and the intervention resource factor itself. The model real-time error factor weight coefficient is typically set to 0.7, and the intervention resource factor weight coefficient is typically set to 0.3, with the sum of the two weight coefficients being 1. This ensures that the threshold is both appropriate for model accuracy and aligns with the hospital's resource capacity. The confidence level stratification rule is as follows: when the probability value of a high-value sample is not lower than the high confidence threshold, it is marked as high-confidence high-value; when the probability value of a high-value sample is not lower than the low confidence threshold but lower than the high confidence threshold, it is marked as low-confidence high-value. The high confidence threshold is typically set to 0.85, and the low confidence threshold is the final decision threshold. This stratification rule is synchronized to the feedback data classification interface of the model iteration module for subsequent model confidence optimization, such as increasing the labeling priority for low-confidence samples. The data acquisition module labels emergency / critical care data, such as emergency hematemesis patients and acute kidney injury patients; feature subset prediction uses only a specific number of core features, usually 10, such as hemoglobin value, chief complaint semantic vector, and past history of severe illness; its inference time does not exceed a specific time, usually 20ms; the prediction result is directly pushed to the emergency intervention interface of the result linkage module, and a second verification is performed after the complete feature vector is completed, and the verification error is calculated, which is the absolute value of the difference between the prediction probability of the feature subset and the prediction probability of the complete feature vector; if the verification error exceeds the error threshold, usually 0.1, it is fed back to the model training module to optimize the emergency feature subset, such as adjusting the selection of core features or increasing the number of features; As an optional embodiment, the specific working steps of the churn risk assessment module are as follows: The LLM diagnosis and treatment trajectory reasoning module outputs diagnosis and treatment deviation labels, triggering the adjustment of indicator weights in the ML quantitative scoring card; when the ML scoring card outputs a high-risk warning, it triggers the deep semantic review of the LLM module; and the risk level threshold of the ML scoring card is corrected based on the specialty feature distribution of the feature construction module.
[0025] It should be noted that the diagnosis and treatment deviation label is generated by the LLM diagnosis and treatment trajectory reasoning module after inputting the patient's chief complaint, examination report, current diagnosis, and registered department. This includes types such as diagnosis omission, departmental deviation, and incomplete treatment plan. For example, when a diagnosis omission is marked as "esophageal and gastric varices not mentioned," the weight of the disease severity dimension of the ML scorecard is adjusted. The adjustment method is as follows: the weight of the disease severity dimension after adjustment is equal to the weight before adjustment multiplied by 1 plus the deviation adjustment coefficient. The weight before adjustment is usually 0.3, the deviation adjustment coefficient is usually 0.5, and the weight after adjustment is usually 0.45. Moreover, the adjustment logic is positively correlated with the deviation severity output by LLM. For example, the deviation adjustment coefficient for diagnosis omission of malignant diseases is greater than the deviation adjustment coefficient for diagnosis omission of benign diseases. High-risk warning means that the risk score output by the ML scoring card is not lower than the ML high-risk threshold, which is usually 35 points. At this time, the deep semantic review of the LLM module focuses on analyzing the patient's imaging reports such as CT / MRI descriptions, medical order execution records such as medication execution rate and examination completion rate, and supplements potential dropout causes such as low medication execution rate due to reimbursement issues, incomplete examinations due to transportation inconvenience. The review results are synchronized to the risk cause remarks interface of the interpretability module as a supplement to the risk judgment basis in the report. The distribution of specialty characteristics is determined by the mean of core characteristics of specialty patients statistically analyzed by the feature construction module, such as the average number of complications in patients with cirrhosis and the average number of complications in patients with stage 4 CKD. The risk level threshold of the ML scorecard is adjusted according to specialty, specifically as follows: the specialty risk level threshold is equal to the base threshold multiplied by 1 plus the sum of the product of the complication adjustment coefficient and the product of 'the average number of complications in the specialty minus the average number of complications in the entire specialty'. The base threshold is usually 30 points, and the complication adjustment coefficient is usually 0.1. For example, the high-risk threshold for cirrhosis is usually 32 points, and the high-risk threshold for stage 4 CKD is usually 35 points. In addition, the system calculates the rate of change of the patient's risk score in the most recent specific number of assessments, which is the ratio of the current score minus the historical score to the historical score. This specific number of assessments is usually 3. If the rate of change exceeds the rate of change threshold, which is usually 0.2, the high-frequency feature update of the feature construction module is triggered. The sliding window length of the time series feature is adjusted from the original length, which is usually 7 days, to a new length, which is usually 3 days. At the same time, the rate of change data is pushed to the model iteration module to optimize the weight of the time series feature in the ML model, such as increasing the weight of recent time series features.
[0026] As an optional embodiment, the specific working steps of the interpretability module are as follows: The system generates explanatory content based on a list of core features of specialties built on the feature construction module; it supplements the explanatory logic with feedback results from the dual-engine system of the churn risk assessment module; and it automatically traces the cause of the error if the doctor's feedback explanation does not match the clinical reality.
[0027] It should be noted that the interpretation of specialty-specific adaptation prioritizes the annotation of the impact of key specialty features. For example, the interpretation of liver cirrhosis focuses on the correlation between the fluctuation range of Hb value and the semantic vector of hematemesis-related chief complaint in LIS. The interpretation of end-stage renal disease stage 4 CKD focuses on the correlation between Cr value and the completion rate of dialysis preparation-related examinations in LIS. The interpretation content must clearly explain the causal relationship between feature values and assessment conclusions. For example, a continuous decrease in Hb value indicates an increased risk of bleeding, which leads to an increased probability of high value. The dual-engine feedback results include diagnostic bias labels for LLM (Less-Low Morbidity), such as an increase in the risk score due to the failure to mention esophageal and gastric varices, where the score change is due to LLM; and risk dimension scores for ML (Multi-Level Modeling), such as the proportion of the patient visit behavior risk dimension score to the total risk score, which is the ratio of the patient visit behavior risk dimension score to the total risk score multiplied by 100%, ensuring consistency between the interpretation and risk assessment logic. Error tracing is triggered by interpretation error markers in doctor feedback, such as incorrect feature association or insufficient basis for conclusions. The system automatically traces two levels: first, the feature layer, which calls the feature generation logs of the feature construction module to check for feature selection biases, such as missing key LIS indicators or unreasonable temporal feature window lengths, and calculates the feature missing rate, i.e., the ratio of missing features to the total number of features; if the feature missing rate exceeds the missing rate threshold (usually 0.1), it is marked as a feature layer error; second, the model layer, which calls the loss function logs of the model training module to check for dual-model collaboration errors, such as LLM semantic loss exceeding the semantic loss threshold (usually 0.3), if present, it is marked as a model layer error. The source tracing results are synchronized to the explanation error library of the model iteration module, which is used to optimize the Prompt template in a targeted manner, such as adding descriptions of the association between specialty features or feature selection rules, such as reducing the importance weight of features with high missing rates. The report format adopts a clinical decision-oriented structure, with the beginning fixed as the user's historical medical behavior record and current medical information analysis. The main text is divided into two first-level headings: value judgment basis and risk judgment basis. Each first-level heading is further divided into two second-level headings: historical dimension and current medical treatment dimension. Key conclusions are marked in bold. Paragraphs are separated by blank lines. The format is compatible with the electronic medical record embedding requirements of the doctor's workstation, such as supporting HTML format export. As an optional embodiment, the specific working steps of the result linkage module are as follows: establishing intervention resource priorities based on the confidence level stratification of the high-value prediction module and the risk change rate of the churn risk assessment module; adjusting the intervention frequency according to the risk score changes tracked by the churn risk assessment module; and collecting post-intervention effect data to calculate the strategy effectiveness.
[0028] It should be noted that intervention resources are prioritized based on a combination of confidence level and risk change rate, following the rules: High confidence high value + risk change rate greater than the risk change rate threshold (usually 0.2) > High confidence high value + risk change rate not greater than the risk change rate threshold > Medium confidence high value + risk change rate greater than the risk change rate threshold > Medium confidence high value + risk change rate not greater than the risk change rate threshold > Low confidence high value + high risk. Priority directly maps to resource allocation. For example, the highest priority is allocated to VIP health manager + multidisciplinary consultation, the medium priority to dedicated customer service + monthly follow-up, and the low priority to basic follow-up. Resource occupancy rate is fed back to the intervention resource factor in the high-value prediction module in real time. If the resource occupancy rate is higher than the high occupancy rate threshold (usually 0.9), the high-value judgment threshold is temporarily increased (usually 0.05) to avoid resource overload. The duration of the increase is positively correlated with the resource occupancy rate minus the high occupancy rate threshold. The rules for adjusting the intervention frequency are as follows: If the churn risk assessment module tracks a risk score decrease and the risk change rate is less than 0, the intervention frequency will be adjusted from the original frequency to the original frequency multiplied by '1 minus the absolute value of the risk change rate', such as changing from weekly follow-up to once every two weeks; if the risk score increase and the risk change rate is greater than 0, the frequency will be adjusted to the original frequency multiplied by '1 plus the risk change rate'; if the secondary verification interval of the high-value prediction module is usually 7 days and the probability value of low-confidence high value turning into non-high value is less than the judgment threshold, the VIP intervention will be terminated and converted to regular follow-up, and the adjustment record will be synchronized to the intervention adjustment log of the model iteration module. The strategy effectiveness data includes the patient follow-up rate after intervention (followed by a follow-up visit within 30 days, considered effective) and the doctor's adoption rate of the intervention recommendations. The strategy effectiveness rate is calculated as follows: the product of the follow-up rate weight coefficient and the follow-up rate, plus the product of the adoption rate weight coefficient and the adoption rate, divided by the sum of the follow-up rate weight coefficient and the adoption rate weight coefficient. The follow-up rate weight coefficient is usually 0.6, and the adoption rate weight coefficient is usually 0.4. The effectiveness rate data is pushed to the model iteration module to optimize the strategy weights in the value-risk matrix. For example, a strategy with low effectiveness will have its resource allocation ratio reduced. The adjustment method is that the new strategy weight is equal to the original strategy weight multiplied by the strategy effectiveness rate.
[0029] As an optional embodiment, the specific working steps of the model iteration module are as follows: perform hierarchical processing on multi-source feedback data and divide it into three levels according to the iteration frequency; implement fluctuation upper limit control on the weight of the core features of the specialty during the iteration process; and ensure system stability through full-link verification after each iteration.
[0030] It should be noted that multi-source feedback data includes the interpretation error of the interpretability module, such as error type and error rate; the strategy effectiveness of the result linkage module; the threshold calibration error of the high-value prediction module, i.e., the absolute value of the difference between the predicted threshold and the actual optimal threshold; the risk change rate distribution and statistical distribution characteristics of the risk change rate of the churn risk assessment module; the feature quality changes of the feature construction module, such as the increase in the feature missing rate and the increase in the proportion of abnormal features; and the data priority scheduling effect of the data acquisition module, such as the collection delay compliance rate of the data marked by the data acquisition module. The tiered processing rules are as follows: The first tier is high-frequency, lightweight iteration, typically with a 7-day cycle. It receives interpretation errors and strategy effectiveness, and is used to fine-tune the LLM Prompt template, such as optimizing specialist feature descriptions, adding descriptions incorporating patient bleeding history analysis, and adjusting the ML model feature weights, such as increasing the weights of features with high strategy effectiveness. The adjustment method is that the new feature weight equals the original feature weight multiplied by 1 plus the product of the effectiveness coefficient and the strategy effectiveness. The effectiveness coefficient is typically set to 0.2. The second tier is medium-frequency, moderate-intensity iteration, typically with a 30-day cycle. It receives threshold calibration errors and risk change rate distribution, and is used for... The dynamic threshold formula of the high-value prediction module is adjusted, such as correcting the weight coefficient of the real-time error factor of the model, the weight coefficient of the intervention resource factor, and recalculating the specialty risk threshold of the loss risk assessment module; the third level is low-frequency deep iteration, with a cycle of 90 days, which receives changes in feature quality and data priority scheduling effects, and is used to retrain the joint evaluation model. The merged dataset of new samples and historical samples is used to update the specialty feature priority library of the feature construction module and recalculate the feature importance score and the data collection priority of the data collection module, such as increasing the collection priority of features with high missing rates. The upper limit of fluctuation control targets core specialty features such as hemoglobin values for cirrhosis and serum creatinine values for stage 4 CKD. The fluctuation range of their weights is such that the absolute value of the difference between the weight of the new feature and the weight of the original feature does not exceed the fluctuation threshold. The fluctuation threshold is usually set to 0.15 to avoid the model’s sensitivity to key specialty indicators decreasing due to iteration. The constraint parameters are derived from the specialty feature importance score of the feature construction module. For example, the fluctuation threshold is positively correlated with the specialty feature importance score. The end-to-end validation uses a historical test set and training set from the data acquisition module that have no overlap. It sequentially goes through the stages of feature construction, model training, high-value prediction, risk assessment, interpretability generation, and result linkage. Validation indicators include: the absolute value of the difference between the current iteration AUC and the previous version AUC does not exceed the iteration AUC fluctuation threshold, which is usually set to 0.03; the physician acceptance rate of the interpretation report is not lower than the report acceptance rate threshold, which is usually set to 0.8; and the effectiveness rate of the intervention strategy is not lower than the strategy effectiveness threshold, which is usually set to 0.7. If the validation fails, it rolls back to the previous version to ensure the stability of the system throughout the entire process.
[0031] This invention also proposes a method for assessing patient value and churn risk based on ML combined with LLM, comprising the following steps: S1: Dynamic priority data collection. Based on the list of core indicators for specialties in the subsequent feature construction steps, the data from the hospital's HIS, LIS, PACS, and EMR systems are subject to three-level priority scheduling. At the same time, layered desensitization processing is adopted. After the collected data undergoes integrity, consistency, and timeliness verification, it is synchronized to the data quality logs of the feature construction and model iteration steps. S2: Specialty Adaptive Feature Construction. Based on specialty diagnosis and treatment guidelines, a specialty feature priority library is built. Core features are selected through feature importance scoring. Missing value imputation, outlier removal, and normalization are performed on structured data. Semantic vectors are generated for text data. Time decay coefficients are introduced to calculate weighted features for time-series behavioral data. The selected multimodal feature vectors are synchronized to the model training step. Feature format standardization is synchronized to the data consistency verification step in S1. Abnormal feature annotation and source tracing labels are pushed to the interpretability step and the model iteration step. S3: Dual-model collaborative training, constructing a joint evaluation model of ML and medical-specific LLM, using a dynamic loss function to calculate the total loss, optimizing hyperparameters through a hyperparameter optimization framework, and adding a specialty error correction term. During the training process, each round is verified through the real-time inference interface of the high-value prediction step, and the finally trained joint model is pushed to the high-value prediction step. S4: Dynamic threshold high-value prediction. It receives the real-time multimodal feature vector of patients from S2 and the joint model of S3. It uses a dual-driving factor dynamic threshold to derive and determine the threshold, outputs the high-value probability value and stratifies it according to confidence level. The prediction results are synchronized to the prediction error log of the churn risk assessment step, the interpretability step and the model iteration step. S5: Dual-engine feedback loss risk assessment. Based on the high-value prediction results of S4, the dual engines of LLM diagnosis and treatment reasoning and ML quantitative scoring are launched. The LLM engine outputs deviation labels to trigger the weight adjustment of the ML scoring card. The ML engine calculates the risk score to trigger the deep semantic review of the LLM engine. Finally, the risk level and cause analysis are synchronized to the interpretability step and the result linkage step. The dual-engine feedback data is pushed to the model iteration step. S6: Specialty-adapted interpretability report generation. It receives the high-value judgment results from S4 and the risk assessment results from S5, generates a report according to clinical decision-making guidance, and synchronizes the report to the results linkage step after privacy compliance verification. The interpretation error reported by the doctor is pushed to the error source library of the model iteration step. S7: Dynamic resource scheduling intervention. Based on the confidence level of S4 and the risk change rate of S5, intervention resource priority is established, and a value-risk four-quadrant strategy is matched. The intervention resource occupancy rate is fed back to the intervention resource factor of S4 in real time, and the post-intervention effect data is pushed to the model iteration steps. S8: Multi-source feedback hierarchical iteration. Based on the feedback data of S1-S7, a three-level iteration is implemented. During the iteration process, the fluctuation upper limit of the weight of the core features of the specialty is controlled. After the iteration, the whole link is verified. If it fails, it is rolled back to the previous version.
[0032] It should be noted that the latency requirements for the three-level priority scheduling in S1 are as follows: Priority 1 emergency / critical care data latency should not exceed 1 second; Priority 2 key indicators of the current visit latency should not exceed 5 seconds; and Priority 3 historical archived data should be batch synchronized every morning at midnight. In the layered desensitization, privacy information is desensitized using SHA-256 hashing. Specifically, the hash value is equal to the result of calculating the combination of the plaintext privacy information and a random salt value using the SHA-256 algorithm. Clinical core terms retain their original expressions. The integrity check of the triple check is determined by the non-empty rate of the core fields being no less than 99%. The consistency check refers to the feature format specification of S2. The timeliness check is determined by the data latency not exceeding a preset threshold.
[0033] In S2, the feature importance score is calculated as follows: the feature importance score equals the product of the first weight coefficient and the feature's SHAP value, plus the product of the second weight coefficient and the feature's semantic relevance, and the sum of the two weight coefficients is 1; the time decay coefficient is calculated as follows: the time decay coefficient equals the product of the negative decay coefficient of the natural exponent and the difference between the current time and the historical consultation time raised to the power of the product; structured data normalization uses the Min-Max method, that is, the normalized value equals the original value minus the minimum value of the feature and the maximum value minus the minimum value of the feature; text data generates semantic vectors through the BioBERT model; The dynamic loss function in S3 is calculated as follows: the total loss equals the product of the ML model loss weight coefficient and the ML model loss, plus 1 minus the product of the ML model loss weight coefficient and the LLM loss; where the ML model loss is the cross-entropy loss, specifically calculated as the cross-entropy loss equals the product of the reciprocal of the negative number of samples multiplied by the logarithm of the true label of all samples and the probability that the model predicts the sample to be of high value, plus the sum of the logarithms of '1 minus the true label' and '1 minus the probability that the model predicts the sample to be of high value'; the LLM loss is the semantic loss, specifically calculated as the semantic loss equals 1 minus the cosine similarity between the semantic vector generated by the LLM parsing of the diagnostic text and the standard semantic vector annotated by clinical experts; hyperparameter optimization uses the Optuna framework; the specialty error correction term adjusts the hyperparameter search range based on the prediction error of the specialty core features; The dynamic threshold in S4 is calculated as follows: the decision threshold is equal to the product of the model real-time error factor weight coefficient and the model real-time error factor, plus the product of the intervention resource factor weight coefficient and the intervention resource factor, and the sum of the two weight coefficients is 1; where the model real-time error factor is calculated by linear interpolation from the real-time AUC, and the intervention resource factor is dynamically adjusted by the resource occupancy rate; confidence level is divided into high confidence (not lower than the high confidence threshold) and low confidence (not lower than the decision threshold but lower than the high confidence threshold). In S5, the LLM engine's bias label triggers ML scorecard weight adjustments as follows: the adjusted disease severity dimension weight equals the original weight multiplied by 1 plus the bias adjustment coefficient; the ML scorecard's specialty risk threshold is calculated as follows: the specialty risk threshold equals the base threshold multiplied by 1 plus the complication adjustment coefficient and the product of 'specialty average complication count minus overall specialty average complication count'; the risk change rate is calculated as follows: the risk change rate equals the current score minus the historical score and the ratio of the historical scores. The S6 report format begins with an analysis of the user's historical medical records and current medical information, divided into two parts: value judgment basis and risk judgment basis, including a feature-conclusion causal chain; privacy compliance verification references the de-identification rules of S1; The S7 Value-Risk Four-Quadrant Strategy includes: High Value + High Risk corresponds to VIP Health Manager + Multidisciplinary Consultation; High Value + Low Risk corresponds to Routine Follow-up + Health Education; Low Value + High Risk corresponds to Basic Follow-up + Cost Control Reminders; and Low Value + Low Risk corresponds to System Archiving. The strategy effectiveness is calculated as follows: Strategy effectiveness equals the product of the follow-up visit rate weight coefficient and the follow-up visit rate, plus the product of the adoption rate weight coefficient and the adoption rate, and then divided by the sum of the follow-up visit rate weight coefficient and the adoption rate weight coefficient. The three-level iteration cycle in S8 is: 7 days for high frequency, 30 days for medium frequency, and 90 days for low frequency; the fluctuation range of the core feature weight is that the absolute value of the difference between the new feature weight and the original feature weight does not exceed the fluctuation threshold; the full-link validation indicators include AUC fluctuation not exceeding 0.03, explanation acceptance rate not less than 0.8, and strategy effectiveness not less than 0.7, to ensure the clinical adaptability and stability of the method throughout the entire process.
[0034] Example: Example 1: Specialist assessment of patients with cirrhosis complicated by upper gastrointestinal bleeding This embodiment takes a patient with cirrhosis complicated by upper gastrointestinal bleeding admitted to the Department of Gastroenterology of a tertiary hospital as an example to illustrate the specific implementation process of the system and method described in this invention. All operations follow the full-link logic of data collection, feature construction, model training, high-value prediction, loss risk assessment, interpretability reporting, result linkage, and model iteration, and the core parameters are deeply adapted to the clinical diagnosis and treatment guidelines of the specialty.
[0035] 1. Data collection; Connect to the hospital's HIS, LIS, PACS, and EMR systems via the HL7 / FHIR standard interface, and implement three-level priority data collection based on the list of core specialty indicators: Priority 1 (Emergency Data): Patient's chief complaint of intermittent hematemesis for 3 hours, hemoglobin (Hb) value detected by LIS in real time, emergency registration record, collection delay ≤1 second, after integrity check (core fields non-empty rate 100%), consistency check (Hb value format is numerical + g / L), and timeliness check (data generation and collection time difference <1 second) are synchronized to the system; Priority 2 (Key indicators for this visit): Abdominal CT report text from PACS system (including description of esophageal and gastric varices), doctor's preliminary diagnosis (decompensated cirrhosis, upper gastrointestinal bleeding), details of this outpatient expenses, collection delay ≤ 5 seconds, original expression of clinical core terms in unstructured text, and patient ID number desensitized by SHA-256 hash; Priority 3 (Historical Archived Data): Patient's hospitalization records for the past 12 months (2 hospitalizations related to cirrhosis), Hb value fluctuation records in previous LIS examinations, and records of canceled appointments. Data such as historical hospitalization frequency and cost range are converted into interval format according to range anonymization rules to ensure privacy compliance.
[0036] After the data collection is completed, the data is synchronized to the feature construction module in real time, and the collection time, data source, and verification results are recorded in the data quality log of the model training module. There is no abnormal data in the verification.
[0037] 2. Feature construction; Based on the "Guidelines for the Diagnosis and Treatment of Liver Cirrhosis", a priority database of specialty features for liver cirrhosis complicated with upper gastrointestinal bleeding was constructed, and specialty adaptive feature selection and multimodal feature vector generation were performed: Feature selection: Based on feature importance scoring (combining LightGBM model SHAP value and medical LLM semantic correlation), 15 core features are retained, including the fluctuation range of Hb value in LIS during the current visit, the semantic vector of hematemesis, the frequency of hospitalization in the past 30 days, the number of matching keywords related to varicose veins in CT reports, the cancellation rate in the past 90 days, and the medication execution rate. The selection results are synchronized to the data acquisition module to optimize the focus of subsequent data acquisition for similar patients. Feature processing: For structured data (such as Hb values and hospitalization frequency), median imputation for missing values (without missing values), IQR method for outlier removal (without outliers), and Min-Max normalization are used to convert Hb values into standardized feature vectors; for text data (such as CT reports and chief complaints), semantic vectors are generated using a BioBERT pre-trained model, and entity relationship triples for hematemesis-3 hours-esophageal varices are extracted; for time-series behavioral data (such as hospitalization frequency in the past 30 days), a time decay coefficient (k=0.03, with higher weight for recent hospitalization behavior) is introduced to calculate weighted features; Anomaly labeling: The collected data did not contain any features that exceeded the clinically reasonable range (although the Hb value was lower than the reference range, it was consistent with the clinical manifestations of upper gastrointestinal bleeding and was judged as a valid feature). No abnormal source labeling was required. Finally, a 52-dimensional multimodal feature vector was generated and synchronized to the model training module and the high-value prediction module.
[0038] 3. Model training; The LightGBM + medical-grade LLM joint assessment model was used, trained based on historical data of patients with cirrhosis and upper gastrointestinal bleeding in this hospital over the past 24 months (including 1200 valid samples, divided into training, validation, and test sets in an 8:1:1 ratio). Dynamic loss function settings: Since the quality score of the specialty features output by the feature construction module (core feature missing rate <1%) is higher than the threshold, α=0.6 (ML model loss weight) is set. L_ML adopts cross-entropy loss (based on high-value patients = the top 30% of pure medical service fees or the annotations of patients hospitalized within 30 days), and L_LLM adopts semantic loss (based on the standard semantic annotations of clinical experts on diagnosis and treatment texts, the cosine similarity deviation between the LLM parsing vector and the expert annotation vector is calculated). Hyperparameter optimization: Hyperparameters were searched using the Optuna framework. Since the prediction error of the ML model for samples with Hb < 70 g / L was below the threshold, the search range of num_leaves was set to [60, 100], and learning_rate was set to 0.08 according to the principle of balancing prediction accuracy and computational efficiency. After each round of training, the AUC was verified through the real-time inference interface of the high-value prediction module. The fluctuation was < 5%, and there was no need to trigger feature re-selection. Model validation: The trained joint model achieved a high-value identification AUC of 0.93 on the validation set and a churn risk prediction accuracy of 86%, meeting clinical accuracy requirements. It was then pushed to the high-value prediction module for backup.
[0039] 4. High-value forecasting; Input the real-time multimodal feature vector of the patient generated in step 2 into the trained joint model: Dynamic threshold calculation: The real-time AUC pushed by the model training module is 0.93, and τ_model=0.6 is obtained by linear interpolation; the resource utilization rate of VIP health manager reported by the result linkage module is 60% (below the threshold), so τ_res=0.57 is obtained. The final judgment threshold τ=0.59 is calculated according to the weights ω_model=0.7 and ω_res=0.3. Prediction results: The model outputs a high-value probability of 0.88 for this patient, which is higher than τ=0.59 and meets the high confidence stratification criteria of P≥0.85. The patient is marked as a high-confidence high-value patient (judgment criteria: low Hb value indicates the need for emergency endoscopic hemostasis, and the previous hospitalization history shows high resource consumption, which meets the definition of a high-value patient). Emergency Validation: Since the patient is an emergency case, feature subset prediction is triggered. Only 10 core features, including Hb value, chief complaint semantic vector, and history of bleeding, are used for rapid inference. The time taken is 18ms, the prediction probability is 0.86, and the error with the full feature prediction result is <10%. The validation is successful, and the prediction result is synchronized to the churn risk assessment module and the interpretability module.
[0040] 5. Attrition risk assessment; Activate the dual-engine approach of LLM diagnostic reasoning and ML quantitative scoring: LLM diagnostic reasoning: Input the patient's chief complaint, CT report, and preliminary diagnosis. LLM reasons based on the "Guidelines for the Diagnosis and Treatment of Liver Cirrhosis" and outputs a diagnostic deviation label - the diagnosis does not clearly indicate esophageal and gastric variceal rupture and bleeding, and there is no recommendation for interventional radiology consultation. This triggers the weight of the disease severity dimension of the ML quantitative scoring card to be increased from 0.3 to 0.45. ML quantitative scoring: Based on the time-weighted features in step 2, the scores for each dimension were calculated as follows: Basic information dimension (age 58 years old, local resident) scored 7 points; medical behavior dimension (return appointment rate of 5% in the past 90 days, follow-up visit interval of 45 days) scored 12 points; treatment adherence dimension (previous medication execution rate of 85%) scored 5 points; disease severity dimension (2 complications, 2 hospitalizations) scored 8 points, with a total basic score of 32 points; combined with the specialty risk threshold (high-risk threshold for cirrhosis specialty = 32 points), it was marked as high risk; Dual-engine integration: LLM re-examination supplements potential churn causes - the patient is a resident of another province, and the cost of medical treatment in another place may affect subsequent follow-up visits. The final risk score is 32×1.05 (LLM bias adjustment coefficient) = 33.6 points, which is determined to be a high churn risk. The results are synchronized to the interpretability module and the result linkage module.
[0041] 6. Generate interpretability reports; Receive high-value prediction results and churn risk assessment results, and generate reports based on clinical decision-making guidelines: Report Structure: The opening is fixed as an analysis of the user's historical medical records and current medical information; under the first-level heading "Value Judgment Criteria," it is divided into historical dimensions (two hospitalizations for cirrhosis in the past 12 months, high resource consumption) and current medical dimension (vomiting blood with significantly reduced Hb levels, requiring emergency intervention, meeting the characteristics of a high-value patient); under the first-level heading "Risk Judgment Criteria," it is divided into treatment deviations (unclear diagnosis of varicose vein rupture and bleeding, which may affect the integrity of treatment) and behavioral risks (residents from other provinces, seeking medical treatment in other places may increase the probability of attrition). Causal chain explanation: The report highlights the core logic of decreased Hb value → demand for emergency intervention → increased probability of high-value residence in other provinces → low convenience of follow-up visits → increased risk of attrition. No original indicators are listed. The report is pushed to the doctor's workstation after privacy compliance verification (no anonymization omissions).
[0042] 7. Results linkage; The four-quadrant matching rule is based on high value and high risk: Resource allocation: VIP health managers are assigned based on priority, high confidence, high value, and high risk; multidisciplinary team (MDT) consultations are coordinated between gastroenterology and interventional radiology departments; intervention timing is determined by combining the golden window of clinical intervention (48 hours for bleeding in cirrhosis) with the patient's treatment cycle, and is set to proactive intervention within 24 hours. Real-time feedback: The VIP health manager resource occupancy rate is fed back to the high-value prediction module in real time. The current occupancy rate has risen to 75%, but the threshold adjustment has not been triggered. Subsequent tracking shows that the patient will have a follow-up visit within 30 days, and the intervention effectiveness rate is recorded in the strategy effect log of the results linkage module.
[0043] 8. Model iteration; After the patient assessment is completed, the relevant data is incorporated into the model iteration feedback: Feedback data collection: The physicians' acceptance rate of the interpretability report was 100% (no error feedback), no patients were lost after the intervention, and the strategy's effectiveness rate met the target; Iteration Trigger: Since the number of new samples has not reached 20% of the historical sample size and the model prediction accuracy has not declined continuously, batch updates will not be triggered for the time being. Instead, the LLMPrompt template will be fine-tuned through high-frequency iterations to add guiding statements that highlight varicose vein-related diagnoses for patients with bleeding cirrhosis, ensuring more accurate assessments of similar patients in the future.
[0044] Example 2: Specialist assessment of a stage 4 CKD patient with end-stage renal disease; This embodiment takes a patient with end-stage renal disease (CKD) stage 4 in the nephrology department of a tertiary hospital as an example to further verify the adaptability of the present invention in different specialties. The implementation process is consistent with the logic of embodiment one, and the core difference lies in the matching of specialty characteristics, model parameters and clinical needs.
[0045] 1. Data collection; Collected according to the list of core indicators for stage 4 CKD (end-stage renal disease): Priority 1 data: Patient's serum creatinine (Cr) value from this LIS test, chief complaint of fatigue and poor appetite, with a collection delay of <1 second; Priority 2 data: Kidney function test report (five items), kidney ultrasound report (including description of bilateral kidney atrophy), doctor's diagnosis of stage 4 chronic kidney disease; Priority 3 data: Records of dialysis preparation-related examinations completed in the past 90 days (2 items not completed), records of 3 previous hospitalizations in the nephrology department, and privacy information is anonymized according to the range (age converted to the 65-70 age range).
[0046] 2. Feature construction; Based on the "Guidelines for the Diagnosis and Treatment of Chronic Kidney Disease", 12 core features were selected, including Cr value, calcium-phosphorus product, completion rate of dialysis preparation examination in the past 90 days, and outpatient frequency in the past 30 days. The time-series features adopted a 90-day sliding window with a time decay coefficient k=0.02. The entity relationship between bilateral renal atrophy and chronic kidney disease was extracted from text data (such as ultrasound reports) to generate a 50-dimensional multimodal feature vector with no abnormal feature annotations.
[0047] 3. Model training; The training samples for the joint model were data from 1000 CKD stage 4 patients in the hospital over the past 24 months. Due to the high feature quality score, α=0.6 was used. The hyperparameter num_leaves was searched within the range of [50,90]. After training, the model's high-value identification AUC was 0.92, and the churn risk prediction accuracy was 84%.
[0048] 4. High-value forecasting; The dynamic threshold τ=0.58, the model output patient high value probability=0.91 (high confidence high value, judgment criteria: high Cr value indicates need for dialysis preparation, previous hospitalization shows high resource dependence), the emergency feature subset prediction time is 20ms, and the validation error is <10%.
[0049] 5. Attrition risk assessment; LLM inference: The output diagnosis did not mention secondary hyperparathyroidism, which is a bias label that requires collaborative evaluation by an endocrinologist. The weight of the ML disease severity dimension was increased to 0.4. ML score: Basic information (residents from other provinces) scored 16 points, medical behavior (cancellation rate 0%) scored 7 points, compliance (drug execution rate 95%) scored 2 points, disease severity (3 complications) scored 13 points, basic total score = 38 points, combined with the CKD stage 4 specialty high risk threshold (35 points), high risk was marked; The final risk score was 38 × 1.03 = 39.14 points. The supplementary examination for LLM re-examination, medical treatment in another province, and dialysis preparation examination were not completed, increasing the risk of attrition.
[0050] 6. Explainability report; The report focuses on explaining the causal chain of abnormal Cr levels → dialysis need → incomplete high-value dialysis examinations + out-of-town medical treatment → high risk. The report format conforms to clinical reading habits and has a 100% adoption rate among doctors.
[0051] 7. Results linkage; Matching high-value + high-risk strategies: assigning a dedicated follow-up nurse, weekly telephone follow-up, scheduling endocrinology consultations, and providing feedback on intervention resource occupancy rates to the prediction module to avoid triggering threshold adjustments. Patients subsequently complete dialysis preparation examinations, and the 30-day follow-up rate is 100%.
[0052] 8. Model iteration; We collected data on the effectiveness of the intervention (strategy effectiveness rate 85%), and fine-tuned the CKD stage 4 specialty risk threshold to 34 points through mid-frequency iterations to ensure that subsequent assessments are more adapted to clinical practice.
[0053] As can be seen from the two specialized embodiments above, the system and method described in this invention can adaptively adjust the core parameters and processes according to the clinical needs of different specialties. The evaluation results are accurate and highly interpretable, and can form a prediction-intervention-feedback closed loop, which fully meets the specialized and refined patient management needs of medical institutions.
[0054] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A patient value and churn risk assessment system based on ML combined with LLM, characterized in that, It includes a data acquisition module, a feature construction module, a model training module, a high-value prediction module, a churn risk assessment module, an interpretability module, a result linkage module, and a model iteration module; The data acquisition module is used to collect patients' historical and real-time multi-source medical data; The feature construction module is used to integrate the data collected by the data acquisition module to generate multimodal feature vectors; The model training module is used to build an evaluation model using ML and LLM. The input is the multimodal feature vector generated by the feature construction module, and the output is the trained joint evaluation model. The high-value prediction module is used to output patient value judgment results based on the joint assessment model. The input is the joint assessment model generated by the model training module and the real-time multimodal feature vector of the patient. The output is the patient's high-value probability value and binary judgment result. The churn risk assessment module is used to comprehensively analyze the possibility of patient churn. The input is the patient's real-time multimodal feature vector, and the output is the churn risk score and risk level. The interpretability module is used to generate clinically understandable assessment criteria. The inputs are the judgment results of the high-value prediction module, the risk score of the churn risk assessment module and the corresponding multimodal feature vector, and the output is a structured interpretation report. The result linkage module is used to match targeted intervention strategies. The inputs are the binary judgment result of the high-value prediction module and the risk level of the churn risk assessment module, and the output is a personalized intervention plan. The model iteration module is used to dynamically optimize the evaluation model. The inputs are doctors' feedback data on the evaluation results, actual patient churn and value conversion results data, and newly added multimodal feature data. The output is the updated joint evaluation model parameters.
2. The patient value and churn risk assessment system based on ML combined with LLM according to claim 1, characterized in that, The specific working steps of the data acquisition module are as follows: The subsequent feature construction module divides the data into three levels of priority. For unstructured text data, field-level dynamic data masking is used; for structured data, range-based data masking is used. Data collection must pass integrity, consistency, and timeliness checks. Data that fails the checks will trigger a second data collection.
3. The patient value and churn risk assessment system based on ML combined with LLM according to claim 1, characterized in that, The specific working steps of the feature construction module are as follows: A specialty feature priority database is constructed based on specialty clinical diagnosis and treatment guidelines, and the feature selection logic is dynamically adjusted through feature importance scoring; A time decay coefficient is introduced into the historical medical visit behavior characteristics to calculate the weighted features; Anomaly traceability labels are used to mark any abnormal features discovered during the cleaning process.
4. The patient value and churn risk assessment system based on ML combined with LLM according to claim 1, characterized in that, The specific working steps of the model training module are as follows: The total loss of the joint evaluation model is calculated by weighting the ML model loss and the LLM loss; a specialized error correction term is added during the hyperparameter optimization process; After each round of training, the results are verified through the real-time inference interface of the high-value prediction module.
5. The patient value and churn risk assessment system based on ML combined with LLM according to claim 1, characterized in that, The specific working steps of the high-value prediction module are as follows: The judgment threshold is derived from the model's real-time error factor and the intervention resource factor; Stratify the confidence levels of the high-value probability values in the output; Predict feature subsets based on data marked by the data acquisition module.
6. The patient value and churn risk assessment system based on ML combined with LLM according to claim 1, characterized in that, The specific working steps of the churn risk assessment module are as follows: The LLM diagnostic trajectory reasoning module outputs diagnostic deviation labels, triggering the adjustment of indicator weights in the ML quantitative scoring card; when the ML scoring card outputs a high-risk warning, it triggers the deep semantic review of the LLM module. Risk level thresholds of ML scoring cards are corrected based on the distribution of specialty features built using feature building modules.
7. The patient value and churn risk assessment system based on ML combined with LLM according to claim 1, characterized in that, The specific working steps of the interpretability module are as follows: The system generates explanatory content based on a list of core features of specialties built on the feature construction module; it supplements the explanatory logic with feedback results from the dual-engine system of the churn risk assessment module; and it automatically traces the cause of the error if the doctor's feedback explanation does not match the clinical reality.
8. The patient value and churn risk assessment system based on ML combined with LLM according to claim 1, characterized in that, The specific working steps of the result linkage module are as follows: establish intervention resource priorities based on the confidence level stratification of the high-value prediction module and the risk change rate of the churn risk assessment module; adjust the intervention frequency according to the risk score changes tracked by the churn risk assessment module; and collect post-intervention effect data to calculate the strategy effectiveness.
9. A patient value and churn risk assessment system based on ML combined with LLM according to claim 1, characterized in that, The specific working steps of the model iteration module are as follows: hierarchical processing of multi-source feedback data, divided into three levels according to the iteration frequency; fluctuation upper limit control of the weight of the core features of the specialty during the iteration process; and system stability is ensured through full-link verification after each iteration.
10. A method for assessing patient value and churn risk based on ML combined with LLM, applicable to the patient value and churn risk assessment system based on ML combined with LLM as described in any one of claims 1 to 9, characterized in that, Includes the following steps: S1: Dynamic priority data collection. Based on the list of core indicators for specialties in the subsequent feature construction steps, the data from the hospital's HIS, LIS, PACS, and EMR systems are subject to three-level priority scheduling. At the same time, layered desensitization processing is adopted. After the collected data undergoes integrity, consistency, and timeliness verification, it is synchronized to the data quality logs of the feature construction and model iteration steps. S2: Specialty Adaptive Feature Construction. Based on specialty diagnosis and treatment guidelines, a specialty feature priority library is built. Core features are selected through feature importance scoring. Missing value imputation, outlier removal, and normalization are performed on structured data. Semantic vectors are generated for text data. Time decay coefficients are introduced to calculate weighted features for time-series behavioral data. The selected multimodal feature vectors are synchronized to the model training step. Feature format standardization is synchronized to the data consistency verification step in S1. Abnormal feature annotation and source tracing labels are pushed to the interpretability step and the model iteration step. S3: Dual-model collaborative training, constructing a joint evaluation model of ML and medical-specific LLM, using a dynamic loss function to calculate the total loss, optimizing hyperparameters through a hyperparameter optimization framework, and adding a specialty error correction term. During the training process, each round is verified through the real-time inference interface of the high-value prediction step, and the finally trained joint model is pushed to the high-value prediction step. S4: Dynamic threshold high-value prediction. It receives the real-time multimodal feature vector of patients from S2 and the joint model of S3. It uses a dual-driving factor dynamic threshold to derive and determine the threshold, outputs the high-value probability value and stratifies it according to confidence level. The prediction results are synchronized to the prediction error log of the churn risk assessment step, the interpretability step and the model iteration step. S5: Dual-engine feedback loss risk assessment. Based on the high-value prediction results of S4, the dual engines of LLM diagnosis and treatment reasoning and ML quantitative scoring are launched. The LLM engine outputs deviation labels to trigger the weight adjustment of the ML scoring card. The ML engine calculates the risk score to trigger the deep semantic review of the LLM engine. Finally, the risk level and cause analysis are synchronized to the interpretability step and the result linkage step. The dual-engine feedback data is pushed to the model iteration step. S6: Specialty-adapted interpretability report generation. It receives the high-value judgment results from S4 and the risk assessment results from S5, generates a report according to clinical decision-making guidance, and synchronizes the report to the results linkage step after privacy compliance verification. The interpretation error reported by the doctor is pushed to the error source library of the model iteration step. S7: Dynamic resource scheduling intervention. Based on the confidence level of S4 and the risk change rate of S5, intervention resource priority is established, and a value-risk four-quadrant strategy is matched. The intervention resource occupancy rate is fed back to the intervention resource factor of S4 in real time, and the post-intervention effect data is pushed to the model iteration steps. S8: Multi-source feedback hierarchical iteration. Based on the feedback data of S1-S7, a three-level iteration is implemented. During the iteration process, the fluctuation upper limit of the weight of the core features of the specialty is controlled. After the iteration, the whole link is verified. If it fails, it is rolled back to the previous version.