Method for predicting and early warning of follow-up risk of chronic disease in combination with asynchronous electronic medical record sequence

CN122552012APending Publication Date: 2026-08-11THE FIRST AFFILIATED HOSPITAL OF ARMY MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0005]然而,现有医疗信息化与人工智能技术在适配慢病随访场景的风险预警需求时,面临着诸多现实瓶颈:一方面,慢病随访数据本身具有显著的“时序性、不规则性、异步性、高缺失性”,与常规结构化数据存在本质差异,传统数据处理与建模方法难以适配;另一方面,临床决策对风险预警的“实时性、可解释性”要求极高,现有技术要么无法实现实时预警,要么无法给出可被临床医护人员理解、采信的解释依据,导致技术与临床工作流脱节,难以真正落地应用

Benefits of technology

本发明针对现有技术无法适配不规则随访序列、异步与高缺失数据的问题,提供一种能够天然适配慢病随访数据特性的处理与建模方式,减少伪观测数据与系统性偏差,提升风险预警的准确性与鲁棒性,使其能够精准捕捉患者病情的动态演变轨迹,适配真实临床随访场景。针对现有技术风险预警滞后、无法实现逐次随访实时更新的问题,设计贴合临床随访工作流的预警模式,实现“每次随访即时触发、风险实时更新”,满足医护人员在随访现场快速获取风险评估结果的需求,为临床及时干预争取时间。针对现有技术解释结果与预警不同步、稳定性不足、临床可读性差的问题,实现风险预警与可解释结果的同步生成,提供“随访级、可追踪”的个体级与群体级解释依据,明确风险驱动因素,提升解释结果的可信度与临床可接受度,辅助医护人员制定个性化干预方案。针对现有技术泛化能力不足、难以跨中心、跨科室推广的问题,设计具有强扩展性与可移植性的技术方案,不依赖特定病种或固定不良结局,适配多中心数据分布差异,便于在不同医疗机构、不同慢病随访场景中推广应用,提升慢病随访管理的整体效率与质量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122552012A_ABST
    Figure CN122552012A_ABST
Patent Text Reader

Abstract

This invention discloses a method for real-time early warning of chronic disease follow-up risks based on asynchronous electronic medical record sequence prediction, belonging to the field of medical management technology. The method involves acquiring asynchronous electronic medical records, extracting dynamic follow-up information to construct a follow-up axis, and extracting static baseline information. For each node on the axis, an input sample is constructed, consisting of the historical prefix sequence up to the current follow-up and the static baseline information. Labels are assigned based on whether a specified event has occurred within the risk window. Dynamic variables in the input sample are grouped by category to obtain dynamic variable groups. A time-series encoder is used to extract variable lexical representations from the dynamic variable groups. The static baseline information is projected into a static lexical representation, which is then concatenated with the lexical representations of each variable and subjected to feature-level attention aggregation to obtain a comprehensive risk representation as input features. Labels are used as supervision signals, and a class-imbalanced loss function is employed to train the risk prediction model. The comprehensive risk representation is constructed using the latest node and input into the model to achieve early warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical management technology, specifically relating to a method for real-time early warning of chronic disease follow-up risks by combining asynchronous electronic medical record sequence prediction. Background Technology

[0002] With the aging population and increased health awareness, chronic diseases such as hypertension, diabetes, coronary heart disease, stroke, and chronic kidney disease have become core challenges facing global healthcare systems. These diseases are characterized by long disease courses, high recurrence rates, the need for long-term follow-up monitoring, and preventable adverse outcomes (such as complication onset, disease progression, hospitalization, and death). According to relevant statistics, the global population of chronic disease patients continues to expand, and the number of chronic disease patients in my country has exceeded 400 million. The efficiency and precision of chronic disease management are directly related to the rational allocation of medical and health resources and the improvement of residents' health levels.

[0003] In the core aspect of chronic disease management—follow-up management—electronic medical records (EMRs) and follow-up data have become the core basis for clinical assessment of patients' conditions and prediction of adverse risks. Currently, medical institutions at all levels in my country have gradually realized the informatization of electronic medical records, and the accumulation of multi-source follow-up data (including outpatient follow-up records, laboratory test results, vital sign monitoring data, medication records, patient self-reported symptoms, etc.) is becoming increasingly rich, providing a data foundation for risk warning using artificial intelligence technology.

[0004] In clinical practice, the core requirement of chronic disease follow-up management is "early warning and precise intervention." This means that medical staff need to quickly grasp the patient's current condition at each follow-up visit and accurately predict the risk of adverse outcomes in the future (e.g., 30 days, 90 days). At the same time, they need to identify the key driving factors of the risk and then develop personalized intervention plans (e.g., adjusting medication, increasing follow-up frequency, strengthening indicator monitoring) to reduce the incidence of adverse events and improve the efficiency of follow-up management.

[0005] However, existing medical informatics and artificial intelligence technologies face many practical bottlenecks when adapting to the risk warning needs of chronic disease follow-up scenarios: On the one hand, chronic disease follow-up data itself has significant "temporal sequence, irregularity, asynchronicity, and high missingness," which is fundamentally different from conventional structured data, making it difficult for traditional data processing and modeling methods to adapt; on the other hand, clinical decision-making has extremely high requirements for the "real-time and interpretability" of risk warnings, and existing technologies either cannot achieve real-time warnings or cannot provide explanatory evidence that can be understood and accepted by clinical medical staff, resulting in a disconnect between technology and clinical workflow, making it difficult to truly implement and apply. Summary of the Invention

[0006] In view of this, the purpose of this invention is to provide a method for real-time early warning of chronic disease follow-up risks that combines asynchronous electronic medical record sequence prediction, in order to solve the above-mentioned technical problems.

[0007] To achieve the above objectives, the present invention provides the following technical solution: A method for real-time early warning of chronic disease follow-up risks based on asynchronous electronic medical record sequence prediction includes: The system acquires asynchronous electronic medical records of target patients, extracts dynamic follow-up information to construct a real follow-up axis of patients sorted by time, and extracts static baseline information. The real follow-up axis of patients contains multiple follow-up nodes, each of which corresponds to one dynamic follow-up information. The dynamic follow-up information includes a variable state tuple constructed from multiple dynamic variables. Using each follow-up node as a trigger point, input samples are constructed based on the historical prefix sequence and static baseline information up to that follow-up. Follow-up-level labels are assigned based on whether the target adverse event occurs within the future preset risk window. Dynamic variables in the input samples are grouped by category to obtain multiple dynamic variable groups. A time encoder is used to extract variable lexical representations from each dynamic variable group. The static baseline information is projected into a static lexical representation and concatenated with the lexical representations of each variable. Feature-level attention aggregation is then performed to obtain a comprehensive risk representation as input features. At the same time, the corresponding follow-up-level labels are used as supervision signals, and a class imbalance-friendly loss function is used to train the risk prediction model. Input samples with unclosed future preset risk windows and insufficient outcomes are not directly trained as ordinary supervision samples. Using the latest follow-up node of the target patient as the early warning point, the input sample is constructed and the comprehensive risk characterization of the input sample is obtained in the same way as in the training phase. The input is then fed into the trained risk prediction model to predict whether a specified adverse event will occur within a preset risk window period in the future. At the same time, the attention weights of each dynamic variable during aggregation are output as the endogenous feature importance synchronized with the early warning, so as to achieve interpretable results at both the individual and group levels.

[0008] Furthermore, static baseline information includes patient age, gender, primary diagnosis, disease stage, comorbidities, history of major events, and long-term medication background; dynamic follow-up information includes outpatient visit records, laboratory results, vital signs, symptom scores, medication adjustments, treatment adherence records, post-hospitalization follow-up results, and data transmitted from outpatient monitoring.

[0009] Furthermore, follow-up-level labels are assigned based on whether the target adverse event occurs within a pre-defined risk window, including: Obtain the endpoint event information for each follow-up node; the endpoint event information includes whether the target adverse event occurred, the time of occurrence, the event type, and the source of event confirmation; Based on the endpoint event information of each follow-up node, if a target adverse event occurs within the preset risk window in the future, the input sample constructed by the corresponding follow-up node will be marked as a high-risk sample. If no target adverse event occurs within the preset risk window in the future and the subsequent follow-up records are continuous and complete, the input sample constructed at the corresponding follow-up node will be marked as a low-risk sample. If a target patient is lost to follow-up, transferred out, or has incomplete follow-up records before the preset risk window closes, the input samples constructed from the corresponding follow-up nodes will be marked as incomplete samples and will not participate in the subsequent model training process.

[0010] Furthermore, before grouping the dynamic variables in the input samples by category to obtain multiple dynamic variable groups, the process also includes missing variable retention and asynchronous state construction for the dynamic variables at each follow-up node. For dynamic variables with missing values ​​at each follow-up node, either of the following two strategies is used to complete the missing value retention: zero-value imputation and adding an observation mask to each dynamic variable to explicitly encode the missing value pattern, or intra-patient forward imputation combined with historical statistics from the same institution for backfilling. For each dynamic variable at each follow-up node, an asynchronous state is constructed. The asynchronous state construction includes at least whether the current observation is a true observation and the time information since the last true observation.

[0011] Furthermore, a time-series encoder is used to extract variable lexical representations for each dynamic variable group, and the static baseline information is projected as a static lexical representation and concatenated with each variable lexical representation. This includes treating each dynamic variable group as an independent time stream, inputting it into a time-series encoder with a length mask, obtaining the variable lexical representation of the corresponding dynamic variable group through mask pooling, and simultaneously projecting the static baseline information of the corresponding target patient as a static lexical representation and concatenating it with each variable lexical representation.

[0012] Furthermore, feature-level attention aggregation is performed on the concatenated lexical set, including: Calculate the attention weight of each word representation in the word set, and use the attention weights to perform a weighted summation of the word representations to obtain a comprehensive risk characterization; Before calculating the attention weight for each lexical representation, a comprehensive evaluation of the current credibility and relevance of each dynamic variable to the current task is required. The current level of credibility is related to at least the following factors: whether the current dynamic variable is observed at the trigger point, the length of time since the last real observation, whether the data source is reliable, whether there is obvious noise recently, and whether the change of the dynamic variable is continuous and consistent. The relevance of the current task refers to the strength of the indication that the current dynamic variable provides for the specified adverse event.

[0013] Furthermore, the risk prediction model includes multiple independent models trained for different adverse events, or a multi-head output model with a shared underlying representation obtained through multi-task joint training; the risk prediction model outputs a risk score for the target patient to experience the target adverse event within a future preset risk window period, and determines whether the target patient will experience the specified adverse event within the future preset risk window period based on the risk score.

[0014] The beneficial effects of this invention are as follows: This invention addresses the limitations of existing technologies in handling irregular follow-up sequences and asynchronous or highly missing data. It provides a processing and modeling method that naturally adapts to the characteristics of chronic disease follow-up data, reducing spurious observations and systematic biases, and improving the accuracy and robustness of risk warnings. This allows for precise capture of the dynamic evolution of a patient's condition, making it suitable for real-world clinical follow-up scenarios. To address the issues of delayed risk warnings and the inability to achieve real-time updates for each follow-up visit in existing technologies, this invention designs a warning mode that aligns with clinical follow-up workflows, enabling "instant triggering and real-time risk updates for each follow-up visit." This meets the needs of medical staff to quickly obtain risk assessment results on-site, allowing for timely clinical intervention. Furthermore, to address the problems of asynchronous interpretation and warning, insufficient stability, and poor clinical readability in existing technologies, this invention achieves the synchronous generation of risk warnings and interpretable results, providing "follow-up-level, traceable" individual and group-level interpretive evidence. It clarifies risk drivers, improves the credibility and clinical acceptability of interpretive results, and assists medical staff in developing personalized intervention plans. To address the issues of insufficient generalization capabilities and difficulty in promoting existing technologies across centers and departments, a highly scalable and portable technical solution is designed. This solution is not dependent on specific diseases or fixed adverse outcomes, adapts to the differences in multi-center data distribution, and is easy to promote and apply in different medical institutions and different chronic disease follow-up scenarios, thereby improving the overall efficiency and quality of chronic disease follow-up management.

[0015] Other advantages, objectives, and features of the invention will be set forth in the following description and will be apparent to those skilled in the art in some respects, or may be learned by practice of the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.

[0016] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1This is a flowchart illustrating a method for calculating the relative velocity and direction of colliding droplets in a colliding chamber, as described in an embodiment of the present invention. Detailed Implementation

[0018] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0019] In hospital chronic disease follow-up management and clinical decision support, it is necessary to output the risk of adverse outcomes / events occurring within a fixed future time window based on electronic medical record history at each follow-up visit, and provide clinically understandable explanations. Current implementations mainly face the following technical challenges: 1. Irregular follow-up sequences: Patient follow-up intervals are not fixed, and the number of follow-ups varies greatly, resulting in inconsistent sequence lengths, making it difficult for models to directly align and model. 2. Asynchronous and missing indicator collection: Different test indicators are issued and entered at different time points, resulting in strongly asynchronous observations; the missing rate is high, and the missing patterns are related to the disease condition, so simple imputation introduces bias. 3. Differences in distribution across multiple centers / departments: Differences in testing systems, dimensional ranges, recording habits, and populations lead to data distribution drift during training and deployment, resulting in insufficient generalization and stability. 4. Difficulty in synchronizing interpretability with early warning: Clinical practice requires "follow-up-level, traceable" explanations (why the risk increases / decreases, which indicators drive it), but existing methods are mostly ex-post interpretations, making it difficult to reproduce stably and integrate into workflows. This technology belongs to the fields of medical informatics and artificial intelligence. Specifically, it involves a risk assessment method for time-series modeling of multi-source electronic medical records / follow-up data in hospitals, as well as a computational method for generating interpretable results while outputting risk warnings. It can be used in scenarios such as outpatient follow-up management, hierarchical early warning, clinical decision support, and quality control auditing.

[0020] To address the problems in existing implementations, the core design idea of ​​the proposed solutions is to "abandon time-series characteristics and transform time-series data into static summary features." Essentially, this simplifies data complexity to reduce modeling difficulty. However, this design naturally ignores the temporal order and dynamic changes of follow-up data, failing to capture early warning signals such as "continuous increases / decreases in indicators," leading to delayed risk warnings and hindering real-time updates for each follow-up. Feature extraction relies on manual operation, and different operators have varying clinical knowledge and data processing experience, potentially extracting different summary features from the same batch of follow-up data (e.g., some extract "3-month average," while others extract "6-month average"). These feature differences directly lead to inconsistent model training results and insufficient stability. Furthermore, the proposed solutions lack a dedicated processing mechanism for the asynchronous and high-missing-value characteristics of chronic disease follow-up data, relying solely on simple methods like mean and forward imputation to fill in missing data. Since missing data in chronic disease follow-up data is often related to the condition itself, simple imputation forces patients to accept non-true observations, introducing systematic bias and affecting model prediction accuracy.

[0021] Secondly, the original design of this scheme was to solve the problem of "irregular follow-up sequences". The core logic is to "transform irregular sequences into regular sequences" to adapt to traditional time series models. However, this transformation requires filling the missing data of fixed time grids through interpolation, forward padding and other methods. These padding data are not the actual test results of patients, but virtual values ​​calculated by humans. Especially when the patient's condition changes suddenly, the pseudo-observation data cannot reflect the real changes in the condition, resulting in the distortion of model prediction. This solution only focuses on "sequence regularization" and does not design solutions for the root causes of data asynchrony and missing data (such as differences in the frequency of indicator collection and disease-related missing data). Interpolation, forward filling and other methods are still simple filling methods. When the missing pattern is related to the patient's condition (such as the patient's condition deteriorating and not being followed up), the feature values ​​when the condition is stable will be filled into the time step when the condition deteriorates, which will introduce systematic bias. Fixed time grids (such as monthly or weekly) are uniformly set and do not take into account the differences in follow-up cycles for different chronic diseases (such as diabetes requiring weekly monitoring, while hypertension can be monitored monthly), nor do they take into account the differences in individual patient conditions (such as severe patients requiring more frequent follow-up than mild patients), lacking flexibility and resulting in decreased model adaptability.

[0022] Meanwhile, conventional RNN / GRU models require "synchronized time steps and consistent feature vectors" as input. To address the issues of asynchronous and missing data, this approach only uses simple methods such as forward padding and mean padding, failing to recognize clinically significant behavioral signals such as "some indicators not being measured for a long time" or "some indicators suddenly being monitored frequently" (e.g., doctors frequently monitoring blood lipids because they suspect the patient has abnormalities). This results in the model's inability to capture such crucial early warning information. In chronic disease follow-up, the collection time and frequency of different indicators naturally differ, forming asynchronous observation sequences. This approach requires forcibly synchronizing the feature vectors of each time step (ensuring consistency in the number and type of features). This forced synchronization loses the asynchronous information from indicator collection, failing to reflect "differences in the clinical significance of different indicators at different time points," thus affecting the model's accurate judgment of the condition. While the model captures temporal dependencies through a recurrent structure, this structure has high requirements for the consistency of follow-up intervals. When the follow-up intervals for patients vary significantly (e.g., from 1 month to 6 months), the model struggles to accurately capture the disease correlation between different time steps, leading to insufficient predictive stability.

[0023] Furthermore, the design logic of this scheme is "predict first, then interpret," meaning that the interpretation results need to be generated through additional complex calculations after the model completes risk prediction (e.g., the SHAP value needs to be decomposed from the model's decision-making process). This "separation of prediction and interpretation" mode results in the interpretation results not being output synchronously with risk warnings, failing to meet the need for "instant risk and interpretation" in clinical follow-up, leading to low efficiency and difficulty in integrating into the closed-loop workflow of "follow-up data entry - risk warning - clinical decision-making." The principle of ex-post interpretation algorithms such as SHAP and LIME is based on local approximation or feature contribution decomposition. In the scenario of "high missing data, strong feature correlation, and dynamic changes over time" in chronic disease follow-up data, data noise and model parameter fluctuations will directly affect the interpretation results, leading to different interpretation results for the same patient with the same risk level. This results in insufficient stability and consistency, making it difficult for medical staff to accept the information. The interpretation results of this scheme are only for a single prediction and are not associated with the specific data of each patient's follow-up. It cannot achieve "follow-up level traceability," that is, it cannot clearly identify "which indicators changed in the previous follow-up that are related to the current increase / decrease in risk," making it difficult to trace the specific reasons for risk changes and failing to meet the actual needs of clinical decision-making.

[0024] Finally, this scheme lacks a dedicated mechanism for handling asynchronous and missing data issues. If the model input is still a strongly synchronized feature sequence or an unencoded missing pattern, the attention mechanism may mistakenly treat "missing patterns, sequence length, and imputation strategies" as important features, leading to abnormal attention weights (e.g., high weights for missing indicators). This results in interpretations that do not match clinical reality and have limited credibility. The core function of the attention mechanism in models like Transformer is to capture the dependencies within a sequence (e.g., the correlation between a time step and other time steps). Its weights are essentially "dependency weights during model building," not "feature importance" in a clinical sense. For example, a high weight for an indicator may be due to its strong correlation with other indicators, rather than its direct impact on the condition itself. This ambiguity in physical meaning makes the interpretation results difficult for medical staff to understand and accept. The attention weights in this scheme are only generated for the prediction results of a single patient and a single follow-up visit. It cannot perform statistical analysis on the interpretation results of multiple patients, nor can it extract common risk drivers among multiple patients (e.g., common high-risk indicators for a certain type of chronic disease patient), making it difficult to use for overall optimization of chronic disease management (e.g., developing targeted follow-up guidelines).

[0025] Based on the objective limitations of the existing technical solutions, and considering the core clinical needs of chronic disease follow-up management—namely, real-time risk warning, accurate risk assessment, synchronous interpretability, and adaptability to multiple scenarios—the purpose of this invention is to provide a method for real-time risk warning in chronic disease follow-up that combines asynchronous electronic medical record sequence prediction. This aims to specifically address the shortcomings of existing technologies while avoiding design flaws in existing solutions. The specific objectives are as follows: To address the limitations of existing technologies in handling irregular follow-up sequences and asynchronous or highly missing data, this paper proposes a processing and modeling approach that is naturally adapted to the characteristics of chronic disease follow-up data. This reduces spurious observations and systematic biases, improving the accuracy and robustness of risk warnings and enabling precise capture of the dynamic evolution of patients' conditions, thus adapting to real-world clinical follow-up scenarios. To address the issues of delayed risk warnings and the inability to achieve real-time updates for each follow-up visit, a warning mode tailored to clinical follow-up workflows is designed, enabling "instant triggering and real-time risk updates for each follow-up visit." This meets the needs of medical staff to quickly obtain risk assessment results on-site, allowing for timely clinical intervention. Furthermore, to address the problems of asynchronous interpretation and warning, insufficient stability, and poor clinical readability in existing technologies, this paper achieves the synchronous generation of risk warnings and interpretable results, providing "follow-up-level, traceable" individual and group-level interpretive evidence. This clarifies risk drivers, enhances the credibility and clinical acceptability of interpretation results, and assists medical staff in developing personalized intervention plans. To address the issues of insufficient generalization capabilities and difficulty in promoting existing technologies across centers and departments, a highly scalable and portable technical solution is designed. This solution is not dependent on specific diseases or fixed adverse outcomes, adapts to the differences in multi-center data distribution, and is easy to promote and apply in different medical institutions and different chronic disease follow-up scenarios, thereby improving the overall efficiency and quality of chronic disease follow-up management.

[0026] like Figure 1 As shown, to achieve the above objectives, this invention proposes a method for real-time early warning of chronic disease follow-up risks based on asynchronous electronic medical record sequence prediction, comprising: S101. Obtain the asynchronous electronic medical records of the target patients, extract dynamic follow-up information to construct the real follow-up axis of patients sorted by time, and extract static baseline information; Among them, the patient real follow-up axis contains multiple follow-up nodes, each follow-up node corresponds to one dynamic follow-up information, and the dynamic follow-up information includes a variable state tuple constructed from multiple dynamic variables; S102. Using each follow-up node as a trigger point, construct input samples of historical prefix sequences and static baseline information up to the current follow-up, and assign follow-up-level labels based on whether the target adverse event occurs within the future preset risk window. S103. Group the dynamic variables in the input sample by category to obtain multiple dynamic variable groups, and use a time encoder to extract variable word representations for each dynamic variable group. S104. Project the static baseline information into a static lexical representation and concatenate it with the lexical representation of each variable. Then perform feature-level attention aggregation to obtain a comprehensive risk representation as input features. At the same time, use the corresponding follow-up level label as a supervision signal and train the risk prediction model using a class imbalance-friendly loss function. Among them, input samples with unclosed future risk windows and insufficient outcomes will not be directly used as ordinary supervised samples for training. S105. Using the latest follow-up node of the target patient as the early warning point, construct the input sample and obtain the comprehensive risk characterization of the input sample in the same way as in the training phase, and input it into the trained risk prediction model. S106. Predict whether a specified adverse event will occur within a preset risk window period in the future, and output the attention weights of each dynamic variable during aggregation as the importance of endogenous features synchronized with the early warning, so as to achieve interpretable results at both the individual and group levels. The working principle and beneficial effects of the above technical solution are as follows: This invention proposes a method for real-time early warning of chronic disease follow-up risks based on asynchronous electronic medical record sequence prediction. It is applicable to chronic disease management scenarios requiring long-term follow-up, long-term monitoring, and dynamic intervention, such as diabetes, chronic kidney disease, hypertension, coronary heart disease, chronic heart failure, and secondary prevention of stroke. This method addresses actual business processes such as hospital outpatient follow-up, disease management, post-discharge follow-up, and regional collaborative management of chronic diseases. It uses each actual follow-up visit, test, examination, prescription adjustment, or outpatient report as a trigger point to reorganize, encode, fuse, provide early warnings, and interpret the asynchronous electronic medical record data, thereby achieving "follow-up as early warning, early warning as interpretation, and traceable results."

[0027] The method of the present invention includes at least the following steps: 1. Data Access and Standardization Processing Steps Patient data is extracted from hospital information systems, laboratory systems, examination systems, electronic medical record systems, prescription systems, nursing follow-up systems, disease management platforms, or regional health record systems. The extracted data must include at least the following categories: The first category is static baseline information, which includes information that is relatively stable in the short term, such as patient age, gender, primary diagnosis, disease stage, comorbidities, history of major events, and long-term medication background. The second category is dynamic follow-up information, including outpatient follow-up records, test results, vital signs, symptom scores, medication adjustments, treatment adherence records, post-hospitalization follow-up results, and data transmitted from outpatient monitoring. The third category is endpoint event information, including whether the target adverse event occurred, when it occurred, the type of event, and the source of event confirmation; The fourth category is source and quality information, including the data source institution, department, equipment, data entry method, report status, and quality control markings.

[0028] Standardize the accessed data. This includes: unifying patient identification, time format, variable naming, and units of measurement; identifying duplicate records; handling obvious data entry errors; retaining abnormal but true high-risk clinical values; and performing structured mapping on text-based conclusions. For synonymous fields across institutions or systems, a field mapping dictionary can be established; for the same indicator with multiple units of measurement, unit conversion should be performed before inclusion in subsequent processing.

[0029] 2. Steps for Real-Time Follow-up Axis Reconstruction For each patient, their actual follow-up trajectory is reconstructed chronologically. This follow-up trajectory is not limited to the date of outpatient registration, but is based on effective medical contacts that reflect updates to their condition. Effective medical contacts include at least outpatient follow-up visits, specific disease-specific examinations, return of test results, major medication adjustments, post-discharge follow-up examinations, and home monitoring reports confirmed and received by the management platform.

[0030] If multiple records are generated on the same day, they are merged according to preset rules to form a single "follow-up snapshot." These preset rules may include: prioritizing the latest reported values ​​for similar test items; prioritizing vital signs based on records closest to the time of visit; using the valid prescription status after the current visit; and using the latest structured text conclusions. After merging, each follow-up node corresponds to a set of real medical information available at that point in time.

[0031] This step establishes an irregular but realistic follow-up timeline for each patient, providing a unified business triggering basis for subsequent real-time early warnings.

[0032] 3. Follow-up triggers sample generation step This invention generates early warning samples based on each real and effective follow-up visit of a patient, rather than waiting until the entire course of the disease is completed before uniformly modeling. Whenever a patient reaches a new effective follow-up point, the system extracts all historical information from the initial inclusion in management to the current point, and combines it with the patient's static baseline information to form an early warning sample.

[0033] In terms of label construction, the system uses a pre-defined risk window to determine whether a target event will occur within a future period. This risk window can be configured to 30 days, 60 days, 90 days, 180 days, 365 days, or other durations depending on different chronic disease management needs. If the target event occurs within the window, the sample is marked as a high-risk sample; if it does not occur within the window and subsequent observation is sufficient, it is marked as a low-risk sample; if the patient is lost to follow-up, transferred out, or has incomplete data before the window closes, the sample is recorded as an incomplete sample and is not directly treated as a regular negative sample.

[0034] This sample generation method ensures that each prediction uses only information already obtained up to the current time point, preventing future information from being leaked into the current prediction process. Furthermore, this mechanism is highly consistent with clinical workflows; that is, when a patient visits for a follow-up appointment, the system provides an updated risk assessment based on the patient's historical trajectory up to the current time point.

[0035] The core of this invention lies not in the act of "predicting every time," but in its sample generation mechanism, which employs "real follow-up anchor points + historical prefix samples + window closure labeling rules." Compared to existing technologies, this invention ensures that only currently available historical information is used at the current time point, preventing future data from being leaked into the current prediction. Simultaneously, it addresses the issues of missed follow-ups and right censoring through window closure rules, reducing bias caused by incorrect labels. In other words, this invention does not protect an abstract "real-time early warning function," but rather a specific sample generation and label construction technology solution tailored to real-world chronic disease follow-up scenarios.

[0036] 4. Missing State Preservation and Asynchronous State Construction Steps This invention does not interpret missing values ​​merely as "empty values," but rather preserves the missing value itself as clinical information. For each dynamic clinical variable, at each follow-up point, not only is its value recorded, but the following status information is also recorded simultaneously: Whether it was actually detected in this test; How long has it been since the last true test of this variable? How long has it been since the patient's last effective follow-up? Does this result exceed the reference range? This variable has recently shown a continuous increase, a continuous decrease, increased volatility, or relative stability; What system, device, or organization does this record originate from, and what is its level of trustworthiness? If no observation is observed this time, it will be further marked as having different reasons for missing information, such as no order placed, order placed but not reported, logically inapplicable, removed by quality control, or not connected to external institutions.

[0037] Regarding numerical closure, this invention allows for various implementation strategies. For example, placeholder values ​​can be used in conjunction with explicit observation markers, or the most recent valid observation within the patient, the statistical median of similar patients in the same institution, or the clinical default safety value can be used for assisted closure. However, regardless of the strategy adopted, the original state of "whether this observation was genuine" and "time information since the last genuine observation" must be retained to avoid mistaking artificial closure values ​​for genuine detection values.

[0038] Unlike simple mean filling or forward filling, the purpose of this step is not to forcibly fill all gaps, but to preserve the true clinical meaning as much as possible while constructing a unified calculation input, so that the system can recognize that "a certain indicator has not been reviewed for a long time" may itself be a risk signal.

[0039] 5. Variable History Summary Coding Steps For each dynamic variable, this invention does not directly concatenate all variables into a single synchronous vector for input. Instead, it first extracts the change trajectory of each variable during the patient's historical follow-up process, and then processes it using a dedicated variable history summarization unit. The core function of this summarization unit is to compress the historical observation rhythm, abnormal persistence, recent and long-term change trends, observation freshness, and stability of a variable into a state summary oriented towards the current follow-up node.

[0040] For example, for blood pressure variables, the summary unit can extract the state of "consistently high for the last three times, with the most recent abnormality worsening, and the interval between the two tests shortening"; for renal function indicators, it can extract "consistently declining for the last six months, with the most recent decline further widening"; for a key test that has not been measured for a long time, it can extract "long-term lack of recent observation, current state is not recent". In this way, what is obtained is not a single numerical value, but a variable state expression containing clinical dynamic meaning.

[0041] The variable history summarization unit can be implemented using gated temporal networks, locally convolutional networks, lightweight self-attention networks, state-space networks, or other encoding components capable of handling temporal dependencies. This invention does not limit itself to a specific publicly available model structure; the key point is that each variable first generates a separate state summary for the current time point before entering the subsequent fusion process.

[0042] The difference between this invention and existing technologies lies not only in whether a time-series model is used, but also in the change of the input object. Existing technologies mostly use "processed synchronous feature vectors" as the modeling object; this invention uses "variable state tuples that retain missing semantics and temporal freshness" as the modeling object. Existing technologies typically eliminate asynchronicity before modeling; this invention treats asynchronicity as part of the usable information. Missing data handling in existing technologies primarily serves numerical closure, while missing data handling in this invention simultaneously serves semantic expression, reliability identification, and risk assessment, thus making it more suitable for the real-world asynchronous observation scenarios in chronic disease follow-up.

[0043] 6. Static information integration and reliability screening fusion steps After obtaining historical summaries of each dynamic variable, the patient's static baseline information is then introduced into the fusion module. Static baseline information is used to provide a relatively stable risk background for the patient, such as age group, major disease type, history of major events, and long-term risk factors.

[0044] The fusion step does not treat all variables equally by averaging them. Instead, it first comprehensively evaluates the "current credibility" and "current task relevance" of each variable. Current credibility is related to at least the following factors: whether the variable was recently observed, the time since the last actual observation, the reliability of the data source, the presence of significant recent noise, and the consistency of the variable's changes. Current task relevance refers to the strength of the variable's indication of the currently set target event.

[0045] For example, if a variable is historically important but has not been measured for a long time recently, its current credibility can be lowered; if a variable has only been observed a few times, but the most recent observation was exceptionally significant and the source is reliable, its current weight can be increased. Through this screening mechanism, this invention achieves a two-layer fusion of "first judging whether it is trustworthy, and then judging whether it is worth paying attention to."

[0046] The fusion yields a comprehensive risk profile of the patient at the current follow-up point. This profile includes the static risk background, the recent and long-term evolution information of each dynamic variable, and reflects the freshness of the observation and the reliability of the data.

[0047] 7. Risk Output and Real-Time Early Warning Classification Steps The system inputs a comprehensive risk profile into the prediction module and outputs a risk score for the patient to experience a target adverse event within a preset risk window. This score can be further mapped to multi-level risk warning results, such as low risk, medium risk, high risk, or different levels like blue warning, yellow warning, orange warning, and red warning.

[0048] Different early warning tasks can be configured for different disease scenarios and clinical goals. For example, in diabetes management, models can be created for the risks of short-term hospitalization, rapid deterioration of renal function, and cardiovascular complications; in chronic kidney disease management, models can be created for the risks of initiating renal replacement therapy, major cardiovascular events, and all-cause mortality. Alternatively, models can be built separately for individual events, or multi-task joint early warning can be achieved by sharing a front-end processing module and setting separate output heads.

[0049] During deployment, the system triggers an early warning calculation immediately after each new follow-up record is added, enabling dynamic updates of risk results, rather than periodic batch updates or updates after manual aggregation.

[0050] 8. Real-time explanation of generation steps The interpretation results of this invention are not generated synchronously during the risk output process, rather than by calling a separate post-hoc interpretation algorithm after prediction is completed. The interpretation module, directed to the current follow-up node, outputs at least one or more of the following interpretation results: The most important risk drivers at present; Does each driving factor increase or decrease risk? The main sources of change that led to an increase or decrease in risk compared to the previous follow-up point; Which indicators have not been reviewed for a long time, thus increasing the uncertainty of the current risk assessment? For group management scenarios, the most common combination of driving factors among similar high-risk patients can be statistically analyzed.

[0051] The explanations should ideally use natural language templates or structured summary outputs that can be directly used for clinical presentations. For example, explanations such as "recent two consecutive increases in systolic blood pressure," "significant recent decline in key renal function indicators," "long-term lack of follow-up for proteinuria-related tests," and "indicators have not improved despite recent medication adjustments" can be generated. This allows healthcare professionals to immediately understand the main reasons for the risk warnings while reviewing the risk values, facilitating decisions on whether to shorten the follow-up period, add key tests, or adjust the treatment plan.

[0052] It should be noted that the interpretation results in this invention are used for risk communication, clinical auxiliary judgment, and management auditing, and are not directly equivalent to causal conclusions. Its key technical point lies in the simultaneous generation of interpretation and early warning, and the traceable correspondence between the interpretation content and the current follow-up node.

[0053] 9. Training, Calibration, and Deployment Procedures During the model training phase, either single-task training or multi-task joint training can be used depending on business needs. To address the issue of a low proportion of positive events in chronic disease follow-up scenarios, a class imbalance-friendly loss design can be employed. To improve the consistency between the output probability and the actual event occurrence rate, threshold optimization and probability calibration steps can be added after training.

[0054] In terms of data partitioning, training, validation, and test sets can be divided by patient, institution, or time period to prevent the leakage of the same patient's information into different datasets. For multi-center data, institution-level standardization and source quality constraints can also be introduced to enhance the stability of cross-center deployment.

[0055] During the online deployment phase, the system includes at least a data access module, a standardization processing module, a follow-up axis update module, a sample generation module, a status summary construction module, a fusion early warning module, an interpretation generation module, a result display module, and a log management module. The system can be deployed on hospital servers, disease management platforms, regional chronic disease collaboration platforms, or cloud service environments. Whenever a new follow-up record is written, the system automatically completes patient history updates, risk inference, result classification, interpretation generation, and message push.

[0056] The following explanation uses outpatient follow-up management of patients with chronic kidney disease complicated by abnormal glucose metabolism as an example, but the present invention is not limited to this disease.

[0057] The static baseline information recorded when a patient was enrolled included: male, 67 years old, long history of abnormal glucose metabolism, hypertension, and a history of adverse cardiovascular events. Subsequently, multiple irregular follow-up records were generated in the management platform, including serum creatinine, estimated glomerular filtration rate, urinary albumin-related indicators, systolic blood pressure, glycated hemoglobin, hemoglobin, albumin, serum potassium, and medication adjustments. These indicators were not collected completely at every follow-up visit; some indicators appeared consecutively multiple times, while others were only measured at specific time points.

[0058] When a patient returns for a follow-up outpatient visit, the system first incorporates the newly added test results, vital signs, and prescription changes into the previous follow-up trajectory. Second, it generates a status summary for each type of dynamic variable, such as "continuous decline in renal function indicators," "persistently low albumin levels," "poor recent control of systolic blood pressure," and "lack of recent follow-up of key urinary protein indicators." Third, it performs reliability screening and comprehensive fusion of each variable in conjunction with the patient's static baseline background. Finally, it outputs a high-risk warning for the patient to experience a target adverse event within a preset time window, and simultaneously generates an interpretation result.

[0059] For clinical users, the interface displays not only the risk level but also the key driving factors leading to the result and their directional explanations. Based on this, physicians can further decide whether to shorten the follow-up period, supplement key tests, or adjust antihypertensive or renal protection treatment plans. This embodiment demonstrates that the present invention can provide real-time alerts and synchronous interpretation of irregular, asynchronous, and highly missing electronic medical record data in a real-world chronic disease follow-up environment.

[0060] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made to it in form and detail without departing from the scope defined by the claims of the present invention.

Claims

1. A method for real-time early warning of chronic disease follow-up risk combining asynchronous electronic medical record sequence prediction, characterized in that, include: The system acquires asynchronous electronic medical records of target patients, extracts dynamic follow-up information to construct a real follow-up axis of patients sorted by time, and extracts static baseline information. The real follow-up axis of patients contains multiple follow-up nodes, each follow-up node corresponds to one dynamic follow-up information, and the dynamic follow-up information includes a variable state tuple constructed from multiple dynamic variables. Using each follow-up node as a trigger point, input samples of historical prefix sequences and static baseline information up to the current follow-up are constructed, and follow-up-level labels are assigned based on whether the target adverse event occurs within a future preset risk window. The dynamic variables in the input samples are grouped by category to obtain multiple dynamic variable groups. A time-series encoder is used to extract variable lexical representations from each dynamic variable group. The static baseline information is projected into a static lexical representation, which is then concatenated with the lexical representations of each variable and subjected to feature-level attention aggregation to obtain a comprehensive risk representation as input features. Simultaneously, the corresponding follow-up-level labels are used as supervision signals, and a class-imbalanced friendly loss function is employed to train the risk prediction model. Input samples with unclosed future preset risk windows and insufficient outcomes are not directly used as ordinary supervision samples for training. Using the latest follow-up node of the target patient as the early warning point, the input sample is constructed and the comprehensive risk characterization of the input sample is obtained in the same way as in the training phase. The input is then fed into the trained risk prediction model to predict whether a specified adverse event will occur within a preset risk window period in the future. At the same time, the attention weights of each dynamic variable during aggregation are output as the endogenous feature importance synchronized with the early warning, so as to achieve interpretable results at both the individual and group levels.

2. The method for real-time early warning of chronic disease follow-up risk combining asynchronous electronic medical record sequence prediction as described in claim 1, characterized in that, The static baseline information includes the patient's age, gender, primary diagnosis, disease stage, comorbidities, history of major events, and long-term medication background; the dynamic follow-up information includes outpatient follow-up records, laboratory results, vital signs, symptom scores, medication adjustments, treatment adherence records, post-hospitalization follow-up results, and data transmitted from outpatient monitoring.

3. The method for real-time early warning of chronic disease follow-up risk combining asynchronous electronic medical record sequence prediction as described in claim 1, characterized in that, Follow-up-level labels are assigned based on whether the target adverse event occurs within a pre-defined risk window, including: Obtain the endpoint event information for each follow-up node; wherein, the endpoint event information includes whether the target adverse event occurred, the time of occurrence, the event type, and the source of event confirmation; Based on the endpoint event information of each follow-up node, if a target adverse event occurs within the preset risk window in the future, the input sample constructed by the corresponding follow-up node will be marked as a high-risk sample. If no adverse event occurs within the preset risk window and the subsequent follow-up records are continuous and complete, the input sample constructed at the corresponding follow-up node will be marked as a low-risk sample. If a target patient is lost to follow-up, transferred out, or has incomplete follow-up records before the preset risk window closes, the input samples constructed from the corresponding follow-up nodes will be marked as incomplete samples and will not participate in the subsequent model training process.

4. The method for real-time early warning of chronic disease follow-up risk combining asynchronous electronic medical record sequence prediction as described in claim 1, characterized in that, Before grouping the dynamic variables in the input samples by category to obtain multiple dynamic variable groups, the process further includes missing variable retention and asynchronous state construction for the dynamic variables at each follow-up node, wherein... For dynamic variables with missing values ​​at each follow-up node, either of the following two strategies is used to complete the missing value retention: zero-value imputation and adding an observation mask to each dynamic variable to explicitly encode the missing value pattern, or intra-patient forward imputation combined with historical statistics from the same institution for backfilling. For each dynamic variable at each follow-up node, an asynchronous state is constructed. The asynchronous state construction includes at least whether the current observation is a true observation and the time information since the last true observation.

5. The method for real-time early warning of chronic disease follow-up risk combining asynchronous electronic medical record sequence prediction according to claim 1, characterized in that, The process involves using a temporal encoder to extract variable lexical representations for each dynamic variable group, and then projecting static baseline information into a static lexical representation and concatenating it with the variable lexical representations. This includes treating each dynamic variable group as an independent time stream, inputting it into a temporal encoder with a length mask, obtaining the variable lexical representation of the corresponding dynamic variable group through mask pooling, and simultaneously projecting the static baseline information of the corresponding target patient into a static lexical representation and concatenating it with the variable lexical representations.

6. The method for real-time early warning of chronic disease follow-up risk combining asynchronous electronic medical record sequence prediction as described in claim 1, characterized in that, Feature-level attention aggregation is performed on the concatenated lexical set, including: Calculate the attention weight of each word representation in the word set, and use the attention weight to perform a weighted summation of each word representation to obtain a comprehensive risk characterization; Before calculating the attention weight for each lexical representation, a comprehensive evaluation of the current credibility and relevance of each dynamic variable to the current task is required. The current level of credibility is related to at least the following factors: whether the current dynamic variable is observed at the trigger point, the length of time since the last real observation, whether the data source is reliable, whether there is obvious noise recently, and whether the change of the dynamic variable is continuous and consistent. The relevance of the current task refers to the strength of the indication of the current dynamic variable to the specified adverse event.

7. The method for real-time early warning of chronic disease follow-up risk combining asynchronous electronic medical record sequence prediction as described in claim 1, characterized in that, The risk prediction model includes multiple independent models trained for different adverse events, or a multi-head output model with a shared underlying representation obtained through multi-task joint training; the risk prediction model outputs a risk score for the target patient to experience a target adverse event within a future preset risk window period, and determines whether the target patient will experience a specified adverse event within the future preset risk window period based on the risk score.