System for evaluating curative effect of end-stage nephropathy and intelligently predicting complication risk
Through incremental learning, interpretability output and feature weighting modules, the problems of poor adaptability of models, opaque and unstable prediction of prediction results in the prior art are solved, and efficient and reliable prediction of end-stage renal disease efficacy evaluation and complication risk are achieved, and doctors’ decision-making is supported.
Patent Information
- Application Number
- CN202510590900.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-15
AI Technical Summary
In the face of dynamic changes in clinical data, the model cannot be adapted, the predicted results lack interpretability, the feature fusion particle size is rough, and the online update process is unstable, which affects the accuracy and reliability of end-stage renal disease efficacy evaluation and complication risk.
The incremental learning module is used to update the model parameters, and the interpretability output module is introduced to provide feature importance analysis. The feature weighting module is deeply integrated with multi-source multi-modal data, and the recurrent neural network modeling time correlation is used, and the model stability is controlled by the online learning module to achieve dynamic adaptation and transparent prediction.
It improves the adaptability and accuracy of the model in actual clinical applications, improves the transparency and trust of the predicted results, ensures the stability and long-term reliability of the model, and supports doctor-assisted decision-making.
Smart Images

Figure CN120496842A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical artificial intelligence technology, and specifically to an intelligent prediction system for end-stage renal disease efficacy evaluation and complication risk. Background Art
[0002] In real-world clinical scenarios, doctors need to assess the progression of a patient's condition in order to develop timely and effective treatment plans. Early warning is crucial, especially when facing high-risk complications. However, faced with massive amounts of multi-source, dynamically changing medical data, doctors often struggle to identify key indicators in a timely manner. Therefore, leveraging intelligent systems to predict future risks and provide reasonable explanations and updates becomes a crucial tool for improving clinical efficiency and reducing the risk of medical delays.
[0003] Some technological advances have made progress in medical AI prediction. For example, some models, based on deep neural network architectures, can be trained on electronic medical record data to predict the probability of certain diseases. These methods are stable on large, static datasets and have demonstrated good accuracy in screening for some chronic diseases. Furthermore, the introduction of attention mechanisms has significantly improved the model's ability to model temporal structures, helping to enhance the model's perception of key time points. Other research is exploring the use of joint analysis of multimodal information to obtain a more comprehensive representation of patient characteristics.
[0004] Despite this, existing technologies still have several key shortcomings when facing real clinical scenarios. First, most models rely on offline training, and their parameters are frozen after deployment. They have no adaptability to new data, and updates rely entirely on manual retraining, which is extremely inefficient. Second, existing methods generally lack a clear explanation of the basis for predictions, only outputting results and lacking quantitative analysis of the impact of key features, making it difficult for doctors to trust them. Furthermore, some methods are rough in feature fusion, important variables are easily obscured, and information loss is serious, affecting the quality of the final prediction. Finally, even if there are attempts at online learning, most of them ignore model stability control, and the update process is uncontrolled, which can easily cause model oscillations and even performance regression. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present invention provides an intelligent prediction system for the efficacy evaluation and complication risk of end-stage renal disease, which solves the problems in the existing technology that the model cannot adapt to the dynamic changes of clinical data, the prediction results lack interpretability, the feature fusion granularity is coarse, and the online update process is unstable.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: an end-stage renal disease efficacy evaluation and complication risk intelligent prediction system, comprising: The data acquisition module acquires multi-source data of patients by establishing communication connections with hospital information systems, physiological monitoring equipment, and user terminals. The multi-source data includes structured physiological indicator data, laboratory test data, and unstructured text data of the patient's medical history; The feature fusion module receives multi-source data from the data acquisition module, standardizes the structured physiological indicator data, converts the unstructured text data into vector form through embedded coding, and then splices it with the structured features to construct a unified time series input feature; The time series modeling module receives the time series input features from the feature fusion module, uses the recurrent neural network structure to model the time correlation, and extracts the sequence features that reflect the evolution of the patient's status; A feature weighting module receives the hidden state output from the time series modeling module and is used to perform attention weighting processing on the time series features to generate context feature representation; The prediction module receives the context feature representation and generates a predicted value representing the risk probability of the patient developing complications within a set time window in the future; The online learning module receives the predicted value, combines it with the historical samples to perform incremental model parameter updates, and sets the parameter change threshold to control the stability of the model; The explainability output module is used to perform feature importance analysis on the predicted values and output the influence of each feature in the current prediction to assist doctors in decision-making.
[0007] Preferably, the data acquisition module includes: An interface establishment unit, used to establish a data interface with the hospital information system, physiological monitoring equipment and user terminals; Data parsing unit, used to convert the format and extract labels of collected structured and unstructured data; The data cache unit is used to temporarily store data organized by time tags and provide it to subsequent modules for access.
[0008] Preferably, the feature fusion module includes: A standardization processing unit is used to normalize structured physiological indicators and laboratory test data; Embedding encoding unit, which converts unstructured text input of patient medical history and symptom description into semantic vector representation; The feature concatenation unit is used to concatenate structured features and text vectors at the same time step to form a unified input format.
[0009] Preferably, the time series modeling module includes: The recurrent neural network unit is used to sequentially receive time series input features and generate a hidden state vector for each time step; the state update unit is used to calculate the current state representation based on the previous hidden state and the current input features at each time step.
[0010] Preferably, the feature weighting module includes: Weight calculation unit, used to calculate the attention weight for each time step; The weighted aggregation unit is used to perform weighted aggregation of the hidden states of all time steps according to the attention weights to obtain the contextual representation.
[0011] Preferably, the prediction module includes: A feature reading unit, configured to receive a context feature vector from a feature weighting module; The risk estimation unit is used to input the context feature vector into the fully connected network and output the predicted probability value of the target variable.
[0012] Preferably, the online learning module includes: A new sample access unit, used to receive the latest input patient data and its feedback labels; Incremental training unit, used to update the current model parameters without retraining the entire model; The stability control unit is used to set the maximum weight change threshold to limit drastic changes in the model structure.
[0013] Preferably, the interpretability output module includes: Feature contribution analysis unit, used to estimate the marginal contribution of each feature to the prediction result based on the model's internal weight and feature sensitivity; The visual output unit is used to convert the interpretation results into a graphical interface form and push it to the doctor's terminal for display.
[0014] Preferably, the standardization processing unit includes: Data screening unit, used to detect and process missing values and outliers in structured input; A numerical normalization unit is used to perform minimum-maximum normalization on structured physiological indicators and laboratory test data according to set standards; The feature alignment unit is used to align and splice structured data at the same time point in different data sources according to timestamps.
[0015] Preferably, the weight calculation unit includes: A feature score generation unit that calculates intermediate score values based on the nonlinear relationship between the hidden state and the trainable weight vector at each time step; The weight normalization unit is used to convert the score values of all time steps into attention distribution probabilities through a normalization function to form weighted weight coefficients for subsequent aggregation processing.
[0016] The present invention provides an intelligent system for evaluating the efficacy of end-stage renal disease and predicting the risk of complications. It has the following beneficial effects: 1. This invention utilizes an online learning module based on incremental learning, enabling real-time updates of model parameters to address changes in patient status and dynamic shifts in data distribution. This achieves the technical effect of improving the model's adaptability and accuracy in actual clinical applications. Compared to existing models that rely solely on offline training and are unable to quickly adapt to new data, this invention effectively addresses the problem of unresponsive models in rapidly changing medical data environments.
[0017] 2. This invention introduces an interpretable output module that analyzes feature importance and provides the specific impact of each feature on the current prediction. This allows doctors to clearly understand the basis for model decisions. Compared to traditional black-box models that lack visual explanations, this invention significantly improves the transparency and trustworthiness of prediction results, assisting clinicians in decision-making.
[0018] 3. By introducing a feature weighting module, this invention deeply integrates multi-source and multi-modal clinical data, optimizing the feature extraction and characterization process. This achieves the technical effect of more accurately capturing the evolution of a patient's health status. Compared to traditional single-data source models, this invention can more comprehensively reflect a patient's actual condition, thereby improving the accuracy and reliability of risk prediction.
[0019] 4. By employing a parameter change threshold control mechanism, this invention ensures model stability during online learning and avoids overfitting caused by excessive adjustments. This achieves the technical effect of improving model stability and long-term reliability. Compared to existing solutions that lack effective constraints on model updates, this invention effectively reduces model instability and bias accumulation caused by frequent updates. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 Construct a diagram for the system of the present invention; Figure 2 This is a data acquisition module framework diagram of the present invention; Figure 3 This is a feature fusion module framework diagram of the present invention; Figure 4 This is a framework diagram of the timing modeling module of the present invention; Figure 5 This is a feature weighting module framework diagram of the present invention; Figure 6This is a framework diagram of the prediction module of the present invention; Figure 7 This is a framework diagram of the online learning module of the present invention; Figure 8 This is a framework diagram of the interpretable output module of the present invention. DETAILED DESCRIPTION
[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the present specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0022] Please see the attached Figure 1 -Attached Figure 8 The embodiment of the present invention provides an end-stage renal disease efficacy evaluation and complication risk intelligent prediction system, comprising: The data acquisition module establishes communication connections with the hospital information system, physiological monitoring equipment, and user terminals to obtain multi-source data of patients. The multi-source data includes structured physiological indicator data, laboratory test data, and unstructured text data of the patient's medical history; The data acquisition module is primarily responsible for communicating with the hospital information system, physiological monitoring equipment, and patient terminals, enabling the automatic acquisition and structured organization of multi-source, heterogeneous patient data. The output of this module serves as direct input to the feature fusion module, so its format and temporal integrity are particularly important.
[0023] In this embodiment, the data acquisition module includes three sub-functional components: communication interface management, data analysis and processing, and cache organization. These components cooperate with each other to ensure stable data input.
[0024] In one possible implementation, the communication interface management component establishes asynchronous communication links with hospital information systems (HIS), laboratory information systems (LIS), and electronic medical records (EMR) systems by setting up a standardized API gateway. These interfaces interact using HL7, FHIR, or a custom JSON-RPC protocol, supporting RESTful or message queue-based data retrieval.
[0025] Typically, data exchange between physiological monitoring devices and this system is achieved via Bluetooth, Wi-Fi, or RS-232 serial port protocols. Device types include, but are not limited to, Holter monitors, blood pressure monitors, respiratory rate sensors, electronic weight scales, and wrist-worn oximeters.
[0026] Specifically, for data collection from physiological monitoring devices, the system collects data every Δt interval, where Δt is a time interval constant in seconds. In some embodiments, Δt can be set to 30 seconds, 60 seconds, or other adjustable values. The collection results form the following structured record: Among them, X t Represents the structured physiological index vector at time t; represents the measured value of the i-th physiological indicator at time t; n represents the total number of physiological characteristics collected.
[0027] As an option, when collecting laboratory test data, the system primarily extracts key indicators such as renal function (such as serum creatinine, urea nitrogen, and eGFR) and electrolytes (such as potassium, sodium, phosphorus, and calcium). Each indicator is accompanied by information about the sampling time, testing institution, and unit. The system uses a standardized mapping dictionary to uniformly map the raw data to an internally defined label coding system.
[0028] In some embodiments, in order to enhance the integrity of the data source, the system can also access user terminal devices (such as mobile applications, wearable device apps) to obtain soft information such as symptom records, medication compliance feedback, and eating habits filled out by patients every day. This information mostly appears in the form of unstructured natural language text and is usually difficult to use directly. To this end, the system uses natural language processing models such as BERT, ERNIE or RoBERTa to semantically encode the text and generate a semantic embedding vector under the corresponding time label: E t =Embed(Text t ); Among them, Text t Represents the text content provided by the patient at time t; Embed represents the text embedding operation, E t is the generated d-dimensional semantic vector; d is the semantic embedding dimension, usually 128 or 256.
[0029] In a possible implementation, structured data and unstructured semantic embeddings are organized in a sliding window manner in the cache organization module. The window size is set to T w , that is, the system takes T w A complete time series input segment is constructed by continuous time steps, ensuring that the subsequent feature fusion module can load sample data in a temporally consistent manner.
[0030] Furthermore, to support multi-task concurrent acquisition, the system uses a bidirectional queue (deque) structure in the data cache unit while maintaining thread-safety for data reading and writing. The cache uses a FIFO strategy to eliminate expired time-step data to avoid redundant accumulation that affects computational efficiency.
[0031] As a supplement, before data cleaning, the system performs timestamp normalization on all collected data, that is, mapping all time information to a unified time domain [0, T], where T is the maximum step size of the current collection time series. For example, for non-uniformly sampled data, the system will use interpolation algorithms (such as linear interpolation or spline interpolation) to fill in the missing time points in the middle and construct a complete time vector {t1, t2, ..., t T}.
[0032] The feature fusion module receives multi-source data from the data acquisition module, standardizes the structured physiological indicator data, converts the unstructured text data into vector form through embedded coding, and then splices it with the structured features to construct a unified time series input feature; After the data acquisition module completes access and caching of multi-source data, the system needs to further structure and integrate the acquired data to form a unified input format for time series modeling. This fusion process is performed by the feature fusion module of the present invention. Its function is to achieve semantic unification, time series synchronization, and numerical normalization among heterogeneous data, providing high-quality, consistent feature input for subsequent time series modeling and prediction.
[0033] The data types received by the feature fusion module include two categories: one is structured physiological indicator data and laboratory test data, and the other is unstructured natural language text data, such as medical history descriptions, main symptoms, self-report records, etc.
[0034] In this embodiment, to ensure consistency in statistical dimensions across different data sources, the system first performs a normalization operation on the structured data. This operation uses a minimum-maximum normalization method to process numerical features. The normalization formula is as follows: in, represents the measured value of the i-th physiological indicator at time t; and Respectively represent the minimum and maximum values of the feature in the historical data; Represents the normalized numerical result, with a value range of [0,1].
[0035] In general, the normalized range can be clipped or truncated according to medical standards. For example, when indicators such as blood pressure and creatinine have physiological upper and lower limits, the present invention supports soft truncation of outliers to prevent extreme values from dominating model learning.
[0036] As an option, to further improve the stability of the semantic representation, the encoder output can be post-processed, such as mean pooling, residual connection, or normalized activation, so that the final embedding vector is stable and comparable in the time dimension.
[0037] In a typical processing cycle, the system converts the normalized structured data X ′ t and the embedded unstructured semantic vector E t Concatenate features by time step to construct the final unified input vector: Z t =X ′ t ⊕E t ; Among them, ⊕ represents the vector-level splicing operation; Z t is the fusion feature vector of time step t; E t is the generated d-dimensional semantic vector; X ′ t Represents the transformed result of the input feature vector at time step t.
[0038] In order to meet the needs of subsequent time series processing, all the spliced features Z t Need to be aligned in time. To this end, the present invention uses linear interpolation to process missing data fragments in the time step. The system defines the target time series window as: Among them, t1, t2, …, t T Represents each specific time step within the time window; Represents the unified aligned time axis set; T represents the total number of time steps in the window; for any Z t If the original data is missing, the average value of the adjacent time steps is used for interpolation.
[0039] In some embodiments, to ensure the consistency of feature dimensions, the system also performs a zero vector filling operation. t ,definition: E t =0 d , Among them, 0 d represents the d-dimensional zero vector, E t is the generated d-dimensional semantic vector, Text t Represents the text content provided by the patient at time t; Represents an empty collection.
[0040] The time series modeling module receives the time series input features from the feature fusion module, uses the recurrent neural network structure to model the time correlation, and extracts the sequence features that reflect the evolution of the patient's status; Building on the integration of multimodal input features by the aforementioned feature fusion module, the present invention introduces a time series modeling module to effectively capture the patient's dynamic state over time. This module receives the structured time series features output by the feature fusion module and aims to deeply explore the temporal dependencies implicit in the sequence, thereby improving the model's ability to characterize disease progression trends.
[0041] Generally speaking, the evolution of a patient's condition has significant temporal correlation, and it is difficult to fully characterize it by relying solely on static or fragmentary features. In this context, building a neural network structure with memory capabilities becomes a reasonable choice.
[0042] In this embodiment, the time series modeling module is implemented based on a recurrent neural network architecture. In one possible implementation, a long short-term memory network is preferably used to enhance the model's ability to model long-term dependency information.
[0043] In some embodiments, the input time series features are assumed to be: X=[x1,x2,…,x T ]; Where X represents a sequence input feature set; x1,x2,…,x T Represent each time step separately.
[0044] In the LSTM network, the state update at each time step satisfies the following formula: f t =σ(W f x t +U f h t-1 +b f ); i t =σ(W i x t +U i h t-1 +b i ); o t =σ(W o x t +U o h t-1 +b o ); h t =o t ⊙tanh(ct ); Among them, f t 、i t 、o t Represent the forget gate, input gate and output gate respectively; is the candidate cell state; c t is the cell state at the current time step; h t is the hidden state of the current time step; σ(·) represents the Sigmoid activation function; tanh(·) represents the hyperbolic tangent function; ⊙ represents element-wise multiplication; h t-1 is the hidden state of the previous time step; W f ,W i , W o , W c are the input weight matrices of the forget gate, input gate, output gate, and candidate state respectively; U f ,U i , U o , U c are the hidden state weight matrices of the forget gate, input gate, output gate, and candidate state respectively; b f ,b i , b o , b c is the bias term corresponding to the gate and candidate state; In one possible implementation, a multi-layer stacked LSTM network structure can be further adopted to improve the model's expressiveness. As an option, the hidden state h t It can be connected to downstream modules, such as classifiers, regressors, or other decision-making units, for subsequent task processing.
[0045] In some embodiments, in order to enhance the network's ability to select features at different time points, an attention mechanism can be introduced based on the LSTM output. Specifically, the time attention weight α is introduced t , calculated as follows: in, is a trainable attention weight vector; tanh(·) represents the hyperbolic tangent function; α t represents the attention weight corresponding to the t-th time step; e t Score the attention at time step t; W a is the trainable weight matrix in the attention mechanism; h t is the hidden state vector at the tth time step; b a is the trainable bias vector in the attention mechanism; exp(·) is the exponential function used to convert the score e t Mapping is positive; To normalize the scores of all time steps; T is the total number of time steps.
[0046] The attention-weighted sequence representation s can be defined as: Where T is the total number of time steps; α t represents the attention weight corresponding to the t-th time step; h t is the hidden state vector of the t-th time step; s is the context representation vector after weighted summation; t is the index of the current time step.
[0047] In a specific embodiment, the s vector is used as the output of the time series modeling module and further participates in tasks such as patient status assessment and risk prediction.
[0048] As an option, to improve the generalization ability of the model in different clinical scenarios, the above-mentioned LSTM module can also be replaced with a gated recurrent unit, which has a relatively simple structure and fewer parameters and is suitable for some resource-constrained scenarios.
[0049] In some other embodiments, for feature sequences with different time resolutions, a multi-scale time series modeling structure can be constructed to model short-term and long-term trends separately, and to obtain a unified sequence representation through feature fusion.
[0050] It should be noted that all parameters in the time series modeling module are optimized through end-to-end training, which can be achieved by minimizing the target loss function, where the loss function can be set according to the specific task, such as cross entropy loss, mean square error, etc.
[0051] In different implementation environments, the module can be adjusted to adapt to changing requirements such as different input dimensions, sequence lengths, network depths, etc., ensuring that the module has good scalability and adaptability.
[0052] The feature weighting module receives the hidden state output from the time series modeling module and is used to perform attention weighting processing on the time series features to generate contextual feature representations; The feature weighting module inherits the output of the time series modeling module within the system architecture and specifically receives the hidden state sequence representation it generates. Typically, each hidden state vector corresponds to a time step in the original input sequence. Because different time points may have varying importance for the prediction target, simply using the average or last time step representation results in information loss. Therefore, it is necessary to perform weighted aggregation of time series features.
[0053] In this embodiment, the feature weighting module uses an attention mechanism to weight all hidden state vectors output by the time series modeling module. Specifically, the module first calculates the attention score corresponding to each time step, which reflects the contribution of that time point to the overall context feature expression.
[0054] In one possible implementation, the scoring process is performed by a trainable feedforward neural network that takes each hidden state vector as input and outputs a corresponding scalar score. The score values are normalized and converted into normalized attention weights.
[0055] Alternatively, normalization can be achieved using a Softmax function, so that the sum of the attention weights across all time steps is 1, ensuring overall semantic consistency. The system then performs a weighted summation of all hidden state vectors according to their corresponding attention weights, ultimately generating a fixed-dimensional context vector.
[0056] The context vector, as the final output of this module, comprehensively integrates the information in the entire time series, especially highlighting the state features contained in the time segment most relevant to the current task.
[0057] Specifically, in some embodiments, the context vector is passed to a discrimination module, a classification module, or a risk scoring module to complete the modeling of the target task.
[0058] In one specific implementation, the attention scoring function can use a two-layer nonlinear transformation, where the first layer uses a hyperbolic tangent activation function, and the second layer uses a linear transformation to output a scalar attention score. This structure can be adjusted to meet different model capacity requirements, such as adjusting the intermediate layer dimensions or activation function type.
[0059] It should be noted that to ensure the interpretability and stability of the attention mechanism during training, this embodiment introduces an entropy regularization constraint on the attention weights. This constraint controls the sparsity of the weight distribution to avoid over-concentration or over-dispersion of the weights.
[0060] In some other embodiments, a multi-head attention structure can also be introduced to divide the hidden state input into multiple subspaces, each subspace learns a set of attention weights separately, and finally concatenates multiple context representations as the final output to enhance the expressive power of the model.
[0061] The feature weighting module, a crucial component of the present invention, has a structural design that directly impacts the sensitivity and discriminative power of the final model across different clinical conditions. The context vector it outputs not only preserves the global trends of temporal evolution but also maintains dimensional compatibility with subsequent prediction structures, ensuring seamless integration between modules.
[0062] The prediction module receives the context feature representation and generates a predicted value representing the risk probability of the patient developing complications within a set time window in the future; After the feature weighting module completes the deep integration of time series information, the system performs a final risk assessment based on the extracted contextual semantic vectors. The prediction module provided by the present invention is designed to receive the contextual feature representation output by the previous module and generate a probability prediction result for a specific clinical event based on this representation.
[0063] Generally, clinical risk prediction tasks require a clear prediction window timeframe and personalized judgments based on the patient's historical multimodal information. Because the feature weighting module already outputs a vector representation that reflects the evolutionary patterns over the entire time period, this module further constructs a prediction function model based on this output to estimate the risk probability of target events (such as complications and disease progression).
[0064] In this embodiment, the prediction module takes the contextual feature vector as input and outputs a value representing the patient's risk of developing a specific complication within a set future time window. The core structure of this module is a feedforward neural network, typically composed of multiple cascades of linear transformation layers, nonlinear activation layers, and normalization layers.
[0065] The prediction function can be expressed as: in, represents the risk probability value of the output, with a value range of [0,1]; W1 and W2 are the weight matrices of the first and second layers respectively; b1 and b2 are the corresponding bias terms; φ represents a nonlinear activation function, such as ReLU or GELU; σ represents the Sigmoid function, which is used to map the output to a probability value; s represents the context representation vector after weighted summation.
[0066] In one possible implementation, the prediction module can output probability distributions for multiple complication categories, corresponding to different clinical risk types. For example, the prediction results for heart failure, acute kidney injury, or infectious complications can be calculated separately using multiple output channels.
[0067] In some embodiments, to prevent the model from overfitting to specific feature dimensions, a Dropout layer may be introduced into the feedforward network structure. This layer randomly blocks some neurons with a set probability, which helps improve the model's generalization ability on unseen samples.
[0068] In another specific implementation, a weighted cross-entropy loss function is introduced during training to guide the model to focus more on high-risk samples. This loss function can effectively alleviate label skew in the context of uneven sample distribution and improve the model's ability to identify rare, high-risk categories.
[0069] It’s important to note that within the framework of this invention, the prediction module not only accepts a single feature vector as input but can also be expanded to handle multi-channel contextual representations. By connecting multiple feedforward channels in parallel, we can achieve multi-perspective concurrent modeling and fuse multiple predictions at the output layer.
[0070] As an option, to enhance the timeliness of the model, the prediction time window can be flexibly adjusted based on clinical needs, such as setting complication prediction tasks for the next 24 hours, 72 hours, or 7 days. This setting must be uniformly processed during the model training phase through a label alignment mechanism to ensure consistency between the input sequence and the prediction target time window.
[0071] The risk probability value output by this module will eventually serve as an externally readable result of the overall system of the present invention, which can be used to drive the early warning interface in the electronic medical record system, or to link the doctor intervention mechanism to improve the response efficiency to high-risk patients.
[0072] The online learning module receives the predicted value, combines it with the historical samples to perform incremental model parameter updates, and sets the parameter change threshold to control the stability of the model; To improve the model's adaptability to diverse patient populations and dynamically reflect the nonstationary nature of data distribution during disease evolution, the present invention introduces an online learning module to continuously update model parameters during system operation. Based on the risk probability values output by the prediction module and combined with historical sample information, this module employs an incremental learning strategy to locally adjust the model. A parameter change control mechanism is also implemented to ensure model stability during evolution.
[0073] In general, data streams in medical environments are constantly updated. Traditional offline training models struggle to adapt to the frequent fluctuations and increasing diversity of patient conditions. Therefore, it is necessary to introduce an online adaptive parameter update mechanism to enable the model to continuously absorb new information and mitigate the impact of concept drift.
[0074] In this embodiment, the online learning module receives the risk probability value and its corresponding actual label information output by the prediction module, and forms a joint training set with the historical samples. The training set serves as the data basis for online incremental updates.
[0075] Specifically, in each round of online learning, the system collects true label feedback samples from the most recent time window. Let the sample set be: Among them, s (i) Represents the context feature vector of the i-th sample; y (i) is the corresponding true label; N represents the number of new samples in the current batch; represents the newly added data set; i is the sample index.
[0076] As an option, to enhance the model's ability to retain historical data, the system also selects a portion of representative samples from the existing training samples to form a historical reference subset, which is denoted as: in, represents the reference subset dataset; s (i) Represents the context feature vector of the i-th sample; y (i) is the corresponding true label; M represents the number of samples in the reference subset, which can be dynamically determined by a fixed window length or importance sampling mechanism; j is the sample index.
[0077] In this embodiment, the online update of the model is achieved by minimizing the target loss function, which is a weighted combination of the current new sample loss and the reference sample loss, expressed as follows: in, represents the standard loss function, such as the cross entropy loss used in binary classification tasks; λ is the weight coefficient, which is used to adjust the proportion of the influence of new and old samples on the update of model parameters; is the total loss function value; To represent the newly added data set; represents the reference subset dataset; In one possible implementation, to prevent unstable behavior caused by frequent model updates, a parameter change threshold mechanism is set in the present invention. After executing the parameter update, the system calculates the Euclidean distance before and after the model parameter update: Δθ=||θ new -θ old ||2; Among them, Δθ represents the update amplitude of model parameters; θ old represents the model parameter vector before updating; θ new Represents the updated model parameter vector.
[0078] Generally, when Δθ exceeds the preset threshold ∈, the system will automatically perform a rollback operation or start the entropy regularization smooth update process to ensure that the parameter update is within a controllable range.
[0079] In some embodiments, to improve update efficiency, the system uses a mini-batch approach to load incremental samples in batches and retains a momentum term in each update to suppress gradient oscillations. During training, the learning rate can also be dynamically adjusted to achieve more fine-grained learning rate annealing.
[0080] As an option, the present invention also introduces a gradient clipping strategy to limit parameter update items whose gradients exceed a predetermined range, thereby preventing large gradients from interfering with the model structure.
[0081] In another specific implementation, the system can set up an online learning frequency control mechanism to trigger parameter updates only when the cumulative error rate or drift index exceeds a threshold, thereby saving computing resources and enhancing model stability.
[0082] It's important to note that the online learning module maintains end-to-end structural consistency during parameter updates, and all updates apply to the parameters of the aforementioned time series modeling module, feature weighting module, and prediction module. Modules seamlessly connect via shared gradients and parameter interfaces, ensuring continuous learning across the entire system.
[0083] The interpretability output module is used to perform feature importance analysis on the predicted values and output the influence of each feature in the current prediction to assist doctors in decision-making. After the prediction module outputs the patient's risk probability of complications within a future time window, further interpreting the prediction results is necessary to ensure the model's comprehensibility and traceability in a clinical setting. The interpretable output module proposed in this paper analyzes the importance of the features underlying the model's predictions and outputs the specific impact of each feature on the current sample's prediction, thereby assisting physicians in risk assessment and formulating intervention strategies.
[0084] Generally speaking, black-box models fail to provide intuitive evidence for clinical decision-making. Even if a model has high predictive accuracy, it may be questioned or abandoned due to its lack of transparency. Therefore, designing interpretable model outputs is particularly critical in medical scenarios.
[0085] In this embodiment, the interpretability output module receives the risk probability value output by the prediction module. After that, the context feature vector s=[s1,s2,...,s d ], calculate the impact of each feature in this prediction through sensitivity analysis, attention attribution, or model response perturbation-based methods.
[0086] Specifically, in one possible implementation, a feature attribution method based on input gradient is used, and the calculation is as follows: Among them, I i Indicates the sensitivity of the i-th input feature to the predicted output (i.e., importance score); Represents the output risk probability value; s i represents the value of the i-th feature; is the partial derivative of the predicted value with respect to the input features.
[0087] The above derivatives represent the responsiveness of the predicted output when the feature value changes slightly, thus reflecting the importance of the feature.
[0088] In another implementation, if the prediction module includes an explicit attention mechanism, the attention weight can be directly used as an estimate of the feature contribution. Assume that the context feature representation is generated by the attention mechanism, and the relevant attention distribution is: Among them, α i represents the attention weight of the i-th feature; e i is the score of the i-th original attention generated by the feature weighting module; d is the total number of feature dimensions; exp is the exponential function; is the sum of the index scores of all features; j is the index of the feature, At this time, the feature importance index can be: I i =α i ; Among them, α i Represents the attention weight of the i-th feature; I i Indicates the sensitivity of the i-th input feature to the predicted output (i.e., importance score).
[0089] In one specific implementation, to prevent the feature importance evaluation results from being highly sensitive to outliers, a normalization mechanism can be introduced after the gradient or attention weight attribution. This mechanism normalizes the feature importance to the [0, 1] interval based on the maximum and minimum scaling method: in, is the normalized feature influence degree; I i It represents the sensitivity of the i-th input feature to the predicted output (i.e., the importance score); min(I) represents the minimum value in the input feature set I; max(I) represents the maximum value in the input feature set I; I represents the input feature set, which contains the values of all features and is used to calculate the maximum and minimum values.
[0090] In some embodiments, the interpretable output module not only outputs the attributed value for each feature but also visualizes these results as a bar chart or heat map, further enhancing the physician's intuitive understanding of the basis for the model's decisions. This graphical interface can be embedded in electronic medical record systems or external analysis platforms.
[0091] As an option, contrastive learning or self-supervised discrimination mechanisms can be introduced during the model training phase to enhance the discrimination of different features between high-risk and low-risk samples, making the final generated feature interpretation results more robust and credible.
[0092] In another specific implementation, in order to avoid redundancy of explanation information caused by high-dimensional feature space, the system can integrate feature selection mechanisms in the explanation module, such as L1 sparsity constraint or principal component projection technology, to only retain the explanation output of key features.
[0093] It's important to note that the feature importance generated by the module is not limited to single prediction samples; it can also be extended to provide a global interpretation at the batch level. By integrating predictions across multiple moments through a sliding window, the system can further identify variables with long-term high impact, providing doctors with more time-continuous clinical references.
[0094] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent system for evaluating the efficacy of end-stage renal disease and predicting the risk of complications, characterized by: include: The data acquisition module acquires multi-source data of patients by establishing communication connections with hospital information systems, physiological monitoring equipment, and user terminals. The multi-source data includes structured physiological indicator data, laboratory test data, and unstructured text data of the patient's medical history; The feature fusion module receives multi-source data from the data acquisition module, standardizes the structured physiological indicator data, converts the unstructured text data into vector form through embedded coding, and then splices it with the structured features to construct a unified time series input feature; The time series modeling module receives the time series input features from the feature fusion module, uses the recurrent neural network structure to model the time correlation, and extracts the sequence features that reflect the evolution of the patient's status; A feature weighting module receives the hidden state output from the time series modeling module and is used to perform attention weighting processing on the time series features to generate context feature representation; The prediction module receives the context feature representation and generates a predicted value representing the risk probability of the patient developing complications within a set time window in the future; The online learning module receives the predicted value, combines it with the historical samples to perform incremental model parameter updates, and sets the parameter change threshold to control the stability of the model; The explainability output module is used to perform feature importance analysis on the predicted values and output the influence of each feature in the current prediction to assist doctors in decision-making.
2. The intelligent prediction system for end-stage renal disease efficacy evaluation and complication risk according to claim 1, characterized in that: The data acquisition module includes: An interface establishment unit, used to establish a data interface with the hospital information system, physiological monitoring equipment and user terminals; Data parsing unit, used to convert the format and extract labels of collected structured and unstructured data; The data cache unit is used to temporarily store data organized by time tags and provide it to subsequent modules for access.
3. The intelligent prediction system for end-stage renal disease efficacy evaluation and complication risk according to claim 1, characterized in that: The feature fusion module includes: A standardization processing unit is used to normalize structured physiological indicators and laboratory test data; Embedding encoding unit, which converts unstructured text input of patient medical history and symptom description into semantic vector representation; The feature concatenation unit is used to concatenate structured features and text vectors at the same time step to form a unified input format.
4. The intelligent prediction system for end-stage renal disease efficacy evaluation and complication risk according to claim 1, characterized in that: The timing modeling module includes: A recurrent neural network unit, which receives time series input features in sequence and generates a hidden state vector for each time step; The state update unit is used to calculate the current state representation based on the previous hidden state and the current input features at each time step.
5. The intelligent prediction system for end-stage renal disease efficacy evaluation and complication risk according to claim 1, characterized in that: The feature weighting module includes: Weight calculation unit, used to calculate the attention weight for each time step; The weighted aggregation unit is used to perform weighted aggregation of the hidden states of all time steps according to the attention weights to obtain the contextual representation.
6. The intelligent prediction system for end-stage renal disease efficacy evaluation and complication risk according to claim 1, characterized in that: The prediction module includes: A feature reading unit, configured to receive a context feature vector from a feature weighting module; The risk estimation unit is used to input the context feature vector into the fully connected network and output the predicted probability value of the target variable.
7. The intelligent prediction system for end-stage renal disease efficacy evaluation and complication risk according to claim 1, characterized in that: The online learning modules include: A new sample access unit, used to receive the latest input patient data and its feedback labels; Incremental training unit, used to update the current model parameters without retraining the entire model; The stability control unit is used to set the maximum weight change threshold to limit drastic changes in the model structure.
8. The intelligent prediction system for end-stage renal disease efficacy evaluation and complication risk according to claim 1, characterized in that: The interpretability output module includes: Feature contribution analysis unit, used to estimate the marginal contribution of each feature to the prediction result based on the model's internal weight and feature sensitivity; The visual output unit is used to convert the interpretation results into a graphical interface form and push it to the doctor's terminal for display.
9. The intelligent prediction system for end-stage renal disease efficacy evaluation and complication risk according to claim 3, characterized in that: The standardization processing unit includes: Data screening unit, used to detect and process missing values and outliers in structured input; A numerical normalization unit is used to perform minimum-maximum normalization on structured physiological indicators and laboratory test data according to set standards; The feature alignment unit is used to align and splice structured data at the same time point in different data sources according to timestamps.
10. The intelligent prediction system for end-stage renal disease efficacy evaluation and complication risk according to claim 5, characterized in that: The weight calculation unit includes: A feature score generation unit that calculates intermediate score values based on the nonlinear relationship between the hidden state and the trainable weight vector at each time step; The weight normalization unit is used to convert the score values of all time steps into attention distribution probabilities through a normalization function to form weighted weight coefficients for subsequent aggregation processing.
Citation Information
Cited By
Chronic kidney disease intelligent follow-up visit management system and method based on multi-source data
CN120809249A
Multi-sign fusion lung respiration monitoring system and method based on LSTM neural network
CN120859475A