Unbiased data stream and secure update method for emergency-oriented model continuous learning

CN122822385APending Publication Date: 2026-09-25THE FIRST AFFILIATED HOSPITAL OF BENGBU MEDICAL COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611037482.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]本发明提供面向急诊模型持续学习的无偏数据流与安全更新方法,用以解决现有技术中急诊数据流存在分布漂移、数据不平衡和质量参差不齐的缺陷

Benefits of technology

[0054]S57:搭建与真实急诊环境一致的模拟诊疗链路,输入包含多患者并发、多模态数据异步到达、紧急抢救工况的复杂场景的测试数据,验证从数据采集、预处理、模型推理到决策输出的全流程因果一致性与响应时效性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122822385A_ABST
    Figure CN122822385A_ABST
Patent Text Reader

Abstract

The application provides an unbiased data stream and safe updating method for emergency model continuous learning, relates to the technical field of unbiased data stream, and comprises the following steps: constructing an original training data stream; performing multi-dimensional unbiased processing and quality control of clinical semantic drift sensing on the original training data stream; performing incremental model training of clinical knowledge anchoring, reserving historical diagnosis and treatment knowledge through causal regularization constraint, and generating local model updating parameters; performing gradient causal desensitization processing and safe aggregation on the local model updating parameters; and performing causal consistency whole-process safe verification and traceable version management on the global updating model. Through the unbiased processing and quality control of clinical semantic drift sensing on the emergency multi-modal data stream, in combination with the causal weight identification, dynamic unbiased sampling and multi-level quality filtering, the objectivity and effectiveness of the training data are effectively guaranteed, and the diagnosis and treatment decision accuracy, generalization ability and safety and reliability of the emergency model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unbiased data stream technology, and more particularly to an unbiased data stream and a secure update method for continuous learning of emergency models. Background Technology

[0002] Emergency medicine is one of the most challenging areas in the healthcare system, characterized by its data's real-time nature, dynamic distribution, multimodal heterogeneity, and high privacy sensitivity. Traditional statically trained AI models experience performance degradation after deployment due to factors such as changes in disease patterns (e.g., seasonal influenza, outbreaks of new infectious diseases), updates to medical equipment, and adjustments to treatment guidelines—a phenomenon known as "model aging." Continuous learning (CL) technology allows models to continuously absorb new knowledge and retain old knowledge without retraining on the entire dataset, representing a key approach to addressing the long-term effectiveness of emergency medicine models.

[0003] Existing emergency data streams suffer from distribution drift, data imbalance, and inconsistent quality. Directly using them for model updates can lead to accumulated biases and performance degradation. Summary of the Invention

[0004] This invention provides an unbiased data stream and a secure update method for continuous learning of emergency models, which addresses the shortcomings of existing emergency data streams, such as distribution drift, data imbalance, and inconsistent quality.

[0005] On the one hand, this invention provides an unbiased data stream and secure update method for continuous learning of emergency models, including:

[0006] S1: Collect multimodal raw data streams in emergency scenarios, synchronously record corresponding clinical conditions, treatment decision chains and timestamp information, and construct raw training data streams containing causal association labels;

[0007] S2: Perform multi-dimensional unbiased processing and quality control on the original training data stream to detect clinical semantic drift, remove biased and low-quality data, and generate an unbiased quality control data stream with causal weight labels;

[0008] S3: Based on unbiased quality control data flow, perform incremental model training anchored to clinical knowledge, retain historical diagnosis and treatment knowledge through causal regularization constraints, and generate local model update parameters;

[0009] S4: Perform gradient causal desensitization and secure aggregation on the local model update parameters to generate a global update model that meets medical privacy compliance requirements;

[0010] S5: Performs full-process security verification of causal consistency and traceable version management for global update models, and completes the safe deployment of models and automatic rollback of failures.

[0011] According to the unbiased data stream and security update method for continuous learning of emergency models provided by the present invention, step S2, the step of generating an unbiased quality control data stream with causal weight labels, includes:

[0012] S21: Standardize and fuse causal features of multi-source heterogeneous emergency data to generate structured causal feature data in a unified format;

[0013] S22: Real-time detection of clinical semantic drift in the data stream, distinguishing between data distribution drift and diagnostic logic drift, performing adaptive causal correction, and obtaining the corrected data stream;

[0014] S23: A cost-aware dynamic unbiased sampling strategy is used to sample the corrected data stream, balancing data distribution, time weight, and clinical risk weight.

[0015] S24: Perform multi-level quality filtering and anomaly detection on the sampled data, retain valid data and label anomalies, forming an unbiased quality control data stream with unique identifiers and causal weights.

[0016] According to the unbiased data stream and secure update method for continuous learning of emergency models provided by the present invention, step S21, the step of generating structured causal feature data, includes:

[0017] S211: Establish a medical terminology mapping system by adopting a pre-defined standard and unified structured data format;

[0018] S212: Based on the medical terminology mapping system, a medical-specific pre-trained model is used to extract unstructured text features from electronic medical records and test reports, while identifying causal paths for diagnosis and treatment decisions.

[0019] S213: Based on the causal path of diagnosis and treatment decision, standardized preprocessing of vital sign time series data and medical image data is performed, different modal features are mapped to a unified causal feature space, and a causal dependency graph between features is constructed to obtain structured causal feature data.

[0020] According to the unbiased data stream and secure update method for continuous learning of emergency models provided by the present invention, step S22, the step of performing adaptive causal correction includes:

[0021] S221: Monitor concept drift from three dimensions: data distribution, model output, and clinical causal consistency. Trigger a drift alarm when the causal consistency deviation exceeds a preset threshold.

[0022] S222: Distinguish between sudden drift, gradual drift, and periodic drift, further subdividing them into data distribution drift and diagnostic logic drift, and determining the start time, scope of impact, and causal root cause of the drift;

[0023] S223: A sliding window update strategy is adopted to address data distribution drift; a correction strategy combining causal graph reconstruction and model fine-tuning is adopted to address diagnostic logic drift; and a seasonal model switching strategy is adopted to address periodic drift.

[0024] According to the unbiased data stream and security update method for continuous learning of emergency models provided by the present invention, step S23, which involves sampling the corrected data stream using a cost-aware dynamic unbiased sampling strategy, includes:

[0025] S231: Calculate the predictive uncertainty and clinical risk level of the sample to increase the sampling weight of rare disease samples;

[0026] S232: Real-time statistical analysis of the distribution of various disease samples, using causal conditional generative enhancement for rare disease samples, and representative undersampling for common disease samples;

[0027] S233: An exponential decay function is used to assign dynamic weights to historical data, while the sampling weights of rare disease samples are adjusted based on the clinical knowledge importance score.

[0028] According to the unbiased data stream and secure update method for continuous learning of emergency models provided by the present invention, step S3, the step of generating local model update parameters, includes:

[0029] S31: Construct a set of clinical knowledge anchors, including core diagnostic and treatment rules, rare disease diagnostic criteria and key causal relationships, and calculate the importance index of model parameters to the anchor set;

[0030] S32: Based on importance indicators, the modification of key parameters of the clinical knowledge anchor set by regularization penalty is reinforced by causal elasticity weights;

[0031] S33: Use a conditional diffusion model to generate pseudo-samples of historical diseases, and perform incremental training in conjunction with new samples. At the same time, maintain a core memory bank, use a prototype-based update strategy to store the most representative historical samples, and generate local model update parameters.

[0032] According to the unbiased data flow and secure update method for continuous learning of emergency models provided by the present invention, step S31, which calculates the importance index of model parameters to the anchor set, includes:

[0033] S311: For each core diagnosis and treatment rule, rare disease diagnosis standard and key causal relationship in the clinical knowledge anchor set, construct a corresponding causal verification subset;

[0034] S312: Using the parameter perturbation method, small Gaussian perturbations are applied to the parameters of each layer of the model in sequence. The changes in the prediction accuracy, causal path matching degree and key node recognition rate of the model before and after the perturbation are calculated on the corresponding causal validation subset, and the performance changes caused by parameter perturbation are obtained.

[0035] S313: Through causal mediation effect analysis, quantify the causal contribution of each model parameter to the activation intensity of key causal nodes in the anchor point set and the efficiency of causal information flow transmission.

[0036] S314: The performance change caused by parameter perturbation is weighted and summed with the causal contribution to obtain a comprehensive importance index of each model parameter relative to the clinical knowledge anchor set.

[0037] According to the unbiased data stream and secure update method for continuous learning of emergency models provided by the present invention, step S33, the step of generating local model update parameters, includes:

[0038] S331: Extract prototype feature vectors of various historical diseases from the core memory bank. For each disease category, retain multiple most representative prototypes. Select features of the prototypes from the causal feature space of historical samples through a clustering algorithm to obtain prototype feature vectors.

[0039] S332: Input the prototype feature vector as a class condition into the pre-trained class conditional diffusion model to generate historical disease pseudo samples that are consistent with the distribution of real clinical data.

[0040] S333: Perform causal consistency verification on historical disease pseudo-samples, calculate the matching degree between the causal path of the historical disease pseudo-samples and the gold standard causal path, remove pseudo-samples with matching degree lower than a preset threshold, and retain high-quality pseudo-samples.

[0041] S334: Mix high-quality pseudo-samples with new samples in the unbiased quality control data stream at a ratio of 1:2 to construct an incremental training dataset, and use the mini-batch gradient descent algorithm for incremental training.

[0042] S335: Calculate the cosine similarity between the new sample and the existing prototypes in the memory bank. Add the new sample with a cosine similarity lower than the preset threshold as the new prototype to the memory bank. At the same time, remove the oldest prototype with the highest similarity in the memory bank to keep the memory bank capacity constant.

[0043] S336: Extract the parameter differences between the trained model and the original model, and combine them with the constraints of the core protected parameter set to perform pruning and normalization on the differences to generate the final local model update parameters.

[0044] According to the unbiased data flow and secure update method for continuous learning of emergency models provided by the present invention, step S4, the step of generating the global update model, includes:

[0045] S41: Employs a horizontal federated learning architecture, allowing medical institutions to complete model training locally;

[0046] S42: Add adaptive causal Gaussian noise to the model gradient to apply high-intensity noise to the causal feature gradients related to patient privacy;

[0047] S43: Use homomorphic encryption algorithm to encrypt model parameters and generate a globally updated model through secure multi-party computation based on causal weights.

[0048] According to the unbiased data flow and secure update method for continuous learning of emergency models provided by the present invention, step S5, the step of performing a full-process security verification of causal consistency of the global update model, includes:

[0049] S51: Construct a multi-dimensional causal verification benchmark set, and label the samples of each causal verification benchmark set with standard diagnosis and treatment decisions, core causal nodes and complete causal reasoning paths;

[0050] S52: Calculate the accuracy, sensitivity, specificity, missed diagnosis rate, and false diagnosis rate of the global update model for various diseases, and evaluate the negative prediction values ​​for acute and critical illnesses and rare diseases; perform incremental knowledge retention validation and test the performance degradation of the global update model on the historical disease validation set.

[0051] S54: Extract the activated causal path and key decision nodes when making diagnosis and treatment decisions in the output of the global update model through the causal attribution algorithm, and perform a structured comparison with the standard causal path in the causal verification benchmark set to calculate the causal path matching degree, the accuracy of key causal nodes and the false causal association rate.

[0052] S55: Verify the stability of causal inference of the global update model under different clinical conditions, different data missing rates, and different time windows. When the global causal consistency score is lower than the preset safety threshold, the model is directly judged to fail the verification.

[0053] S56: Verify the effectiveness of gradient causal desensitization and homomorphic encryption mechanisms through member inference attack, model inversion attack and attribute inference attack tests;

[0054] S57: Build a simulated diagnosis and treatment chain consistent with the real emergency environment. Input test data containing complex scenarios including multiple concurrent patients, asynchronous arrival of multimodal data, and emergency rescue conditions to verify the causal consistency and response timeliness of the entire process from data acquisition, preprocessing, model inference to decision output.

[0055] This invention provides an unbiased data stream and secure update method for continuous learning of emergency models. By implementing unbiased processing and quality control of the multimodal emergency data stream with clinical semantic drift awareness, combined with causal weight labeling, dynamic unbiased sampling, and multi-level quality filtering, the objectivity and validity of training data are effectively guaranteed. Incremental training anchored to clinical knowledge, pseudo-sample generation using a conditional diffusion model, and maintenance of the core memory bank enable efficient model updates and prevent the forgetting of historical diagnostic and treatment knowledge. Utilizing gradient causal desensitization, homomorphic encryption, and a lateral federated learning architecture, global model secure aggregation is achieved while strictly ensuring medical privacy compliance. Through multi-dimensional causal consistency verification and simulated emergency scenario testing, the model is ensured to possess excellent clinical adaptability, inference stability, and response timeliness. Ultimately, this significantly improves the accuracy, generalization ability, and security of emergency model diagnostic and treatment decisions, providing efficient, accurate, and secure technical support for emergency clinical diagnosis and treatment, helping to optimize emergency treatment processes and reduce the risk of missed or misdiagnosed diagnoses. Attached Figure Description

[0056] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0057] Figure 1 This is a flowchart of the unbiased data flow and secure update method for continuous learning of emergency models provided in this embodiment of the invention;

[0058] Figure 2 This is a flowchart of generating an unbiased quality control data stream with causal weight identifiers in an embodiment of the present invention;

[0059] Figure 3 This is a flowchart of the full-process security verification of causal consistency of the global update model in an embodiment of the present invention. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0061] Example:

[0062] The following is combined with Figures 1-3This invention describes an unbiased data flow and secure update method for continuous learning of emergency model.

[0063] like Figures 1-3 As shown in the embodiment of the present invention, the unbiased data stream and secure update method for continuous learning of emergency models includes:

[0064] S1: Collect multimodal raw data streams in emergency scenarios, synchronously record corresponding clinical conditions, treatment decision chains, and timestamp information, and construct a raw training data stream containing causal association labels. The multimodal raw data stream includes structured data (vital signs, laboratory indicators, medication records), semi-structured data (electronic medical record template fields), and unstructured data (medical records, medical images, ECG waveforms). Clinical conditions are divided into four categories: general emergency care, critical care resuscitation, mass casualty treatment, and nighttime emergency care. The treatment decision chain is recorded in the form of a directed acyclic graph, containing six core nodes: symptom collection → preliminary diagnosis → auxiliary examinations → confirmed diagnosis → treatment plan → efficacy evaluation. Causal association labels are calculated using the following formula:

[0065]

[0066] Among them, L causal (x i ,y i ) represents the causal relationship label, x i For the i-th sample, y i Here, m represents the corresponding diagnostic and treatment decision label, and w represents the number of causal features. j f represents the clinical importance weight of the j-th causal feature. j (·) is the feature mapping function for the j-th causal feature, and I(·) is the indicator function. When feature f j Is it a decision y i The value is 1 if the direct cause is a given, and 0 otherwise. The timestamp uses the UTC standard format with millisecond precision to ensure the time synchronization of multimodal data.

[0067] S2: Perform multi-dimensional unbiased processing and quality control on the original training data stream to detect clinical semantic drift, remove biased and low-quality data, and generate an unbiased quality control data stream with causal weight labels.

[0068] Step S2, which involves generating an unbiased quality control data stream with causal weights, includes:

[0069] S21: Standardize and fuse causal features of multi-source heterogeneous emergency data to generate structured causal feature data in a unified format.

[0070] Step S21, the step of generating structured causal feature data, includes:

[0071] S211: A medical terminology mapping system is established using a pre-defined standard unified structured data format. The pre-defined standards include the HL7FHIRR4 data exchange standard, the SNOMEDCT clinical terminology standard, the ICD-10 disease coding standard, and the LOINC test item standard. The medical terminology mapping system adopts a bidirectional mapping structure, mapping free text terms from the original electronic medical records to the standard terminology set, while retaining the original terms as supplementary information. Term similarity calculation uses a weighted fusion method of cosine similarity and edit distance, expressed by the formula:

[0072]

[0073]

[0074] Among them, t1 and t2 are two medical terms. The overall similarity score for two medical terms is calculated as follows: Emb(·) is the term embedding vector generated by the medical pre-training model, EditDist(·) is the Levinstein edit distance, and α is the fusion weight coefficient, ranging from [0.6, 0.8]. CosSim(·) is the cosine similarity calculation function, where Emb(·) represents the term embedding vector generated by the medical pre-training model, and EditDist(·) represents the Levinstein edit distance, characterizing the degree of difference between the term texts. len(·) is the character length of the medical term text, and max(·) is the maximum of the two term text lengths. Terms with a similarity score greater than or equal to 0.85 are considered to be the same term.

[0075] S212: Based on a medical terminology mapping system, a medical-specific pre-trained model is used to extract unstructured text features from electronic medical records and laboratory reports, while simultaneously identifying causal paths in diagnosis and treatment decisions. The medical-specific pre-trained model uses the MedBERT-based model, taking the text sequence after terminology mapping as input and outputting a 768-dimensional text feature vector. The causal path identification for diagnosis and treatment decisions employs a causal extraction method based on an attention mechanism. First, the causal attention score for each entity pair in the text is calculated:

[0076]

[0077]

[0078] Among them, A(e) i ,e j ) for medical entity e i For e j The causal attention normalization score, e i e j Let be two medical entities in the text, exp(·) be the natural exponential function, and score(e) be the medical entity in the text. i ,ej ) represents the original association score for the entity pair, h i h j Let be the embedding vector corresponding to the entity, W be the trainable weight matrix, b be the bias term, and n be the total number of entities in the text. T is the transpose operator. When the causal attention score is greater than or equal to 0.6, e is determined. i It is e j The reasons are used to construct a causal path for diagnosis and treatment decisions.

[0079] S213: Based on the causal path of diagnosis and treatment decisions, standardized preprocessing is performed on vital sign time-series data and medical image data to map different modal features to a unified causal feature space, construct a causal dependency graph between features, and obtain structured causal feature data. Standardized preprocessing of vital sign time-series data includes outlier removal (using the 3σ criterion), missing value imputation (using a combination of linear interpolation and causal interpolation; for missing values ​​at key nodes in the causal path, a causal regression model is used for prediction and imputation), and normalization (using Z-score standardization). Standardized preprocessing of medical image data includes size normalization (uniformly adjusted to 224×224 pixels), grayscale normalization, and noise reduction (using Gaussian filtering). The unified causal feature space mapping uses a cross-modal attention fusion network to map text features, time-series features, and image features to a unified feature space of dimension 512. The causal dependency graph is constructed as G=(V,E), where the vertex set V represents all causal features, the edge set E represents the causal relationships between features, and the edge weights represent the causal strength, calculated using the Pearson causality coefficient, expressed by the formula:

[0080]

[0081] in, Let E represent the expected value of the random variable, and let X and Y be two causal characteristic variables, with μ as the causal strength. X μ Y Let σ be the mean of the variable. X σ Y Let be the standard deviation of the variable. When |ρ xy When |>0.3 and passes the significance test (p<0.05), add an edge between the two features, with the edge weight being |ρ. xy |

[0082] S22: Real-time detection of clinical semantic drift in the data stream, distinguishing between data distribution drift and diagnostic logic drift, performing adaptive causal correction, and obtaining the corrected data stream.

[0083] In step S22, the steps for performing adaptive causal correction include:

[0084] S221: Concept drift is monitored across three dimensions: data distribution, model output, and clinical causal consistency. A drift alarm is triggered when the causal consistency deviation exceeds a preset threshold. Data distribution drift monitoring uses a combination of the KS test and the AD test to calculate the distribution difference between the current data window and the baseline data window. Model output drift monitoring uses a sliding window statistical method for prediction accuracy; an alert is triggered when the accuracy decreases by more than 10%. Clinical causal consistency deviation D... causal The formula is expressed as:

[0085]

[0086] Where N is the number of validation samples, P i P represents the causal path output by the model for the i-th sample. std Let |·| be the standard causal path corresponding to this sample, and |·| be the cardinality of the set. A preset threshold θ is used. drift The value range is [0.2, 0.3], and in this embodiment, it is taken as 0.25. When D causal >θ drift A drift alarm is triggered at that time.

[0087] S222: Differentiate between sudden drift, gradual drift, and periodic drift, further subdividing them into data distribution drift and diagnostic logic drift, determining the start time, scope of impact, and causal root cause of the drift. The drift type determination method is as follows:

[0088] Calculate the rate of change of drift within a continuous time window , where ΔD causal Δt represents the change in causal consistency deviation between adjacent time windows, where Δt is the time interval between adjacent time windows. A sudden drift is defined as r > 0.5 / h, a gradual drift is defined as 0.05 / h ≤ r ≤ 0.5 / h, and a periodic drift is defined as the drift exhibiting periodic fluctuations with a period greater than 7 days.

[0089] If a drift is detected only in the data distribution dimension but not in the clinical causal consistency dimension, then it is a data distribution drift.

[0090] If a drift is detected in the clinical causal consistency dimension, it is considered a drift in the diagnostic and treatment logic.

[0091] The drift start time is determined by a sliding window mutation detection algorithm (such as the PELT algorithm).

[0092] The scope of impact is determined by calculating the degree of influence of drift on different disease categories and different clinical conditions.

[0093] The causal root cause analysis uses the intervention method of causal graphs to locate the key feature nodes that lead to a decline in causal consistency.

[0094] S223: A sliding window update strategy is used to address data distribution drift; a correction strategy combining causal graph reconstruction and model fine-tuning is used to address diagnostic logic drift; and a seasonal model switching strategy is used to address periodic drift. The sliding window update strategy uses an adaptive window size, which is adjusted according to the degree of drift. Dynamic adjustments are made, where W0 is the baseline window size, typically 1000 samples. Causal dependencies between features are recalculated based on the drifted data stream, updating the edges and weights of the causal graph while retaining core causal edges consistent with the clinical knowledge anchor set. Model fine-tuning uses a small learning rate (1e−5) to update only the parameters of the last two layers of the model, with 5-10 fine-tuning rounds. A seasonal model switching strategy pre-trains sub-models corresponding to different seasons. When a periodic drift is detected and the current season matches a sub-model, the model automatically switches to the corresponding sub-model while retaining the core parameters of the global model.

[0095] S23: Use a cost-aware, dynamic, unbiased sampling strategy to sample the corrected data stream, balancing data distribution, time weights, and clinical risk weights.

[0096] Step S23, which involves sampling the corrected data stream using a cost-aware dynamic unbiased sampling strategy, includes:

[0097] S231: Calculate the prediction uncertainty and clinical risk level of the sample to increase the sampling weight of rare disease samples. The prediction uncertainty of the sample is calculated using the Monte Carlo dropout method. T prediction results are obtained through T forward propagations, and the entropy of the prediction results is calculated as the uncertainty, expressed by the formula:

[0098]

[0099] in, The entropy of the prediction result is the prediction uncertainty of a single emergency room sample x. K is the number of treatment decision categories, and p k This represents the average predicted probability for the k-th category. Clinical risk levels are categorized into four levels: low, medium, high, and very high, based on the patient's vital signs, disease type, and complications. A scoring system is used: a total score of 0-10 points, with 0-3 indicating low risk, 4-6 indicating medium risk, 7-8 indicating high risk, and 9-10 indicating very high risk. Rare diseases are defined as those with an incidence rate of less than 1 / 5000. The baseline sampling weight for rare disease samples is set to 10 times that of common disease samples.

[0100] S232: Real-time statistical analysis of the distribution of various disease samples. Causal conditional generative augmentation is used for rare disease samples, while representative undersampling is performed on common disease samples. Real-time statistics employ a sliding window method with a window size of 10,000 samples, updating the distribution statistics every 1,000 samples. Causal conditional generative augmentation uses a causal conditional variational autoencoder to generate rare disease samples based on disease type and key causal features. The number of generated samples is determined by the ratio of rare to common disease samples, aiming to increase the proportion of rare disease samples to 5%-10% of the total sample size. Representative undersampling uses a clustering-based method, performing K-means clustering on common disease samples. A certain number of samples are retained at each cluster center, with the retention number inversely proportional to the cluster size, thus ensuring the diversity of common disease samples.

[0101] S233: An exponential decay function is used to assign dynamic weights to historical data, while the sampling weights for rare disease samples are adjusted based on clinical knowledge importance scores. The dynamic weights for historical data are calculated using an exponential decay function, expressed by the following formula:

[0102]

[0103] in, Let t be the time-dynamic weight of the i-th historical sample. i t represents the sample collection time. current The current time is λ, which is the decay coefficient, ranging from [0.001, 0.01] / day. The clinical knowledge importance score is assessed by domain experts based on the severity of the disease, the difficulty of diagnosis and treatment, and its clinical value, with a score range of 1-10. The final sampling weight for rare disease samples... The formula is expressed as:

[0104]

[0105] Among them, w base Based on the sampling weights, w t For time weighting, S clinical Assess the importance of clinical knowledge.

[0106] S24: Perform multi-level quality filtering and anomaly detection on the sampled data, retaining valid data and annotating anomalies to form an unbiased quality control data stream with unique identifiers and causal weights. Multi-level quality filtering includes: Level 1 integrity filtering, removing samples missing more than 30% of key causal features; Level 2 logicality filtering, removing samples with obvious clinical logical contradictions (e.g., diagnosed with hypertension but with normal blood pressure); Level 3 consistency filtering, removing samples inconsistent with standard clinical guidelines. Anomaly detection uses the Isolation Forest algorithm, identifying anomalous samples with anomaly scores greater than 0.7. Unique identifiers are generated using UUIDv4 format, with each sample corresponding to a globally unique identifier. Causal weights. The calculation is expressed as:

[0107]

[0108] Where U(x) represents the prediction uncertainty of the sample, and R(x) represents the clinical risk level score of the sample. Causal weights The value range is (0,1), and the higher the weight, the greater the value of the sample for model training.

[0109] S3: Based on unbiased quality control data flow, it performs incremental model training anchored to clinical knowledge, retains historical diagnosis and treatment knowledge through causal regularization constraints, and generates local model update parameters.

[0110] Step S3, the steps for generating local model update parameters, include:

[0111] S31: Construct a set of clinical knowledge anchors, including core diagnostic and treatment rules, diagnostic criteria for rare diseases, and key causal relationships, and calculate the importance of model parameters to the anchor set.

[0112] Step S31, which involves calculating the importance index of the model parameters to the anchor set, includes:

[0113] S311: For each core clinical practice rule, rare disease diagnostic standard, and key causal relationship in the clinical knowledge anchor set, a corresponding causal validation subset is constructed. The core clinical practice rules are derived from the *Emergency Medicine Guidelines* and the *Clinical Practice Guidelines*, selecting 100 of the most commonly used core rules. The rare disease diagnostic standards are derived from the *Rare Disease Diagnosis and Treatment Guidelines (2019 Edition)*, containing diagnostic standards for 121 rare diseases. Key causal relationships are extracted from the clinical causal graph by domain experts, selecting 200 key causal edges. Each causal validation subset contains 50-100 samples that match the corresponding anchor point; samples are randomly selected from historical data to ensure the diversity and representativeness of the subset.

[0114] S312: Using the parameter perturbation method, small Gaussian perturbations are applied to the parameters of each layer of the model sequentially. The changes in prediction accuracy, causal path matching degree, and key node recognition rate of the model before and after the perturbation on the corresponding causal validation subset are calculated to obtain the performance change caused by the parameter perturbation. For the l-th layer parameter θ of the model... l Apply Gaussian perturbation Δθ l ~N(0,σ 2 I), where σ is the disturbance intensity, ranging from [0.001, 0.01], and N(·) is a Gaussian normal distribution. The formula for calculating the prediction accuracy Acc is expressed as:

[0115]

[0116] Where TP represents a true positive, TN a true negative, FP a false positive, and FN a false negative. The causal path matching degree PM is calculated in the same way as the path matching degree in S221. The formula for calculating the key node identification rate KR is as follows:

[0117]

[0118] Among them, TP node To correctly identify the number of critical nodes, FN node This represents the number of unidentified critical nodes. The formula for the performance change caused by parameter perturbation is as follows:

[0119]

[0120] in, The parameter denoted by Acc0 represents the overall performance degradation after the perturbation of the parameters of the l-th layer of the model. Acc0, PM0, and KR0 are the performance indicators before the perturbation. l PM l KR l The perturbation performance index is represented by α, β, and γ, which are weighting coefficients. In this embodiment, they are taken as 0.3, 0.4, and 0.3, respectively.

[0121] S313: Through causal mediation effect analysis, quantify the causal contribution of each model parameter to the activation intensity of key causal nodes and the efficiency of causal information flow transmission in the anchor point set. The causal mediation effect analysis adopts the Baron-Kenny method, using the model parameter θ as the independent variable, the activation intensity A of the key causal nodes as the mediating variable, and the final output Y of the model as the dependent variable. The activation intensity of the key causal node is the output value of the corresponding neuron. The efficiency of causal information flow transmission is expressed as the change in information entropy:

[0122]

[0123] Where IE represents the information flow transmission efficiency between causal nodes, and H(·) represents the information entropy. The final diagnostic decision output of the model is represented by A, where A is the neuron activation intensity of the key causal node, do(·) is the intervention operation, a0 is the normal activation value of the node, and a1 is the activation value of the node after intervention. The causal contribution of parameter θ to the key causal node is:

[0124]

[0125] in, β represents the causal contribution of model parameter θ to the clinical causal node. Aθ Let θ be the regression coefficient of the parameter θ on the node activation intensity A, and β be the regression coefficient of θ on the node activation intensity A. YA Let β be the regression coefficient of node activation intensity A on model output Y. Yθ ′ represents the direct regression coefficient of parameter θ after controlling for A on the model output Y.

[0126] S314: The performance changes caused by parameter perturbations are weighted and summed with the causal contribution to obtain a comprehensive importance index for each model parameter relative to the clinical knowledge anchor set. The formula is as follows:

[0127]

[0128] in, Let θ be the overall importance index of the model parameter relative to the clinical knowledge anchor set, ΔP(θ) be the performance change caused by parameter perturbation, C(θ) be the causal contribution of the parameter, and ω1 and ω2 be weighting coefficients, which are set to 0.5 and 0.5 respectively in this embodiment. A comprehensive importance index greater than or equal to the threshold θ is considered. imp The parameters are determined as core protection parameters, and the threshold θ imp The value is taken as the 75th percentile of the importance index of all parameters.

[0129] S32: Based on importance metrics, the key parameters of the clinical knowledge anchor set are modified using Causal Elastic Weight Consolidation (CEWC) regularization penalties. The loss function for CEWC regularization is:

[0130]

[0131] Among them, L total L is the total loss function for incremental training of the model. task Let λ be the cross-entropy loss function for the current task, and λ be the regularization coefficient, with a value range of

[10] . 3 10 4 ], n1 is the total number of model parameters, θ i Let θ be the i-th parameter of the current model.i ∗ Let I(θ) be the i-th parameter of the original model. i Let be the overall importance index of the i-th parameter. This regularization term limits the changes of parameters with high importance during incremental training by imposing a greater penalty on them, thereby preserving historical diagnostic knowledge.

[0132] S33: Use a conditional diffusion model to generate pseudo-samples of historical diseases, and perform incremental training together with new samples. At the same time, maintain a core memory bank, use a prototype-based update strategy to store the most representative historical samples, and generate local model update parameters.

[0133] Step S33, the step of generating local model update parameters includes:

[0134] S331: Extract prototype feature vectors for various historical diseases from the core memory. For each disease category, retain multiple most representative prototypes. Use a clustering algorithm to select features from the causal feature space of historical samples to obtain prototype feature vectors. The clustering algorithm uses K-means++. For each disease category, the number of clusters k is determined based on the number of samples in that category: k=5 when the number of samples is less than 100; k=10 when the number of samples is between 100 and 1000; and k=20 when the number of samples is greater than 1000. The prototype feature vector is the center vector of each cluster, calculated using the following formula:

[0135]

[0136] in, Let C be the prototype feature vector of the j-th cluster. j For the j-th cluster, |C j | represents the number of samples in the cluster, and f(x) is the causal feature vector of sample x. The total capacity of the core memory is 10,000 prototype feature vectors.

[0137] S332: The prototype feature vector is used as class-conditional input to a pre-trained class-conditional diffusion model to generate historical disease pseudo-samples consistent with the distribution of real clinical data. The class-conditional diffusion model adopts the DDPM architecture, which includes an encoder, decoder, and time-step embedding module. A class-conditional embedding module is also introduced, using the prototype feature vector as conditional input. The forward propagation formula for the diffusion process is:

[0138]

[0139]

[0140] Where, x t Let x be the noise sample at step t. t−1 For the previous sample, αt Let α be the diffusion scheduling coefficient. t =1−β t . β t For noise scheduling, the value range is [0.0001, 0.02]. t For standard Gaussian noise, the backpropagation process predicts the noise ϵ using a neural network. θ (x t The process involves generating pseudo-samples by progressively denoising the vectors ,t,c, where c is the class conditional embedding vector. The formula is as follows:

[0141]

[0142]

[0143]

[0144]

[0145] in, Here, c represents the step-by-step noise predicted by the neural network, and c is the conditional embedding vector for the disease class. The cumulative product of the coefficients α over the first t steps. Let z be the standard deviation of the reverse process noise, z be the reverse-added Gaussian noise, and s be the intermediate time step of the diffusion process.

[0146] S333: Perform causal consistency verification on historical disease pseudo-samples, calculate the matching degree between the causal path of the historical disease pseudo-sample and the gold standard causal path, remove pseudo-samples with a matching degree lower than a preset threshold, and retain high-quality pseudo-samples. First, the causal path extraction model is used to extract the causal path of the pseudo-sample, and then it is compared with the gold standard causal path corresponding to the disease to calculate the causal path matching degree. The gold standard causal path is constructed by domain experts based on clinical guidelines. The preset threshold range is [0.6, 0.7], and 0.65 is used in this embodiment. At the same time, the clinical rationality of the pseudo-samples is verified, and pseudo-samples with obvious clinical errors are removed. The proportion of high-quality pseudo-samples retained in the final product should not be less than 70% of the total number of generated pseudo-samples.

[0147] S334: High-quality pseudo-samples are mixed with new samples from the unbiased quality control data stream at a ratio of 1:2 to construct an incremental training dataset. Incremental training is performed using the mini-batch gradient descent algorithm. The parameters of the mini-batch gradient descent algorithm are set as follows: batch size of 32, learning rate of 1e−4, AdamW optimizer, weight decay coefficient of 1e−5, 20 training epochs, and an early stopping strategy: training stops when the validation set loss does not decrease for 5 consecutive epochs. The loss function is the CEWC regularized loss function from step S32. During training, dynamic data augmentation techniques are used: synonym replacement, random insertion, and random deletion are performed on text data; time warping and amplitude scaling are performed on time-series data; and random rotation, flipping, and cropping are performed on image data.

[0148] S335: Calculate the cosine similarity between the new sample and the existing prototypes in the memory bank. Add the new sample with a cosine similarity lower than a preset threshold as a new prototype to the memory bank, while removing the oldest prototype with the highest similarity from the memory bank to maintain a constant memory bank size. The formula for calculating the cosine similarity is:

[0149]

[0150] in, Let be the cosine similarity between feature vectors x and y, where x is the causal feature vector of the new sample and y is the prototype feature vector in the memory. A preset threshold θ is used. sim The value range is [0.7, 0.8]. ||·|| represents the L2 norm of the vector, where the cosine similarity between the new sample and all prototypes in the memory is less than θ. sim When the new sample is full, it is added to the memory as a new prototype. If the memory is full, the oldest prototype with the highest similarity to the new prototype is removed from the memory, thus ensuring that the memory always stores the most representative historical samples.

[0151] S336: Extract the parameter differences between the trained model and the original model. Combined with the constraints of the core protected parameter set, prune and normalize the differences to generate the final local model update parameters. Among these, the parameter differences... The calculation formula is:

[0152]

[0153] Where, where θ new θ represents the parameters of the trained model. old These are the parameters of the original model. The core protection parameter set consists of comprehensive importance indices determined in step S314 that are greater than or equal to θ. imp The parameter set. The parameter differences are pruned to limit the difference between core protection parameters to no more than a threshold θ. clip θ clipThe value range is [0.001, 0.01]. L2 normalization is used, and the formula is as follows:

[0154]

[0155] in, The normalized parameter difference. Let be the L2 norm of the parameter difference, and γ be the normalization coefficient, ranging from [0.1, 0.5]. The final local model update parameter is the difference Δθ between the clipped and normalized parameters. final .

[0156] S4: Perform gradient causal desensitization and secure aggregation on the local model update parameters to generate a global update model that meets medical privacy compliance requirements.

[0157] Step S4, the steps for generating the global update model include:

[0158] S41: Employs a horizontal federated learning architecture, where healthcare institutions train their models locally. This architecture consists of a central server and multiple client healthcare institutions. The central server is responsible for global model distribution and aggregation, while client institutions are responsible for training their local models and uploading updated parameters. All communication uses the TLS 1.3 encryption protocol to ensure data transmission security. Asynchronous communication is used between the clients and the central server. After completing local training, the clients proactively upload updated parameters, and the central server performs a global aggregation after collecting a certain number of client updates.

[0159] S42: Add adaptive causal Gaussian noise to the model gradient, applying high-intensity noise to the gradient of causal features related to patient privacy. The intensity of the adaptive causal Gaussian noise is determined based on the privacy sensitivity of the causal features corresponding to the gradient, and the formula for calculating the privacy sensitivity S(g) is as follows:

[0160]

[0161] Where g is the model gradient, m is the number of causal features, and w i F represents the weight of the i-th causal feature. private This represents a set of privacy-sensitive causal features (including patient personal information, medical history, genetic information, etc.). I(⋅) is an indicator function; it is 1 if the feature belongs to the privacy set, and 0 otherwise. Noise intensity σ g It is directly proportional to the privacy sensitivity S(g), expressed as:

[0162]

[0163] Where σ0 is the basic noise intensity, with a value range of [0.01, 0.1]. The gradient after adding noise is:

[0164]

[0165]

[0166] in, The gradient after adding privacy-preserving noise. The variable is Gaussian noise. Let I be the noise variance, and I be the identity matrix.

[0167] S43: The model parameters are encrypted using a homomorphic encryption algorithm. A globally updated model is generated through secure multi-party computation based on causal weights. The homomorphic encryption algorithm uses the CKKS algorithm, which supports homomorphic operations on floating-point numbers. The encryption keys are generated by each client; the public key is uploaded to the central server, and the private key is stored locally on the client. The secure multi-party computation based on causal weights uses an improved version of the federated averaging algorithm. The aggregation formula for the global model update parameters is:

[0168]

[0169] in, Update parameters for the aggregated global model, where K is the number of clients participating in the aggregation, and n k w represents the number of local training samples for the k-th client. k Let Δθ be the causal weight for the k-th client. k Update parameters for the encrypted local model uploaded by the k-th client. Causal weight w k The value was determined based on the causal consistency score and clinical value of the client data, with a range of [0.5, 1.5].

[0170] S5: Performs full-process security verification of causal consistency and traceable version management for global update models, and completes the safe deployment of models and automatic rollback of failures.

[0171] Step S5, which involves performing a full-process security verification of causal consistency for the global update model, includes the following steps:

[0172] S51: Construct a multi-dimensional causal validation benchmark set, and annotate the samples in each benchmark set with standard treatment decisions, core causal nodes, and complete causal reasoning paths. The multi-dimensional causal validation benchmark set includes validation sets for general diseases, acute and critical illnesses, rare diseases, historical knowledge retention, and adversarial examples. All samples were annotated by at least three experts at the associate chief physician level or above, and the annotation results underwent cross-validation, achieving an annotation consistency rate of over 95%.

[0173] S52: Calculate the accuracy, sensitivity, specificity, false negative rate, and false negative rate of the global update model for various diseases, and evaluate the negative prediction values ​​for acute, critical, and rare diseases. Perform incremental knowledge retention validation to test the performance degradation of the global update model on the historical disease validation set. The formulas for calculating each performance indicator are as follows:

[0174]

[0175]

[0176]

[0177]

[0178]

[0179]

[0180] Where Acc represents accuracy, Sen represents sensitivity, Spe represents specificity, MR represents the false negative rate, ER represents the false positive rate, NPV represents the negative predictive value, P represents true positive, TN represents true negative, FP represents false positive, and FN represents false negative. For acute and critical illnesses, the sensitivity requirement is no less than 95%, and the false negative rate requirement is no more than 5%. For rare diseases, the negative predictive value requirement is no less than 90%. The performance degradation of incremental knowledge preservation validation is also considered. The calculation formula is:

[0181]

[0182] Among them, Acc old Acc represents the accuracy of the original model on the historical disease validation set. new This is used to globally update the model's accuracy on the historical disease validation set. The performance degradation should not exceed 5%.

[0183] S54: Extract the activated causal paths and key decision nodes from the global updated model output during diagnosis and treatment decisions using a causal attribution algorithm. Perform a structured comparison with the standard causal paths in the causal validation benchmark set, and calculate the causal path matching degree, key causal node accuracy, and spurious causal association rate. The Grad-CAM++ algorithm is used to extract the activated causal paths and key decision nodes from the model. The formula for calculating the causal path matching degree is the same as in step S221. The formula for calculating the key causal node accuracy is:

[0184]

[0185] Where KCA represents the accuracy of the model's inference at key causal nodes, and TP... node To correctly identify the number of critical nodes, FPnode This represents the number of critical nodes that were incorrectly identified.

[0186] The formula for calculating the spurious causal rate (FCR) is as follows:

[0187]

[0188] Among them, FP edge TP represents the number of spurious causal edges output by the model. edge This represents the number of correct causal edges output by the model. The causal path matching accuracy should be no less than 80%, the accuracy of key causal nodes should be no less than 85%, and the false causal association rate should be no higher than 15%.

[0189] S55: Verify the stability of causal inference of the global update model under different clinical conditions, different data missing rates, and different time windows. When the global causal consistency score is lower than the preset safety threshold, the model is directly judged as failing validation. Different clinical conditions include routine emergency care, emergency and critical care, mass casualty treatment, and nighttime emergency care. Different data missing rates include 10%, 20%, and 30%. Different time windows include 1 hour, 6 hours, 12 hours, and 24 hours. The formula for calculating the global causal consistency score (GCS) is expressed as follows:

[0190]

[0191] Where PM represents the causal path matching degree, KCA represents the accuracy of key causal nodes, FCR represents the false causal association rate, and ω1, ω2, and ω3 are weighting coefficients, which are set to 0.4, 0.3, and 0.3 respectively in this embodiment. A preset safety threshold θ is used. safe The value range is [0.7, 0.8]. When GCS < θ safe If the model fails validation, it will be immediately deemed unusable and prohibited from going live.

[0192] S56: The effectiveness of gradient causal desensitization and homomorphic encryption mechanisms is verified through member inference attack, model inversion attack, and attribute inference attack tests. Member inference attack uses the shadow model method to train multiple shadow models to simulate the behavior of the target model, and then uses the attack model to determine whether a sample is in the training set of the target model. Model inversion attack attempts to recover sensitive information of training samples from the output of the model. Attribute inference attack attempts to infer a certain attribute (such as gender or age) of the training sample. The evaluation index of privacy protection effectiveness is the attack success rate, which requires that the success rate of all attacks is not higher than 55%, that is, close to 50% of random guessing. If the attack success rate is higher than 60%, the privacy protection mechanism is deemed invalid and gradient desensitization and parameter encryption need to be performed again.

[0193] S57: Build a simulated treatment process consistent with a real emergency room environment. Input test data from complex scenarios including concurrent multi-patient scenarios, asynchronous arrival of multimodal data, and emergency resuscitation conditions. Verify the causal consistency and response timeliness of the entire process from data acquisition, preprocessing, model inference to decision output. The simulated treatment process includes a data acquisition module, a data preprocessing module, a model inference module, a decision output module, and an evaluation module, simulating the workflow of a real emergency department. The concurrent multi-patient scenario simulates 10-20 patients seeking treatment simultaneously. The asynchronous multimodal data arrival scenario simulates different modalities of data (such as vital signs, laboratory reports, and medical images) arriving at different times. The emergency resuscitation condition simulates emergency situations such as cardiac arrest and shock. The causal consistency of the entire process is verified by comparing the causal path consistency between the model output decision and the expert decision. The evaluation index for response timeliness is end-to-end latency, requiring the average latency from data input to decision output to not exceed 1 second, and the 99th percentile latency to not exceed 2 seconds.

[0194] In summary, the unbiased data stream and secure update method for continuous learning of emergency models effectively ensures the objectivity and validity of training data by implementing unbiased processing and quality control of clinical semantic drift perception on multimodal emergency data streams, combined with causal weight labeling, dynamic unbiased sampling, and multi-level quality filtering. Incremental training anchored to clinical knowledge, pseudo-sample generation using a conditional diffusion model, and maintenance of the core memory bank enable efficient model updates while avoiding the forgetting of historical diagnostic and treatment knowledge. Gradient causal desensitization, homomorphic encryption, and a horizontal federated learning architecture ensure global model security aggregation while strictly guaranteeing medical privacy compliance. Multi-dimensional causal consistency verification and simulated emergency scenario testing ensure the model possesses excellent clinical adaptability, inference stability, and response timeliness. Ultimately, this significantly improves the accuracy, generalization ability, and security of emergency model diagnostic and treatment decisions, providing efficient, accurate, and safe technical support for emergency clinical diagnosis and treatment, helping to optimize emergency treatment processes and reduce the risk of missed or misdiagnosed diagnoses.

[0195] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0196] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An unbiased data flow and safe update method for continuous learning of emergency medicine models, characterized in that, include: S1: Collect multimodal raw data streams in emergency scenarios, synchronously record corresponding clinical conditions, treatment decision chains and timestamp information, and construct raw training data streams containing causal association labels; S2: Perform multi-dimensional unbiased processing and quality control on the original training data stream to detect clinical semantic drift, remove biased data and low-quality data, and generate an unbiased quality control data stream with causal weight labels; S3: Based on the unbiased quality control data stream, perform incremental model training anchored to clinical knowledge, retain historical diagnosis and treatment knowledge through causal regularization constraints, and generate local model update parameters; S4: Perform gradient causal desensitization and secure aggregation on the local model update parameters to generate a global update model that meets medical privacy compliance requirements; S5: Perform causal consistency security verification and traceable version management on the global update model to complete the safe online deployment and automatic rollback of the model in case of failure.

2. The unbiased data flow and secure update method for continuous learning of emergency models according to claim 1, characterized in that, Step S2, which involves generating an unbiased quality control data stream with causal weights, includes: S21: Standardize and fuse causal features of multi-source heterogeneous emergency data to generate structured causal feature data in a unified format; S22: Real-time detection of clinical semantic drift in the data stream, distinguishing between data distribution drift and diagnostic logic drift, performing adaptive causal correction, and obtaining the corrected data stream; S23: The corrected data stream is sampled using a cost-aware dynamic unbiased sampling strategy to balance data distribution, time weight, and clinical risk weight. S24: Perform multi-level quality filtering and anomaly detection on the sampled data, retain valid data and mark anomalies, forming the unbiased quality control data stream with unique identifiers and causal weights.

3. The unbiased data flow and secure update method for continuous learning of emergency models according to claim 2, characterized in that, Step S21, the step of generating structured causal feature data, includes: S211: Establish a medical terminology mapping system by adopting a pre-defined standard and unified structured data format; S212: Based on the aforementioned medical terminology mapping system, a medical-specific pre-trained model is used to extract unstructured text features from electronic medical records and test reports, while simultaneously identifying causal paths for diagnosis and treatment decisions; S213: Based on the causal path of diagnosis and treatment decision, the time series data of vital signs and medical imaging data are standardized and preprocessed, different modal features are mapped to a unified causal feature space, and a causal dependency graph between features is constructed to obtain the structured causal feature data.

4. The unbiased data flow and secure update method for continuous learning of emergency models according to claim 2, characterized in that, In step S22, the steps for performing adaptive causal correction include: S221: Monitor concept drift from three dimensions: data distribution, model output, and clinical causal consistency. Trigger a drift alarm when the causal consistency deviation exceeds a preset threshold. S222: Distinguish between sudden drift, gradual drift, and periodic drift, further subdividing them into data distribution drift and diagnostic logic drift, and determining the start time, scope of impact, and causal root cause of the drift; S223: A sliding window update strategy is adopted for the data distribution drift, a correction strategy combining causal graph reconstruction and model fine-tuning is adopted for the diagnosis and treatment logic drift, and a seasonal model switching strategy is adopted for the periodic drift.

5. The unbiased data flow and secure update method for continuous learning of emergency models according to claim 2, characterized in that, Step S23, which involves sampling the corrected data stream using a cost-aware dynamic unbiased sampling strategy, includes: S231: Calculate the predictive uncertainty and clinical risk level of the sample to increase the sampling weight of rare disease samples; S232: Real-time statistical analysis of the distribution of various disease samples, using causal conditional generative enhancement for rare disease samples, and representative undersampling for common disease samples; S233: An exponential decay function is used to assign dynamic weights to historical data, while the sampling weights of the rare disease samples are adjusted based on the clinical knowledge importance score.

6. The unbiased data flow and secure update method for continuous learning of emergency models according to claim 1, characterized in that, Step S3, the steps for generating local model update parameters, include: S31: Construct a set of clinical knowledge anchors, including core diagnostic and treatment rules, rare disease diagnostic criteria and key causal relationships, and calculate the importance index of model parameters to the anchor set; S32: Based on the importance index, the modification of key parameters of the clinical knowledge anchor set is reinforced by causal elasticity weighting through regularization penalty; S33: Use a conditional diffusion model to generate pseudo-samples of historical diseases, and perform incremental training in conjunction with new samples. At the same time, maintain a core memory bank, use a prototype-based update strategy to store the most representative historical samples, and generate the local model update parameters.

7. The unbiased data flow and secure update method for continuous learning of emergency models according to claim 6, characterized in that, Step S31, which involves calculating the importance index of the model parameters to the anchor point set, includes: S311: For each core diagnosis and treatment rule, rare disease diagnosis standard and key causal relationship in the clinical knowledge anchor set, construct a corresponding causal verification subset; S312: Using the parameter perturbation method, small Gaussian perturbations are applied to the parameters of each layer of the model in sequence. The changes in the prediction accuracy, causal path matching degree and key node recognition rate of the model before and after the perturbation are calculated on the corresponding causal validation subset, and the performance changes caused by parameter perturbation are obtained. S313: Through causal mediation effect analysis, quantify the causal contribution of each model parameter to the activation intensity of key causal nodes in the anchor point set and the efficiency of causal information flow transmission. S314: The performance change caused by the parameter perturbation is weighted and summed with the causal contribution to obtain a comprehensive importance index of each model parameter relative to the clinical knowledge anchor set.

8. The unbiased data flow and secure update method for continuous learning of emergency models according to claim 6, characterized in that, Step S33, the step of generating the local model update parameters includes: S331: Extract prototype feature vectors of various historical diseases from the core memory bank. For each disease category, retain multiple most representative prototypes. Select features of the prototypes from the causal feature space of historical samples through a clustering algorithm to obtain prototype feature vectors. S332: The prototype feature vector is used as a class condition input to a pre-trained class conditional diffusion model to generate historical disease pseudo samples that are consistent with the distribution of real clinical data. S333: Perform causal consistency verification on the historical disease pseudo-samples, calculate the matching degree between the causal path of the historical disease pseudo-samples and the gold standard causal path, remove false samples with matching degree lower than a preset threshold, and retain high-quality pseudo-samples. S334: Mix the high-quality pseudo-samples with the new samples in the unbiased quality control data stream at a ratio of 1:2 to construct an incremental training dataset, and use the mini-batch gradient descent algorithm for incremental training. S335: Calculate the cosine similarity between the new sample and the existing prototypes in the memory bank, add the new sample with the cosine similarity lower than the preset threshold as the new prototype to the memory bank, and remove the oldest prototype with the highest similarity in the memory bank to keep the memory bank capacity constant. S336: Extract the parameter differences between the trained model and the original model, and combine them with the constraints of the core protected parameter set to perform pruning and normalization on the differences to generate the final local model update parameters.

9. The unbiased data flow and secure update method for continuous learning of emergency models according to claim 1, characterized in that, Step S4, the steps for generating the global update model include: S41: Employs a horizontal federated learning architecture, allowing medical institutions to complete model training locally; S42: Add adaptive causal Gaussian noise to the model gradient to apply high-intensity noise to the causal feature gradients related to patient privacy; S43: Encrypt the model parameters using a homomorphic encryption algorithm, and generate the global update model through secure multi-party computation based on causal weights.

10. The unbiased data flow and secure update method for continuous learning of emergency models according to claim 1, characterized in that, Step S5, the steps for performing a full-process security verification of causal consistency on the global update model, include: S51: Construct a multi-dimensional causal verification benchmark set, and label the standard diagnosis and treatment decisions, core causal nodes and complete causal reasoning paths for each sample in the causal verification benchmark set; S52: Calculate the accuracy, sensitivity, specificity, missed diagnosis rate and false diagnosis rate of the global update model on various diseases, and evaluate the negative prediction value of acute and critical illnesses and rare diseases; perform incremental knowledge retention verification to test the performance degradation of the global update model on the historical disease validation set; S54: Extract the activated causal path and key decision nodes when the global update model outputs diagnosis and treatment decisions using the causal attribution algorithm, and perform a structured comparison with the standard causal paths in the causal verification benchmark set to calculate the causal path matching degree, the accuracy of key causal nodes, and the false causal association rate. S55: Verify the stability of the global update model's causal inference under different clinical conditions, different data missing rates, and different time windows. When the global causal consistency score is lower than the preset safety threshold, the model verification is directly determined to be unsuccessful. S56: Verify the effectiveness of gradient causal desensitization and homomorphic encryption mechanisms through member inference attack, model inversion attack and attribute inference attack tests; S57: Build a simulated diagnosis and treatment chain consistent with the real emergency environment. Input test data containing complex scenarios including multiple concurrent patients, asynchronous arrival of multimodal data, and emergency rescue conditions to verify the causal consistency and response timeliness of the entire process from data acquisition, preprocessing, model inference to decision output.