Method and system for infectious disease monitoring based on big data deep causal learning network
By constructing a data resource standard library and evidence-based causal graph through a big data deep causal learning network, the problem of data silos in infectious disease surveillance systems has been solved, enabling high-precision and interpretable infectious disease surveillance, and improving the model's generalization ability and decision support for public health prevention and control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG UNIV
- Filing Date
- 2026-06-23
- Publication Date
- 2026-07-21
AI Technical Summary
Existing infectious disease surveillance systems suffer from data silos and a lack of causal linkage mechanisms, resulting in difficulties in integrating cross-system data, poor model generalization ability, and a lack of interpretability and accuracy, making it difficult to provide a reliable basis for public health prevention and control.
By employing a big data deep causal learning network, and constructing a data resource standard library, an evidence-based causal graph, and a measurement-diagnosis-treatment-evaluation graph brain model, we can achieve multimodal data alignment and fusion, and perform causal inference and prediction, thereby breaking down data silos and improving the interpretability and accuracy of the monitoring system.
It achieves high-precision and interpretable predictions for infectious disease surveillance, maintains high generalization ability when facing heterogeneous data or emerging infectious diseases, and provides decision support based on causal evidence.
Smart Images

Figure CN122436263A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical big data and artificial intelligence technology, and in particular relates to a method and system for infectious disease monitoring based on big data deep causal learning networks. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] Currently, infectious disease surveillance systems primarily rely on structured data from Hospital Information Systems (HIS), Laboratory Information Systems (LIS), Electronic Medical Records (EMR), Picture Archiving and Communication Systems (PACS), and public health reporting systems for statistical analysis and risk warning. With the continuous advancement of medical informatization and public health digitalization, medical institutions, disease control and prevention agencies, and regional health platforms have accumulated a wealth of clinical diagnosis and treatment data, laboratory testing data, medical imaging data, vital sign data, and epidemiological data, providing a data foundation for dynamic monitoring of infectious diseases, transmission risk analysis, and public health decision-making.
[0004] However, existing infectious disease surveillance technologies still have the following problems: On the one hand, there are significant differences in data structure, coding rules, and storage methods among different medical information systems, and a lack of a unified data association and fusion mechanism between systems makes it difficult to standardize and integrate cross-system data. Existing monitoring schemes typically analyze data based on a single data source or limited structured fields, such as making rule judgments based on chief complaint information, positive laboratory results, or public health reporting information. This makes it difficult to effectively integrate multimodal heterogeneous information such as text records, medical images, laboratory indicators, vital sign time-series data, and epidemiological data, thus limiting the completeness and accuracy of infectious disease monitoring.
[0005] On the other hand, existing infectious disease surveillance models based on machine learning or deep learning mostly rely on statistical feature fitting of historical case data to achieve risk prediction, lacking constraints on the correlations between medical events. When these models face different regions, time periods, or emerging and variant infectious disease scenarios, they are easily affected by data distribution bias, regional heterogeneity, and spurious correlations, leading to a decline in model generalization ability. Furthermore, existing models typically lack interpretable analytical capabilities regarding disease transmission pathways, symptom evolution relationships, and factors influencing diagnostic and treatment interventions, making it difficult to provide reliable evidence for public health prevention and control and clinical interventions. Summary of the Invention
[0006] To overcome the shortcomings of the existing technologies, this invention provides an infectious disease monitoring method and system based on big data deep causal learning networks. By breaking down data silos through a standardization system and introducing causal deep learning networks, the infectious disease monitoring system is endowed with deep evidence-based interpretability and high-precision predictive decision-making capabilities.
[0007] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: The first aspect of this invention provides a method for infectious disease monitoring based on big data deep causal learning networks.
[0008] Infectious disease surveillance methods based on deep causal learning networks using big data include: By acquiring multi-source heterogeneous data from medical information systems and performing edge computing, a data resource standard library is constructed. Based on the big data queue standard system and the queue general data model, the data resource standard library is cleaned and reorganized to create a medical science queue data warehouse. Using an evidence-based knowledge triplet extraction toolkit, relationships are extracted from multiple evidence-based medical knowledge sources to construct a standard system for evidence-based causal graphs. Based on a medical graph brain database engine, evidence-based causal graphs and measurement-diagnosis-treatment-evaluation graph brain models are generated. A deep learning network based on causal representation is used to perform multimodal alignment and fusion of the cohort data in the medical science cohort data warehouse with the evidence-based causal graph, and to perform preliminary infectious disease monitoring and causal inference prediction. The obtained infectious disease monitoring and causal inference prediction results are deployed to the epidemic prevention and control terminal under the one-brain-multiple-terminal architecture to achieve large-scale infectious disease monitoring.
[0009] Furthermore, the construction of the data resource standard library includes: adopting a data acquisition and integration general data model to collect multi-source heterogeneous data from the target health and medical information system and performing edge computing; then, through mirror library and standard library intelligent tools, the automatic construction of the data resource standard library is realized; wherein, the multi-source heterogeneous data includes microscopic data and macroscopic data.
[0010] Furthermore, in the process of constructing the data resource standard library, the collected multi-source heterogeneous data is initially filtered, outlier removed, missing value imputed, and desensitized based on the data acquisition and integration general data model. The heterogeneous medical feature vectors corresponding to the multi-source heterogeneous data are mapped into a unified latent space representation through the feature preprocessing encoder.
[0011] Furthermore, a feature preprocessing encoder maps the heterogeneous medical feature vectors corresponding to multi-source heterogeneous data into a unified latent space representation, thereby constructing a highly standardized data resource standard library; the latent space representation is expressed by the following formula: ; in, The latent space representation obtained by the mapping; This indicates the types of multimodal data sources, specifically including images, text records, and structured test data; and They represent the first Projection weights and biases for each modality of data; Represents a non-linear activation function; Indicates the first The sample at the th Input feature vectors under various modalities.
[0012] Furthermore, the general data model for the queue utilizes an evidence-based knowledge triple extraction toolkit and extracts entity and causal relationship networks through natural language processing technology to construct an evidence-based causal graph standard system. Data reconstruction is then performed based on the general data model for the queue: for any patient in the queue, their longitudinal medical trajectory is represented as follows: The weight distribution of historical medical events on the current susceptibility to infectious diseases is calculated using a temporal attention mechanism; the weight distribution and the finally aggregated patient-level queue normalization vector are respectively represented as follows: ; ; in, For the first Attention weights for events at each time step. A learnable query matrix, This is the context representation vector for the global queue; This represents the normalized vector of the resulting queue; Indicates the first The medical event feature vector corresponding to each time step This represents the total number of time steps contained in the longitudinal medical trajectory of the corresponding patient.
[0013] Furthermore, the nodes in the evidence-based causal graph and measurement-diagnosis-treatment-assessment graph brain model include four types of medical entities: measurement, diagnosis, treatment, and assessment. The update formula for node features is expressed as follows: ; in, express In the Hidden layer representation of a layer, For nodes The set of neighboring nodes; The causal edge weight coefficients calculated by the medical graph brain engine. The layer transition matrix; Represents a node The neighboring nodes, Neighboring nodes In the Hidden feature representation in layered graph neural networks; Representation layer normalization, This represents the activation function.
[0014] Furthermore, the deep learning network removes background noise and extracts the causal features of infectious disease pathogenesis and transmission by jointly optimizing the prediction loss and the causal representation identifiability loss.
[0015] A second aspect of this invention provides an infectious disease monitoring system based on a deep causal learning network using big data.
[0016] An infectious disease surveillance system based on deep causal learning networks using big data includes: The data resource standard library construction module is configured to: construct a data resource standard library by acquiring multi-source heterogeneous data from medical information systems and performing edge computing. The medical science cohort data warehouse construction module is configured to: clean and reorganize the data resource standard library based on the big data cohort standard system and the general data model of the cohort, so as to create a medical science cohort data warehouse; The graph brain model construction module is configured to: use the evidence-based knowledge triplet extraction toolkit to extract relationships from multiple evidence-based medical knowledge sources, construct a standard system of evidence-based causal graphs, and generate evidence-based causal graphs and measurement-diagnosis-treatment-evaluation graph brain models based on the medical graph brain database engine. The preliminary prediction module is configured to: use a deep learning network based on causal representation to perform multimodal alignment and fusion of the queue data in the medical science queue data warehouse and the evidence-based causal graph, and perform preliminary infectious disease monitoring and causal inference prediction; The large-scale prediction module is configured to deploy the obtained infectious disease monitoring and causal inference prediction results to the epidemic prevention and control terminal under the one-brain-multi-terminal architecture to realize large-scale infectious disease monitoring.
[0017] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of the infectious disease monitoring method based on big data deep causal learning networks as described in the first aspect of the present invention.
[0018] The fourth aspect of the present invention provides an electronic device including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the infectious disease monitoring method based on big data deep causal learning network as described in the first aspect of the present invention.
[0019] The above one or more technical solutions have the following beneficial effects: This invention first acquires multi-source heterogeneous data from medical information systems and performs edge computing. Combined with mirror and standard library intelligent tools, it automatically constructs a data resource standard library. Based on this, and using a big data cohort standard system and a general cohort data model, it performs deep cleaning and reorganization of the data resource standard library to create a medical science cohort data warehouse. Through this complete "acquisition-edge computing-standardization-cohort reconstruction" chain, this invention breaks down data barriers between existing hospital information systems, laboratory information systems, image archiving systems, and public health systems. It can map multimodal data such as images, text records, and structured test data into a unified latent spatial representation, thereby fully utilizing all data information, significantly improving the breadth and depth of infectious disease monitoring data, and overcoming the information gaps caused by single data sources or a few structured fields.
[0020] This invention employs an evidence-based knowledge triplet extraction toolkit to extract relationships from multiple evidence-based medical knowledge sources, constructing a standard system for evidence-based causal graphs. It generates evidence-based causal graphs and a measurement-diagnosis-treatment-assessment graph-brain model, embedding medical prior causal logic into the model's underlying layer in a structured form. Simultaneously, it utilizes a deep learning network based on causal representation to perform multimodal alignment and fusion of cohort data and the evidence-based causal graph. By jointly optimizing prediction loss and causal representation identifiability loss, it effectively removes background noise and extracts the true causal characteristics of infectious disease pathogenesis and transmission. Therefore, the monitoring and prediction results of this invention no longer rely on statistical spurious correlations but are built upon interpretable causal chains. Even when facing heterogeneous medical data or novel, mutated infectious diseases, it maintains high generalization ability and robustness, providing causally based decision support for public health interventions.
[0021] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0022] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0023] Figure 1 This is a flowchart of the infectious disease monitoring method based on big data deep causal learning network in Embodiment 1 of the present invention. Detailed Implementation
[0024] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0025] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.
[0026] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0027] Example 1 This embodiment discloses an infectious disease monitoring method based on big data deep causal learning networks.
[0028] like Figure 1 As shown, the infectious disease surveillance method based on big data deep causal learning networks includes: Step S1: By acquiring multi-source heterogeneous data from the medical information system and performing edge computing, a data resource standard library is constructed; Step S2: Based on the big data queue standard system and the queue general data model, the data resource standard library is cleaned and reorganized to create a medical science queue data warehouse; Step S3: Using the evidence-based knowledge triplet extraction toolkit, relationships are extracted from multiple evidence-based medical knowledge sources to construct a standard system for evidence-based causal graphs. Based on the medical graph brain database engine, evidence-based causal graphs and measurement-diagnosis-treatment-evaluation graph brain models are generated. Step S4: Using a deep learning network based on causal representation, the queue data in the medical science queue data warehouse and the evidence-based causal graph are multimodal aligned and fused, and preliminary infectious disease monitoring and causal inference prediction are performed. Step S5: Deploy the obtained infectious disease monitoring and causal inference prediction results to the epidemic prevention and control terminal under the one-brain-multi-terminal architecture to achieve large-scale infectious disease monitoring.
[0029] Based on the above process, this invention breaks down data silos through a standardization system and, by introducing a causal deep learning network, endows the infectious disease monitoring system with deep evidence-based interpretability and high-precision predictive decision-making capabilities. To facilitate understanding of the technical solution of this invention, the specific implementation methods of this invention will be further explained and described below.
[0030] In step S1, a data resource standard library is constructed by acquiring multi-source heterogeneous data from the medical information system and performing edge computing.
[0031] A general data acquisition and fusion model is adopted to collect multi-source heterogeneous data from the target health and medical information system and perform edge computing on the data source side. The general data acquisition and fusion model is a data access and standardization processing model constructed by this invention, which is used to uniformly collect, map fields, convert formats, verify quality, desensitize privacy, and encode features from hospital information systems, laboratory information systems, image archiving systems, electronic medical record systems, public health reporting systems, medical insurance systems, and accessible social and environmental data sources.
[0032] Furthermore, the data acquisition and integration general data model includes a data source adaptation layer, a field mapping layer, a quality control layer, a privacy desensitization layer, and a feature encoding layer. The data source adaptation layer connects different health and medical information systems and extracts raw data. The field mapping layer maps patient identifiers, consultation times, diagnostic codes, laboratory indicators, imaging examinations, medication records, and public health reporting fields from different systems to a unified field system. The quality control layer performs duplicate record identification, outlier removal, missing value imputation, and timestamp correction. The privacy desensitization layer de-identifies sensitive fields such as names, ID numbers, and contact information. The feature encoding layer converts text, images, structured laboratory data, and vital sign time-series data into heterogeneous medical feature vectors that can be processed by the subsequent feature preprocessing encoder.
[0033] When performing edge computing, the system deploys a data acquisition edge computing toolkit on the data source sides of hospital information systems, laboratory information systems, image archiving systems, electronic medical record systems, public health reporting systems, and medical insurance systems to perform localized preprocessing on the collected raw multi-source heterogeneous data.
[0034] Furthermore, edge computing includes at least data format parsing, field standardization mapping, duplicate record identification, outlier removal, missing value imputation, timestamp correction, patient identifier desensitization, data quality scoring, and lightweight feature encoding. Among these, structured test data and vital sign data are converted into standardized numerical features, text records are segmented, entity recognition, and keyword extraction to generate text features, and image data is normalized in size, standardized in grayscale, and encoded in lightweight image to generate image features.
[0035] After edge computing is completed, the system writes the processed standard fields, data quality identifiers, de-identified indexes, time tags, modality type identifiers, and preliminary feature vectors into the mirror library. The standard library intelligent tool then aggregates these data into a data resource standard library according to a unified field system and modality index rules.
[0036] A mirror repository is a data cache that creates a synchronized copy of the original data in a target healthcare information system without altering the original business system's data structure and operational logic. It stores anonymized copies of the original data, field mapping records, timestamps, data source identifiers, data versions, and data quality identifiers. The mirror repository does not directly replace the original business database; instead, it is established as a synchronized cache on the data source side or within a secure area of the hospital. This avoids direct manipulation of the business system while ensuring that the original data source, version, and quality status can be traced during subsequent standardization processes.
[0037] Furthermore, the standard library's intelligent tools include data source connection tools, field mapping tools, data cleaning tools, quality verification tools, privacy desensitization tools, missing value imputation tools, outlier detection tools, standard encoding conversion tools, and modal feature indexing tools. Specifically, the data source connection tool connects Hospital Information Systems (HIS), Laboratory Information Systems (LIS), Electronic Medical Records (EMR), public health reporting systems, and medical insurance systems; the field mapping tool maps patient numbers, consultation times, diagnosis names, test items, examination results, imaging pathways, and medication information from different systems to a unified field system; the data cleaning tool removes duplicate records, corrects formatting errors, and standardizes time formats; the quality verification tool generates integrity, consistency, and validity identifiers; the privacy desensitization tool de-identifies sensitive information such as names, ID numbers, and mobile phone numbers; the standard encoding conversion tool converts diagnostic, testing, drug, and operational records into a unified code; and the modal feature indexing tool marks the modal types of text, images, structured test data, and vital sign time-series data, as well as their subsequent feature encoding entries. Through the above processing, the multi-source heterogeneous data in the mirror library is transformed into a data resource standard library with a unified field structure, unified encoding rules, unified quality identifiers, and unified modal indexes.
[0038] In the process of constructing the data resource standard library, the collected multi-source heterogeneous data is initially filtered, outlier removed, missing value imputed, and desensitized based on the data collection and integration general data model. The heterogeneous medical feature vectors corresponding to the multi-source heterogeneous data are mapped into a unified latent space representation through the feature preprocessing encoder.
[0039] In the specific implementation process, data acquisition integrates a general data model. For multi-source, heterogeneous, and isolated data sources, input data is processed through feature mapping and multi-omics / multi-modal data alignment mechanisms. A "data acquisition edge computing toolkit" is deployed at each business system to perform preliminary filtering of raw social, economic, environmental, and medical system data, including outlier removal, missing value imputation, and anonymization. Medical data includes text (chief complaint), images (CT / X-ray films), and time-series values (vital signs). The system uses a mirror library intelligent tool to extract copies of the original data from each system. Assuming the input heterogeneous medical feature vector is... The feature preprocessing encoder maps it into a unified latent space representation. : ; in, This indicates the types of multimodal data sources, specifically including images, text records, and structured test data; and They represent the first The projection weights and biases of the modal data were initially set to 0.2 and 0.33, respectively. Represents a non-linear activation function; Indicates the first The sample at the th Heterogeneous medical feature vectors under various modalities. The unified latent space representation. As a standardized feature layer in the data resource standard library, it together with standard fields, data source identifiers, time tags, quality identifiers, and modal indexes constitute the data resource standard library, which is used to provide unified input for subsequent construction of medical science cohort data warehouses, evidence-based causal graph fusion, and infectious disease monitoring and prediction.
[0040] In step S2, based on the big data queue standard system and the queue general data model, the data resource standard library is cleaned and reorganized to create a medical science queue data warehouse.
[0041] Based on the big data queue standard system and the queue general data model, the standardized data fields, modal feature indexes, time tags, data source identifiers and data quality identifiers in the data resource standard library are cleaned, queued and reorganized, and vertical trajectory is constructed to create a medical science queue data warehouse.
[0042] First, the system uses the patient's unique desensitization identifier, consultation time, medical institution code, and event type to correlate and match outpatient records, inpatient records, laboratory test results, image indexes, diagnostic information, medication records, vital sign time series data, and public health reporting data in the data resource standard library, forming a multi-source medical event set centered on the patient.
[0043] Subsequently, the system screens patient samples based on cohort inclusion and exclusion criteria. The inclusion criteria include at least symptoms related to the target infectious disease, diagnostic codes, laboratory test results, or epidemiological exposure records. The exclusion criteria include at least samples with severely missing key fields, abnormal time sequences, duplicate visit records that cannot be merged, and samples that do not meet the follow-up time requirements. For the screened patient samples, the system reconstructs their historical medical events into a longitudinal medical trajectory at a uniform time granularity, and encodes the symptoms, diagnoses, tests, imaging, medications, vital signs, and public health reporting events within each time step into a cohort event vector.
[0044] During the cohort reorganization process, the system further performs event time correction, cross-institutional duplicate record merging, missing field imputation, outlier event removal, unified diagnostic and laboratory coding, image and text modal index binding, and label generation. The label generation process includes generating infectious disease status labels, susceptibility status labels, exposure status labels, and outcome labels for patients within different time windows based on the target infectious disease diagnosis results, positive laboratory test results, syndrome reporting status, or public health confirmation records. Finally, the system jointly writes patient basic information, longitudinal medical trajectory, cohort event vectors, time window labels, follow-up information, data quality identifiers, and latent spatial representations into the medical science cohort data warehouse, providing standardized cohort input for subsequent evidence-based causal graph fusion, temporal attention modeling, and deep causal learning network prediction.
[0045] In step S3, the evidence-based knowledge triplet extraction toolkit is used to extract relationships from multiple evidence-based medical knowledge sources, construct a standard system for evidence-based causal graphs, and generate evidence-based causal graphs and measurement-diagnosis-treatment-evaluation graph models based on the medical graph brain database engine.
[0046] A general data model for cohorts is used to reconstruct cross-sectional, discrete medical records into longitudinal tracking data that meets the input requirements of epidemiology and deep learning. After obtaining a standard library of data resources, cross-sectional data must be transformed into long-term longitudinal tracking data for epidemiological causal inference. By introducing external clinical practice guidelines and other evidence-based medicine knowledge sources, utilizing an evidence-based knowledge triplet extraction toolkit, and employing natural language processing techniques to extract entity and causal relationship networks, a standard system of evidence-based causal graphs is constructed.
[0047] The construction of the evidence-based causal graph standard system includes: obtaining evidence-based medicine knowledge sources from clinical practice guidelines, expert consensus, infectious disease prevention and control standards, evidence-based medicine databases, and publicly available medical literature; performing text cleaning, entity recognition, relation extraction, and terminology standardization on the knowledge sources to obtain a set of medical entities including symptoms, test indicators, imaging manifestations, pathogens, disease diagnoses, treatment measures, and prognostic assessment indicators; using an evidence-based knowledge triple extraction toolkit to represent the causal or evidentiary relationship between entities as a "head entity-relationship-tail entity" triple; calculating the reliability of the relationship based on the knowledge source level, evidence level, recommendation strength, and entity co-occurrence frequency; and finally generating an evidence-based causal graph standard system containing standardized entity nodes, causal relationship edges, evidence levels, and source identifiers.
[0048] Finally, data reconstruction is performed based on the general data model of the queue: for any patient in the queue, their longitudinal medical trajectory is represented as follows: The weight distribution of historical medical events on the current susceptibility to infectious diseases is calculated using a temporal attention mechanism; the weight distribution and the finally aggregated patient-level queue normalization vector are respectively represented as follows: ; ; in, For the first Attention weights for events at each time step. A learnable query matrix, This is the context representation vector for the global queue; This represents the normalized vector of the resulting queue.
[0049] The medical graph brain database engine includes a medical entity standardization module, a causal relationship input module, a queue event mapping module, a graph structure generation module, an edge weight calculation module, and a graph neural network update module. Specifically, the medical entity standardization module maps symptoms, test indicators, imaging findings, diagnoses, treatment measures, and assessment indicators in the evidence-based causal graph to unified medical entities. The queue event mapping module maps longitudinal medical events of patients in the medical science queue data warehouse to corresponding medical entity nodes. The graph structure generation module generates MDTE node sets according to four types of entities: measurement, diagnosis, treatment, and assessment. The engine calculates the causal edge weights between nodes based on the causal relationships, evidence levels, event time order, and co-occurrence frequency in the evidence-based knowledge triplet, and establishes directed causal edges between measurement nodes, diagnosis nodes, treatment nodes, and assessment nodes, generating an evidence-based causal graph and a measurement-diagnosis-treatment-assessment graph brain model. The update formula for node features is as follows: ; in, express In the Hidden layer representation of a layer, For nodes The set of neighboring nodes; The causal edge weight coefficients calculated for the medical graph brain engine are initially set to 0.21; The layer transition matrix; Represents a node The neighboring nodes, Neighboring nodes In the Hidden feature representation in layered graph neural networks. This formula can be used to deeply mine evidence-based prior knowledge of infectious diseases.
[0050] In step S4, a deep learning network based on causal representation is used to perform multimodal alignment and fusion of the queue data in the medical science queue data warehouse with the evidence-based causal graph, and to perform preliminary infectious disease monitoring and causal inference prediction.
[0051] The deep learning network based on causal representation includes a queue feature encoding module, a graph feature encoding module, a cross-modal causal alignment module, a causal latent variable extraction module, a prediction output module, and a joint loss optimization module.
[0052] The system comprises the following modules: a queue feature encoding module, a graph ...
[0053] Specifically, for the first Let there be a number of patients or monitoring subjects, and let the standardized vector of their medical science cohort data warehouse be denoted as . The features of the evidence-based causal graph extracted by the graph brain model are: First, the queue representation and the graph representation are obtained through two encoders: ; ; in, Indicates the first Characterization of cohort data from individual patients Represents the queue feature encoder. Represents the learnable parameters of the queue feature encoder; Indicates the first Evidence-based causal profiles for each patient. This represents a spectral feature encoder. This represents the learnable parameters of the graph feature encoder.
[0054] Subsequently, the queue representation With spectral characterization Perform cross-modal fusion to obtain joint representation : ; in, The projection weight matrix represents the queue representation. The projection weight matrix represents the spectral representation. Indicates the fusion bias term. Represents a non-linear activation function. This represents the fused multimodal causal candidate representation. Through this fusion process, statistical features from real-world patient cohort data and medical prior knowledge from evidence-based causal graphs are mapped to a unified representation space.
[0055] To extract risk representations with causal interpretability, the network further integrates representations. Input Causal Latent Variable Extraction Module: ; in, Indicates the first Causal latent variables of individual patients or monitored subjects are used to characterize potential causal factors that are directly related to the occurrence, spread, progression or outcome of infectious diseases; This represents a causal latent variable extraction network. This indicates its learnable parameters. This module is used to remove background noise and spurious correlation features from the fused representation, retaining key features that are causally related to infectious disease risk.
[0056] The prediction output module is based on causal latent variables. Generate infectious disease surveillance and prediction results: ; in, Indicates the first The prediction results for an individual patient or monitored subject can be classified as infection risk category, transmission risk level, or warning status. This represents the prediction layer weight matrix. Indicates the prediction layer bias term. Used to output the probability distribution of different risk categories.
[0057] During training, the deep learning network updates its parameters by jointly optimizing the prediction loss and the causal representation identifiability loss, and its joint loss function is expressed as: ; in, Represents the total loss function. This represents the monitoring and prediction loss, used to measure the actual infectious disease status label. Compared with the prediction results The differences between them; Indicates the true label, Indicates the predicted label or predicted probability; This represents the loss of identifiability of causal representations, used to constrain latent causal variables. Characteristics of Evidence-Based Causal Map Maintain consistency; This represents the balance coefficient between prediction loss and causal characterization identifiability loss.
[0058] In one implementation, the causal characterization identifiability loss can take the following form: ; in, Denotes KL divergence, Indicates the fusion features The distribution of causal latent variables obtained through inference, This indicates the characteristics of the evidence-based causal graph. It provides a constrained prior causal distribution. Through this constraint, when making infectious disease surveillance and predictions, the network not only relies on the statistical correlation in the data, but also incorporates causal priors from evidence-based medicine, thereby reducing the influence of spurious correlation features and improving the robustness and interpretability of the model in different regions, seasons, and emerging infectious disease scenarios.
[0059] Through the aforementioned network structure, the medical science cohort data warehouse provides real-world longitudinal medical data on patients, while the evidence-based causal graph provides medical causal priors. After cross-modal causal alignment, these two data points jointly generate causal latent variables, which are then used by the prediction output module to perform infectious disease monitoring and causal inference prediction. This network can provide causal explanations related to symptoms, testing, diagnosis, treatment, and assessment while maintaining prediction accuracy, thus providing a basis for subsequent risk warnings, transmission chain tracing, and intervention strategy evaluation in epidemic prevention and control.
[0060] A deep learning network based on causal representation, by jointly optimizing the prediction loss and the causal representation identifiability loss, aims to remove background noise and extract the true causal features of infectious disease pathogenesis and transmission. The joint loss function is defined as follows: ; in, To monitor the predicted cross-entropy loss function, and These are labels for actual and predicted infectious disease states, respectively. Characteristics of the patient cohort The graph brain spectrum features extracted in step S3; KL divergence is used to constrain the extracted causal latent variables. Follows prior independent distribution , This is the balance coefficient.
[0061] In step S5, the obtained infectious disease monitoring and causal inference prediction results are deployed to the epidemic prevention and control terminal under the one-brain-multi-terminal architecture to achieve large-scale infectious disease monitoring.
[0062] The system's underlying data and analysis engine are uniformly encapsulated into a multi-terminal intelligent medical graph brain central hub, providing high-concurrency API call support through a microservice architecture. The data resource standard library, medical science cohort data warehouse, evidence-based causal graph, evidence-based causal graph, measurement-diagnosis-treatment-assessment (MDTE) graph brain model, and causal representation-based deep learning network are all uniformly encapsulated into this multi-terminal intelligent medical graph brain central hub. This central hub includes a data service layer, a graph service layer, a model inference layer, a task scheduling layer, an interface service layer, and a permission and security layer. Through a microservice architecture, the central hub encapsulates data querying, queue construction, graph inference, model prediction, risk warning, report generation, and user permission management as independent services, and provides infectious disease risk prediction, transmission chain tracing, causal explanation paths, and intervention target assessment results to the epidemic prevention and control end, clinical support end, scientific research analysis end, and management decision-making end through a unified API. First, combined with the measurement-diagnosis-treatment-assessment (MDTE) graph brain model, real-time regional epidemic monitoring, transmission chain tracing, and risk warning are achieved. Second, it provides researchers with convenient big data real-world and randomized clinical design tools to accelerate research output. For suspected or confirmed patients, personalized medication recommendations and treatment inferences are provided based on graph reasoning. Ultimately, as the "medical graph brain center," it outputs externally to the epidemic prevention and control end of the "one brain, multiple terminals" system. Relevant personnel can use this port to view the infectious disease transmission chain, predictive trend charts, and intervention target assessments provided by the graph brain in real time.
[0063] The infectious disease monitoring method provided by this invention achieves several technological breakthroughs compared to existing technologies: 1) By using the RCDM data collection standard system and tools, it breaks down the silos of systems within and outside the industry, realizing the automatic construction of a "data resource standard library" and the fusion of multimodal data; 2) By designing the CCDM big data cohort standard system, it realizes the creation of a "medical science cohort data warehouse," providing structured, large-scale, high-quality basic data for complex epidemiological tracing and tracking; 3) Based on evidence-based medicine knowledge sources, it realizes the construction of the GCDM evidence-based causal graph standard system and the "medical graph brain" database engine. At the same time, it introduces MDTE prior guidance into deep networks, effectively solving the black-box drawback of existing artificial intelligence that "knows what but not why"; 4) By using a deep causal inference and predictive decision-making method system, combined with a "one brain, multiple terminals" smart medical architecture, it realizes large-scale and industrialized applications from bottom-level data processing to top-level epidemic prevention and control, clinical research, precision medicine, and other scenarios.
[0064] Example 2 This embodiment discloses an infectious disease monitoring system based on a deep causal learning network using big data.
[0065] An infectious disease surveillance system based on deep causal learning networks using big data includes: The data resource standard library construction module is configured to: construct a data resource standard library by acquiring multi-source heterogeneous data from medical information systems and performing edge computing. The medical science cohort data warehouse construction module is configured to: clean and reorganize the data resource standard library based on the big data cohort standard system and the general data model of the cohort, so as to create a medical science cohort data warehouse; The graph brain model construction module is configured to: use the evidence-based knowledge triplet extraction toolkit to extract relationships from multiple evidence-based medical knowledge sources, construct a standard system of evidence-based causal graphs, and generate evidence-based causal graphs and measurement-diagnosis-treatment-evaluation graph brain models based on the medical graph brain database engine. The preliminary prediction module is configured to: use a deep learning network based on causal representation to perform multimodal alignment and fusion of the queue data in the medical science queue data warehouse and the evidence-based causal graph, and perform preliminary infectious disease monitoring and causal inference prediction; The large-scale prediction module is configured to deploy the obtained infectious disease monitoring and causal inference prediction results to the epidemic prevention and control terminal under the one-brain-multi-terminal architecture to realize large-scale infectious disease monitoring.
[0066] Example 3 The purpose of this embodiment is to provide a computer-readable storage medium.
[0067] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the infectious disease monitoring method based on big data deep causal learning networks as described in Embodiment 1 of this disclosure.
[0068] Example 4 The purpose of this embodiment is to provide an electronic device.
[0069] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the infectious disease monitoring method based on big data deep causal learning networks as described in Embodiment 1 of this disclosure.
[0070] The steps and methods involved in the apparatuses of Embodiments 2, 3, and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.
[0071] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.
[0072] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method for infectious disease surveillance based on deep causal learning networks using big data, characterized in that, Includes the following steps: By acquiring multi-source heterogeneous data from medical information systems and performing edge computing, a data resource standard library is constructed. Based on the big data queue standard system and the queue general data model, the data resource standard library is cleaned and reorganized to create a medical science queue data warehouse. Using an evidence-based knowledge triplet extraction toolkit, relationships are extracted from multiple evidence-based medical knowledge sources to construct a standard system for evidence-based causal graphs. Based on a medical graph brain database engine, evidence-based causal graphs and measurement-diagnosis-treatment-evaluation graph brain models are generated. A deep learning network based on causal representation is used to perform multimodal alignment and fusion of the cohort data in the medical science cohort data warehouse with the evidence-based causal graph, and to perform preliminary infectious disease monitoring and causal inference prediction. The obtained infectious disease monitoring and causal inference prediction results are deployed to the epidemic prevention and control terminal under the one-brain-multiple-terminal architecture to achieve large-scale infectious disease monitoring.
2. The infectious disease surveillance method based on big data deep causal learning networks as described in claim 1, characterized in that, The construction of the data resource standard library includes: using a data acquisition and integration general data model to collect multi-source heterogeneous data from the target health and medical information system and performing edge computing; then, through mirror library and standard library intelligent tools, the automatic construction of the data resource standard library is realized; wherein, the multi-source heterogeneous data includes microscopic data and macroscopic data.
3. The infectious disease monitoring method based on big data deep causal learning networks as described in claim 1, characterized in that, In the process of constructing the data resource standard library, the collected multi-source heterogeneous data is initially filtered, outlier removed, missing value imputed, and desensitized based on the data acquisition and integration general data model. The heterogeneous medical feature vectors corresponding to the multi-source heterogeneous data are mapped into a unified latent space representation through the feature preprocessing encoder.
4. The infectious disease monitoring method based on big data deep causal learning networks as described in claim 3, characterized in that, A feature preprocessing encoder maps heterogeneous medical feature vectors corresponding to multi-source heterogeneous data into a unified latent space representation, thereby constructing a highly standardized data resource standard library; the latent space representation is expressed by the following formula: ; in, The latent space representation obtained by the mapping; This indicates the types of multimodal data sources, specifically including images, text records, and structured test data; and They represent the first Projection weights and biases for each modality of data; Represents a non-linear activation function; Indicates the first The sample at the th Input feature vectors under various modalities.
5. The infectious disease surveillance method based on big data deep causal learning networks as described in claim 1, characterized in that, The general data model for the queue utilizes an evidence-based knowledge triple extraction toolkit and employs natural language processing techniques to extract entity and causal relationship networks. This is used to construct a standard system for evidence-based causal graphs, and the data is reconstructed based on the general data model for the queue: for any patient in the queue, their longitudinal medical trajectory is represented as follows: The weight distribution of historical medical events on the current susceptibility to infectious diseases is calculated using a temporal attention mechanism; the weight distribution and the finally aggregated patient-level queue normalization vector are respectively represented as follows: ; ; in, For the first Attention weights for events at each time step. A learnable query matrix, This is the context representation vector for the global queue; This represents the normalized vector of the resulting queue; Indicates the first The medical event feature vector corresponding to each time step This represents the total number of time steps contained in the longitudinal medical trajectory of the corresponding patient.
6. The infectious disease surveillance method based on big data deep causal learning networks as described in claim 1, characterized in that, The nodes in the evidence-based causal graph and the measurement-diagnosis-treatment-assessment graph brain model include four types of medical entities: measurement, diagnosis, treatment, and assessment. The update formula for the features of each node is expressed as follows: ; in, express In the Hidden layer representation of a layer, For nodes The set of neighboring nodes; The causal edge weight coefficients calculated by the medical graph brain engine. The layer transition matrix; Represents a node The neighboring nodes, Neighboring nodes In the Hidden feature representation in layered graph neural networks; Representation layer normalization, This represents the activation function.
7. The infectious disease surveillance method based on big data deep causal learning networks as described in claim 1, characterized in that, The deep learning network extracts the pathogenic and causal features of infectious diseases by jointly optimizing the prediction loss and the causal representation identifiability loss.
8. An infectious disease surveillance system based on big data deep causal learning networks, characterized in that, include: The data resource standard library construction module is configured to: construct a data resource standard library by acquiring multi-source heterogeneous data from medical information systems and performing edge computing. The medical science cohort data warehouse construction module is configured to: clean and reorganize the data resource standard library based on the big data cohort standard system and the general data model of the cohort, so as to create a medical science cohort data warehouse; The graph brain model construction module is configured to: use the evidence-based knowledge triplet extraction toolkit to extract relationships from multiple evidence-based medical knowledge sources, construct a standard system of evidence-based causal graphs, and generate evidence-based causal graphs and measurement-diagnosis-treatment-evaluation graph brain models based on the medical graph brain database engine. The preliminary prediction module is configured to: use a deep learning network based on causal representation to perform multimodal alignment and fusion of the queue data in the medical science queue data warehouse and the evidence-based causal graph, and perform preliminary infectious disease monitoring and causal inference prediction; The large-scale prediction module is configured to deploy the obtained infectious disease monitoring and causal inference prediction results to the epidemic prevention and control terminal under the one-brain-multi-terminal architecture to realize large-scale infectious disease monitoring.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the infectious disease monitoring method based on big data deep causal learning network as described in any one of claims 1-7.
10. An electronic device, comprising a memory and a processor, characterized in that, The memory stores a computer program; when the processor executes the computer program, it implements the infectious disease monitoring method based on big data deep causal learning network as described in any one of claims 1-7.