A Public Health Case Database Management System
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-08-14
AI Technical Summary
[0007]其四,现有案例库仅提供被动浏览功能,无法支持用户对案例中干预措施进行交互式因果推演与反事实分析,教学与决策辅助能力严重受限
[0028](1)通过内嵌动态病原体知识库并调用SMT求解器进行时序逻辑约束一致性验证,本发明突破了现有系统仅能做格式校验的局限,能够自动识别案例数据中深层的流行病学专业逻辑矛盾并给出基于专业知识的修正建议。同时,对关键时间节点缺失的案例,利用贝叶斯结构学习模型,以同地区同期病例的潜伏期分布作为专业先验进行最大后验估计补全,相较于通用插值方法,补全结果具有流行病学专业依据,显著提升了案例数据的完整性和可信度。
Smart Images

Figure CN122575763A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical data management technology, specifically a public health case database management system. Background Technology
[0002] Public health case databases are crucial foundational tools for epidemic monitoring, emergency response, scientific research, and teaching and training in the public health field. By systematically recording and archiving real public health events, case databases provide critical data support for epidemiological investigations, prevention and control strategy formulation, and medical education. However, existing public health case database management systems remain at the level of traditional relational databases in their architecture design. Their core functions are limited to case entry, storage, and keyword retrieval, essentially making them passive data archiving tools. In terms of data quality control, existing systems can only perform basic format checks on case data and cannot identify and correct inconsistencies in the epidemiological logic within cases. Regarding data utilization, cases are only provided as static archives for reference, failing to meet the growing needs of dynamic debriefing, causal analysis, and teaching and training in the public health field.
[0003] The aforementioned shortcomings manifest themselves in the following technical aspects:
[0004] First, the entered cases often contain professional logical errors due to recall bias or reporting negligence, such as infection date being later than death date or missing key time nodes. The existing system lacks an automatic verification and intelligent completion mechanism based on epidemiological expertise, making it difficult to guarantee data quality.
[0005] Secondly, under the traditional storage architecture, privacy desensitization and source tracing are conflicting. For example, to protect patient privacy, key source tracing fields must be removed or obscured, but this also cuts off the discovery clues of implicit homologous associations across cases, and cannot meet the dual needs of privacy compliance and chain tracing.
[0006] Third, existing systems all adopt a static final version archiving mode, retaining only the final version of a case. However, real public health prevention and control is a dynamic game process of repeated trials and multiple rounds of adjustments. All information about the process of trial and error and strategy iteration in the early stage is lost, and the review and reuse value of the case is extremely low.
[0007] Fourth, the existing case library only provides a passive browsing function and cannot support users to conduct interactive causal inference and counterfactual analysis of the intervention measures in the cases, which severely limits its teaching and decision support capabilities. Summary of the Invention
[0008] (a) Technical problems to be solved
[0009] To address the shortcomings of existing technologies, this invention provides a public health case database management system, which solves the problems mentioned in the background section.
[0010] (II) Technical Solution
[0011] To achieve the above objectives, the present invention provides the following technical solution: a public health case database management system, including an epidemiological model verification and completion module, used to perform logical consistency verification and missing data inference and completion on the entered public health case data, and output standardized case data;
[0012] The dynamic knowledge network construction module is used to perform multi-granular semantic annotation on the standardized case data, extract entities and relationships, and construct a public health event knowledge network.
[0013] The implicit homology association clustering and desensitization storage module is used to generate hierarchical materialized snapshots of the original fields of the core traceability of cases and simultaneously generate an irreversible privacy desensitization derivative index layer. Based on the three implicit dimensions of pathogen gene traceability fragments, population flow intersection trajectory and environmental point overlap features, the module automatically clusters and stores the cross-temporal implicit homology association case chain on the desensitization derivative index layer.
[0014] The case version differential closed-loop storage module is used to perform version differential storage of case data generated by multiple rounds of prevention and control interventions under the same public health event, generate multiple versions of dynamic case differential copies carrying the loss coefficient of intervention effectiveness, and form a closed-loop storage of the whole-link evolution trajectory of prevention and control strategy trial, optimization and finalization.
[0015] The counterfactual deduction module is used to perform counterfactual outcome deductions based on causal models and visualize the results, based on user modifications to intervention measures in a case.
[0016] Furthermore, the epidemiological model validation and completion module embeds a dynamic pathogen knowledge base, storing the incubation period intervals, passage intervals, and effective reproduction numbers for various infectious diseases; it generates temporal logic constraints for case timelines and location information, calls the SMT solver to verify consistency, marks data items that do not meet the constraints, and provides correction suggestions; for cases with missing key time nodes, it uses a Bayesian structured learning model, taking the incubation period distribution of cases in the same region during the same period as prior, and obtains the maximum posterior estimate of the missing node by solving the following formula. and confidence interval:
[0017] ;
[0018] in, This is a key time point that needs to be filled in. For the time series data of the observed cases, Let be the likelihood function. This is a prior probability density constructed based on the distribution of incubation periods of cases in the same region during the same period, and the completed data is labeled as inferred values.
[0019] Furthermore, the underlying layer of the dynamic knowledge network construction module adopts a native fusion index kernel for heterogeneous data anchored by cross-modal case vital signs: taking the core public health treatment vital signs as the anchoring benchmark, the heterogeneous data of different modalities are not converted into formats, and the corresponding vital sign anchor points are directly attached in their native form to generate a cross-modal parallel dedicated index queue; the kernel automatically aligns the timestamps, spatial locations and population association three-dimensional benchmarks of multi-source data, and natively solidifies the coupling and association logic of heterogeneous data.
[0020] Furthermore, the specific implementation of the implicit homology clustering and desensitization storage module is as follows: a hierarchical materialized snapshot is generated and sealed for the original core traceability fields of the case, and an irreversible privacy desensitization derivative index layer is generated simultaneously; on the desensitization derivative index layer, based on three implicit dimensions—pathogen gene traceability fragments, population flow intersection trajectories, and environmental point overlap features—the case chain of implicit homology association across time and space is automatically clustered and sealed.
[0021] Furthermore, the specific implementation of the case version differential closed-loop storage module is as follows: following multiple rounds of prevention and control intervention actions, it captures real-time data such as on-site operation logs, resource scheduling receipts, and population management feedback every second, and automatically generates multiple versions of dynamic case differential copies under the same public health event; it embeds a health intervention time-series game weight model, marks the effectiveness loss coefficient of intervention actions version by version, and stores the entire evolution trajectory of closed-loop retention strategy trial, optimization, and finalization; the health intervention time-series game weight model is specifically as follows:
[0022] ;
[0023] in, For the first The effectiveness loss coefficient of the intervention action. and These represent the effective regeneration numbers before and after implementing this version of the intervention. This is the normalized value of the resource input cost for this intervention; this coefficient is annotated in the corresponding version of the differential copy. The larger the value, the higher the effectiveness of epidemic control achieved per unit cost.
[0024] Furthermore, the counterfactual deduction teaching module uses a causal discovery algorithm to extract a causal graph of exposure, infection, transmission, intervention, and control from a complete event chain case, and fits structural equation model parameters based on historical data; it provides an interactive interface to receive user modifications to the type, intensity, or time of intervention measures on the case timeline, calls the structural equation model, and performs probabilistic intervention calculations on relevant variables in the causal graph with the intervention operation as a condition, and displays the counterfactual outcome.
[0025] Furthermore, the counterfactual deduction teaching module also records the user's deduction operation sequence, compares it with the preset expert deduction path by editing distance, evaluates the rationality of the decision and outputs feedback; when performing deduction, the counterfactual deduction teaching module calls the multiple versions of dynamic case differential copies stored in the case version differential closed-loop storage module as a benchmark reference.
[0026] Furthermore, it also includes a tiered computing power scheduling module, which dynamically divides the core tracing case dedicated computing power pool and the ordinary archive case rate limiting pool according to the public health event response level; during the first level emergency response, it triggers the hardware-level offline instantaneous sealing of core cases to avoid data damage caused by network peak congestion.
[0027] (III) Beneficial Effects
[0028] (1) By embedding a dynamic pathogen knowledge base and calling the SMT solver to verify the consistency of temporal logic constraints, this invention breaks through the limitation of existing systems that can only perform format verification. It can automatically identify deep-seated epidemiological professional logical contradictions in case data and provide correction suggestions based on professional knowledge. At the same time, for cases with missing key time nodes, a Bayesian structure learning model is used to perform maximum a posteriori estimation to complete the data, using the incubation period distribution of cases in the same region and at the same time as a professional prior. Compared with general interpolation methods, the completion results have epidemiological professional basis, which significantly improves the completeness and credibility of case data.
[0029] (2) This invention uses an incremental desensitization materialized sharding storage architecture to physically isolate and seal the core traceability original fields in materialized snapshots, and simultaneously generate an irreversible privacy desensitization derivative index layer. On the desensitization index layer, automatic homology clustering is performed around the three implicit dimensions of pathogen genes, population trajectory and environmental location. For the first time, the parallel design of privacy desensitization and chain traceability is realized at the storage physical architecture level, breaking through the long-standing dilemma that the two cannot be achieved at the same time in traditional systems.
[0030] (3) This invention generates multiple differential copies carrying the loss coefficient of intervention effectiveness by capturing multi-round prevention and control intervention action data in real time, thus retaining the complete evolution trajectory of prevention and control strategies from early trial and error, mid-term optimization to late-term finalization in a closed loop. The case is no longer just an archived file that records the results, but a dynamic data asset containing the complete decision-making process, which can support the full-process accurate review and strategy benchmarking and reproduction in the scenario of new and sudden epidemics.
[0031] (4) This invention embeds a structural causal model into a case database management system, providing an interactive counterfactual reasoning interface. Users can arbitrarily modify the type, intensity, and time of intervention measures on the case timeline. The system automatically calls the causal model to perform probability intervention calculations and displays the comparison between the counterfactual outcome and the actual outcome in real time. At the same time, the module calls multiple versions of real intervention control materials in the version difference closed-loop storage module as benchmark references and automatically evaluates the editing distance between the student's reasoning path and the expert's path, realizing a paradigm upgrade from passive case reading to active causal reasoning training. Attached Figure Description
[0032] Figure 1 This is a system structure block diagram of the present invention;
[0033] Figure 2 This is a flowchart of the epidemiological model validation and completion module of the present invention;
[0034] Figure 3 This is a flowchart of the differential closed-loop storage module of the present invention. Detailed Implementation
[0035] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0036] Please see Figures 1 to 3 As shown, the embodiments of the present invention provide the following technical solutions:
[0037] A public health case database management system has been systematically reconstructed from the underlying storage architecture to the upper-level application functions. It consists of six major functional modules: an epidemiological model validation and completion module, a dynamic knowledge network construction module, a latent homology clustering and desensitized storage module, a case version differential closed-loop storage module, a counterfactual inference teaching module, and a hierarchical computing power scheduling module. The specific implementation methods of each module are described in detail below:
[0038] 1. Epidemiological model validation and completion module
[0039] The epidemiological model validation and completion module is the data quality control entry point of the system. It is used to perform logical consistency verification and missing data inference and completion on the entered public health case data, and output the validated and completed standardized case data. This module solves the shortcomings of the existing system, which can only perform format verification and cannot identify logical contradictions in epidemiology.
[0040] Specifically, the epidemiological model validation and completion module embeds a dynamic pathogen knowledge base, storing the incubation period range, passage interval, and effective reproduction number for various infectious diseases. This knowledge base stores core epidemiological parameters for multiple infectious diseases in a structured format, including but not limited to: reference ranges for the incubation period range (composed of the shortest and longest incubation periods), passage interval, and effective reproduction number. Data sources for the knowledge base include authoritative textbooks such as *Infectious Diseases* and *Epidemiology*, technical guidelines issued by the World Health Organization, prevention and control plans issued by the National Center for Disease Control and Prevention, and the latest peer-reviewed research literature. The knowledge base supports regular online updates; when a new infectious disease is identified or the epidemiological parameters of a known infectious disease are re-estimated, the administrator can enter or import new parameter ranges into the knowledge base.
[0041] When a new public health case is entered into the system, the verification and completion module first extracts information from the case, identifying the time sequence and geographical location information it contains. The time sequence includes, but is not limited to: exposure time, onset time, first visit time, diagnosis time, hospitalization time, recovery or death time, etc. For cases with a complete time sequence, the module automatically generates temporal logical constraints based on epidemiological parameters stored in the pathogen knowledge base. For example, for a confirmed case, the system generates constraints such as: onset time - exposure time = incubation period interval; death time - onset time > 0, etc. When multiple time points form a complete logical chain, the system constructs all constraints into a constraint satisfaction problem and calls the SMT solver to determine satisfiability.
[0042] The SMT solver is an automated reasoning tool that can determine the satisfiability of first-order logic formulas. In this embodiment, the Z3 solver is preferred. For case data items that the SMT solver determines do not satisfy the constraints, the module automatically marks them as logical contradictions and provides correction suggestions based on the reverse derivation of the constraints. For example, if the solver determines that the onset date in a case is later than the death date, it provides a correction suggestion to check the onset date or death date and lists the most recent feasible value that satisfies the constraints as a reference.
[0043] For cases lacking key time points, such as missing the exact exposure date but knowing the onset date, the module initiates a missing data inference and completion mechanism. This mechanism utilizes a Bayesian structured learning model, using the incubation period distribution of cases in the same region and during the same period as prior, and obtains the maximum posterior estimate of the missing node by solving the following formula. and confidence interval:
[0044] ;
[0045] in, This is a key time point that needs to be filled in. For the time series data of the observed cases, Let be the likelihood function, representing the likelihood function in a given situation. Data observed under these circumstances The probability, This is a prior probability density constructed based on the distribution of incubation periods of cases in the same region during the same period. All data items that have been inferred and completed are labeled as inferred values to distinguish them from the directly collected raw data, making it easier for subsequent users to understand their uncertainty when referencing them. Through the above verification and completion process, the epidemiological logical consistency and completeness of the case data are ensured, providing a high-quality data foundation for subsequent knowledge network construction and inferential analysis.
[0046] 2. Dynamic Knowledge Network Construction Module
[0047] The dynamic knowledge network construction module performs multi-granular semantic annotation on the standardized case data output by the epidemiological model validation and completion module, extracts entities and their relationships, and constructs and dynamically maintains a public health event knowledge network. This module solves the technical problem of isolated entries and lack of deep semantic connections in traditional case databases.
[0048] Specifically, the underlying layer of the dynamic knowledge network construction module adopts a native fusion index kernel for heterogeneous data anchored to cross-modal case characteristics. Public health case data sources are extremely diverse, including but not limited to: structured infectious disease reporting forms from CDC centers, electronic medical records from hospital information systems, on-site audio and video recordings collected during epidemiological investigations, scanned copies of handwritten paper investigation ledgers, and sensor data streams from environmental disinfection records. Existing conventional solutions require these heterogeneous data to be pre-converted into a unified format. This post-processing inevitably loses detailed features of the original data; for example, modification marks in handwritten ledgers may reflect the verification process of frontline personnel for certain information, and this information will be permanently lost after format conversion.
[0049] The native fusion index kernel proposed in this invention adopts a completely different technical approach. First, based on the characteristics of public health operations, a set of core treatment signs are predefined as anchoring benchmarks. These core treatment signs include, but are not limited to, the following six dimensions: case demographic characteristics, clinical manifestations, exposure factors, time course, spatial distribution, and intervention measures. For any modal raw data, the system does not perform format conversion, but directly attaches it in its native form to the most relevant treatment sign anchor point. Taking a field investigation recording as an example, after extracting the timestamp and spatial location information, the system natively attaches the complete audio file to the corresponding case's time course and spatial distribution anchor points. At the same time, an index record is generated under this anchor point, containing the data modality type, time range, spatial coordinates, and population association identifier. The index records under all anchor points are aggregated into a cross-modal parallel dedicated index queue. During queries, users can quickly locate all related native data through any anchor point.
[0050] In addition, the index kernel automatically aligns the three benchmark dimensions of timestamps, spatial locations, and population associations of multi-source data, enabling temporal matching and spatial registration of different modal data in terms of collection time and location. It natively solidifies the coupling and association logic between heterogeneous data, enabling high-fidelity fusion to be completed as soon as the data is entered into the database, thus eliminating distortion problems caused by post-processing.
[0051] At the knowledge network construction level, the module performs semantic annotation on standardized case data at three granularities: word segmentation annotation, which identifies medical entities in the text, including but not limited to pathogen names, drug names, symptom descriptions, and population attributes; sentence-level annotation, which classifies each paragraph by function, identifying whether the paragraph belongs to a description of clinical manifestations, epidemiological investigation records, laboratory test results, or description of prevention and control measures; and chapter-level annotation, which generates a structured summary of the case and extracts the core event chain of the case.
[0052] Based on the above annotation results, the module automatically extracts the temporal and causal relationships between entities and constructs a multi-dimensional knowledge graph with four types of nodes: case, event, intervention, and outcome. Each node in the graph carries a timestamp, spatial coordinates, and population attribute dimension labels. When a new case is entered and verified, the module calculates the semantic similarity between the new case and existing nodes in the knowledge graph. If the similarity exceeds a preset threshold, the new case is fused with similar nodes to strengthen the weight of the corresponding edges. If a new entity type or a new relationship type is detected, the graph pattern layer is automatically expanded to achieve self-growth and updating of the knowledge network.
[0053] 3. Latent homology association clustering and desensitized storage module
[0054] Used to generate hierarchical materialized snapshots of the original fields of the core traceability of cases and simultaneously generate an irreversible privacy-de-identified derived index layer. Based on the three implicit dimensions of pathogen gene traceability fragments, population flow intersection trajectory and environmental point overlap features, the de-identified derived index layer is automatically clustered and sealed to preserve cross-temporal implicit homology case chains.
[0055] Specifically, this module performs hierarchical storage processing on the standardized case data processed by the epidemiological model validation and completion module. First, it extracts the core original fields from the cases related to personal privacy and source tracing analysis, generating a separate materialized snapshot. This materialized snapshot is stored in a separate storage partition using physical isolation, logically and physically isolated from the front-end query interface. This means that regular front-end queries cannot access the original content of this materialized snapshot. Access to the materialized snapshot is subject to strict permission control and audit log monitoring, and it can only be unlocked under specific authorized scenarios.
[0056] While generating the materialized snapshot, the module performs irreversible privacy-de-identifying processing on the aforementioned core original fields, generating a de-identified derived index layer. This derived index layer has two key characteristics: first, irreversibility, meaning the de-identification process uses a one-way hash function or differential privacy injection mechanism to ensure that the specific values of the original sensitive data cannot be derived from the derived index layer; second, query availability, meaning the de-identified data retains certain statistical characteristics and similarity calculation capabilities. The derived index layer is stored in a storage area that can be directly accessed by regular queries, allowing other modules of the system to call it.
[0057] Specifically, for gene source tracing fragments, the desensitization process extracts single nucleotide polymorphism (SNP) feature vectors from the whole genome sequence and maps them to irreversible hash signatures, enabling the system to calculate the homology distance between pathogens in different cases without exposing the original sequence. For overlapping trajectories of population flows, the desensitization process converts precise GPS coordinates into statistical region codes, such as mapping latitude and longitude coordinates to 500m × 500m grid cell numbers, and adding noise perturbations that satisfy differential privacy to the grid numbers to generate desensitized trajectory sequences. For overlapping environmental locations, the desensitization process replaces specific building and location names with functional type codes, retaining location type information but removing unique identifiers.
[0058] The module's backend continuously runs a latent homology association clustering engine, which operates on the de-identified derived index layer and performs automatic clustering analysis around three latent dimensions:
[0059] In terms of pathogen gene tracing, the pathogen hash feature codes of different cases are compared to calculate the homology distance, and cases with a homology distance below the threshold are grouped into the same potential homology cluster.
[0060] Based on the dimension of population flow intersection trajectory, spatiotemporal co-occurrence analysis is performed on the desensitized trajectory sequences of different cases to identify whether there is spatial grid overlap within the time window or whether the same grid is visited successively.
[0061] By comparing the functional type coding sequences of places visited by different cases, environmental sites with overlapping dimensions can be identified as commonly exposed environmental factors.
[0062] When a group of cases simultaneously meets clustering criteria across at least two latent dimensions, the engine automatically marks them as a spatiotemporal latent homology chain and creates an index record for this chain in the desensitized derived index layer. The chain index record includes the de-identified ID of each case within the chain, the type of association dimension, a comprehensive score of association strength, and the timestamp when the association was discovered. Thus, during data entry and storage, the system can automatically discover and preserve latent homology chains spanning years, districts, and even diseases without relying on manual annotation or prior assumptions, truly achieving a balance between privacy protection and source tracing analysis capabilities.
[0063] 4. Case version differential closed-loop storage module
[0064] It is used to perform version-by-version differential storage of case data generated from multiple rounds of prevention and control interventions under the same public health event, generate multiple versions of dynamic case differential copies carrying the loss coefficient of intervention effectiveness, and form a closed-loop storage of the entire chain evolution trajectory of prevention and control strategy trial, optimization and finalization.
[0065] Specifically, the implementation of the case version differential closed-loop storage module is as follows: This module establishes real-time data channels with the frontline health emergency command system, hospital information system, and grassroots grid management platform through data interfaces. When a public health event triggers an emergency response, the module initiates an event-level version tracking mechanism: it creates a version chain container for the event, initializing the version number to the original state snapshot at the time of the event's first report. Subsequently, as prevention and control work progresses, the module automatically triggers the generation of a version snapshot whenever a prevention and control intervention action is detected.
[0066] The data collection for version snapshots is multi-dimensional: the system captures on-site operational logs, resource scheduling receipts, and population control feedback data second by second. These multi-source real-time data are then timestamped and aligned with the epidemiological data of the corresponding case to form the complete content of that version snapshot. It is worth noting that a differential storage strategy is used between adjacent version snapshots, meaning only data fields that have changed compared to the previous version are recorded, rather than generating a complete copy of the data each time. This strategy significantly reduces storage resource consumption, and is particularly suitable for major epidemic cases with long durations and frequent version iterations.
[0067] Each time a snapshot is generated, the module invokes an embedded health intervention time-series game weight model to quantitatively evaluate the effectiveness of the intervention actions in that version. This model treats a single round of intervention as a game between the control measures and the spread of the epidemic: the intervention invests certain social resources in exchange for a reduction in the intensity of epidemic transmission. The model uses the following formula for calculation:
[0068] ;
[0069] in, For the first The effectiveness loss coefficient of the intervention action. and These represent the effective reproduction number before and after implementing this version of the intervention. The effective reproduction number can be estimated using the classic EpiEstim method or the maximum likelihood estimation method based on generational time distribution. This is the normalized value of the resource input cost for this intervention. Taking into account direct economic costs, social costs, and human resource input, all cost dimensions are weighted, summed, and normalized to the [0, 1] interval to facilitate horizontal comparison between different versions.
[0070] The numerical value has a clear practical meaning: when When this occurs, it indicates that the intervention has achieved positive results and the intensity of the epidemic's spread has decreased; The higher the value, the higher the effectiveness of epidemic control achieved per unit of cost input; when When the value is close to 0 or negative, it indicates that the intervention was not effective or that the epidemic was rising instead of falling, suggesting that the strategy needs to be adjusted.
[0071] Through the aforementioned mechanism, the case version differential closed-loop storage module retains the complete prevention and control process of a public health event, from the initial response to the final closure, in a closed-loop manner as multiple differential copies with effectiveness annotations. After the final closure, the entire version chain forms a clear strategy evolution trajectory: from early exploratory responses to mid-term strategy corrections and optimizations, and finally to the final strategy formulation. This full-process retention design provides crucial process data that traditional static case libraries cannot offer for subsequent dynamic benchmarking and replication of emergency situations and for the debriefing and analysis of emerging and re-emerging epidemics.
[0072] 5. Counterfactual reasoning teaching module
[0073] It is used to perform counterfactual outcome deduction based on causal models and visualize the results, based on the user's modification of the intervention measures in the case.
[0074] Specifically, this module first automatically extracts the causal structure of complete event chain cases from the case version differential closed-loop storage module. It employs a PC causal discovery algorithm that incorporates prior knowledge from the public health domain: the domain prior knowledge is input into the algorithm in the form of a whitelist and a blacklist; the whitelist specifies the causal edge directions that must exist according to epidemiological principles, while the blacklist specifies impossible causal edges that violate medical common sense. Under the constraints of the domain prior knowledge, the algorithm performs conditional independence tests on the variables in the case data, gradually constructing a causal graph containing five types of nodes: exposure, infection, transmission, intervention, and control. Based on the structure of this causal graph, the module uses historical data to fit the structural equation parameters corresponding to each causal edge, obtaining a complete structural equation model and quantifying the effect of each variable on downstream nodes.
[0075] This module also provides a visual, interactive simulation interface. The core interactive area of this interface is a horizontal timeline, marking the occurrence times of key events in the real-world case and the initiation times of various intervention measures. Below the timeline, the actual epidemic development is presented as an epidemiological curve. Users can select any intervention node on the timeline by dragging and dropping with the mouse or using touch controls, modifying its three attributes: intervention type (e.g., changing advisory health education to mandatory school closures); intervention intensity (e.g., changing the closure of 50% of schools to the closure of all schools); and intervention initiation time. After the user completes the modifications, the module invokes a structural equation model, using the user's intervention modifications as conditions, to perform probabilistic intervention calculations on the intervention nodes in the causal graph. Specifically, the intervention variable is forcibly set to the user-specified value, the dependency relationship between the variable and its original parent node is broken, and the probability distribution of each downstream variable is recalculated on the modified causal graph, ultimately obtaining an estimate of the epidemic development under this counterfactual intervention condition. The system displays the real-world outcome curve and the counterfactual outcome curve side-by-side in the same coordinate system, distinguished by different colors or line types, providing a visually appealing comparison.
[0076] In public health teaching scenarios, this module also features automatic evaluation. Teachers can pre-define one or more expert deduction paths as evaluation benchmarks based on expert assessments, storing them in the system's expert path library. An expert deduction path is defined as a sequence of intervention parameter combinations selected by experts at several key decision nodes along the case timeline. After a student completes a set of deduction operations, the module records the student's entire operation sequence, including the intervention type, intensity, and time selection at each decision node, forming the student's deduction path. The module uses an edit distance algorithm to compare the student's deduction path with the pre-defined expert deduction paths. Edit distance is defined as the minimum number of operations required to convert the student's path into an expert path; operation types include modifying intervention parameters, adding intervention nodes, and deleting intervention nodes. Based on the edit distance, the module generates evaluation feedback: if the edit distance is below a preset threshold, the decision logic is deemed reasonable; if the edit distance is high, the module identifies the decision node with the greatest discrepancy, provides targeted feedback, and guides the student to reflect on their decision-making bias.
[0077] During simulation and evaluation, the counterfactual simulation teaching module uses multiple versions of dynamically differentiated case copies stored in the case version differential closed-loop storage module as benchmarks. Specifically, the system uses the intervention actions and their effectiveness loss coefficients for each version in the real case version chain. Marked at corresponding positions on the timeline, students can browse various versions of strategies and their quantified effectiveness during the actual epidemic prevention and control process before the simulation, and then compare their own simulation strategies with them. This feature allows students not only to see what ultimately happened, but also to see how decision-makers in the real world went through trial and error and adjustment processes, thus gaining an immersive public health decision-making training experience.
[0078] 6. Tiered computing power scheduling module
[0079] The core source tracing cases are dynamically divided into dedicated computing power pools and ordinary archived cases rate limiting pools according to the public health emergency response level; during the first level of emergency response, the core cases are triggered to be instantly sealed offline at the hardware level to avoid data damage caused by network peak congestion.
[0080] This module maintains a mapping table between public health emergency response levels and computing power scheduling strategies. Response levels are classified according to the grading standards of the "National Emergency Response Plan for Public Health Emergencies," and are categorized into four levels: general, relatively serious, major, and extremely serious, based on the scope of the event, the number of cases, and the degree of harm. Response levels can be manually set by the system administrator or automatically adjusted by monitoring external authoritative data sources.
[0081] The module divides the system's underlying computing resources into two logical resource pools: a dedicated computing power pool for core source tracing cases and a rate-limited pool for ordinary archived cases. The dedicated computing power pool for core source tracing cases is allocated a fixed percentage of the system's total computing power, specifically ensuring the computing power needs of core tasks such as SMT solving operations for the epidemiological model verification and completion module, background comparison calculations for the implicit homology clustering engine, and causal model reasoning for the counterfactual inference teaching module. The rate-limited pool for ordinary archived cases is used for daily case queries, routine statistical analysis, and non-urgent research tasks; its computing power limit is dynamically constrained in emergency situations.
[0082] During normal operation, dynamic borrowing is allowed between the two computing power pools. When the core pool has sufficient computing power, surplus computing power can be temporarily allocated to the ordinary pool to accelerate archiving task processing. When the response level is upgraded to Level 1, the module triggers an emergency scheduling strategy: First, it automatically disconnects all low-priority access links that are not for scientific research or analysis; second, it reduces the computing power limit of the ordinary archiving case rate-limiting pool to 20% of the normal level, and all the released computing power is supplemented to the dedicated computing power pool for core source tracing cases, ensuring that the computing power of the entire domain is tilted towards the core business of source tracing analysis, emergency analysis, and command and dispatch. At the same time, the module triggers the hardware-level offline instantaneous sealing mechanism for core cases: the materialized snapshot data of core source tracing cases is synchronously solidified from the online storage medium and written to the non-erasable offline storage medium. This operation is completed before the peak network traffic surge. The sealed core case copy cannot be tampered with or crawled in batches by network attacks at the physical level, effectively avoiding the risk of database congestion, downtime, or even data corruption caused by high-concurrency access in emergency situations. When the response level is downgraded from Level 1, the module automatically releases its offline storage state, gradually restores the computing power limit of the ordinary pool, and the system returns to normal operation mode.
[0083] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0084] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
[0085] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A public health case database management system, characterized in that, include: The epidemiological model validation and completion module is used to perform logical consistency checks and missing data inference and completion on the entered public health case data, and output standardized case data. The dynamic knowledge network construction module is used to perform multi-granular semantic annotation on the standardized case data, extract entities and relationships, and construct a public health event knowledge network. The implicit homology association clustering and desensitization storage module is used to generate hierarchical materialized snapshots of the original fields of the core traceability of cases and simultaneously generate an irreversible privacy desensitization derivative index layer. Based on the three implicit dimensions of pathogen gene traceability fragments, population flow intersection trajectory and environmental point overlap features, the module automatically clusters and stores the cross-temporal implicit homology association case chain on the desensitization derivative index layer. The case version differential closed-loop storage module is used to perform version differential storage of case data generated by multiple rounds of prevention and control interventions under the same public health event, generate multiple versions of dynamic case differential copies carrying the loss coefficient of intervention effectiveness, and form a closed-loop storage of the whole-link evolution trajectory of prevention and control strategy trial, optimization and finalization. The counterfactual deduction teaching module is used to perform counterfactual outcome deduction based on causal models and visualize the results, based on the user's modification of the intervention measures in the case.
2. The public health case database management system according to claim 1, characterized in that, The epidemiological model validation and completion module embeds a dynamic pathogen knowledge base, storing the incubation period intervals, passage intervals, and effective reproduction numbers for various infectious diseases. It generates temporal logic constraints for case timelines and location information, calls the SMT solver to verify consistency, marks data items that do not meet the constraints, and provides correction suggestions. For cases with missing key time nodes, a Bayesian structured learning model is used, with the incubation period distribution of cases in the same region and period as prior, to obtain the maximum posterior estimate of the missing node by solving the following formula. and confidence interval: ; in, This is a key time point that needs to be filled in. For the time series data of the observed cases, Let be the likelihood function. This is a prior probability density constructed based on the distribution of incubation periods of cases in the same region during the same period, and the completed data is labeled as inferred values.
3. The public health case database management system according to claim 1, characterized in that, The underlying layer of the dynamic knowledge network construction module adopts a native fusion index kernel for heterogeneous data anchored by cross-modal case vital signs: taking the core public health treatment vital signs as the anchoring benchmark, it does not perform format conversion on heterogeneous data of different modalities, but directly attaches the corresponding vital sign anchor points in the native form to generate a cross-modal parallel dedicated index queue. The kernel automatically aligns the timestamps, spatial locations, and population-related three-dimensional benchmarks of multi-source data, and natively solidifies the coupling and association logic of heterogeneous data.
4. The public health case database management system according to claim 1, characterized in that, The specific implementation of the implicit homology clustering and desensitization storage module is as follows: a hierarchical materialized snapshot is generated and sealed for the original core traceability fields of the case, and an irreversible privacy desensitization derivative index layer is generated simultaneously; on the desensitization derivative index layer, based on three implicit dimensions—pathogen gene traceability fragments, population flow intersection trajectories, and environmental point overlap features—the case chain of implicit homology associations across time and space is automatically clustered and sealed.
5. The public health case database management system according to claim 1, characterized in that, The specific implementation of the case version differential closed-loop storage module is as follows: following multiple rounds of prevention and control intervention actions, it captures real-time data such as on-site operation logs, resource scheduling receipts and population control feedback every second, and automatically generates multiple versions of dynamic case differential copies under the same public health event. An embedded health intervention time-series game weight model is used to annotate the effectiveness loss coefficient of intervention actions version by version, and to record the full-link evolution trajectory of closed-loop retention strategy trial and error, optimization and finalization; the specific health intervention time-series game weight model is as follows: ; in, For the first The effectiveness loss coefficient of the intervention action. and These represent the effective reproduction numbers before and after implementing this version of the intervention. This is the normalized value of the resource input cost for this intervention; this coefficient is annotated in the corresponding version of the differential copy. The larger the value, the higher the effectiveness of epidemic control achieved per unit cost.
6. The public health case database management system according to claim 1, characterized in that, The counterfactual deduction teaching module uses a causal discovery algorithm to extract a causal graph of exposure, infection, transmission, intervention, and control from a complete event chain case, and fits structural equation model parameters based on historical data; it provides an interactive interface to receive user modifications to the type, intensity, or time of intervention measures on the case timeline, calls the structural equation model, and performs probabilistic intervention calculations on relevant variables in the causal graph with the intervention operation as a condition, and displays the counterfactual outcome.
7. The public health case database management system according to claim 6, characterized in that, The counterfactual deduction teaching module also records the user's deduction operation sequence, compares the edit distance with the preset expert deduction path, evaluates the rationality of the decision, and outputs feedback; when performing deduction, the counterfactual deduction teaching module calls the multiple versions of dynamic case differential copies stored in the case version differential closed-loop storage module as a benchmark reference.
8. The public health case database management system according to claim 1, characterized in that, It also includes a tiered computing power scheduling module, which dynamically divides the core tracing case dedicated computing power pool and the ordinary archive case rate limiting pool according to the public health event response level; during the first level emergency response, it triggers the hardware-level offline instantaneous sealing of core cases to avoid data damage caused by network peak congestion.