A method and system for tracing the source of infection in the population
By integrating multiple data sources and intelligent algorithms, the problem of low efficiency in traditional infectious disease tracing methods has been solved, enabling efficient and accurate epidemic monitoring and resource optimization, rapid location of potential sources of infection, and reduction of the risk of epidemic spread.
Patent Information
- Application Number
- CN202510329488.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-03-20
AI Technical Summary
Traditional methods of tracing the source of infectious diseases rely on manual investigation, which is inefficient and prone to information delays and analytical errors, making it difficult to comprehensively track complex transmission paths.
By integrating data from multiple sources, such as medical institutions, public health departments, and social media, and performing data cleaning and standardization, machine learning and infectious disease dynamics models are used to calculate the number of infections and high-risk areas. By combining spatiotemporal weighting factors and multidimensional condition screening, potential sources of infection can be accurately identified.
It improves the accuracy and efficiency of infectious disease tracing, enables real-time monitoring of the epidemic situation, accurate prediction of the number of infections and high-risk areas, rapid location of potential sources of infection, optimization of resource allocation, and reduction of the risk of epidemic spread.
Smart Images

Figure CN120221123B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of epidemic surveillance technology, and more specifically, to a method and system for tracing the source of infection in a population. Background Technology
[0002] In modern society, the prevention and control of infectious diseases has increasingly become a key issue in global public health. Especially in the early stages of emerging infectious diseases or outbreaks, rapid and effective source tracing is crucial for preventing the spread of the epidemic and protecting public health. However, traditional source tracing methods rely on manual investigation and on-site sampling, which is not only inefficient but also presents numerous challenges. First, manual investigation requires a significant amount of human resources, which are often limited in the face of large-scale outbreaks. Second, manual data processing is prone to information delays and analytical errors. Furthermore, the transmission routes of many infectious diseases are complex, making comprehensive tracing difficult solely through manual methods.
[0003] In recent years, the rapid development of big data technology has provided new solutions for tracing the origins of infectious diseases. By integrating multiple data sources, including electronic health records from medical institutions, public health reporting systems, social media activity, and geolocation data from mobile devices, more detailed and real-time epidemic information can be obtained. Such data integration requires efficient data cleaning and standardization processes to ensure data quality and consistency. Simultaneously, the application of artificial intelligence and machine learning technologies makes it possible to extract valuable patterns and trends from massive amounts of data, providing strong support for accurate epidemic prediction and the identification of high-risk areas.
[0004] In mathematical modeling, with the development of infectious disease dynamics models, model calibration and parameter estimation combined with real data can more accurately reflect the dynamics of epidemic transmission. Such models can not only be used to predict the number of infections and epidemic trends, but also to assess the potential impact of different intervention measures by simulating them, providing a scientific basis for decision-makers. Furthermore, network-based transmission models can be used to analyze complex transmission chains and identify key transmission nodes, thereby assisting in the rapid location of potential sources of infection. Summary of the Invention
[0005] In view of this, the present invention proposes a method and system for tracing the source of infection in the population, aiming to solve the problem of low efficiency of current manual analysis and tracing methods.
[0006] This invention proposes a method for tracing the source of infection in a population, the method comprising:
[0007] Data is collected from the data source, cleaned, and then sent to the traceability module and stored in the data warehouse.
[0008] The number of infected people at the current and future time points is calculated using the cleaned data, and high-risk areas are identified. The most critical high-risk areas are set as super-risk areas, and the super-risk areas are divided into multiple equally divided preset nodes. The probability of each preset node being a potential source of infection is calculated, and the probability is corrected by a spatiotemporal weighting factor. The corrected probability is compared with the preset probability, and the node where the potential source of infection is located is determined based on the comparison result.
[0009] After obtaining the node where the potential source of infection is located, multi-dimensional conditions are used to screen for potential sources of infection, and the potential source of infection is determined based on the screening results and the order of onset of the population at the node where the potential source of infection is located.
[0010] Furthermore, the process of collecting data from the data source, cleaning the data, and then sending it to the traceability module specifically includes:
[0011] Data is collected from data sources, including hospitals, public health departments, social media, and mobile devices;
[0012] Clean the data, identify and process missing values, duplicate data and outliers, and convert the data into a consistent format and units;
[0013] Record background information for the data, including data source, update time, and processing history.
[0014] Furthermore, the data warehouse includes: raw data, the raw data cleaning process, and data sent to the traceability module.
[0015] Furthermore, the identification of high-risk areas, including designating the most critical high-risk areas as super-risk areas, includes:
[0016] The infection growth rate is calculated based on the number of infections at the current and future time points. The region is divided into high, medium and low risk levels based on the infection growth rate. Information on infected persons and contacts in all high-risk areas is entered into the source tracing module. Each individual high-risk area is treated as an infection node, and contact events are treated as edges to determine super-risk areas.
[0017] Furthermore, the formula for calculating the number of infections at the future time point is as follows:
[0018] ;
[0019] Where I(t) represents the number of infected people at time t, β is the transmission coefficient, S(t) is the number of susceptible people, γ is the recovery coefficient, and N is the total population.
[0020] Furthermore, the method for determining the super-risk zone is as follows:
[0021] Calculate the frequency with which an infected node appears on the shortest path between all other pairs of infected nodes; the formula is as follows:
[0022] ;
[0023] Where, σ st σ is the total number of shortest paths from infected node s to infected node t. st (v) is through infected nodes The number of paths, C B (v) is the C of node v. B value;
[0024] Compare the C values of the infection nodes described in all high-risk areas. B The value of C B The high-risk area represented by the highest infection node is the super-risk area.
[0025] Furthermore, calculating the probability of each preset node being a potential source of infection specifically includes:
[0026] ;
[0027] Where P(j) represents the probability that node j is the source of infection, and W ij Let T be the weight from node i to node j. ij W represents the propagation time between node i and node j. ij For custom settings; n is the total number of nodes, T i Let be the propagation time of the i-th node.
[0028] Furthermore, the above probabilities are corrected using a spatiotemporal weighting factor, and the corrected probabilities are compared with preset probabilities. Based on the comparison results, the specific nodes where potential sources of infection are located are determined, including:
[0029] ;
[0030] Where P(j) represents the probability that node j is the source of infection. Let represent the corrected probability that node j is the source of infection. It is a time weighting factor. This is the spatial weighting factor; both the temporal weighting factor and the spatial weighting factor are custom settings.
[0031] In all nodes, The node with the largest value is the potential source of infection.
[0032] Furthermore, the step of screening potential sources of infection using multidimensional conditions and determining potential sources of infection based on the screening results and the order of onset of the disease among the population at the node where the potential source of infection is located specifically includes:
[0033] Screening is based on the nature of the case type, with pre-set weights for confirmed cases (k1), suspected cases (k2), asymptomatic infections (k3), and ordinary contacts (k4).
[0034] The results of the property-based screening based on case type are as follows:
[0035] ;
[0036] Screening is based on close contact chains, and the strength of close contact relationships, Q, is defined. ij Weights are assigned based on contact frequency and contact duration; where Q ij =f(F ij E ij );
[0037] The screening results based on close contact chains are as follows:
[0038] ;
[0039] Screening is based on epidemiological investigation and contact history. Di=1 is pre-defined to indicate direct contact with confirmed cases, and Di<1 indicates indirect contact with confirmed cases.
[0040] The screening results based on contact history are as follows:
[0041] ;
[0042] Therefore, the result of the multi-dimensional screening is S=P1∩P2∩P3;
[0043] Where, k x k represents the case type weight. min Indicates the minimum threshold; F ij Indicates contact frequency; E ij Indicates the duration of contact;
[0044] In the multidimensional screening results S, the results are arranged according to the order of onset, and the earliest onset time is identified as the potential source of infection.
[0045] On the other hand, the present invention also proposes an infection tracing system, applied to the above-mentioned infection tracing method, the system comprising:
[0046] The data acquisition module is configured to collect data from the data source, clean the data, and then send it to the data traceability module and store it in the data warehouse.
[0047] The source tracing module is configured to use cleaned data to calculate the number of infected people at the current and future time points, identify high-risk areas, set the most critical high-risk areas as super-risk areas, divide the super-risk areas into multiple equally divided preset nodes, calculate the probability of each preset node as a potential source of infection, and correct the above probability through a spatiotemporal weight factor. The corrected probability is compared with the preset probability, and the node where the potential source of infection is located is determined based on the comparison result.
[0048] The allocation module is configured to, after obtaining the node where the potential source of infection is located, perform multi-dimensional screening of the potential source of infection and determine the potential source of infection based on the screening results and the order of onset of the population at the node where the potential source of infection is located.
[0049] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention provides a method and system for tracing the source of infected populations based on multi-source data fusion and intelligent algorithms, which can effectively improve the accuracy and efficiency of infectious disease tracing. By collecting data from multiple data sources such as hospitals, public health departments, social media, and mobile devices, and then cleaning and analyzing it, the dynamics of the epidemic can be monitored in real time. Based on this, using transmission mathematical models and intelligent algorithms, the number of current and future infections can be accurately predicted, and high-risk areas can be identified. Through precise division and node analysis of super-risk areas, potential sources of infection can be quickly located, thus providing an important basis for timely epidemic control. In addition, based on the geographical location of potential sources of infection, the system can select appropriate resource allocation schemes, optimize the allocation and use of public health resources, and reduce the risk of epidemic spread. This comprehensive method not only improves the scientific nature and accuracy of tracing, but also provides a more intelligent tool for public health management. Moreover, this invention innovatively combines multi-source data fusion with intelligent algorithms, providing an efficient and accurate solution for the tracing and control of infectious diseases. First, by integrating multiple data sources, the system can obtain comprehensive epidemic information in real time, covering a full range of data perspectives from detailed information on individual infected individuals to macro-epidemic trends. This integration not only improves the timeliness of information but also significantly enhances the depth and breadth of data analysis. Secondly, the system utilizes advanced machine learning algorithms and mathematical models to accurately predict the development of an epidemic. Through training with historical data and dynamic model adjustments, it can precisely identify high-risk areas and potential outbreak points. This accurate predictive capability enables public health departments to deploy prevention and control measures in advance to prevent the spread of the epidemic. Furthermore, the system can guide targeted intervention strategies by identifying key nodes and super-spreaders in the transmission chain, thereby reducing the overall risk of transmission. In terms of resource allocation, the system can optimize the allocation of public health resources based on the geographical location of potential sources of infection and the severity of the epidemic. This intelligent resource management not only improves the response speed to sudden outbreaks but also effectively reduces unnecessary resource waste. Attached Figure Description
[0050] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0051] Figure 1 This is a flowchart of an infection tracing method according to an embodiment of the present invention.
[0052] Figure 2 This is a functional block diagram of an infected population tracing system according to an embodiment of the present invention. Detailed Implementation
[0053] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the disclosure to those skilled in the art. It should be noted that, unless otherwise specified, embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0054] In modern society, the prevention and control of infectious diseases has increasingly become a key issue in global public health. Especially in the early stages of emerging infectious diseases or outbreaks, rapid and effective source tracing is crucial for preventing the spread of the epidemic and protecting public health. However, traditional source tracing methods rely on manual investigation and on-site sampling, which is not only inefficient but also presents numerous challenges. First, manual investigation requires a significant amount of human resources, which are often limited in the face of large-scale outbreaks. Second, manual data processing is prone to information delays and analytical errors. Furthermore, the transmission routes of many infectious diseases are complex, making comprehensive tracing difficult solely through manual methods.
[0055] In recent years, the rapid development of big data technology has provided new solutions for tracing the origins of infectious diseases. By integrating multiple data sources, including electronic health records from medical institutions, public health reporting systems, social media activity, and geolocation data from mobile devices, more detailed and real-time epidemic information can be obtained. Such data integration requires efficient data cleaning and standardization processes to ensure data quality and consistency. Simultaneously, the application of artificial intelligence and machine learning technologies makes it possible to extract valuable patterns and trends from massive amounts of data, providing strong support for accurate epidemic prediction and the identification of high-risk areas.
[0056] In mathematical modeling, with the development of infectious disease dynamics models, model calibration and parameter estimation combined with real data can more accurately reflect the dynamics of epidemic transmission. Such models can not only be used to predict the number of infections and epidemic trends, but also to assess the potential impact of different intervention measures by simulating them, providing a scientific basis for decision-makers. Furthermore, network-based transmission models can be used to analyze complex transmission chains and identify key transmission nodes, thereby assisting in the rapid location of potential sources of infection.
[0057] See Figure 1 As shown in the figure, this embodiment of the invention provides a method for tracing the source of an infected population, the method comprising:
[0058] S1: Collect data from multiple data sources, clean the data, send it to the traceability module, and store it in the data warehouse;
[0059] S2: Calculate the number of infected people at the current and future time points using the cleaned data, identify high-risk areas, set the most critical high-risk areas as super-risk areas, divide the super-risk areas into multiple equally divided preset nodes, calculate the probability of each preset node as a potential source of infection, and correct the above probability through a spatiotemporal weight factor. Compare the corrected probability with the preset probability, and determine the node where the potential source of infection is located based on the comparison results.
[0060] Understandably, the initial probability is adjusted to account for changes in time and space. Spatiotemporal weighting factors can include the following aspects: Time factors: adjusting the probability based on activity characteristics at different times (e.g., rush hours, holidays, etc.). Spatial factors: considering distances between nodes, transportation connectivity, etc., adjusting the propagation potential of nodes in different locations.
[0061] S3: After obtaining the node where the potential source of infection is located, perform multi-dimensional screening of potential sources of infection and determine the potential source of infection based on the screening results and the order of onset of the population at the node where the potential source of infection is located.
[0062] In this preferred embodiment, collecting data from multiple data sources, cleaning the data, and sending it to the traceability module specifically includes:
[0063] Data was collected from multiple data sources, including hospitals, public health departments, social media, and mobile devices.
[0064] Clean the data, identify and process missing values, duplicate data and outliers, and convert the data into a consistent format and units;
[0065] Record background information for the data, including data source, update time, and processing history.
[0066] Understandably, the first step involves collecting relevant data from multiple sources, including but not limited to hospital records, public health department reports, social media platform information, and mobile device location information. The collected data may contain noise and inconsistencies, necessitating data cleaning to remove duplicates, incompleteness, or errors. This step ensures high-quality input data, forming the foundation for subsequent analysis. The cleaned data is then transmitted to the source tracing module and stored in a dedicated data warehouse for easy access and analysis. After data cleaning, the system uses statistical models and predictive algorithms to calculate the number of infections at current and future points in time. This prediction helps identify high-risk areas—regions with relatively high infection rates or rapid growth. For further precise targeting, the most critical high-risk areas are designated as super-risk zones. These super-risk zones are divided into several preset nodes for subsequent refined analysis; for each preset node, a mathematical model is established to calculate its probability of being a potential source of infection. This calculation considers multi-dimensional factors such as historical infection data, population movement information, and social behavior. Based on this, a spatiotemporal weighting factor is introduced to correct the calculated probability, ensuring that dynamic changes in time and space are taken into account. This correction improves the accuracy of the results and helps to better identify potential sources of infection. The corrected probabilities are compared with preset reference probabilities to determine the likelihood of each node being a potential source of infection. Through this comparative analysis, the system can locate the node most likely to be a potential source of infection. This judgment result will indicate the node's geographical location and related risk level, providing guidance for further action. After identifying potential sources of infection, on-site verification is a necessary step to verify the accuracy of the system's judgment. The verification process may involve on-site investigation or collaboration with relevant departments. Based on the verification results, the system will recommend appropriate preset resource allocation plans, including personnel scheduling, medical resource allocation, and the implementation of prevention and control measures to effectively address potential infection risks. Furthermore, data cleaning is a crucial part of data processing, involving the identification and handling of missing values, duplicate data, and outliers to ensure data accuracy and consistency. When handling missing values, records containing missing data can be deleted, imputed using the mean or median, or imputed using a predictive model. For duplicate data, unique identifiers must be used to identify and merge them to eliminate data redundancy. Regarding outlier handling, statistical methods can be used for detection, and the decision to delete or correct outliers depends on their source. Furthermore, data standardization is essential, unifying the format and units of data from different sources to improve compatibility and comparability. Simultaneously, recording data metadata such as source, update time, and processing history is crucial for traceability and verification, ensuring compliance. These steps not only improve data quality, making it more suitable for analysis and decision-making, but also enhance the transparency of data management, supporting deeper business insights and innovation.Through comprehensive data cleansing and management, enterprises can better utilize data resources and improve overall operational efficiency and competitiveness.
[0067] In this preferred embodiment, the data warehouse includes: raw data, the raw data cleaning process, and data sent to the traceability module. The rationale behind this structure—comprising raw data, the raw data cleaning process, and data sent to the traceability module—lies in providing a systematic and efficient data management framework. First, the storage of raw data ensures data integrity, preserving the original records and providing a comprehensive foundation for subsequent analysis. This is crucial for tracing data sources, verifying data authenticity, and conducting historical data backtracking. Second, the raw data cleaning process occupies a key position in the data warehouse, improving data quality and consistency through a series of data cleaning, transformation, and standardization operations. This not only enhances data reliability but also provides more accurate evidence for data analysis and decision support. The transparency and traceability of the cleaning process also help to quickly identify and correct potential errors in the data. Furthermore, the cleaned data is sent to the traceability module for further tracking and analysis. The traceability module records the data processing path and evolution history, providing an effective way to understand the data lifecycle and changes. This design is essential for meeting compliance requirements and enhancing data governance capabilities. Overall, this data warehouse architecture not only optimizes data processing workflows and improves data utilization efficiency, but also enhances the enterprise's flexibility and responsiveness in data management. This facilitates more accurate and effective business decisions, driving innovation and growth. Through this comprehensive and integrated data management strategy, enterprises can better unlock the value of their data and enhance their competitive advantage.
[0068] Understandably, to ensure data integrity and traceability, a data warehouse is designed as a multi-layered storage structure. Specifically, a data warehouse includes the following components:
[0069] Raw Data Layer: This layer stores unprocessed raw data collected from various data sources. This data retains its original format and content, providing a foundation for re-analysis and validation when needed. Raw data may include hospital reports, location data, social media content, etc.
[0070] Data Cleaning Process Layer: This layer records the entire process of cleaning the raw data, including the cleaning algorithms used, the noise data removed, and the detailed steps for data format standardization. This record ensures the transparency and repeatability of each cleaning operation, facilitating auditing and improvement of data processing.
[0071] Post-processed data layer (data sent to the traceability module): Data that has been cleaned and pre-processed is stored in this layer. Redundancy and errors have been removed from this data, and it has been standardized for later use in the traceability module. Processed data is often more structured and has greater analytical value.
[0072] Through this multi-layered data storage design, the data warehouse not only provides high-quality data support for traceability analysis but also enhances the system's flexibility and reliability. This architecture ensures that all analysis is based on verifiable data and can be traced and re-evaluated as needed.
[0073] In this preferred embodiment, identifying high-risk areas includes:
[0074] The infection growth rate is calculated based on the number of infections at the current and future time points. Based on the infection growth rate, the region is divided into high, medium and low risk levels. Information on infected persons and contacts in all high-risk areas is entered into the source tracing module. Each individual high-risk area is treated as an infection node, and contact events are treated as edges to determine super-risk areas.
[0075] Understandably, identifying high-risk areas involves the following steps: First, the infection rate growth rate for each region is calculated based on the current and future infection numbers. The growth rate is a crucial indicator for assessing the regional epidemic's development trend; by comparing infection numbers at different points in time, the speed and trend of the epidemic's spread can be identified. Based on the calculated infection rate growth rate, each region is divided into high, medium, and low-risk levels. This classification helps decision-makers quickly identify areas requiring immediate attention and action. For identified high-risk areas, information on infected individuals and close contacts is entered into the source tracing module. This module records and manages detailed information on personnel activities related to the epidemic for subsequent analysis and tracking. Each high-risk area is considered an infection node, and contact events are considered edges connecting different nodes. By analyzing these nodes and edges, super-risk areas that may become super-spreading pathways can be identified. These super-risk areas are the focus of epidemic prevention and control, requiring stricter control and intervention measures.
[0076] In this preferred embodiment, the formula for calculating the number of infections at future points in time is:
[0077] ;
[0078] Where I(t) represents the number of infected people at time t, β is the transmission coefficient, S(t) is the number of susceptible people, γ is the recovery coefficient, and N is the total population.
[0079] Understandably, after identifying high-risk areas, further analysis of the nodes within these areas can help determine specific sources of infection or transmission routes. This helps refine public health intervention measures and enable targeted isolation or resource allocation.
[0080] In this preferred embodiment, the method for determining the super-risk zone is as follows:
[0081] Calculate the frequency with which an infected node appears on the shortest path between all other pairs of infected nodes; the formula is as follows:
[0082] ;
[0083] Where, σ st σ is the total number of shortest paths from infected node s to infected node t. st (v) is through infected nodes The number of paths;
[0084] Comparing the C-level infection nodes in all high-risk areas B The value of C B The high-risk area represented by the highest infection node is the super-risk area.
[0085] Understandably, machine learning models (such as random forests and Bayesian networks) are used to refine their models based on historical data and new evidence. This approach can capture more complex nonlinear relationships. The model can automatically adjust its prediction probabilities through the training process. The above formula can identify infected nodes that act as key bridges in the spread of information or infection networks. The influence of these infected nodes often extends beyond their direct connections, affecting the overall propagation dynamics of the network. When using the above formula to determine the most important high-risk areas, the analysis can be performed as follows: Calculate the betweenness centrality value of each infected node using network analysis tools or software. Betweenness centrality measures the number of paths an infected node acts as between other infected node pairs in the network. Rank the infected nodes according to the calculated betweenness centrality values. Infected nodes with higher values indicate that they act as bridges on more shortest paths and are key infected nodes in the spread of information or infection in the network. The infected node ranked first is the one with the most prominent betweenness centrality. This infected node can be considered the most important high-risk area because its "bridging" role in the network is the greatest, and it may have a significant impact on the spread of information or influence. Further analysis should be conducted based on domain knowledge and practical situations, building upon the numerical ranking alone. Considering factors such as the geographical location, economic center status, and population density of the infection node, its rationale for being designated as a high-risk area is confirmed.
[0086] In this preferred embodiment, calculating the probability of each preset node being a potential source of infection specifically includes:
[0087] ;
[0088] Where P(j) represents the probability that node j is the source of infection, and W ij Let T be the weight from node i to node j. ij Wij is the propagation time between node i and node j; Wij is a custom setting, such as W. 12 This represents the weight from node 1 to node 2; n is the total number of nodes, and Ti is the propagation time of the i-th node.
[0089] In this preferred embodiment, the above probability is corrected by a spatiotemporal weighting factor, and the corrected probability is compared with a preset probability. Determining the node where the potential source of infection is located based on the comparison result specifically includes:
[0090] ;
[0091] Where P(j) represents the probability that node j is the source of infection. Let represent the corrected probability that node j is the source of infection. It is a time weighting factor. This is the spatial weighting factor; both the temporal weighting factor and the spatial weighting factor are custom settings.
[0092] In all nodes, The node with the largest value is the potential source of infection.
[0093] Understandably, a custom-defined time weighting factor is used to adjust probabilities to account for the impact of the timing of events on transmission. For example, nodes with earlier cases may be assigned higher weights. A custom-defined spatial weighting factor is used to adjust probabilities to account for the spatial distribution characteristics of nodes. Nodes geographically closer to the epicenter of the outbreak may have a higher adjusted probability. Based on the comparative results, key nodes containing potential sources of infection are identified. Special attention is paid to nodes whose probabilities increase significantly after adjustment, as these nodes, after considering spatiotemporal factors, may become new high-risk sources of infection. Through such systematic analysis, potential sources of infection within super-risk areas can be identified and monitored more accurately, facilitating more effective prevention and control measures.
[0094] Please refer to Table 1 below. Table 1 shows the specific process for determining potential sources of infection after multidimensional adjustment screening in this application embodiment. In this preferred embodiment, the process of screening potential sources of infection under multidimensional conditions and determining potential sources of infection based on the screening results and the disease progression of the population at the node where the potential source of infection is located specifically includes:
[0095] Screening is based on the nature of the case type, with pre-set weights for confirmed cases (k1), suspected cases (k2), asymptomatic infections (k3), and ordinary contacts (k4).
[0096] The results of the property-based screening based on case type are as follows:
[0097] ;
[0098] Screening is based on close contact chains, and the strength of close contact relationships, Q, is defined. ij Weights are assigned based on contact frequency and contact duration; where Q ij =f(F ij E ij );
[0099] The screening results based on close contact chains are as follows:
[0100] ;
[0101] Screening is based on epidemiological investigation and contact history. Di=1 is pre-defined to indicate direct contact with confirmed cases, and Di<1 indicates indirect contact with confirmed cases.
[0102] The screening results based on contact history are as follows:
[0103] ;
[0104] Therefore, the result of the multi-dimensional screening is S=P1∩P2∩P3;
[0105] Where, k x k represents the case type weight. min Indicates the minimum threshold; F ij Indicates contact frequency; E ij The duration of contact is indicated. Based on the results of the multidimensional screening and the order of onset of the disease among the potential sources of infection, the earliest onset of the disease is identified as the potential source of infection.
[0106] Table 1. Specific procedures for identifying potential sources of infection after multidimensional regulatory screening.
[0107] Operation Name Function Description filter Obtain and acquire information on potential sources of infection control Set potential sources of infection to a controlled state. Sure The source of infection is determined based on the onset time of the potential source of infection.
[0108] In this preferred embodiment, a preset allocation scheme can also be set; wherein the preset allocation scheme may include:
[0109] If the potential source of infection is located in the city center, public transportation and business activities in that area will be restricted.
[0110] If the potential source of infection is located in a rural or remote area, a medical team will be dispatched to that area to provide medical resources to the local community.
[0111] Furthermore, it can be divided into high-population-density area schemes and low-population-density area schemes based on population density:
[0112] Plan for high-density population areas: Establish temporary isolation points: Set up isolation points in public facilities to prevent transmission within households. Optimize the distribution of supplies: Ensure the rapid delivery of essential supplies to reduce residents' need to go out.
[0113] Low-population-density area program: Personalized health guidance: Provide personalized health guidance and support to enhance residents' self-protection awareness.
[0114] It can also be categorized according to the severity of the epidemic:
[0115] High-risk area plan: Complete lockdown: Implement strict lockdown measures in the area, restricting the movement of people. Centralize medical resources: Allocate more medical resources, such as intensive care facilities and specialized medical teams, to the area. Enhanced disinfection measures: Increase the frequency of disinfection in public places to ensure environmental hygiene.
[0116] Medium-risk area plan: Strengthen testing: Expand testing scope, conduct large-scale nucleic acid testing, and promptly identify and isolate cases. Community patrols: Organize community staff to conduct daily patrols to ensure residents comply with epidemic prevention regulations.
[0117] Low-risk area plan: Routine prevention and control: Maintain daily prevention and control measures, such as wearing masks and maintaining social distancing. Epidemic monitoring: Continuously monitor the epidemic situation and prepare emergency plans.
[0118] Furthermore, the pre-set deployment plan can be further expanded to improve its comprehensiveness and flexibility: big data analytics and artificial intelligence technologies can be used to predict the development trend of the epidemic, identify potential hotspots, and thus formulate prevention and control strategies in advance. Residents are encouraged to use health monitoring applications to self-report their health status and location information for real-time tracking and risk assessment. Community education activities can be conducted through a combination of online and offline methods to improve residents' awareness of epidemic prevention and their self-protection capabilities; a community volunteer network can be established to assist in the distribution of supplies, health promotion, and psychological support, reducing the workload of professionals. Psychological hotlines and online counseling services can be set up to provide professional mental health support and help residents cope with the psychological stress brought about by the epidemic. Online mental health lectures and activities can be organized to promote a positive mindset among residents during the epidemic. The frequency of cleaning and disinfection of public facilities can be increased to ensure the hygiene and safety of the public environment. In remote areas and areas with scarce medical resources, infrastructure construction can be improved, such as increasing mobile medical vehicles and temporary medical points.
[0119] In summary, this invention provides a method and system for tracing the source of infected populations based on multi-source data fusion and intelligent algorithms, which can effectively improve the accuracy and efficiency of infectious disease tracing. By collecting data from multiple data sources such as hospitals, public health departments, social media, and mobile devices, and then cleaning and analyzing it, the dynamics of the epidemic can be monitored in real time. Based on this, using transmission mathematical models and intelligent algorithms, the number of current and future infections can be accurately predicted, and high-risk areas can be identified. Through precise division and node analysis of super-risk areas, potential sources of infection can be quickly located, thus providing an important basis for timely epidemic control. In addition, based on the geographical location of potential sources of infection, the system can select appropriate resource allocation schemes, optimize the allocation and use of public health resources, and reduce the risk of epidemic spread. This comprehensive method not only improves the scientific nature and accuracy of tracing, but also provides a more intelligent tool for public health management. Moreover, this invention innovatively combines multi-source data fusion with intelligent algorithms, providing an efficient and accurate solution for the tracing and control of infectious diseases. First, by integrating multiple data sources, the system can obtain comprehensive epidemic information in real time, covering a full range of data perspectives from detailed information on individual infected individuals to macro-epidemic trends. This integration not only improves the timeliness of information but also significantly enhances the depth and breadth of data analysis. Secondly, the system utilizes advanced machine learning algorithms and mathematical models to accurately predict the development of an epidemic. Through training with historical data and dynamic model adjustments, it can precisely identify high-risk areas and potential outbreak points. This accurate predictive capability enables public health departments to deploy prevention and control measures in advance to prevent the spread of the epidemic. Furthermore, the system can guide targeted intervention strategies by identifying key nodes and super-spreaders in the transmission chain, thereby reducing the overall risk of transmission. In terms of resource allocation, the system can optimize the allocation of public health resources based on the geographical location of potential sources of infection and the severity of the epidemic. This intelligent resource management not only improves the response speed to sudden outbreaks but also effectively reduces unnecessary resource waste.
[0120] See Figure 2 As shown, this embodiment of the invention provides an infection tracing system, the system comprising:
[0121] The data acquisition module is configured to collect data from multiple data sources, clean the data, and then send it to the data traceability module and store it in the data warehouse.
[0122] The source tracing module is configured to use cleaned data to calculate the number of infected people at the current and future time points, identify high-risk areas, set the most critical high-risk areas as super-risk areas, divide the super-risk areas into multiple equally divided preset nodes, calculate the probability of each preset node as a potential source of infection, and correct the above probability through a spatiotemporal weight factor. The corrected probability is compared with the preset probability, and the node where the potential source of infection is located is determined based on the comparison result.
[0123] The allocation module is configured to, after obtaining the node where the potential source of infection is located, perform multi-dimensional screening of the potential source of infection and determine the potential source of infection based on the screening results and the order of onset of the population at the node where the potential source of infection is located.
[0124] Understandably, the data collection module is responsible for acquiring relevant data from various data sources. These sources may include medical records, travel histories, social media data, etc. After collection, the data undergoes cleaning to ensure accuracy and consistency. The data cleaning process involves removing duplicate data, correcting errors, and filling in missing data to improve data quality. Using the cleaned data, the source tracing module enables the system to calculate the number of infections at current and future points in time. This can be achieved through epidemiological models and predictive algorithms. The system analyzes spatial and temporal data to identify potentially high-risk areas and further designates them as super-risk zones. The division of super-risk zones helps to concentrate resources and measures. Super-risk zones are divided into multiple equally spaced pre-defined nodes, and the probability of each node being a potential source of infection is calculated by analyzing historical data, the degree of close contact, and other factors. The probabilities of the nodes are corrected using spatiotemporal weighting factors, which takes into account the effects of time and space, such as the speed of infectious disease transmission and distance attenuation effects. The corrected probabilities are compared with the pre-defined probabilities to determine which node might be a potential source of infection, allowing for further action. After the allocation module obtains the node where the potential source of infection is located, the system will arrange on-site investigations to confirm the accuracy of the source of infection. Based on the confirmation results, a suitable pre-set resource allocation plan is selected. Resource allocation can include medical resource allocation, personnel scheduling, and material transportation to quickly respond to the epidemic. This system aims to improve the efficiency and effectiveness of epidemic prevention and control by using a data-driven approach to achieve accurate tracing of infected individuals and optimized resource allocation.
[0125] In summary, this invention provides a method and system for tracing the source of infected populations based on multi-source data fusion and intelligent algorithms, which can effectively improve the accuracy and efficiency of infectious disease tracing. By collecting data from multiple data sources such as hospitals, public health departments, social media, and mobile devices, and then cleaning and analyzing it, the dynamics of the epidemic can be monitored in real time. Based on this, using transmission mathematical models and intelligent algorithms, the number of current and future infections can be accurately predicted, and high-risk areas can be identified. Through precise division and node analysis of super-risk areas, potential sources of infection can be quickly located, thus providing an important basis for timely epidemic control. In addition, based on the geographical location of potential sources of infection, the system can select appropriate resource allocation schemes, optimize the allocation and use of public health resources, and reduce the risk of epidemic spread. This comprehensive method not only improves the scientific nature and accuracy of tracing, but also provides a more intelligent tool for public health management. Moreover, this invention innovatively combines multi-source data fusion with intelligent algorithms, providing an efficient and accurate solution for the tracing and control of infectious diseases. First, by integrating multiple data sources, the system can obtain comprehensive epidemic information in real time, covering a full range of data perspectives from detailed information on individual infected individuals to macro-epidemic trends. This integration not only improves the timeliness of information but also significantly enhances the depth and breadth of data analysis. Secondly, the system utilizes advanced machine learning algorithms and mathematical models to accurately predict the development of an epidemic. Through training with historical data and dynamic model adjustments, it can precisely identify high-risk areas and potential outbreak points. This accurate predictive capability enables public health departments to deploy prevention and control measures in advance to prevent the spread of the epidemic. Furthermore, the system can guide targeted intervention strategies by identifying key nodes and super-spreaders in the transmission chain, thereby reducing the overall risk of transmission. In terms of resource allocation, the system can optimize the allocation of public health resources based on the geographical location of potential sources of infection and the severity of the epidemic. This intelligent resource management not only improves the response speed to sudden outbreaks but also effectively reduces unnecessary resource waste.
[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for tracing the source of infection in a population, characterized in that, The method includes: Data is collected from the data source, cleaned, and then sent to the traceability module and stored in the data warehouse. The number of infected people at the current and future time points is calculated using the cleaned data, and high-risk areas are identified. The most critical high-risk areas are set as super-risk areas, and the super-risk areas are divided into multiple equally divided preset nodes. The probability of each preset node being a potential source of infection is calculated, and the probability is corrected by a spatiotemporal weighting factor. The corrected probability is compared with the preset probability, and the node where the potential source of infection is located is determined based on the comparison result. After obtaining the node where the potential source of infection is located, multi-dimensional conditions are used to screen for potential sources of infection, and the potential source of infection is determined based on the screening results and the order of onset of the population at the node where the potential source of infection is located.
2. The method for tracing the source of infection in a population according to claim 1, characterized in that, The process of collecting data from the data source, cleaning the data, and sending it to the traceability module specifically includes: Data is collected from data sources, including hospitals, public health departments, social media, and mobile devices; Clean the data, identify and process missing values, duplicate data and outliers, and convert the data into a consistent format and units; Record background information for the data, including data source, update time, and processing history.
3. The method for tracing the source of infection in a population according to claim 2, characterized in that, The data warehouse includes: raw data, the raw data cleaning process, and data sent to the traceability module.
4. The method for tracing the source of infection in a population according to claim 3, characterized in that, The identification of high-risk areas, and the designation of the most critical high-risk areas as super-risk areas, includes: The infection growth rate is calculated based on the number of infections at the current and future time points. The region is divided into high, medium and low risk levels based on the infection growth rate. Information on infected persons and contacts in all high-risk areas is entered into the source tracing module. Each individual high-risk area is treated as an infection node, and contact events are treated as edges to determine super-risk areas.
5. The method for tracing the source of infection in a population according to claim 4, characterized in that, The formula for calculating the number of infections at the future time point is: ; Where I(t) represents the number of infected people at time t, β is the transmission coefficient, S(t) is the number of susceptible people, γ is the recovery coefficient, and N is the total population.
6. The method for tracing the source of infection in a population according to claim 5, characterized in that, The method for determining the super-risk zone is as follows: Calculate the frequency with which an infected node appears on the shortest path between all other pairs of infected nodes; the formula is as follows: ; Where, σ st σ is the total number of shortest paths from infected node s to infected node t. st (v) is through infected nodes The number of paths, C B (v) is the C of node v. B value; Compare the C values of the infection nodes described in all high-risk areas. B The value of C B The high-risk area represented by the highest infection node is the super-risk area.
7. The method for tracing the source of infection in a population according to claim 6, characterized in that, The calculation of the probability that each preset node is a potential source of infection specifically includes: ; Where P(j) represents the probability that node j is the source of infection, and W ij Let T be the weight from node i to node j. ij W represents the propagation time between node i and node j. ij For custom settings; n is the total number of nodes, T i Let be the propagation time of the i-th node.
8. The method for tracing the source of infection in a population according to claim 7, characterized in that, The probabilities are corrected using a spatiotemporal weighting factor. The corrected probabilities are then compared with preset probabilities. Based on the comparison results, the specific nodes where potential sources of infection are located are determined, including: ; Where P(j) represents the probability that node j is the source of infection. Let represent the corrected probability that node j is the source of infection. It is a time weighting factor. This is the spatial weighting factor; both the temporal weighting factor and the spatial weighting factor are custom settings. In all nodes, The node with the largest value is the potential source of infection.
9. An infection tracing system, applied to the infection tracing method according to any one of claims 1-8, characterized in that, The system includes: The data acquisition module is configured to collect data from multiple data sources, clean the data, and then send it to the data traceability module and store it in the data warehouse. The source tracing module is configured to use cleaned data to calculate the number of infected people at the current and future time points, identify high-risk areas, set the most critical high-risk areas as super-risk areas, divide the super-risk areas into multiple equally divided preset nodes, calculate the probability of each preset node as a potential source of infection, and correct the above probability through a spatiotemporal weight factor. The corrected probability is compared with the preset probability, and the node where the potential source of infection is located is determined based on the comparison result. The allocation module is configured to, after obtaining the node where the potential source of infection is located, perform multi-dimensional screening of the potential source of infection and determine the potential source of infection based on the screening results and the order of onset of the population at the node where the potential source of infection is located.
Citation Information
Patent Citations
Early-warning and tracing method for unknown infectious diseases
CN111403048A
Major infectious disease propagation risk early warning and prevention and control analysis system for COVID-19
CN111863271A