Infected population tracing method and system

Through multi-source data fusion and intelligent algorithms, the problem of inefficiency of traditional traceability methods is solved, the rapid and accurate traceability and control of infectious disease epidemics is achieved, and the scientificity and efficiency of public health management is improved.

CN120221123AActive Publication Date: 2025-06-27SECOND MEDICAL CENT OF CHINESE PLA GENERAL HOSPITAL

Patent Information

Application Number
CN202510329488.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-27
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

Traditional infectious disease traceability methods are inefficient and rely on manual investigation and on-site sampling, making it difficult to quickly and effectively identify and track complex transmission paths.

Method used

By collecting data from a variety of data sources (such as hospitals, public health departments, social media and mobile devices), cleaning and analysis, using dissemination mathematical models and intelligent algorithms to calculate the number of infected people, identify high-risk areas, and modify the probability of potential infection sources through spatiotemporal weight factors to quickly locate potential infection sources.

Benefits of technology

It improves the accuracy and efficiency of infectious disease tracing, can grasp the epidemic dynamics in real time, accurately predict the number of infected people and epidemic trends, quickly locate potential sources of infection, optimize the allocation of public health resources, and reduce the risk of epidemic spread.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120221123A_ABST
    Figure CN120221123A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of epidemic disease monitoring, and discloses an infected population tracing method and system. The system comprehensively utilizes various data sources such as medical records, public health reports, social media dynamic and geographic position data and the like to realize real-time monitoring and accurate prediction of infectious disease epidemic situations. The system identifies high-risk areas and key propagation nodes through an advanced machine learning algorithm and a mathematical model, and provides accurate prevention and control measure suggestions. The method improves the traceability efficiency, optimizes the resource allocation, effectively supports public health decisions, and provides technical support for epidemic prevention and control. The system can guide a targeted intervention strategy by identifying key nodes and super propagators in a propagation chain, so that the overall propagation risk is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of epidemic monitoring. Specifically, it relates to a method and system for tracing the source of infected people. Background Art

[0002] In modern society, the prevention and control of infectious diseases has increasingly become a key issue in global public health. Especially in the early stage of emerging infectious diseases or outbreaks, rapid and effective source tracing work is crucial for preventing the spread of the epidemic and ensuring public health. However, traditional source tracing methods rely on manual investigations and on-site sampling. This approach is not only inefficient but also poses many challenges. First, manual investigations require a large amount of human resources, and such resources are often limited when faced with large-scale epidemics. Second, manual data processing is prone to information delays and analysis errors. In addition, the transmission paths of many infectious diseases are complex, and it is difficult to comprehensively track them manually.

[0003] In recent years, the rapid development of big data technology has provided new solutions for infectious disease source tracing. By integrating multiple data sources, including electronic health records of medical institutions, public health reporting systems, social media dynamics, geographical location data of mobile devices, etc., more detailed and real-time epidemic information can be obtained. Such data integration requires an efficient data cleaning and standardization process to ensure the quality and consistency of the data. At the same time, the application of artificial intelligence and machine learning technologies makes it possible to extract valuable patterns and trends from the vast amount of data, which provides strong support for accurate epidemic prediction and high-risk area identification.

[0004] In terms of mathematical modeling, with the development of infectious disease dynamics models, model calibration and parameter estimation combined with actual data can more accurately reflect the spread dynamics of the epidemic. Such models can not only be used to predict the number of infected people and epidemic trends but also evaluate the potential impact by simulating different intervention measures, providing a scientific basis for decision-makers. In addition, network-based transmission models can be used to analyze complex transmission chains and identify key transmission nodes, thus assisting in quickly locating potential sources of infection. Summary of the Invention

[0005] In view of this, the present invention proposes a method and system for tracing the source of infected people, aiming to solve the problem of low efficiency in the current manual analysis and source tracing method.

[0006] A method for tracing the source of infected people proposed by the present invention, the method includes:

[0007] Collect data from data sources, clean the data and then send it to the source tracing module and store it in the data warehouse;

[0008] Calculate the number of infected people at the current and future time points using the cleaned data, identify high-risk areas, set the most critical high-risk area as the super-risk area, divide the super-risk area into multiple equally divided preset nodes, calculate the probability of each preset node being a potential source of infection, correct the probability using the spatio-temporal weight factor, compare the corrected probability with the preset probability, and determine the node where the potential source of infection is located based on the comparison result;

[0009] After obtaining the node where the potential source of infection is located, perform multi-dimensional conditions to screen the potential source of infection and determine the potential source of infection based on the screening result in combination with the order of onset of the population at the node where the potential source of infection is located.

[0010] Further, the collecting data from the data source and cleaning the data and then sending it to the traceability module specifically includes:

[0011] Collect data from the data source, where the data source includes hospitals, public health departments, social media, and mobile devices;

[0012] Clean the data, identify and process missing values, duplicate data, and outliers, and convert the data into a consistent format and unit;

[0013] Record the background information of the data, including the data source, update time, and processing history.

[0014] Further, the data warehouse includes: raw data, the raw data cleaning process, and the data sent to the traceability module.

[0015] Further, the identifying high-risk areas and setting the most critical high-risk area as the super-risk area includes:

[0016] Calculate the growth rate of the number of infected people based on the number of infected people at the current and future time points, divide the regions into high, medium, and low risk levels according to the growth rate of the number of infected people, enter the information of the infected and contacts in all high-risk areas into the traceability module, take each individual high-risk area as an infection node, and the contact event as an edge to determine the super-risk area.

[0017] Further, the calculation formula for the number of infected people at the future time point is:

[0018] ;

[0019] where I(t) represents the number of infected people at time t, β is the transmission coefficient, S(t) is the number of susceptible people, γ is the recovery coefficient, and N is the total population.

[0020] Further, the method for determining the super-risk area is:

[0021] Calculate the frequency of an infected node appearing on the shortest path between all other pairs of infected nodes; the calculation formula is:

[0022] ;

[0023] where σ st is the total number of shortest paths from the infected node s to the infected node t, and σ st (v) is the number of paths passing through the infected node , and C B (v) is the C B value of the node v;

[0024] Compare the C B values of the infected nodes described in all high-risk areas. The high-risk area represented by the infected node with the highest C B is the super-risk area.

[0025] Furthermore, the specific calculation of the probability of each of the preset nodes as a potential source of infection includes:

[0026] ;

[0027] where P(j) represents the probability of node j as a source of infection, W ij is the weight from node i to node j, and T ij is the transmission time between node i and node j; W ij is a custom setting; n is the total number of nodes, and T i is the transmission time of the i-th node.

[0028] Furthermore, the above probability is corrected by the spatio-temporal weight factor, and the corrected probability is compared with the preset probability. Judging the node where the potential source of infection is located according to the comparison result specifically includes:

[0029] ;

[0030] where P(j) represents the probability of node j as a source of infection, represents the corrected probability of node j as a source of infection, is the time weight factor, is the space weight factor, and both the time weight factor and the space weight factor are custom settings;

[0031] Among all the nodes, the node with the largest value is the potential source of infection.

[0032] Furthermore, the specific steps of screening the potential source of infection by multi-dimensional conditions and determining the potential source of infection according to the screening results combined with the onset order of the population at the node where the potential source of infection is located include:

[0033] Screening is carried out based on the nature of the case type. The weight of confirmed cases is preset as k1, the weight of suspected cases is k2, the weight of asymptomatic infected persons is k3, and the weight of general contacts is k4;

[0034] The screening result based on the nature of the case type is:

[0035] ;

[0036] Screening is carried out based on the close contact chain, and the intensity Q of the close contact relationship is defined ij , and weights are assigned according to the contact frequency and contact duration; where Q ij =f(F ij , E ij );

[0037] The screening result based on the close contact chain is:

[0038] ;

[0039] Screening is carried out based on the epidemiological contact history. It is preset that Di = 1 indicates direct contact with a confirmed case, and Di < 1 indicates indirect contact with a confirmed case;

[0040] The screening result based on the epidemiological contact history is:

[0041] ;

[0042] Then the screening result according to the multi-dimensional conditions is S = P1 ∩ P2 ∩ P3;

[0043] Among them, k x represents the case type weight, and k min represents the lowest threshold; F ij represents the contact frequency; E ij represents the contact duration;

[0044] Arrange according to the onset order in the multi-dimensional condition screening result S, and determine the one with the earliest onset time as the potential source of infection.

[0045] On the other hand, the present invention also proposes an infection population tracing system, which is applied to the above-mentioned infection population tracing method. The system includes:

[0046] A collection module, configured to collect data from a data source, clean the data and then send it to the tracing module and store it in a data warehouse;

[0047] The traceability module is configured to calculate the number of infected people at current and future time points by using the cleaned data, identify high-risk areas, set the most critical high-risk area as the super-risk area, divide the super-risk area into multiple equally divided preset nodes, calculate the probability of each preset node as a potential source of infection, correct the above probability through a spatio-temporal weight factor, compare the corrected probability with a preset probability, and judge the node where the potential source of infection is located according to the comparison result;

[0048] The deployment module is configured to, after obtaining the node where the potential source of infection is located, perform multi-dimensional conditions to screen the potential source of infection and determine the potential source of infection according to the screening result in combination with the onset order of the population at the node where the potential source of infection is located.

[0049] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention provides a method and system for tracing infected people based on multi-source data fusion and intelligent algorithms, which can effectively improve the accuracy and efficiency of infectious disease tracing. By collecting data from multiple data sources such as hospitals, public health departments, social media, and mobile devices, and performing cleaning and analysis, the epidemic situation can be grasped in real time. On this basis, by using a propagation mathematical model and intelligent algorithms, the number of current and future infected people can be accurately predicted, and high-risk areas can be identified. Through the precise division of the super-risk area and node analysis, the potential source of infection can be quickly located, providing an important basis for timely controlling the epidemic. In addition, according to the geographical location of the potential source of infection, the system can select a suitable resource deployment plan, optimize the allocation and use of public health resources, and reduce the risk of epidemic spread. This comprehensive method not only improves the scientificity and accuracy of tracing, but also provides a more intelligent tool for public health management. Moreover, the present invention innovatively combines multi-source data fusion with intelligent algorithms, providing an efficient and accurate solution for the tracing and control of infectious diseases. First, by integrating multiple data sources, the system can obtain comprehensive epidemic information in real time, covering a full range of data perspectives from the detailed information of individual infected people to the macroscopic epidemic trend. This integration not only improves the timeliness of information, but also significantly enhances the depth and breadth of data analysis. Second, the system uses advanced machine learning algorithms and mathematical models to accurately predict the development of the epidemic. Through the training of historical data and the dynamic adjustment of the model, high-risk areas and potential outbreak points can be accurately identified. This accurate prediction ability enables public health departments to deploy prevention and control measures in advance to prevent the spread of the epidemic. In addition, the system can guide targeted intervention strategies by identifying key nodes and super spreaders in the transmission chain, thereby reducing the overall transmission risk. In terms of resource deployment, the system can optimize the allocation of public health resources according to the geographical location and severity of the epidemic of the potential source of infection. This intelligent resource management not only improves the response speed to sudden epidemics, but also effectively reduces unnecessary resource waste. Description of the Drawings

[0050] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0051] Figure 1 It is a flowchart of a method for tracing the source of infected people according to an embodiment of the present invention.

[0052] Figure 2 It is a functional block diagram of a system for tracing the source of infected people according to an embodiment of the present invention. Detailed Embodiments

[0053] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.

[0054] In modern society, the prevention and control of infectious diseases have increasingly become a key issue in global public health. Especially in the early stage of emerging infectious diseases or outbreaks, rapid and effective source tracing work is crucial for preventing the spread of the epidemic and ensuring public health. However, traditional source tracing methods rely on manual investigations and on-site sampling, which are not only inefficient but also pose many challenges. First of all, manual investigations require a large amount of human resources, and in the face of large-scale epidemics, such resources are often limited. Secondly, manual data processing is prone to information delays and analysis errors. In addition, the transmission paths of many infectious diseases are complex, and it is difficult to comprehensively track them manually.

[0055] In recent years, the rapid development of big data technology has provided new solutions for infectious disease source tracing. By integrating multiple data sources, including electronic health records of medical institutions, public health reporting systems, social media dynamics, geographical location data of mobile devices, etc., more detailed and real-time epidemic information can be obtained. Such data integration requires an efficient data cleaning and standardization process to ensure the quality and consistency of the data. At the same time, the application of artificial intelligence and machine learning technologies makes it possible to extract valuable patterns and trends from the vast amount of data, which provides strong support for accurate epidemic prediction and high-risk area identification.

[0056] In mathematical modeling, with the development of infectious disease dynamics models, model calibration and parameter estimation combined with actual data can more accurately reflect the spread dynamics of the epidemic. Such models can not only be used to predict the number of infected people and the epidemic trend, but also evaluate their potential impact by simulating different intervention measures, providing a scientific basis for decision-makers. In addition, network-based transmission models can be used to analyze complex transmission chains and identify key transmission nodes, thus assisting in quickly locating potential sources of infection.

[0057] Refer to Figure 1 As shown, an embodiment of the present invention provides a method for tracing the source of infected people, the method including:

[0058] S1: Collect data from multiple data sources, clean the data and send it to the tracing module and store it in the data warehouse;

[0059] S2: Use the cleaned data to calculate the number of infected people at current and future time points, identify high-risk areas, set the most critical high-risk area as the super-risk area, divide the super-risk area into multiple equally divided preset nodes, calculate the probability of each preset node as a potential source of infection, and correct the above probability through a spatio-temporal weight factor, compare the corrected probability with the preset probability, and judge the node where the potential source of infection is located according to the comparison result;

[0060] It can be understood that considering the changes in time and space, the initial probability is corrected. The spatio-temporal weight factor can include the following aspects: Time factor: Adjust the probability according to the activity characteristics of different periods (such as rush hours, holidays, etc.). Space factor: Consider the distance between nodes, traffic connectivity, etc., and adjust the transmission potential of nodes in different positions.

[0061] S3: After obtaining the node where the potential source of infection is located, conduct multi-dimensional conditions to screen the potential source of infection and determine the potential source of infection according to the screening result combined with the onset order of the people in the node where the potential source of infection is located.

[0062] In this preferred embodiment, collecting data from multiple data sources and cleaning the data and sending it to the tracing module specifically includes:

[0063] Collect data from multiple data sources, and the data sources include hospitals, public health departments, social media and mobile devices;

[0064] Clean the data, identify and process missing values, duplicate data and outliers, and convert the data into a consistent format and unit;

[0065] Record the background information of the data, including the data source, update time, and processing history.

[0066] It is understandable that, first of all, relevant data is collected from multiple data sources, which include but are not limited to hospital records, reports from public health departments, information on social media platforms, and location information of mobile devices. The collected data may contain noise and inconsistencies, so data cleaning is required to remove duplicate, incomplete, or incorrect data. This step ensures the high quality of the input data and is the basis for subsequent analysis. The cleaned data is transmitted to the traceability module and stored in a dedicated data warehouse for easy access and analysis at any time. After completing data cleaning, the system uses statistical models and prediction algorithms to calculate the number of infected people at current and future time points. This prediction helps to identify high-risk areas, that is, areas where the number of infected people is relatively high or the growth rate is relatively fast. For further precise positioning, the most critical high-risk areas are set as super-risk areas. These super-risk areas are divided into several preset nodes for subsequent refined analysis; for each preset node, a mathematical model is established to calculate the probability of it being a potential source of infection. This calculation takes into account multi-dimensional factors such as historical infection data, population mobility information, and social behavior. On this basis, a spatio-temporal weight factor is introduced to correct the calculated probability to ensure that dynamic changes in time and space are taken into account. This correction improves the accuracy of the results and helps to better identify potential sources of infection. The corrected probability is compared with a preset reference probability to judge the possibility of each node being a potential source of infection. Through this comparative analysis, the system can locate the node where the most likely potential source of infection is located. This judgment result will indicate the geographical location of the node and the relevant risk level, providing guidance for further actions. After determining the potential source of infection, on-site confirmation is a necessary step to verify the accuracy of the system's judgment. The confirmation process may involve on-site investigations or cooperation with relevant departments. According to the confirmation results, the system will recommend appropriate preset resource allocation plans, and the specific plans include personnel scheduling, medical resource allocation, and implementation of prevention and control measures to effectively respond to possible infection risks. Moreover, data cleaning is an important part of data processing, which involves identifying and handling missing values, duplicate data, and outliers to ensure the accuracy and consistency of the data. When dealing with missing values, records containing missing data can be deleted, filled with the mean or median, or filled through a prediction model. For duplicate data, it needs to be identified and merged through unique identifiers to eliminate data redundancy. In terms of outlier handling, statistical methods can be used for detection, and it is decided whether to delete or correct them according to their sources. In addition, data standardization is necessary to unify the formats and units of data from different sources to improve their compatibility and comparability. At the same time, it is crucial to record the meta-information of the data such as the source, update time, and processing history for traceability and verification to meet compliance requirements. These steps not only improve the data quality, making it more suitable for analysis and decision-making, but also enhance the transparency of data management, supporting in-depth insights and innovation in business.Through comprehensive data cleaning and management, enterprises can make better use of data resources and enhance overall operational efficiency and competitiveness.

[0067] In this preferred embodiment, the data warehouse includes: raw data, the raw data cleaning process, and the data sent to the traceability module. Moreover, the design of the data warehouse includes raw data, the raw data cleaning process, and the data sent to the traceability module. The rationality of this structure lies in providing a systematic and efficient data management framework. First, the storage of raw data ensures data integrity, retains the original records, and provides a comprehensive basis for subsequent analysis. This is of great significance for tracing data sources, verifying data authenticity, and conducting historical data backtracking. Second, the raw data cleaning process occupies a key position in the data warehouse. Through a series of data cleaning, transformation, and standardization operations, it improves the quality and consistency of the data. This not only enhances data reliability but also provides a more accurate basis for data analysis and decision support. The transparency and traceability of the cleaning process also help quickly identify and correct potential errors in the data. In addition, the cleaned data is sent to the traceability module for further tracking and analysis. The traceability module can record the processing path and evolution history of the data, providing an effective way to understand the data life cycle and changes. Such a design is crucial for meeting compliance requirements and enhancing data governance capabilities. Overall, this data warehouse structure not only optimizes the data processing flow, improves data utilization efficiency, but also enhances the flexibility and responsiveness of enterprises in data management, helps make more accurate and effective business decisions, and promotes enterprise innovation and development. Through this comprehensive and integrated data management strategy, enterprises can better explore data value and enhance competitive advantages.

[0068] It can be understood that to ensure data integrity and traceability, the data warehouse is designed as a multi-level storage structure. Specifically, the data warehouse includes the following parts:

[0069] Raw data layer: This layer stores the unprocessed raw data collected from various data sources. These data retain their original format and content, providing a basis for re-analysis and verification when needed. Raw data may include hospital reports, location data, social media content, etc.

[0070] Data cleaning process layer: This layer records the entire process of cleaning the raw data, including the applied cleaning algorithms, the removed noise data, and the detailed steps of data format standardization. The recording of this process ensures the transparency and repeatability of each cleaning operation, facilitating the auditing and improvement of data processing.

[0071] Processed Data Layer (Data Sent to the Traceability Module): The data after cleaning and preprocessing is stored in this layer. This data has had redundant and error information removed and has been standardized for subsequent use in the traceability module. The processed data is often more structured and has higher analytical value.

[0072] Through this multi-level data storage design, the data warehouse not only provides high-quality data support for traceability analysis but also enhances the flexibility and reliability of the system. This architecture ensures that all data for analysis is verifiable and can be traced and re-evaluated as needed.

[0073] In this preferred embodiment, identifying high-risk areas includes:

[0074] Calculate the growth rate of the number of infected people based on the number of infected people at the current and future time points, divide the regions into high, medium, and low-risk levels according to the growth rate of the number of infected people, enter the information of the infected and contacts in all high-risk areas into the traceability module, take each individual high-risk area as an infection node, and the contact event as an edge to determine the super-risk area.

[0075] It can be understood that the process of identifying high-risk areas involves the following steps: First, calculate the growth rate of the number of infected people in each region based on the number of infected people at the current and future time points. The growth rate is an important indicator for evaluating the development trend of the epidemic in a region. By comparing the number of infected people at different time points, the speed and trend of the spread of the epidemic can be identified. Based on the calculated growth rate of the number of infected people, each region is divided into high, medium, and low-risk levels. This division helps decision-makers quickly identify areas that require immediate attention and action. For the identified high-risk areas, the information of the infected and close contacts is entered into the traceability module. This module is used to detailed record and manage the information of personnel activities related to the epidemic for subsequent analysis and tracking. Each high-risk area is regarded as an infection node, and the contact event is regarded as the edge connecting different nodes. By analyzing these nodes and edges, super-risk areas that may become super-spreading routes can be identified. These super-risk areas are the key points of epidemic prevention and control and require more stringent control and intervention measures.

[0076] In this preferred embodiment, the calculation formula for the number of infected people at the future time point is:

[0077] ;

[0078] where I(t) represents the number of infected people at time t, β is the transmission coefficient, S(t) is the number of susceptible people, γ is the recovery coefficient, and N is the total population.

[0079] It is understandable that after identifying high-risk areas, the nodes within these areas can be further analyzed to determine specific sources of infection or transmission routes. This helps to refine public health intervention measures and target isolation or resource allocation accordingly.

[0080] In this preferred embodiment, the method for determining the super high-risk area is as follows:

[0081] Calculate the frequency of an infected node appearing on the shortest paths between all other pairs of infected nodes; its calculation formula is:

[0082] ;

[0083] where σ st is the total number of shortest paths from the infected node s to the infected node t, and σ st (v) is the number of paths passing through the infected node .

[0084] Compare the C B values of the infected nodes in all high-risk areas. The high-risk area represented by the infected node with the highest C B value is the super high-risk area.

[0085] It is understandable that machine learning models (such as random forests, Bayesian networks) are used for calibration based on historical data and new evidence. This method can capture more complex non-linear relationships. The model can automatically adjust the prediction probability through the training process. The above formula can identify those infected nodes that play a key bridging role in the information or infection transmission network. The influence of these infected nodes often extends beyond their directly connected range and can affect the overall transmission dynamics of the network. When using the above formula to determine the most important high-risk areas, the following steps can be taken for analysis: Use network analysis tools or software to calculate the betweenness centrality value of each infected node. Betweenness centrality measures the number of paths of an infected node between other pairs of infected nodes in the network. Sort the infected nodes according to the calculated betweenness centrality values. The higher the value of an infected node, the more it indicates that it plays a bridging role on more shortest paths and is a key infected node for the spread of information or infection in the network. The infected node ranked first is the one with the most prominent betweenness centrality. This infected node can be regarded as the most important high-risk area because its "bridging" role in the network is the greatest and it may have a significant impact on the spread of information or influence. On the basis of simply relying on numerical sorting, further analysis is combined with domain knowledge and actual situations. Consider factors such as the geographical location, economic central status, and population density of this infected node to confirm the rationality of its being a high-risk area.

[0086] In this preferred embodiment, calculating the probability of each preset node as a potential source of infection specifically includes:

[0087] ;

[0088] Among them, P(j) represents the probability that node j is the source of infection, and W ij is the weight from node i to node j, and T ij is the transmission time between node i and node j; Wij is a custom setting, such as W 12 represents the weight from node 1 to node 2; n is the total number of nodes, and Ti is the transmission time of the i-th node.

[0089] In this preferred embodiment, the above probability is corrected by the spatio-temporal weight factor, and the corrected probability is compared with the preset probability. Judging the node where the potential source of infection is located according to the comparison result specifically includes:

[0090] ;

[0091] Among them, P(j) represents the probability that node j is the source of infection, represents the corrected probability that node j is the source of infection, is the time weight factor, is the space weight factor, and both the time weight factor and the space weight factor are custom settings;

[0092] Among all the nodes, the node with the largest

[0093] value is the potential source of infection.

[0094] It can be understood that the custom time weight factor is used to adjust the probability to consider the impact of the time of event occurrence on transmission. For example, nodes with earlier cases may be given higher weights. The custom space weight factor is used to adjust the probability to consider the characteristics of the nodes in the spatial distribution. Nodes close to the epidemic center may have a higher corrected probability. According to the comparison result, identify the key nodes where the potential source of infection is located. Pay special attention to those nodes with a significant increase in the corrected probability, because these nodes may become new high-risk sources of infection after considering spatio-temporal factors. Through such systematic analysis, potential sources of infection in the super-risk area can be identified and monitored more accurately, which helps to implement more effective prevention and control measures.

[0095] Screening is carried out based on the nature of case types. The weight of confirmed cases is preset as k1, the weight of suspected cases is k2, the weight of asymptomatic infected persons is k3, and the weight of general contacts is k4;

[0096] The screening result based on the nature of case types is:

[0097] ;

[0098] Screening is carried out based on the close contact chain. Define the intensity of the close contact relationship Q ij , and assign weights according to the contact frequency and contact duration; where Q ij =f(F ij , E ij );

[0099] The screening result based on the close contact chain is:

[0100] ;

[0101] Screening is carried out based on the epidemiological contact history. It is preset that Di = 1 indicates direct contact with a confirmed case, and Di < 1 indicates indirect contact with a confirmed case;

[0102] The screening result based on the epidemiological contact history is:

[0103] ;

[0104] Then the screening result according to the multi-dimensional conditions is S = P1 ∩ P2 ∩ P3;

[0105] Among them, k x represents the case type weight, and k min represents the lowest threshold; F ij represents the contact frequency; E ij represents the contact duration; Screening is carried out according to the obtained multi-dimensional condition screening result in combination with the onset order of the population in the potential source of infection, and the earliest onset is determined as the potential source of infection.

[0106] Table 1 Specific process for determining the potential source of infection after multi-dimensional adjustment screening

[0107] Operation Name Function Description Filter Obtain and get potential source of infection information Control Set the potential source of infection to the controlled state Determine Determine the source of infection based on the onset time of the potential source of infection

[0108] In this preferred embodiment, a preset deployment plan can also be set; the preset deployment plan can include:

[0109] If the potential source of infection is located in the central urban area, public transportation management and commercial activity restrictions will be imposed on this area;

[0110] If the potential source of infection is located in rural or remote areas, a medical team will be dispatched to this area for the sinking of medical resources.

[0111] Moreover, it can also be classified into a high population density area plan and a low population density area plan according to population density:

[0112] High population density area plan: Set up temporary quarantine points: Set up quarantine points in public facilities to avoid in-house transmission. Optimize material distribution: Ensure the rapid distribution of living materials to reduce the need for residents to go out.

[0113] Low population density area plan: Provide personalized health guidance: Provide personalized health guidance and support to enhance residents' self-protection awareness.

[0114] It can also be classified according to the severity of the epidemic:

[0115] High-risk area plan: Implement a full lockdown: Implement strict lockdown measures in this area to restrict personnel movement. Concentrate medical resources: Allocate more medical resources, such as intensive care facilities and professional medical teams, to this area. Strengthen disinfection measures: Increase the disinfection frequency in public places to ensure environmental hygiene.

[0116] Medium-risk area plan: Strengthen detection: Expand the detection scope, conduct large-scale nucleic acid testing, and promptly detect and isolate cases. Community patrol: Organize community workers to conduct daily patrols to ensure that residents comply with epidemic prevention regulations.

[0117] Low-risk area plan: Implement routine prevention and control: Maintain daily prevention and control measures, such as wearing masks and maintaining social distancing. Monitor the epidemic situation: Continuously monitor the epidemic situation and prepare emergency plans.

[0118] Moreover, when formulating a preset deployment plan, further expansion can be carried out to improve the comprehensiveness and flexibility of the plan: Use big data analysis and artificial intelligence technology to predict the development trend of the epidemic, identify potential hotspots, and thus formulate prevention and control strategies in advance. Encourage residents to use health monitoring applications to report their health status and location information by themselves for real-time tracking and risk assessment. Carry out community education activities through a combination of online and offline methods to improve residents' awareness of epidemic prevention and self-protection capabilities; Establish a community volunteer network to assist in material distribution, health promotion, and psychological support to relieve the work pressure of professionals. Set up a psychological hotline and online consultation services to provide professional mental health support to help residents cope with the psychological pressure brought by the epidemic. Organize online mental health lectures and activities to promote residents to maintain a positive attitude during the epidemic. Increase the cleaning and disinfection frequency of public facilities to ensure the health and safety of the public environment. In remote areas and areas with scarce medical resources, improve infrastructure construction, such as increasing mobile medical vehicles and temporary medical points.

[0119] In summary, the present invention provides a method and system for tracing the source of infected populations based on multi-source data fusion and intelligent algorithms, which can effectively improve the accuracy and efficiency of infectious disease source tracing. By collecting data from multiple data sources such as hospitals, public health departments, social media, and mobile devices, and cleaning and analyzing the data, the epidemic situation can be grasped in real time. On this basis, using the propagation mathematical model and intelligent algorithms, the current and future number of infected people can be accurately predicted, and high-risk areas can be identified. By accurately dividing the super-risk areas and analyzing the nodes, potential sources of infection can be quickly located, providing an important basis for timely controlling the epidemic. In addition, according to the geographical location of the potential source of infection, the system can select a suitable resource allocation plan, optimize the allocation and use of public health resources, and reduce the risk of epidemic spread. This comprehensive method not only improves the scientific nature and accuracy of source tracing, but also provides a more intelligent tool for public health management. Moreover, the present invention innovatively combines multi-source data fusion with intelligent algorithms, providing an efficient and accurate solution for the source tracing and control of infectious diseases. First, by integrating multiple data sources, the system can obtain comprehensive epidemic information in real time, covering a full range of data perspectives from the detailed information of individual infected persons to the macro-epidemic trend. This integration not only improves the timeliness of information, but also significantly enhances the depth and breadth of data analysis. Second, the system uses advanced machine learning algorithms and mathematical models to accurately predict the development of the epidemic. Through the training of historical data and the dynamic adjustment of the model, high-risk areas and potential outbreak points can be accurately identified. This accurate prediction ability enables public health departments to deploy prevention and control measures in advance to prevent the spread of the epidemic. In addition, the system can guide targeted intervention strategies by identifying key nodes and super spreaders in the transmission chain, thereby reducing the overall transmission risk. In terms of resource allocation, the system can optimize the allocation of public health resources according to the geographical location of the potential source of infection and the severity of the epidemic. This intelligent resource management not only improves the response speed to sudden epidemics, but also effectively reduces unnecessary resource waste.

[0120] Referring to Figure 2 as shown, an embodiment of the present invention provides an infected population source tracing system, which includes:

[0121] A collection module, configured to collect data from multiple data sources, clean the data, and then send it to the source tracing module and store it in the data warehouse;

[0122] The traceability module is configured to calculate the number of infected people at current and future time points using the cleaned data, identify high-risk areas, set the most critical high-risk area as the super-risk area, divide the super-risk area into multiple equally divided preset nodes, calculate the probability of each preset node being a potential source of infection, correct the above probability through the spatio-temporal weight factor, compare the corrected probability with the preset probability, and judge the node where the potential source of infection is located according to the comparison result;

[0123] The deployment module is configured to, after obtaining the node where the potential source of infection is located, perform multi-dimensional conditions to screen the potential source of infection and determine the potential source of infection according to the screening result in combination with the onset order of the population at the node where the potential source of infection is located.

[0124] It can be understood that the collection module is responsible for obtaining relevant data from various data sources. These data sources may include medical records, movement trajectories, social media data, etc. After collection, the data will be cleaned to ensure the accuracy and consistency of the data. The data cleaning process involves removing duplicate data, correcting error information, filling in missing data, etc. to improve the data quality. Using the cleaned data, the traceability module enables the system to calculate the number of infected people at current and future time points. This can be achieved through epidemiological models and prediction algorithms. The system will analyze spatial and temporal data, identify possible high-risk areas, and further set them as super-risk areas. The division of super-risk areas helps to concentrate resources and measures. The super-risk area is divided into multiple equally divided preset nodes, and the probability of each node being a potential source of infection is calculated by analyzing factors such as historical data and degree of close contact. The probability of the node is corrected through the spatio-temporal weight factor, which can consider the influence of time and space, such as the speed of infectious disease transmission and distance attenuation effect. The corrected probability is compared with the preset probability to judge which node may be the potential source of infection for further action. After the deployment module obtains the node where the potential source of infection is located, the system will arrange on-site investigations to confirm the accuracy of the source of infection. According to the confirmation result, a suitable preset resource deployment plan will be selected. Resource deployment can include medical resource allocation, personnel scheduling, material transportation, etc. to quickly respond to the epidemic. The system aims to achieve precise traceability of infected people and optimal allocation of resources through data-driven methods, thereby improving the efficiency and effectiveness of epidemic prevention and control.

[0125] In summary, the present invention provides a method and system for tracing the source of infected people based on multi-source data fusion and intelligent algorithms, which can effectively improve the accuracy and efficiency of infectious disease source tracing. By collecting data from various data sources such as hospitals, public health departments, social media, and mobile devices, and cleaning and analyzing it, the epidemic situation can be grasped in real time. On this basis, using the propagation mathematical model and intelligent algorithms, the current and future number of infected people can be accurately predicted, and high-risk areas can be identified. Through the precise division and node analysis of super-risk areas, potential sources of infection can be quickly located, providing an important basis for timely controlling the epidemic. In addition, according to the geographical location of potential sources of infection, the system can select a suitable resource allocation plan, optimize the allocation and use of public health resources, and reduce the risk of epidemic spread. This comprehensive method not only improves the scientificity and accuracy of source tracing, but also provides a more intelligent tool for public health management. Moreover, the present invention innovatively combines multi-source data fusion with intelligent algorithms, providing an efficient and accurate solution for the source tracing and control of infectious diseases. First, by integrating multiple data sources, the system can obtain comprehensive epidemic information in real time, covering a full range of data perspectives from the detailed information of individual infected people to the macro-epidemic trend. This integration not only improves the timeliness of information, but also significantly enhances the depth and breadth of data analysis. Second, the system uses advanced machine learning algorithms and mathematical models to accurately predict the development of the epidemic. Through the training of historical data and the dynamic adjustment of the model, high-risk areas and potential outbreak points can be accurately identified. This precise prediction ability enables public health departments to deploy prevention and control measures in advance to prevent the spread of the epidemic. In addition, the system can guide targeted intervention strategies by identifying key nodes and super spreaders in the transmission chain, thereby reducing the overall transmission risk. In terms of resource allocation, the system can optimize the allocation of public health resources according to the geographical location and severity of the epidemic of potential sources of infection. This intelligent resource management not only improves the response speed to sudden epidemics, but also effectively reduces unnecessary resource waste.

[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A method for tracing the source of an infected population, characterized in that: The method comprises: Collect data from the data source, clean the data, send it to the traceability module and store it in the data warehouse; The cleaned data is used to calculate the number of infections at current and future time points, and to identify high-risk areas. The most critical high-risk areas are set as super-risk areas, and the super-risk areas are divided into a plurality of equally divided preset nodes. The probability of each preset node being a potential source of infection is calculated, and the probability is corrected by a spatiotemporal weight factor, and the corrected probability is compared with the preset probability, and the node where the potential source of infection is located is determined based on the comparison result; After obtaining the node where the potential source of infection is located, multi-dimensional conditions are used to screen the potential source of infection and the potential source of infection is determined based on the screening results and the order of onset of the population at the node where the potential source of infection is located.

2. A method for tracing the source of an infected population according to claim 1, characterized in that: The collecting of data from the data source and sending the data to the traceability module after cleaning specifically includes: Collecting data from data sources including hospitals, public health departments, social media, and mobile devices; Clean the data, identify and handle missing values, duplicate data, and outliers, and convert the data into consistent formats and units; Record the background information of the data, including data source, update time, and processing history.

3. A method for tracing the source of an infected population according to claim 2, characterized in that: The data warehouse includes: original data, original data cleaning process and data sent to the traceability module.

4. A method for tracing the source of an infected population according to claim 3, characterized in that: The identification of high-risk areas and setting the most critical high-risk areas as super-risk areas includes: The growth rate of the number of infections is calculated based on the number of infections at the current and future time points, and the areas are divided into high, medium, and low risk levels based on the growth rate of the number of infections. The information of infected persons and contacts in all high-risk areas is entered into the traceability module, and each individual high-risk area is taken as an infection node, and the contact event is taken as an edge to determine the super-risk area.

5. A method for tracing the source of an infected population according to claim 4, characterized in that: The calculation formula for the number of infected people at the future time point is: ; Among them, I(t) represents the number of infected people at time t, β is the transmission coefficient, S(t) is the number of susceptible people, γ is the recovery coefficient, and N is the total population.

6. A method for tracing the source of an infected population according to claim 5, characterized in that: The method for determining the super risk area is: Calculate the frequency of an infected node appearing on the shortest path between all other pairs of infected nodes; the calculation formula is: ; Among them, σ st is the total number of shortest paths from infected node s to infected node t, σ st (v) Through infected nodes The number of paths, C B (v) is the C of node v B value; Compare the C of infected nodes in all high-risk areas B The value of C B The high-risk area represented by the highest infection node is the super-risk area.

7. A method for tracing the source of an infected population according to claim 6, characterized in that: The calculating the probability of each of the preset nodes being a potential infection source specifically includes: ; Among them, P(j) represents the probability that node j is the source of infection, W ij is the weight from node i to node j, T ij is the propagation time between node i and node j; W ij is a custom setting; n is the total number of nodes, T i is the propagation time of the i-th node.

8. A method for tracing the source of an infected population according to claim 7, characterized in that: The above probability is corrected by the spatiotemporal weight factor, and the corrected probability is compared with the preset probability. According to the comparison result, the nodes where the potential infection source is located are determined, including: ; Among them, P(j) represents the probability that node j is the source of infection, represents the corrected probability of node j being the source of infection, is the time weighting factor, is a spatial weight factor, wherein both the temporal weight factor and the spatial weight factor are user-defined settings; In all nodes, The node with the largest value is a potential source of infection.

9. A method for tracing the source of an infected population according to claim 8, characterized in that: The screening of potential infection sources by multi-dimensional conditions and determining the potential infection sources according to the screening results combined with the order of onset of the population at the node where the potential infection sources are located specifically include: Screening is performed based on the nature of the case type, with the weight of confirmed cases set to k1, the weight of suspected cases set to k2, the weight of asymptomatic infections set to k3, and the weight of common contacts set to k4; The results of the screening based on the nature of the case type are: ; Screening based on close contact chains and defining the close contact relationship strength Q ij , weights are assigned according to contact frequency and contact duration; where Q ij =f(F ij , E ij ); The screening results based on close contact chains are: ; Screening is based on the epidemiological contact history, with Di=1 pre-set to indicate direct contact with confirmed cases, and Di<1 to indicate indirect contact with confirmed cases; The screening results based on the epidemiological contact history are: ; Then the screening result according to multi-dimensional conditions is S=P1∩P2∩P3; Among them, k x represents the case type weight, k min Indicates the lowest threshold; F ij Indicates contact frequency; E ij Indicates the duration of contact; The multi-dimensional condition screening results S are arranged according to the order of onset, and the one with the earliest onset time is determined as the potential source of infection.

10. An infected population tracing system, applied to an infected population tracing method according to any one of claims 1 to 9, characterized in that: The system comprises: The acquisition module is configured to collect data from various data sources, clean the data, send it to the traceability module and store it in the data warehouse; The tracing module is configured to use the cleaned data to calculate the number of infections at current and future time points, identify high-risk areas, set the most critical high-risk areas as super-risk areas, and divide the super-risk areas into a plurality of equally divided preset nodes, calculate the probability of each preset node being a potential source of infection, and correct the above probability by a spatiotemporal weight factor, compare the corrected probability with the preset probability, and determine the node where the potential source of infection is located according to the comparison result; The deployment module is configured to obtain the node where the potential infection source is located, screen the potential infection source under multi-dimensional conditions, and determine the potential infection source based on the screening results combined with the order of onset of the population at the node where the potential infection source is located.

Citation Information

Patent Citations

  • Early-warning and tracing method for unknown infectious diseases

    CN111403048A

  • Major infectious disease propagation risk early warning and prevention and control analysis system for COVID-19

    CN111863271A

  • Structured new coronal pneumonia epidemic situation prediction and evaluation method and device, equipment and medium

    CN112382405A

  • Generative AI emotion propagation prediction and guidance large model construction method and system

    CN119047512A

  • Epidemic prediction method and electronic device

    WO2022135197A1

Cited By

  • Big data analysis-based infectious disease transmission risk prediction method and system

    CN121528580A