Multi-dimensional Disease Monitoring and Early Warning Method

By building a multi-dimensional data set and a real-time disease source network, combining weighted fusion algorithms and geographic information systems, the monitoring and early warning problems in the initial stage of infectious diseases are solved, and accurate identification and early prevention and control are achieved.

CN119128040BActive Publication Date: 2025-08-01SUZHOU HEALTH & FAMILY PLANNING STATISTICS INFORMATION CENT +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411621652.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-14
Publication Date
2025-08-01
Estimated Expiration
2044-11-14

AI Technical Summary

Technical Problem

It is difficult for the existing technology to effectively monitor and early warning of infectious diseases in the initial stage of occurrence, resulting in delays in prevention and control time.

Method used

By constructing a multi-dimensional data set, a real-time disease source network and a historical disease source network are established, and the hierarchical comparison principle and weighted fusion algorithm are adopted to identify the most similar historical disease source network, formulate prevention and control measures, and conduct real-time monitoring and labeling of key areas in combination with the geographical information system to identify suspected disease source nodes and infection sources.

Benefits of technology

It has achieved accurate identification and precise prevention and control in the initial stage of infectious diseases, improved the accuracy and robustness of disease warning, and can promptly capture signs of early transmission and formulate targeted measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119128040B_ABST
    Figure CN119128040B_ABST
Patent Text Reader

Abstract

The present application discloses a multi-dimensional disease monitoring and early warning method, which relates to the technical field of data processing, and includes: obtaining multi-dimensional data to generate a multi-dimensional data set; establishing a historical disease source network according to the multi-dimensional data set, and obtaining existing case data to establish a real-time disease source network; comparing the real-time disease source network with multiple historical disease source networks respectively according to the hierarchical comparison principle to obtain the most similar historical disease source; marking key areas around each real-time network node according to geographic information system technology to form punctuation data; continuously monitoring the punctuation data, and defining an abnormal period and an initial onset period; preliminarily determining the disease source area and suspected disease source nodes, and cross-identifying the infection source; through the integration and weighted fusion algorithm of multi-dimensional data, the accuracy and robustness of disease early warning are improved, the pattern difference from existing diseases is accurately identified, and prevention and control measures are formulated based on the most similar historical disease source network, achieving the effects of accurate identification and targeted prevention and control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a multi-dimensional disease monitoring and early warning method. Background Art

[0002] With the rapid development of intelligent healthcare, monitoring and early warning systems play an increasingly important role in the field of epidemiology. Traditional disease monitoring methods rely on manual data collection and analysis, which are not only inefficient but also prone to false early warnings due to information delays or errors. To overcome these limitations, intelligent healthcare monitoring and early warning systems have been proposed to achieve automated monitoring and timely early warning of disease transmission.

[0003] The Chinese invention patent with the application number 202311679085.6 provides a method for intelligent healthcare monitoring and early warning based on disease diagnosis data. By using the epidemiological trend heat map blocks of the first x - 1 case epidemiological investigation events, a trend heat matching vector is determined for the xth case epidemiological investigation event in the epidemiological trend heat map stream. Based on the feature differences between the P trend heat embedding vectors and the trend heat matching vector of the xth case epidemiological investigation event, the epidemiological trend heat map block of the xth case epidemiological investigation event is determined, and then a judgment result is obtained.

[0004] In the prior art patents, the target knowledge vectors generated based on existing cases and historical data are used to identify and compare new case epidemiological investigation events, and a discrimination result is obtained for prevention and control, which solves the problem of data delay in traditional manual reporting. However, the accuracy of its discrimination result needs to be based on a large amount of existing case data, resulting in the inability to effectively monitor and early warn of infectious diseases in the initial stage, delaying the prevention and control time. Summary of the Invention

[0005] By providing a multi-dimensional disease monitoring and early warning method, the present application solves the problem in the prior art that infectious diseases cannot be effectively monitored and early warned in the initial stage, and achieves the technical effect of accurately identifying the initial stage of infectious diseases for monitoring and early warning.

[0006] The present application provides a multi-dimensional disease monitoring and early warning method, and the method includes:

[0007] S100: Obtain multi-dimensional data to generate a multi-dimensional data set, and the multi-dimensional data set includes a number of multi-dimensional feature vectors;

[0008] S200: Establish a historical disease source network based on the cube, obtain existing case data to establish a real-time disease source network. In the historical disease source network, the fusion feature vectors of multiple cases are used as historical network nodes, and the disease change differences are used as edges for connection; in the real-time disease source network, the multi-dimensional feature vectors of a single case are used as real-time network nodes, and the change differences between different cases are used as edges for connection; the fusion feature vector is generated by fusing the multi-dimensional feature vectors of multiple cases.

[0009] S300: Compare the real-time disease source network with multiple historical disease source networks respectively according to the hierarchical comparison principle to obtain the most similar historical disease source, set corresponding prevention and control measures, and conduct monitoring and early warning.

[0010] S400: Mark key areas around each real-time network node according to geographic information system technology to form punctuation data, assign a unique identifier to each punctuation data, and record its basic information; continuously monitor the punctuation data, and define the abnormal period and the initial period; preliminarily determine the disease source area and suspected disease source nodes, and cross-identify the infection source.

[0011] Further, the hierarchical comparison principle is to compare the real-time disease source network with multiple historical disease source networks one by one through different levels, calculate the comprehensive similarity, and select the historical disease source network with the highest comprehensive similarity.

[0012] Further, the hierarchical comparison principle also includes: sorting the historical disease source networks in descending order according to the comprehensive similarity, presetting a similarity threshold, conducting spatio-temporal overlap analysis on the historical disease source networks higher than the similarity threshold and the real-time disease source network to obtain several suspected paths of precursor key nodes, screening the suspected paths according to the pre-set condition filtering mechanism, and calculating the difference value between multiple suspected paths and the actual propagation path; calculating the estimated path from the terminal propagation node to the next real-time network node according to the difference value, and presetting prevention and control measures.

[0013] Further, the condition filtering mechanism includes mandatory constraint conditions and flexible evaluation conditions. All suspected paths are preliminarily screened according to the mandatory constraint conditions, and the preliminarily screened suspected paths are evaluated item by item according to the flexible evaluation conditions to give an evaluation score. Multiply the evaluation score by its corresponding weight to obtain the weighted score of this condition. Add up all the weighted scores to obtain the final weight of this suspected path; the mandatory constraint conditions include time consistency, regional rationality, and type matching; the flexible evaluation conditions include quantity similarity, speed consistency, and reference of other factors.

[0014] Further, the spatio-temporal overlap analysis regards the real-time source network and the historical source network as graph structures, selects the most similar real-time network nodes and historical network nodes as the comparison starting nodes, and uses a graph matching algorithm for overlap comparison according to the comparison starting nodes to identify the overlapping parts between the real-time source network and the historical source network.

[0015] Further, the precursor key node is the penultimate real-time network node of the real-time source network, and the terminal propagation node is the last real-time network node of the real-time source network.

[0016] Further, the suspected path refers to the development path from multiple historical network nodes overlapping with the precursor key node to the next historical network node; the actual propagation path is the propagation change path between the precursor key node and the terminal propagation node, and the difference value is the comprehensive calculation value of multiple suspected paths and the actual propagation path.

[0017] Further, the abnormal period refers to the abnormal situation in the punctuation data monitoring in a short period of time, and the initial stage refers to the situation of continuous abnormality and the increase of case data in a period of time; monitoring rules are set for the abnormal period and the initial stage respectively.

[0018] Further, step S400 also includes: performing a cut-off screening on the areas between multiple different real-time network nodes around the suspected source node. The cut-off screening is to calculate the distances between different real-time network nodes and the distances of punctuation data between real-time network nodes to form multiple radiation areas for secondary screening; identifying the overlapping areas of multiple radiation areas, calculating the intersection points of adjacent radiation areas, connecting the intersection points to form an overlapping area, and performing in-depth screening on the overlapping area; finally determining the disease area and the disease source.

[0019] Further, the secondary screening is to perform a preliminary screening on the real-time network nodes in the radiation area, and evaluate the preliminarily screened real-time network nodes according to the pre-determined node risk level evaluation index to obtain high-risk nodes and high-risk areas; the in-depth screening is to use a clustering algorithm to cluster the real-time network nodes in the overlapping area, divide them into different groups according to the relevance and similarity between the real-time network nodes, and apply association rule mining technology to identify the potential connections between different groups, and finally determine the disease source and the actual transmission method.

[0020] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0021] Through the integration and weighted fusion algorithm of multi-dimensional data, the accuracy and robustness of disease early warning are improved. Constructing a real-time disease source network and a historical disease source network can monitor the disease transmission status in real time and accurately identify the pattern differences from existing diseases. Based on the most similar historical disease source network, prevention and control measures are formulated, achieving the effects of accurate identification and targeted prevention and control. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is a schematic diagram of the overall process of the multi-dimensional disease monitoring and early warning method in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] To facilitate the understanding of the present invention, the present application will be described more comprehensively with reference to the relevant drawings; the drawings show preferred embodiments of the present invention. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein; on the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs; the terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention; the term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0025] Example 1: As Figure 1 shown, a multi-dimensional disease monitoring and early warning method, the method includes:

[0026] S100: Obtain multi-dimensional data to generate a multi-dimensional data set, and the multi-dimensional data set includes a number of multi-dimensional feature vectors.

[0027] In some embodiments, the multi-dimensional data is obtained relying on a health care big data platform and a municipal government affairs exchange and sharing platform. The obtained multi-dimensional data uses machine learning algorithms to automatically identify and correct data errors, improving the efficiency and accuracy of data cleaning. The standardization method is dynamically adjusted according to data characteristics to ensure the consistency and comparability of data at different times and in different situations. Key indicators related to diseases are extracted from the data, and multi-dimensional features such as incidence rate, transmission speed, public attention, and meteorological factors are extracted by combining multi-source data. An early warning model based on multi-source data fusion is constructed, and a weighted fusion algorithm is used to improve the accuracy and robustness of early warning.

[0028] In some embodiments, the multidimensional dataset includes medical treatment data, personnel information and drug data, Internet data, disease monitoring subject library data information, etc. The Internet data is the network search data, discussion data, meteorological data, etc. generated in social media, and the disease monitoring subject library data information refers to syndrome monitoring, fever clinic monitoring, risk population monitoring, drug monitoring and disease epidemic intensity monitoring data.

[0029] In some embodiments, statistical analysis is conducted based on data such as the number of fever clinic visits, main symptoms, and epidemiological history. The data is plotted using a combination of averages and control charts, using control chart-related algorithms. Control rules are then used to identify outliers in the statistical charts. Simultaneously, intelligent early warning analysis models are generated based on different symptoms, different moving average methods, and different rules to enable real-time monitoring of clusters and high-incidence populations and symptoms. Early warning of infectious diseases and monitoring of epidemic intensity are achieved through monitoring of drug retail information at pharmacies. Analysis is conducted on shared daily sales data of medical insurance drugs at retail pharmacies, specifically focusing on daily sales of cold and fever-reducing drugs, cough suppressants, antidiarrheal drugs, and anti-inflammatory drugs. The EARS-3Cs early warning model is used to monitor and warn based on sales over the previous seven days. At the same time, the prescription volume of the top drugs in hospital prescriptions is monitored, with a focus on the ranking and usage trends of hormones, anticoagulants, and antibiotics.

[0030] In some embodiments, the multidimensional dataset includes several multidimensional feature vectors. A multidimensional feature vector is a vector containing multiple features (or attributes) that collectively describe multiple aspects of an entity or event. For infectious disease monitoring and early warning systems, multidimensional feature vectors contain various information related to disease spread, such as incidence rate, transmission rate, public attention, and meteorological factors. For example, a multidimensional feature vector is constructed containing the following features: date xxx; region xxx; incidence rate xxx; transmission rate xxx (relative value); public attention xxx (number of social media mentions); average temperature xxx; relative humidity xxx; cold medicine sales xxx; number of hospital visits on the same day xxx; top medication prescription volume xxx. In this example, the multidimensional feature vector contains 11 features, each with a corresponding eigenvalue, which collectively describe the disease development status in the region on a specific date.

[0031] S200: Establish a historical pathogenicity network based on a multidimensional dataset, and obtain existing case data to establish a real-time pathogenicity network. In the historical pathogenicity network, the fused feature vectors of multiple cases are used as historical mesh nodes, and the differences in disease changes are used as edges for connection; in the real-time pathogenicity network, the multidimensional feature vectors of a single case are used as real-time mesh nodes, and the differences in changes between different cases are used as edges for connection.

[0032] In some embodiments, the historical disease source network is a network constructed based on historical case data for analyzing the historical patterns and trends of disease transmission, which includes complete records of all past disease transmissions from the start, development to the end. Relevant case data is collected from the historical case database, including patients' basic information, diagnosis information, treatment information, etc. The collected data is cleaned and standardized to ensure data quality and consistency. Features related to disease transmission, such as the location of the patient, onset time, contact history, etc., are extracted from the cleaned data. Based on the extracted features, a historical disease source network is constructed using graph theory or network science methods, and each disease generates a corresponding historical disease source network.

[0033] In the historical disease source network, the fusion feature vectors of multiple cases are used as historical network nodes. The fusion feature vector is generated by fusing the multi-dimensional feature vectors of multiple cases. Suppose there are multi-dimensional feature vectors of n cases, and each vector contains m features, then the fusion feature vector F can be expressed as:

[0034]

[0035] Where, represents the j-th eigenvalue of the i-th case. The average value of each feature in the multi-dimensional feature vectors of multiple cases is obtained and combined to get the fusion feature vector. The historical network node represents the comprehensive feature status of multiple cases at a specific time node. The comprehensive feature vector obtained based on the relevant disease conditions of multiple cases under this node is connected with the disease change difference as the edge. The disease change difference refers to one or more change differences in the comprehensive feature vectors of two historical network nodes. The change difference can be determined by calculating metrics such as the Euclidean distance and cosine similarity between two fusion feature vectors. Analyze the historical disease source network to identify the key nodes and paths of disease transmission and discover potential transmission rules and patterns. In the real-time disease source network, the information feature vector of a single case is used as a real-time network node, and the change differences between different cases are used as edges for connection. A real-time network node represents the information of an actual case data, which should at least include date, location, onset time, contact history, basic disease condition information, medical record, adverse reactions, etc. The connection line between different real-time network nodes, i.e., the edge, represents the change difference between the two, and the change difference can also be determined by calculating the similarity or distance between two multi-dimensional feature vectors. For example, calculate the difference between two feature vectors using the Euclidean distance:

[0036]

[0037] Where, and represent feature vector 1 and feature vector 2 respectively. and respectively represent the j-th eigenvalue in two vectors, and m represents the total number of features. Both the historical mesh nodes and the real-time mesh nodes are connected with multiple edges.

[0038] S300: Compare the real-time disease source network with multiple historical disease source networks respectively according to the hierarchical comparison principle to obtain the most similar historical disease source, set corresponding prevention and control measures, and conduct monitoring and early warning.

[0039] In the prior art, several multi-dimensional feature vectors are obtained from the historical disease source network, and the feature differences are directly obtained by comparing multiple multi-dimensional feature vectors with the case feature vector generated by the currently newly determined case. The feature difference refers to the difference in each feature dimension between different data entities. According to different differences, the most similar known disease pattern is determined, and then the discrimination result is obtained for prevention and control.

[0040] The hierarchical comparison principle is to compare the real-time disease source network with multiple historical disease source networks one by one through different levels, calculate the comprehensive similarity, and select the historical disease source network with the highest similarity.

[0041] In some embodiments, it is set to be divided into a time layer, a region layer, and a feature layer to ensure that the data of all historical disease source networks are complete and standardized. From the time layer aspect: according to the timestamp of the real-time disease source network, a reasonable time window is determined (such as the past year, half year, or a specific season), and the historical disease source networks within this time window are selected as the preliminary candidate set, and further screen the historical disease source networks with the same or similar disease types as those in the real-time disease source network; from the region layer aspect: according to the region where the real-time disease source network is located, the geographical region is divided into several sub-regions, and in the preliminary candidate set, the historical disease source networks with the same region as the real-time disease source network or similar geographical features are retained; from the feature layer aspect: according to the requirements of disease monitoring and early warning, key features are selected for comparison, such as incidence rate, transmission speed, public attention, meteorological factors, etc. For each feature, the cosine similarity can be used to calculate the similarity between the real-time disease source network and the candidate historical disease source network. Assume that the feature vector of the real-time disease source network is , and the feature vector of the historical disease source network is , and the direction similarity of the two vectors in the multi-dimensional space The calculation formula is as follows:

[0042]

[0043] Taking into account the similarities of all features, a weighted average method is used to obtain a comprehensive similarity score. According to the comprehensive similarity score, the candidate historical source networks are sorted, and the historical source network with the highest similarity is selected as the most similar historical source. The disease transmission patterns and trends of the most similar historical source network are analyzed in depth. Based on the analysis results and combined with the current actual situation, targeted prevention and control measures are formulated.

[0044] The technical solutions in the embodiments of the present application at least have the following technical effects or advantages:

[0045] Through the integration of multi-dimensional data and the weighted fusion algorithm, the present application improves the accuracy and robustness of disease early warning. Constructing a real-time source network and a historical source network can monitor the disease transmission status in real time and accurately identify the pattern differences from existing diseases. Based on the most similar historical source network, prevention and control measures are formulated, achieving the effects of accurate identification and targeted prevention and control.

[0046] Embodiment 2: In Embodiment 1, by constructing a multi-dimensional feature vector and adopting a weighted fusion algorithm, the accuracy and reliability of disease early warning are improved. This embodiment makes further improvements on the basis of the above embodiment.

[0047] The hierarchical comparison principle further includes: sorting the historical source networks in descending order according to the comprehensive similarity, presetting a similarity threshold, performing spatio-temporal overlap analysis on the historical source networks higher than the similarity threshold and the real-time source network to obtain several suspected paths of the precursor key nodes, filtering the suspected paths according to a pre-set condition filtering mechanism, and calculating the difference values between the multiple suspected paths and the actual propagation path; calculating the estimated path from the terminal propagation node to the next real-time network node according to the difference value, and presetting prevention and control measures.

[0048] In some embodiments, several historical source networks after hierarchical comparison are obtained according to the above content, sorted in descending order according to the comprehensive similarity of each historical source network, a similarity threshold is preset according to historical experience and experimental data, and spatio-temporal overlap analysis is performed on the historical source networks higher than the similarity threshold and the real-time source network. The spatio-temporal overlap analysis regards the real-time source network and the historical source network as graph structures, selects the most similar real-time network node and historical network node as the comparison starting nodes. Here, the similarity calculation is performed according to the specific vector data of the real-time network node and the historical network node, including the analysis and comparison of different levels (time level, region level, and feature level), and the most similar one is selected. In this embodiment, the real-time network node is the first node in the real-time source network, but the selected historical network node is not necessarily the first node in the historical source network, and only needs to be selected according to the similarity; using a graph matching algorithm for overlap comparison based on the comparison starting nodes to identify the overlapping part between the real-time source network and the historical source network.

[0049] In some embodiments, the method for setting the similarity threshold is as follows: Select multiple historical disease source networks. Since the disease information of the historical disease source networks is known information data, use the same method to judge the similarity between multiple historical disease source networks to obtain the comprehensive similarity of the historical disease source networks. Set an appropriate similarity threshold through the existing similarity calculation to ensure that effective historical disease source networks for comparison can be selected. To ensure the analysis efficiency of the historical disease source networks and the real-time disease source network, calculate the number of historical disease source networks with a similarity higher than the similarity threshold. If the number is large, for example, exceeding the preset maximum selection number, then adopt a dynamic adjustment method. For example, select the top 25% of the total number; the maximum selection number is set in the same way as the above similarity threshold method, and the number is set using the analysis situation of the historical disease source networks to improve the comparison efficiency while ensuring the comparison effect.

[0050] In some embodiments, the suspected path refers to the development path from multiple historical network nodes overlapping with the precursor key node to the next historical network node; the actual propagation path is the propagation change path between the precursor key node and the terminal propagation node, and the difference value is the comprehensive calculation value of multiple suspected paths and the actual propagation path; the precursor key node is the penultimate real-time network node of the real-time disease source network, and the terminal propagation node is the last real-time network node of the real-time disease source network. Assume that the actual development path of the real-time disease source network is , where represents the i-th node on the path, and the suspected path obtained from the historical disease source network is , where represents the first suspected path, represents the i-th node on the suspected path. The difference value adopts a calculation method based on node similarity. Define the node similarity function , which represents the similarity between nodes and . The path difference value D is defined as the negative average value of the corresponding node similarities, and the calculation formula is as follows: , where ranges from (0, 1), indicating from completely dissimilar to completely identical. The smaller the difference value D, the more similar the two paths are.

[0051] In some embodiments, a difference value threshold is preset to judge the similarity degree between the suspected path and the actual path. Select the paths with a difference value less than from all the suspected paths as candidate paths. For each candidate path, check the next node after its corresponding terminal propagation node, obtain the change trend of the front and back nodes, and correspondingly estimate the next change trend of the real-time disease source network.

[0052] In this application, another method for screening suspected paths is also provided. Using a conditional screening method, a conditional filtering mechanism is set up. The conditional filtering mechanism includes mandatory constraint conditions and flexible evaluation conditions. The mandatory constraint conditions include time consistency, regional rationality, and type matching; the flexible evaluation conditions include quantity similarity, speed consistency, and reference to other factors. All suspected paths are initially screened according to the mandatory constraint conditions, and the suspected paths after the initial screening are evaluated item by item according to the flexible evaluation conditions, and an evaluation score is given. Multiply the evaluation score by its corresponding weight to obtain the weighted score of this condition. Add up all the weighted scores to obtain the final weight of this suspected path. The mandatory constraint conditions represent the inviolable basic principles and laws in disease transmission analysis, which are obtained based on scientific research and actual data and must be strictly observed to ensure the accuracy and reliability of the analysis results. For example, time consistency, geographical location rationality, and virus type matching all belong to the mandatory constraint conditions; the flexible evaluation conditions are auxiliary indicators used to further evaluate the rationality and possibility of suspected paths on the basis of meeting the mandatory constraint conditions, providing additional information and enabling a more comprehensive understanding of the potential patterns and factors of disease transmission. For example, case quantity similarity, transmission speed consistency, and consideration of environmental factors can all be regarded as flexible evaluation conditions. Suppose the suspected path is P, the set of mandatory constraint conditions is , and the set of flexible evaluation conditions is . The corresponding weight settings are and respectively, and the condition scores are and respectively; the calculation formula for the final weight is:

[0053]

[0054] where represents the total weighted score of the mandatory constraint conditions, and represents the total weighted score of the flexible evaluation conditions.

[0055] In some embodiments, an estimated path from the terminal transmission node to the next real-time mesh node is calculated according to the difference value, and prevention and control measures are preset. The estimated path is the estimation of the next node of the current real-time disease source network (i.e., a new case or a new symptom change of the current case). According to the estimation result, prevention and control means and related measures are implemented.

[0056] The technical solutions in the above embodiments of this application have at least the following technical effects or advantages:

[0057] Through spatio-temporal overlap analysis and path difference value calculation, this application can more accurately identify historical disease source networks similar to the real-time disease source network, improving the accuracy of disease early warning; based on the estimated path and conditional filtering mechanism, it can formulate prevention and control measures more precisely, improving the prevention and control effect; by setting the similarity threshold and the maximum selection quantity, it reduces ineffective comparison and improves the analysis efficiency, achieving the effect of accurately judging the disease transmission mode and making precise predictions for early prevention.

[0058] Embodiment 3: In the above embodiment, the identification and monitoring of diseases are realized, and the transmission mode and onset characteristics of the diseases are determined. This embodiment makes further improvements on the basis of the above embodiment.

[0059] The method further includes: S400: Mark key areas around each real-time network node according to geographic information system technology to form punctuation data, assign a unique identifier to each punctuation data, and record its basic information; continuously monitor the punctuation data, and define the abnormal period and the initial onset period; preliminarily determine the disease source area and suspected disease source nodes, and cross-identify the infection source.

[0060] In some embodiments, the key areas refer to pharmacies, clinics, hospitals, etc. existing around a case. The basic information of each punctuation data includes information data such as the geographical location, type, scale, and number of floating population. The distance selection of the punctuation data should be based on a comprehensive consideration of the actual possibility of disease transmission and monitoring efficiency. In the above embodiments, the monitoring of the current disease pattern and the prediction of the development path are given, and the basis is still the data foundation that needs to have cases. When the case data is very few, basic monitoring and prevention and control should be the main focus. In this embodiment, the key areas are marked to form punctuation data, and through the real-time monitoring of the punctuation data, the occurrence of abnormal situations is timely detected. The abnormal period refers to the stage when abnormal conditions occur in the monitoring of punctuation data in a short period of time, but the emergence of an infectious disease has not yet been confirmed. Abnormal phenomena include a sharp increase in drug sales, an increase in the number of cases with specific symptoms, etc. Among them, according to specific experiments, the short period of time refers to that within a continuous week, the punctuation data shows continuous abnormalities, which often serves as the initial stage of the occurrence of a disease; the initial onset period refers to the situation where continuous abnormalities occur within a period of time and the case data increases. The initial onset period is determined after the abnormal period has occurred for a period of time. The period of time is about one to three months, and with a significant increase in case data, it is basically determined that the infectious disease has emerged, such as the appearance of a large number of cases with the same symptoms, a sharp increase in the number of hospital admissions, etc.; the method for distinguishing the abnormal period and the initial onset period is to set monitoring thresholds. According to historical data and expert experience, monitoring thresholds are set for different types of punctuation data (such as the growth rate of drug sales, the number of cases, etc.); real-time monitoring and comparison are carried out. By real-time monitoring the monitoring situation of the punctuation data and comparing it with the set thresholds, the abnormal period and the initial onset period are distinguished. The normal sales volume of specific drugs and the number of cases at ordinary times are set as the safe period, and the difference is distinguished through comparison; the monitoring situations of multiple punctuation data are combined for comprehensive analysis to improve the accuracy of distinction.

[0061] In some embodiments, monitoring rules are set respectively for the abnormal period and the initial onset period. For the monitoring of the abnormal period, the monitoring frequency of the punctuation data should be strengthened, the development and changes of abnormal phenomena should be closely monitored in real time, a preliminary investigation and analysis of the abnormal phenomena should be carried out, and emergency preparations should be made to be ready to enter the monitoring state of the initial onset period at any time; for the monitoring of the initial onset period, the monitoring frequency should be further increased to ensure the real-time nature and accuracy of the data, the suspected disease source nodes should be monitored key points, and in combination with the monitoring methods of Embodiment 1 and Embodiment 2 above, the characteristics and infection patterns of the current disease are determined according to the historical disease source network, the development path of the next stage is predicted, and further precise prevention and control are carried out.

[0062] In some embodiments, the preliminary determination of the disease source area and suspected disease source nodes is set for the initial stage. During the initial stage, based on the monitoring of punctuation data and combined with the data of real-time mesh nodes, the trend and path of disease transmission are analyzed. Combining the above-described technical solutions, the patterns and characteristics of the disease are determined, and prevention and control strategies are set accordingly. Through GIS technology, spatial overlay analysis is performed on the punctuation data and real-time mesh nodes to identify the disease source area and suspected disease source nodes. The disease source area generally refers to the area where the disease transmission is most concentrated, while the suspected disease source node may be the location where the disease first appears. Combining the data information of the historical disease source network, further analysis is performed on the suspected disease source nodes. By comparing the monitoring data at different time points, the transmission patterns and rules of infectious diseases are identified, and multi-dimensional information such as geographical location, population flow, and environmental factors is comprehensively analyzed to cross-identify the infection source.

[0063] The technical solutions in the embodiments of the present application at least have the following technical effects or advantages:

[0064] By using GIS technology to mark key areas and continuously monitor, the present application can more accurately capture the early signs of disease transmission, improve the monitoring accuracy and efficiency; by defining the abnormal period and the initial stage, different stages of disease transmission can be identified in a timely manner, achieving the effects of early warning and pre-prevention; through GIS spatial overlay analysis and comprehensive analysis of multi-dimensional information, the effects of accurately positioning the disease source area and suspected disease source nodes are achieved.

[0065] Embodiment 4: In the above embodiment, real-time monitoring is performed by setting punctuation data based on multi-dimensional data information, which can timely detect disease abnormalities and then carry out prevention and control. This embodiment makes further improvements on the basis of the above embodiment.

[0066] Step S400 further includes: performing cut-off screening on the areas between multiple different real-time mesh nodes around the suspected disease source node. The cut-off screening is to calculate the distances between different real-time mesh nodes and the distances of punctuation data between real-time mesh nodes to form multiple radiation areas for secondary screening; identifying the overlapping areas of the multiple radiation areas, calculating the intersection points of adjacent radiation areas, connecting the intersection points to form an overlapping area, and performing in-depth screening on the overlapping area; finally determining the disease area and the disease source.

[0067] In some embodiments, based on Embodiment III, the key regions are marked and punctuation data is formed, and its basic information is recorded. Meanwhile, in the network structure, the suspected disease source node and multiple different network nodes around it are determined, the punctuation data existing between the suspected disease source node and the remaining nodes is calculated, and then the distances between the suspected disease source node and the punctuation data are calculated in sequence. Taking the distance from the suspected disease source node as the radius, a radiation area is formed, and the radiation area is screened secondly; the overlapping regions of multiple radiation areas are identified, the intersection points of adjacent radiation areas are calculated, the intersection points are connected to form an overlapping area, and the overlapping area is screened deeply; for example, there are multiple pharmacies in the area between the suspected disease source node and node a. When obtaining the relevant disease drugs sold by the pharmacies, it will be found that there are suspected pharmacies 1, 2, 3, and 4. Calculate the distance between the suspected disease source node and pharmacy 1, and use pharmacy 1 as the origin and the distance as the radius to search to form radiation area 1, and secondly screen the network nodes in the radiation area; similarly, calculate the distance between pharmacy 1 and pharmacy 2, and use pharmacy 2 as the origin and the distance as the radius to search to form radiation area 2. The intersection area of the front and back radiation areas 1 and radiation area 2 is set as the overlapping area, the intersection points on both sides of the overlapping area are connected, and the data information in the overlapping area is screened deeply to finally determine the disease source.

[0068] In some embodiments, the secondary screening is to initially screen the real-time network nodes in the radiation area. The initial screening is to exclude the real-time network nodes that have nothing to do with the disease or have little impact, and evaluate the real-time network nodes after the initial screening according to the pre-determined node risk level evaluation index to obtain high-risk nodes and high-risk areas; the deep screening is to use the clustering algorithm to cluster the real-time network nodes in the overlapping area, divide them into different groups according to the relevance and similarity between the real-time network nodes, apply the association rule mining technology to identify the potential connections between different groups, and finally determine the disease source and the actual transmission method, which helps to identify the nodes with similar characteristics and may be the key path of disease transmission. The secondary screening is mainly based on the weighted scoring method to evaluate the risk level of the nodes, which is an extension of the preliminary screening, while the deep screening pays more attention to the relevance and similarity between the nodes, as well as the change trend of the data, and is a more in-depth analysis based on the secondary screening. The deep screening is based on the LSTM neural network to evaluate the data change trend. The time series data (such as the number of daily cases, population flow data, etc.) of each node in the overlapping area is pre-processed, including data cleaning, normalization, etc. Determine the number of layers of the LSTM network and the number of neurons in each layer, select the activation function (such as ReLU, Sigmoid, etc.) and the loss function (such as mean square error MSE), use the training data to train the LSTM model, optimize the model parameters through the backpropagation algorithm, use the trained LSTM model to predict the test data, evaluate the prediction performance of the model, and analyze the change trend of the data of each node according to the prediction results to identify the nodes with abnormal changes or trend reversals; its calculation formula is:

[0069]

[0070] wherein represents the hidden state at time step t, which is a vector containing the internal state information of the LSTM cell at the current time step; represents the activation function, and different activation functions are used for calculation according to different actual situations; represents the weight matrix from the hidden state to the hidden state, which determines how the hidden state of the previous time step affects the hidden state of the current time step; represents the hidden state of the previous time step; represents the weight matrix from the input data to the hidden state, which determines how the input data of the current time step affects the hidden state; represents the input data at time step t, which is a vector containing the external information received by the LSTM cell at the current time step; represents the bias term, which is used to adjust the calculation of the hidden state.

[0071] In some embodiments, by combining multi-dimensional data information (such as geographical location, population flow, environmental factors, etc.) with the analysis of mesh node data, a screening mode is set. In the overlapping area, through data comparison and trend analysis, the source of the disease can be quickly located. According to the located source of the disease, the disease area can be locked. Combining with GIS technology, targeted prevention and control measures can be formulated, such as resource allocation, setting up quarantine areas, controlling personnel flow, etc., continuously monitoring the prevention and control effect, and timely adjusting the prevention and control strategy according to the monitoring results.

[0072] The technical solutions in the embodiments of the present application at least have the following technical effects or advantages:

[0073] Through cut-off screening, secondary screening and in-depth screening, and combining with the evaluation of the data change trend of the LSTM neural network, the present application can more accurately locate the source of the disease and the transmission path; by combining multi-dimensional data information and GIS technology, more accurate and effective prevention and control measures can be formulated, improving the prevention and control efficiency, and achieving the effect of accurately locating the source of the disease and the disease source area.

[0074] The above is only the preferred embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A multi-dimensional disease monitoring and early warning method, characterized in that The method includes: S100: Obtain multi-dimensional data to generate a multi-dimensional data set, where the multi-dimensional data set includes a number of multi-dimensional feature vectors; S200: Establish a historical disease source network based on the multi-dimensional data set, and obtain existing case data to establish a real-time disease source network. In the historical disease source network, the fusion feature vectors of multiple cases are used as historical network nodes, and the disease change differences are used as edges for connection; in the real-time disease source network, the multi-dimensional feature vectors of a single case are used as real-time network nodes, and the change differences between different cases are used as edges for connection; the fusion feature vector is generated by fusing the multi-dimensional feature vectors of multiple cases; S300: Compare the real-time disease source network with multiple historical disease source networks respectively according to the hierarchical comparison principle to obtain the most similar historical disease source, set corresponding prevention and control measures, and conduct monitoring and early warning; S400: Mark key areas around each real-time network node according to geographic information system technology to form punctuation data, assign a unique identifier to each punctuation data, and record its basic information; continuously monitor the punctuation data, and define the abnormal period and the initial period; preliminarily determine the disease source area and suspected disease source nodes, and cross-identify the infection source; The hierarchical comparison principle also includes: sorting the historical disease source networks in descending order according to the comprehensive similarity, presetting a similarity threshold, conducting spatio-temporal overlap analysis on the historical disease source networks higher than the similarity threshold and the real-time disease source network to obtain several suspected paths of precursor key nodes, filtering the suspected paths according to a preset condition filtering mechanism, and calculating the difference value between multiple suspected paths and the actual propagation path; calculating the estimated path from the terminal propagation node to the next real-time network node according to the difference value, and presetting prevention and control measures; the precursor key node is the second-to-last real-time network node of the real-time disease source network, and the terminal propagation node is the last real-time network node of the real-time disease source network; the suspected path refers to the development path from multiple historical network nodes overlapping with the precursor key node to the next historical network node; the actual propagation path is the propagation change path between the precursor key node and the terminal propagation node, and the difference value is the comprehensive calculation value of multiple suspected paths and the actual propagation path; Step S400 also includes: performing a cut-off type screening on the areas between multiple different real-time network nodes around the suspected disease source node. The cut-off type screening is to calculate the distances between different real-time network nodes and the distances of punctuation data between real-time network nodes to form multiple radiation areas for secondary screening; identifying the overlapping areas of multiple radiation areas, calculating the intersection points of adjacent radiation areas, connecting the intersection points to form an overlapping area, and performing in-depth screening on the overlapping area; finally determining the disease area and the disease source; among them, the distances between the suspected disease source node and the punctuation data are calculated in turn as the radii to form radiation areas.

2. The multi-dimensional disease monitoring and early warning method according to claim 1, wherein The hierarchical comparison principle is to compare the real-time disease source network with multiple historical disease source networks one by one through different levels, calculate the comprehensive similarity, and select the historical disease source network with the highest comprehensive similarity.

3. The multi-dimensional disease monitoring and early warning method according to claim 1, wherein The above-mentioned conditional filtering mechanism includes mandatory constraint conditions and flexible evaluation conditions. All suspected paths are preliminarily screened according to the mandatory constraint conditions, and the suspected paths after preliminary screening are evaluated item by item according to the flexible evaluation conditions, and an evaluation score is given. The evaluation score is multiplied by its corresponding weight to obtain the weighted score of this condition. All weighted scores are added together to obtain the final weight of this suspected path; the mandatory constraint conditions include time consistency, regional rationality, and type matching; the flexible evaluation conditions include quantity similarity, speed consistency, and reference of other factors.

4. The multi-dimensional disease monitoring and early warning method according to claim 1, wherein The above-mentioned spatio-temporal overlap analysis regards the real-time pathogen network and the historical pathogen network as graph structures, selects the most similar real-time network nodes and historical network nodes as the comparison starting nodes, and uses the graph matching algorithm for overlap comparison according to the comparison starting nodes to identify the overlap between the real-time pathogen network and the historical pathogen network.

5. The multi-dimensional disease monitoring and early warning method according to claim 1, characterized in that The abnormal period refers to the abnormal situation that occurs in the punctuation data monitoring within a short period of time, and the initial period refers to the situation where continuous abnormalities occur within a period of time and the case data increases; monitoring rules are set for the abnormal period and the initial period respectively.

6. The multi-dimensional disease monitoring and early warning method according to claim 1, wherein The secondary screening is to conduct a preliminary screening on the real-time network nodes in the radiation area, and evaluate the real-time network nodes after preliminary screening according to the pre-determined node risk level evaluation indicators to obtain high-risk nodes and high-risk areas; The in-depth screening is to use the clustering algorithm to cluster the real-time network nodes in the overlapping area, divide them into different groups according to the relevance and similarity between the real-time network nodes, apply the association rule mining technology to identify the potential connections between different groups, and finally determine the disease source and the actual transmission method.

Citation Information

Patent Citations

  • Intelligent medical monitoring and early warning method and system based on disease diagnosis data

    CN118098625A

  • Food-borne disease outbreak identification method and system based on link prediction

    CN114049966A

  • Early warning method and system for infectious diseases and readable storage medium

    CN114141385A