A respiratory infectious disease monitoring method, device, equipment and storage medium
By employing principal component analysis and a dynamic threshold generation mechanism, the problems of noise interference and static thresholds in existing respiratory infectious disease monitoring methods are solved, enabling intelligent cleaning of multidimensional data and timely and accurate early warning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 四川国际旅行卫生保健中心(成都海关口岸门诊部)
- Filing Date
- 2026-02-03
- Publication Date
- 2026-05-01
AI Technical Summary
Existing methods for monitoring respiratory infectious diseases rely on a single core indicator, are susceptible to noise interference, have static thresholds that are difficult to respond sensitively to changes in the epidemic, and lack multi-dimensional data intelligent identification and cleaning mechanisms, leading to false alarms, missed reports, and delayed early warnings.
Principal component analysis is used to reduce and reconstruct the dimensionality of multidimensional monitoring time series data, identify and remove abnormal cycles, and generate dynamic early warning thresholds by combining recent data clustering analysis. The early warning is triggered by continuous verification logic.
It improved the stability of monitoring methods and the timeliness of early warning, reduced the risk of false alarms, and enabled early and sensitive response to the epidemic and accurate early warning during the peak period.
Smart Images

Figure CN121617659B_ABST
Abstract
Description
A method, device, equipment and storage medium for monitoring respiratory infectious diseases Technical Field
[0001] This invention relates to the field of infectious disease surveillance technology, and more specifically, to a method, apparatus, equipment, and storage medium for monitoring respiratory infectious diseases. Background Technology
[0002] Respiratory infectious diseases (such as influenza) spread rapidly and pose a continuous threat to public health security. Early monitoring and warning are crucial for prevention and control. Current mainstream monitoring methods mainly rely on historical threshold comparisons of a single core indicator (such as the percentage of influenza-like illness cases). While simple and easy to implement, these methods have significant drawbacks: First, single-indicator models are susceptible to occasional noise interference such as holidays and abnormal data reporting, leading to false alarms or missed reports. Second, static thresholds based on long-term history or fixed rules are difficult to respond sensitively to dynamic changes in the early stages of an epidemic or abnormal growth caused by new pathogens, resulting in delayed warnings. Third, they fail to effectively integrate multi-dimensional auxiliary information such as outpatient volume and patient demographics, limiting the comprehensiveness of risk perception and early identification capabilities.
[0003] Current monitoring technologies still face two major bottlenecks in practical applications: First, there is a lack of intelligent identification and cleaning mechanisms for non-epidemic anomalies in multidimensional data, which can contaminate the model and affect judgment. Second, most early warning thresholds are still static or based on historical data, which cannot be quickly and adaptively adjusted according to the recent spread, resulting in insufficient sensitivity of early warning when the epidemic changes rapidly.
[0004] Therefore, there is an urgent need for an intelligent monitoring method for respiratory infectious diseases that can automatically clean multidimensional monitoring data and dynamically generate early warning thresholds based on recent data patterns, in order to improve the timeliness, accuracy and practicality of early warnings. Summary of the Invention
[0005] The purpose of this invention is to provide a method, apparatus, device, and storage medium for monitoring respiratory infectious diseases, so as to improve the above-mentioned problems.
[0006] To achieve the above objectives, the embodiments of this application provide the following technical solutions:
[0007] On one hand, embodiments of this application provide a method for monitoring respiratory infectious diseases, the method comprising:
[0008] Acquire historical multidimensional monitoring time series data, which includes at least one core monitoring indicator sequence and at least one auxiliary monitoring indicator sequence. The core monitoring indicator sequence and the auxiliary monitoring indicator sequence correspond one-to-one in terms of time period. The core monitoring indicator includes the percentage of influenza-like cases, and the auxiliary monitoring indicators include the total number of outpatient visits and the proportion of pediatric outpatient visits.
[0009] Based on historical multidimensional monitoring time series data, abnormal time periods are identified; from the core monitoring indicator sequence and its corresponding historical year-on-year sequence, the data corresponding to the abnormal time periods are removed to obtain the cleaned core monitoring indicator sequence and the cleaned historical year-on-year sequence.
[0010] Based on the cleaned core monitoring indicator sequence and the cleaned historical year-on-year sequence, the dynamic early warning threshold is calculated; early warning information is generated based on the cleaned core monitoring indicator sequence and the dynamic early warning threshold.
[0011] Secondly, embodiments of this application provide a respiratory infectious disease monitoring device, the device comprising:
[0012] The acquisition module is used to acquire historical multidimensional monitoring time series data. The historical multidimensional monitoring time series data includes at least one core monitoring indicator sequence and at least one auxiliary monitoring indicator sequence. The core monitoring indicator sequence and the auxiliary monitoring indicator sequence correspond one-to-one in terms of time period. The core monitoring indicator includes the percentage of influenza-like cases, and the auxiliary monitoring indicators include the total number of outpatient visits and the proportion of pediatric outpatient visits.
[0013] The identification module is used to identify abnormal time periods based on historical multidimensional monitoring time series data; from the core monitoring indicator sequence and its corresponding historical year-on-year sequence, the data corresponding to the abnormal time periods are removed to obtain the cleaned core monitoring indicator sequence and the cleaned historical year-on-year sequence.
[0014] The early warning module is used to calculate the dynamic early warning threshold based on the cleaned core monitoring indicator sequence and the cleaned historical year-on-year sequence; and to generate early warning information based on the cleaned core monitoring indicator sequence and the dynamic early warning threshold.
[0015] Thirdly, embodiments of this application provide a respiratory infectious disease monitoring device, the device including a memory and a processor. The memory is used to store a computer program; the processor is used to execute the computer program to implement the steps of the above-described respiratory infectious disease monitoring method.
[0016] Fourthly, embodiments of this application provide a readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described respiratory infectious disease monitoring method.
[0017] The beneficial effects of this invention are as follows:
[0018] This invention applies principal component analysis to anomaly detection in multidimensional monitoring time-series data. Instead of examining each indicator in isolation, this method reduces and reconstructs the overall pattern formed by multiple indicators over time, identifying global anomalies by calculating the reconstruction error at each time point. This approach can keenly capture "pseudo-normal" points where multiple indicators exhibit coordinated anomalies (such as a sudden drop in outpatient volume due to a holiday in a particular week), as well as subtle anomaly patterns that are difficult to detect with a single indicator. Removing these identified anomalous cycles from the core analysis sequence effectively purifies the data foundation used for early warning calculations, significantly reduces the risk of false alarms caused by data noise and non-epidemic interference factors, and enhances the stability and reliability of the entire monitoring method.
[0019] This invention abandons the traditional approach of relying on long-term fixed thresholds or simple statistics, and designs a dynamic threshold generation mechanism based on comparative learning between recent data and historical data from the same period. The method first focuses on core indicator data within a recent time window (e.g., 15 weeks) and its historical data from the same period, forming data pairs for cluster analysis. The clustering process can unsupervisedly discover clusters of different growth patterns in the current data compared to the historical period. Then, by comprehensively considering the growth rate, occurrence time, and cluster size of each cluster, the most noteworthy target cluster categories are selected. Finally, based on the statistical characteristics (mean and standard deviation) and adjustable sensitivity coefficients of the data within these target categories, candidate warning thresholds are generated, and the most sensitive one is selected as the dynamic warning threshold. This mechanism allows the warning threshold to closely align with the latest epidemic spread, providing more sensitive early signals in the early stages of an epidemic and providing warning lines more consistent with the current level during the peak of the epidemic, achieving scenario-adaptive warning sensitivity.
[0020] This invention does not simply treat a single instance of exceeding the threshold as an early warning, but instead introduces a verification logic for continuous exceedances. When the latest monitored value exceeds the dynamic threshold, the system traces back several consecutive periods to check the persistence of the exceedance. Only when the threshold is continuously breached within a certain number of adjacent periods will a high-level early warning be triggered; otherwise, only a low-level alert may be triggered. This design effectively filters out accidental exceedances caused by weekly data fluctuations, making early warning triggering more rigorous and more consistent with the epidemiological characteristics of infectious disease cluster outbreaks.
[0021] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing embodiments of the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 is a schematic flowchart of the respiratory infectious disease monitoring method according to an embodiment of the present invention;
[0024] Figure 2 is a schematic diagram of the respiratory infectious disease monitoring device described in an embodiment of the present invention;
[0025] Figure 3 is a schematic diagram of the respiratory infectious disease monitoring device described in an embodiment of the present invention. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0027] It should be noted that similar reference numerals or letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0028] Example 1
[0029] As shown in Figure 1, this embodiment provides a method for monitoring respiratory infectious diseases, which includes steps S1, S2 and S3.
[0030] Step S1: Obtain historical multidimensional monitoring time series data. The historical multidimensional monitoring time series data includes at least one core monitoring indicator sequence and at least one auxiliary monitoring indicator sequence. The core monitoring indicator sequence and the auxiliary monitoring indicator sequence correspond one-to-one in terms of time period. The core monitoring indicator includes the percentage of influenza-like cases, and the auxiliary monitoring indicators include the total number of outpatient visits and the proportion of pediatric outpatient visits.
[0031] This embodiment can be applied to a specific city, specifically for monitoring respiratory infectious diseases within that city. In this step, the core monitoring indicator includes the percentage of influenza-like illness cases, and auxiliary monitoring indicators include the total number of outpatient visits and the proportion of pediatric outpatient visits.
[0032] Percentage of influenza-like illness cases: The percentage of patients diagnosed with influenza-like illness by attending physicians in all hospitals in the city each week outpatient and emergency department visits; the criteria for attending physician diagnosis of influenza-like illness are: meeting two of the following symptoms: 1) Fever (body temperature ≥38℃); 2) At least one respiratory symptom (cough or sore throat).
[0033] Total outpatient visits: The total number of outpatient visits per week across all hospitals in the city;
[0034] Pediatric outpatient visits percentage: The percentage of pediatric outpatient visits to all hospitals in the city each week.
[0035] In this step, a rolling time window mechanism is used to acquire historical multidimensional monitoring time-series data. The historical data covers a continuous periodic sequence that traces back from the past, and its duration can be configured according to actual monitoring needs and application scenarios.
[0036] For example, the backtracking period can be set to the most recent N time periods (where N is a positive integer). In a typical implementation configuration, N can be set to 52, representing the data backtracking to the most recent full year (52 weeks). When the system runs, it will automatically extract N consecutive periods of data from the previous period based on the current analysis time to form the input dataset for this analysis.
[0037] For example: Assume the current system is performing analysis on January 18, 2026, and the system's configured backtracking length N is 52 weeks. The system will then retrieve data from that date back 52 weeks. For each of these 52 weeks, the system will obtain the following three metrics:
[0038] Percentage of flu-like cases that week;
[0039] Total outpatient visits this week;
[0040] The percentage of pediatric outpatient visits this week.
[0041] Therefore, the system will generate three sequences, each 52 bytes long: the percentage of influenza-like illness cases, the total number of outpatient visits, and the percentage of pediatric outpatient visits. These three sequences are strictly aligned in terms of time period, with the same index position corresponding to three different indicator values in the same week, together forming the historical multidimensional monitoring time series data used for subsequent analysis.
[0042] Step S2: Identify abnormal time periods based on historical multidimensional monitoring time series data; remove the data corresponding to abnormal time periods from the core monitoring indicator sequence and its corresponding historical year-on-year sequence to obtain the cleaned core monitoring indicator sequence and the cleaned historical year-on-year sequence.
[0043] The specific implementation steps of this step include steps S21 and S22;
[0044] Step S21: Standardize each sequence in the historical multidimensional monitoring time series data to obtain multiple standardized sequences; construct a standardized data matrix based on the standardized sequences, where the number of rows in the standardized data matrix is the number of time periods and the number of columns is the number of monitoring indicators; perform principal component analysis on the standardized data matrix to form a projection matrix; multiply the standardized data matrix and the projection matrix to obtain a dimension reduction matrix.
[0045] In the step of standardizing each sequence in the historical multidimensional monitoring time series data to obtain multiple standardized sequences, the sequences include core monitoring indicator sequences and auxiliary monitoring indicator sequences; that is, each core monitoring indicator sequence and each auxiliary monitoring indicator sequence are standardized separately to obtain multiple standardized sequences; in the step of constructing a standardized data matrix based on the standardized sequences, the number of rows in the standardized data matrix is the number of time periods, and the number of columns is the number of monitoring indicators, the monitoring indicators include core monitoring indicators and auxiliary monitoring indicators. If the core monitoring indicators include the percentage of influenza-like cases, and the auxiliary monitoring indicators include the total number of outpatient visits and the proportion of pediatric outpatient visits, then the number of monitoring indicators is 3.
[0046] After standardization, each sequence yields a standardized sequence with a mean of 0 and a standard deviation of 1. For example, Z-score standardization can be used. After standardization, all standardized sequences are aligned by time to construct a standardized data matrix. The number of rows in the standardized data matrix corresponds to the number of time periods (e.g., 52 weeks); the number of columns corresponds to the number of monitoring indicators (e.g., the percentage of influenza-like illness cases, total outpatient visits, and the percentage of pediatric outpatient visits); and the elements of the standardized data matrix are the standardized values of the i-th time period and the j-th monitoring indicator.
[0047] Principal component analysis (PCA) is performed on the standardized data matrix to construct the projection matrix. Specifically, the covariance matrix is calculated and eigenvalues are decomposed to obtain eigenvalues arranged in descending order and their corresponding eigenvectors (i.e., principal component directions). The number M of principal components to retain is determined based on the cumulative variance contribution rate of the eigenvalues. The smallest M value that satisfies the condition that the cumulative contribution rate first exceeds a preset threshold (e.g., 90%) is selected. The projection matrix is constructed from these first M eigenvectors (principal component directions).
[0048] Step S22: Multiply the dimensionality reduction matrix by the transpose of the projection matrix to obtain the reconstructed data matrix; for each time period i, calculate the reconstruction error value between the i-th row of the standardized data matrix and the i-th row of the reconstructed data matrix; based on the distribution of all reconstruction error values, set an error threshold, and mark the time periods corresponding to reconstruction error values greater than the error threshold as abnormal time periods; construct a historical year-on-year sequence, which is of the same length as the core monitoring indicator sequence, and the elements correspond one-to-one; for any time period in the core monitoring indicator sequence, its year-on-year value is the core monitoring indicator value of the same time period in the previous year; remove the data corresponding to the abnormal time periods from the core monitoring indicator sequence, and also remove the data at the corresponding positions in the historical year-on-year sequence to obtain the cleaned core monitoring indicator sequence and the cleaned historical year-on-year sequence.
[0049] In this step, for each time period i (corresponding to the i-th row of the normalized data matrix and the reconstructed data matrix), its reconstruction error value is calculated. This value is defined as the Euclidean distance between the data in the i-th row of the normalized data matrix and the data in the i-th row of the reconstructed data matrix. The reconstruction error value sequence is calculated for each row separately, and the 95th quantile of the reconstruction error value sequence is taken as the error threshold. Then, all time periods i are traversed. If the reconstruction error value is greater than the error threshold, the time period i is marked as an abnormal time period.
[0050] A historical year-on-year sequence is constructed, with the same length as the core monitoring indicator sequence and a one-to-one correspondence between elements. For any time period in the core monitoring indicator sequence, its year-on-year value is the core monitoring indicator value for the same time period in the previous year. Data corresponding to abnormal time periods are removed from the core monitoring indicator sequence, and the corresponding data in the historical year-on-year sequence are also removed, resulting in a cleaned core monitoring indicator sequence and a cleaned historical year-on-year sequence. This can be understood as:
[0051] For example, if the core monitoring indicator sequence contains data from week Ws of year Y to week We of year Y+1, then its historical year-on-year sequence contains corresponding data from week Ws of year Y−1 to week We of year Y; assuming that week 3 is an abnormal time period, then the data of week 3 in the core monitoring indicator sequence will be removed, and the data of week 3 in the historical year-on-year sequence will also be removed, resulting in the cleaned core monitoring indicator sequence and the cleaned historical year-on-year sequence.
[0052] Step S3: Calculate the dynamic early warning threshold based on the cleaned core monitoring indicator sequence and the cleaned historical year-on-year sequence; generate early warning information based on the cleaned core monitoring indicator sequence and the dynamic early warning threshold.
[0053] In this step, the specific implementation steps for calculating the dynamic early warning threshold based on the cleaned core monitoring indicator sequence and the cleaned historical year-on-year sequence include steps S31 and S32.
[0054] Step S31: Extract data of a preset time length from the end of the cleaned core monitoring indicator sequence to form the current observation subsequence; extract data corresponding to the time of the current observation subsequence from the cleaned historical year-on-year sequence to form the year-on-year observation subsequence;
[0055] In this step, the preset time length can be 15 weeks, which means extracting data from the end to the beginning for a total of 15 weeks;
[0056] Step S32: Normalize the current observation subsequence and the same period last year observation subsequence respectively, and pair the data points corresponding to the positions in the two normalized subsequences to form sample data pairs, and count the total number of sample data pairs; record the data from the current observation subsequence in the sample data pairs as the first data, and the data from the same period last year observation subsequence as the second data; perform cluster analysis on all sample data pairs to obtain multiple cluster categories, and obtain the dynamic early warning threshold based on the generated cluster categories.
[0057] In this step, the min-max normalization method can be used to normalize the current observation subsequence and the same period last year observation subsequence respectively; at the same time, in this step, cluster analysis is performed on all sample data pairs to obtain multiple cluster categories. The specific implementation steps for obtaining the dynamic early warning threshold based on the generated cluster categories include step S321:
[0058] Step S321: Use a density-based clustering algorithm to cluster all sample data pairs to obtain multiple cluster categories; determine multiple target cluster categories among all cluster categories; for each target cluster category, obtain the mean and standard deviation of all first data in all sample data pairs contained therein; multiply the standard deviation by a preset sensitivity coefficient to obtain the third data; add the third data to the mean to obtain the candidate warning threshold corresponding to the target cluster category; select the candidate warning threshold with the smallest value from all candidate warning thresholds corresponding to the target cluster categories as the dynamic warning threshold.
[0059] In this step, the sensitivity coefficient can be adjusted within the range of 1.0 to 4.0 according to actual monitoring needs; for example, it can be set to 2.
[0060] Meanwhile, in this step, the specific implementation steps for determining multiple target clustering categories among all clustering categories include steps S3211, S3212, and S3213.
[0061] Step S3211: Establish a time location index for each core monitoring indicator in the current observation subsequence; subtract the second data from the first data in each sample data pair to obtain the difference value;
[0062] In this step, the time position index refers to a consecutive integer number starting from 0 assigned to each data point in the current observation subsequence, used to represent the relative position order of the data point in the current observation subsequence. The time position index = the index of the point in the current observation subsequence - 1. Assuming that the core indicator monitoring data of the most recent 15 weeks are extracted, then the time position index corresponding to the first week is 0, and the time position index corresponding to the last week is 15.
[0063] Step S3212: For each cluster category: calculate the mean of all the difference values corresponding to all sample data pairs contained in it, and obtain the average difference value; calculate the mean of the time position indices corresponding to all the first data in all sample data pairs contained in it, and obtain the average time position; count the logarithm of the sample data pairs contained in it; calculate the comprehensive score of each cluster category according to formula (1), which is:
[0064] (1)
[0065] In formula (1), P is the comprehensive score for each cluster category. The average difference is weighted (e.g., 0.5); H is the average difference. G is the time proximity weight (e.g., 0.3); G is the average time position; Y is the maximum time position index (the maximum value among the time position indices corresponding to the current observation subsequence. Assuming the time position index corresponding to the current observation subsequence is 0-15, then the maximum time position index is 15). B is the sample number weight (e.g., 0.2); B is the number of sample data pairs contained in each cluster category; F is the total number of sample data pairs, which is the total number of sample data pairs calculated in step S32.
[0066] The average time position is calculated by averaging the time position indices of all the first data points in all the sample data pairs contained in the cluster. This can be understood as follows: if the cluster contains 3 first data points with time position indices of 3, 5, and 10 respectively, then the average time position is (3+5+10) / 3=6.
[0067] Step S3213: Among all cluster categories, select a preset number of cluster categories as target cluster categories according to the order of comprehensive scores from high to low.
[0068] In this step, the preset number can be half of the total number of cluster categories. If it is not divisible, it will be rounded up.
[0069] In step S3, the specific implementation steps for generating early warning information based on the cleaned core monitoring indicator sequence and dynamic early warning threshold include step S33;
[0070] Step S33: Determine the latest time period, which is the latest time period in the cleaned core monitoring indicator sequence. Compare the core monitoring indicator value corresponding to the latest time period with the dynamic early warning threshold. If the core monitoring indicator value does not exceed the dynamic early warning threshold, generate a safety warning message. If the core monitoring indicator value exceeds the dynamic early warning threshold, trace back L consecutive time periods from the latest time period in the cleaned core monitoring indicator sequence. If the number of time periods in which the core monitoring indicator value is greater than the dynamic early warning threshold exceeds a preset threshold, generate a first early warning message; otherwise, generate a second early warning message. The warning level of the first early warning message is higher than that of the second early warning message, where L is a positive integer greater than 1.
[0071] In this step, L can be customized, tracing back L consecutive time periods from the latest time period, excluding the latest time period, for example, it can be 4; the preset quantity threshold can be L / 2, and if it is not divisible, it is rounded up; the warning level of the first warning information is higher than that of the second warning information, which can be understood as the first warning information representing a higher degree of severity.
[0072] Example 2
[0073] As shown in Figure 2, this embodiment provides a respiratory infectious disease monitoring device, which includes an acquisition module 1, an identification module 2, and an early warning module 3.
[0074] Module 1 is used to acquire historical multidimensional monitoring time series data. The historical multidimensional monitoring time series data includes at least one core monitoring indicator sequence and at least one auxiliary monitoring indicator sequence. The core monitoring indicator sequence and the auxiliary monitoring indicator sequence correspond one-to-one in terms of time period. The core monitoring indicator includes the percentage of influenza-like cases, and the auxiliary monitoring indicators include the total number of outpatient visits and the proportion of pediatric outpatient visits.
[0075] The identification module 2 is used to identify abnormal time periods based on historical multidimensional monitoring time series data; and to remove the data corresponding to the abnormal time periods from the core monitoring indicator sequence and its corresponding historical year-on-year sequence to obtain the cleaned core monitoring indicator sequence and the cleaned historical year-on-year sequence.
[0076] The early warning module 3 is used to calculate the dynamic early warning threshold based on the cleaned core monitoring indicator sequence and the cleaned historical year-on-year sequence; and to generate early warning information based on the cleaned core monitoring indicator sequence and the dynamic early warning threshold.
[0077] In one specific embodiment of this disclosure, the identification module 2 further includes a processing unit 21 and a cleaning unit 22.
[0078] Processing unit 21 is used to standardize each sequence in the historical multidimensional monitoring time series data to obtain multiple standardized sequences; construct a standardized data matrix based on the standardized sequences, where the number of rows in the standardized data matrix is the number of time periods and the number of columns is the number of monitoring indicators; perform principal component analysis on the standardized data matrix to form a projection matrix; and multiply the standardized data matrix and the projection matrix to obtain a dimension reduction matrix.
[0079] The cleaning unit 22 is used to multiply the dimensionality reduction matrix by the transpose of the projection matrix to obtain the reconstructed data matrix; for each time period i, the reconstruction error value between the i-th row of the standardized data matrix and the i-th row of the reconstructed data matrix is calculated; according to the distribution of all reconstruction error values, an error threshold is set, and the time periods corresponding to the reconstruction error values greater than the error threshold are marked as abnormal time periods; a historical year-on-year sequence is constructed, which is of the same length as the core monitoring indicator sequence, and the elements correspond one-to-one; for any time period in the core monitoring indicator sequence, its year-on-year value is the core monitoring indicator value of the same time period in the previous year; the data corresponding to the abnormal time periods are removed from the core monitoring indicator sequence, and the data at the corresponding positions in the historical year-on-year sequence are also removed, resulting in the cleaned core monitoring indicator sequence and the cleaned historical year-on-year sequence.
[0080] In one specific embodiment of this disclosure, the early warning module 3 further includes an interception unit 31 and a clustering unit 32.
[0081] The extraction unit 31 is used to extract data of a preset time length from the end of the cleaned core monitoring indicator sequence to form the current observation subsequence; and to extract data corresponding to the time of the current observation subsequence from the cleaned historical year-on-year sequence to form the year-on-year observation subsequence.
[0082] Clustering unit 32 is used to normalize the current observation subsequence and the same-year observation subsequence respectively, and pair the data points corresponding to the positions in the two normalized subsequences to form sample data pairs. The data from the current observation subsequence in the sample data pair is recorded as the first data, and the data from the same-year observation subsequence is recorded as the second data. Cluster analysis is performed on all sample data pairs to obtain multiple cluster categories, and a dynamic early warning threshold is obtained based on the generated cluster categories.
[0083] In one specific embodiment of this disclosure, the clustering unit 32 further includes a selection unit 321.
[0084] Unit 321 is selected to cluster all sample data pairs using a density-based clustering algorithm to obtain multiple cluster categories; multiple target cluster categories are determined among all cluster categories; for each target cluster category, the mean and standard deviation of all first data in all sample data pairs contained therein are obtained; the standard deviation is multiplied by a preset sensitivity coefficient to obtain third data, and the third data is added to the mean to obtain the candidate warning threshold corresponding to the target cluster category; from the candidate warning thresholds corresponding to all target cluster categories, the candidate warning threshold with the smallest value is selected as the dynamic warning threshold.
[0085] It should be noted that the specific manner in which each module performs its operation in the apparatus described in the above embodiments has been described in detail in the embodiments of the method, and will not be elaborated here.
[0086] Example 3
[0087] Corresponding to the above method embodiments, this disclosure also provides a respiratory infectious disease monitoring device. The respiratory infectious disease monitoring device described below can be referred to in correspondence with the respiratory infectious disease monitoring method described above.
[0088] Figure 3 is a block diagram illustrating a respiratory infectious disease monitoring device 300 according to an exemplary embodiment. As shown in Figure 3, the respiratory infectious disease monitoring device 300 may include a processor 301 and a memory 302. The respiratory infectious disease monitoring device 300 may also include one or more of a multimedia component 303, an I / O interface 304, and a communication component 305.
[0089] The processor 301 controls the overall operation of the respiratory infectious disease monitoring device 300 to complete all or part of the steps in the aforementioned respiratory infectious disease monitoring method. The memory 302 stores various types of data to support the operation of the respiratory infectious disease monitoring device 300. This data may include, for example, instructions for any application or method used on the respiratory infectious disease monitoring device 300, as well as application-related data such as contact data, sent and received messages, images, audio, video, etc. The memory 302 can be implemented using any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 303 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory 302 or transmitted via the communication component 305. The audio component also includes at least one speaker for outputting audio signals. I / O interface 304 provides an interface between processor 301 and other interface modules, such as keyboards, mice, and buttons. These buttons can be virtual or physical. Communication component 305 is used for wired or wireless communication between the respiratory infectious disease monitoring device 300 and other devices. Wireless communication includes Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination thereof. Therefore, the corresponding communication component 305 may include a Wi-Fi module, a Bluetooth module, and an NFC module.
[0090] In an exemplary embodiment, the respiratory infectious disease monitoring device 300 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the respiratory infectious disease monitoring method described above.
[0091] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the respiratory infectious disease monitoring method described above. For example, the computer-readable storage medium may be the memory 302 including program instructions, which may be executed by the processor 301 of the respiratory infectious disease monitoring device 300 to complete the respiratory infectious disease monitoring method described above.
[0092] Example 4
[0093] Corresponding to the above method embodiments, this disclosure also provides a readable storage medium, and the readable storage medium described below can be referred to in conjunction with the respiratory infectious disease monitoring method described above.
[0094] A readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the respiratory infectious disease monitoring method described in the above method embodiments.
[0095] Specifically, the readable storage medium can be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or any other readable storage medium capable of storing program code.
[0096] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for monitoring respiratory infectious diseases, characterized in that, This includes: acquiring historical multidimensional monitoring time-series data, which includes at least one core monitoring indicator sequence and at least one auxiliary monitoring indicator sequence, with a one-to-one correspondence between the core and auxiliary monitoring indicator sequences in terms of time periods; the core monitoring indicator includes the percentage of influenza-like illness cases, and the auxiliary monitoring indicators include the total number of outpatient visits and the proportion of pediatric outpatient visits; identifying abnormal time periods based on the historical multidimensional monitoring time-series data; removing data corresponding to abnormal time periods from the core monitoring indicator sequences and their corresponding historical year-on-year sequences to obtain cleaned core monitoring indicator sequences and cleaned historical year-on-year sequences; calculating dynamic early warning thresholds based on the cleaned core monitoring indicator sequences and cleaned historical year-on-year sequences; and further... The system generates early warning information based on the cleaned core monitoring indicator sequence and dynamic early warning threshold. Specifically, the dynamic early warning threshold is calculated based on the cleaned core monitoring indicator sequence and the cleaned historical year-on-year sequence. This includes: extracting data of a preset time length from the end of the cleaned core monitoring indicator sequence to form the current observation subsequence; extracting data corresponding to the time of the current observation subsequence from the cleaned historical year-on-year sequence to form the year-on-year observation subsequence; normalizing both the current and year-on-year observation subsequences, and pairing corresponding data points in the two normalized subsequences to form sample data pairs. Data from the current observation subsequence is designated as the first data point, and data from the year-on-year observation subsequence is designated as the second data point. The second data involves clustering all sample data pairs to obtain multiple cluster categories, and then determining a dynamic warning threshold based on these cluster categories. This process includes: using a density-based clustering algorithm to cluster all sample data pairs to obtain multiple cluster categories; identifying multiple target cluster categories within these cluster categories; for each target cluster category, obtaining the mean and standard deviation of all first data points from all sample data pairs it contains; multiplying the standard deviation by a preset sensitivity coefficient to obtain third data points; adding the third data points to the mean to obtain the candidate warning threshold corresponding to that target cluster category; and then... (The sentence is incomplete and requires further context to translate accurately.) Among the candidate warning thresholds corresponding to the category, the candidate warning threshold with the smallest value is selected as the dynamic warning threshold; among them, multiple target cluster categories are determined in all cluster categories, including: establishing a time position index for each core monitoring indicator in the current observation subsequence; subtracting the second data from the first data in each sample data pair to obtain the difference value; for each cluster category: calculating the average difference value for all the difference values corresponding to all the sample data pairs it contains; calculating the average time position for all the time position indices corresponding to the first data in all the sample data pairs it contains; counting the number of logarithms of the sample data pairs it contains; calculating the comprehensive score of each cluster category according to formula (1), which is: (1) In formula (1), P is the comprehensive score for each cluster category. The average difference is the weight; H is the average difference. G represents the time proximity weight; G is the average time position; Y is the index of the largest time position. B is the sample size weight; F is the number of sample data pairs contained in each cluster category; F is the total number of sample data pairs; among all cluster categories, a preset number of cluster categories are selected as target cluster categories according to the order of comprehensive scores from high to low.
2. The method for monitoring respiratory infectious diseases according to claim 1, characterized in that, Based on historical multidimensional monitoring time-series data, abnormal time periods are identified. From the core monitoring indicator sequences and their corresponding historical year-on-year sequences, data corresponding to abnormal time periods are removed, resulting in cleaned core monitoring indicator sequences and cleaned historical year-on-year sequences. This includes: standardizing each sequence in the historical multidimensional monitoring time-series data to obtain multiple standardized sequences; constructing a standardized data matrix based on the standardized sequences, where the number of rows in the standardized data matrix corresponds to the number of time periods and the number of columns corresponds to the number of monitoring indicators; performing principal component analysis on the standardized data matrix to construct a projection matrix; multiplying the standardized data matrix with the projection matrix to obtain a dimensionality-reduced matrix; and multiplying the dimensionality-reduced matrix with the transpose of the projection matrix to obtain a reconstructed data matrix. For each time period i, calculate the reconstruction error between the i-th row of the standardized data matrix and the i-th row of the reconstructed data matrix; based on the distribution of all reconstruction error values, set an error threshold, and mark the time periods corresponding to reconstruction error values greater than the error threshold as abnormal time periods; construct a historical year-on-year sequence, which is of equal length to the core monitoring indicator sequence, and the elements correspond one-to-one; for any time period in the core monitoring indicator sequence, its year-on-year value is the core monitoring indicator value of the same time period in the previous year; remove the data corresponding to the abnormal time periods from the core monitoring indicator sequence, and also remove the data at the corresponding positions in the historical year-on-year sequence, to obtain the cleaned core monitoring indicator sequence and the cleaned historical year-on-year sequence.
3. A respiratory infectious disease monitoring device, characterized in that, include: The acquisition module is used to acquire historical multidimensional monitoring time-series data, which includes at least one core monitoring indicator sequence and at least one auxiliary monitoring indicator sequence. The core monitoring indicator sequence and the auxiliary monitoring indicator sequence correspond one-to-one in terms of time period. The core monitoring indicator includes the percentage of influenza-like illness cases, and the auxiliary monitoring indicators include the total number of outpatient visits and the proportion of pediatric outpatient visits. The identification module is used to identify abnormal time periods based on the historical multidimensional monitoring time-series data. From the core monitoring indicator sequence and its corresponding historical year-on-year sequence, the data corresponding to the abnormal time periods are removed to obtain the cleaned core monitoring indicator sequence and the cleaned historical year-on-year sequence. The early warning module is used to calculate the dynamic early warning threshold based on the cleaned core monitoring indicator sequence and the cleaned historical year-on-year sequence. Early warning information is generated based on the cleaned core monitoring indicator sequence and dynamic early warning threshold. The early warning module includes: a truncation unit, used to truncate data of a preset time length from the end of the cleaned core monitoring indicator sequence to form the current observation subsequence; and to truncate data corresponding to the time of the current observation subsequence from the cleaned historical year-on-year sequence to form the year-on-year observation subsequence; a clustering unit, used to normalize the current observation subsequence and the year-on-year observation subsequence respectively, and to pair data points corresponding to positions in the two normalized subsequences to form sample data pairs, designating data from the current observation subsequence as the first data and data from the year-on-year observation subsequence as the second data; performing cluster analysis on all sample data pairs to obtain multiple cluster categories, and obtaining the dynamic early warning threshold based on the generated cluster categories; the clustering unit includes: a selection unit, used to cluster all sample data pairs using a density-based clustering algorithm to obtain multiple cluster categories; and determining multiple target clusters among all cluster categories. Category; for each target cluster category, obtain the mean and standard deviation of all first data in all sample data pairs contained therein; multiply the standard deviation by the preset sensitivity coefficient to obtain the third data, add the third data to the mean to obtain the candidate warning threshold corresponding to the target cluster category; select the candidate warning threshold with the smallest value from all candidate warning thresholds corresponding to all target cluster categories as the dynamic warning threshold; wherein, multiple target cluster categories are determined in all cluster categories, including: establishing a time position index for each core monitoring indicator in the current observation subsequence; subtracting the second data from the first data in each sample data pair to obtain the difference value; for each cluster category: calculate the mean of all difference values corresponding to all sample data pairs contained therein to obtain the average difference value; calculate the mean of the time position index corresponding to all first data in all sample data pairs contained therein to obtain the average time position; count the number of logarithms of the sample data pairs contained therein; calculate the comprehensive score of each cluster category according to formula (1), formula (1) is: (1) In formula (1), P is the comprehensive score for each cluster category. The average difference is the weight; H is the average difference. G represents the time proximity weight; G is the average time position; Y is the index of the largest time position. B is the sample size weight; F is the number of sample data pairs contained in each cluster category; F is the total number of sample data pairs; among all cluster categories, a preset number of cluster categories are selected as target cluster categories according to the order of comprehensive scores from high to low.
4. The respiratory infectious disease monitoring device according to claim 3, characterized in that, The identification module includes: a processing unit, used to standardize each sequence in the historical multidimensional monitoring time series data to obtain multiple standardized sequences; constructing a standardized data matrix based on the standardized sequences, where the number of rows in the standardized data matrix is the number of time periods and the number of columns is the number of monitoring indicators; performing principal component analysis on the standardized data matrix to construct a projection matrix, and multiplying the standardized data matrix with the projection matrix to obtain a dimension-reduced matrix; a cleaning unit, used to multiply the dimension-reduced matrix with the transpose of the projection matrix to obtain a reconstructed data matrix; and for each time period i, calculating the ratio of the i-th row in the standardized data matrix to the i-th row in the reconstructed data matrix. The reconstruction error value between the time periods is determined. Based on the distribution of all reconstruction error values, an error threshold is set, and the time periods corresponding to reconstruction error values greater than the error threshold are marked as abnormal time periods. A historical year-on-year sequence is constructed, which is of the same length as the core monitoring indicator sequence, and the elements correspond one-to-one. For any time period in the core monitoring indicator sequence, its year-on-year value is the core monitoring indicator value of the same time period in the previous year. The data corresponding to the abnormal time periods are removed from the core monitoring indicator sequence, and the data at the corresponding positions in the historical year-on-year sequence are also removed to obtain the cleaned core monitoring indicator sequence and the cleaned historical year-on-year sequence.
5. A respiratory infectious disease monitoring device, characterized in that, include: Memory, used to store computer programs; A processor for executing a computer program to implement the steps of the respiratory infectious disease monitoring method as claimed in any one of claims 1 to 2.
6. A storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, implements the steps of the respiratory infectious disease monitoring method as claimed in any one of claims 1 to 2.
Citation Information
Patent Citations
Heart and cerebral vessel health monitoring and early warning method and system based on individual physiological features
CN118177766A