Infectious disease early warning method and device based on spatiotemporal heterogeneity, medium and equipment

By constructing seasonal and distance weight matrices and combining Fourier transform and Gaussian decay function, a seasonally weighted case count matrix is ​​generated, which solves the problem that existing technologies fail to consider the heterogeneity of infectious disease data and achieves more accurate infectious disease early warning.

CN117747128BActive Publication Date: 2026-07-24厦门畅享信息技术有限公司 +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
厦门畅享信息技术有限公司
Filing Date
2023-12-26
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing forward-looking spatiotemporal scanning algorithms fail to effectively consider the heterogeneity of infectious disease data in time and space, resulting in inaccurate early warning results.

Method used

By collecting and processing infectious disease data, a seasonal weight matrix and a distance weight matrix are constructed. By combining Fourier transform and Gaussian decay function, a seasonally weighted case count matrix is ​​generated, and spatiotemporal factors are integrated for early warning.

Benefits of technology

It improves the accuracy of infectious disease early warning, better reflects the correlation in time and space, and provides early warning results that are more in line with actual patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117747128B_ABST
    Figure CN117747128B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of infectious disease early warning, and particularly relates to an infectious disease early warning method based on space-time heterogeneity, comprising the following steps: collecting first-time infection data of a prediction area to obtain a first seasonal weight matrix; performing spatial weighting on geographical data of the first-time infection area to obtain a distance weight matrix; collecting second-time infection data of the first-time infection area at a second time to generate a case number matrix; arranging the first seasonal weight matrix corresponding to the second-time infection data time sequence to obtain a second seasonal weight matrix; and fusing the second seasonal weight matrix, the distance weight matrix and the case number matrix to obtain a seasonal weighted case number matrix. The seasonal weight matrix, the distance weight matrix and the case number matrix are obtained by weighting processing of the weight coefficients in time and space calculated from historical infection data, and the seasonal weighted case number matrix is obtained by fusing the three matrices, so that the distribution of the collected data is more consistent with the correlation in time and space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of infectious disease early warning technology, and in particular to an infectious disease early warning method, device, medium and equipment based on spatiotemporal heterogeneity. Background Technology

[0002] In infectious disease early warning models, the prospective spatiotemporal rearrangement scanning statistics method can be used for early prediction of disease outbreaks. Its advantage is that it only uses the number of cases and time location information, without the need for data on high-risk populations in the region. This model is a spatiotemporal scanning statistics method based on a dynamically sized circular (or elliptical if projected coordinates) moving window.

[0003] The scan is based on each cell, using circles of different radii for each scan of the surrounding area. Statistics are superimposed over time with the scan window as the base, until the statistic is customized. The scan statistic is defined as the Generalized Likelihood Ratio (GLR) of the scan window. A larger GLR value indicates a more statistically significant difference and a stronger tendency for abnormal clustering within the window. The window with the largest statistic is selected as the window with the highest infectious disease clustering, used to determine if there are any anomalies in the number of cases within the window. Monte Carlo hypothesis testing is used to test the confidence level of the non-random distribution of elements within the cluster. This hypothesis test is performed on the largest and smallest clusters scanned, thus anchoring the space of the highest clustering area as the Most Likely Cluster (MLC), the second highest clustering area as the second most likely cluster, and so on. Since the study is based on a null hypothesis—that is, calculating the p-value by comparing the likelihood of the random dataset with the real dataset—it does not need to consider extremely complex probability distribution issues.

[0004] Prospective spatiotemporal scanning algorithms assume that the data is homogeneous, but in reality, many disease data are heterogeneous in time and space. For example, many infectious diseases have strong seasonal correlations in time, and two adjacent regions will influence each other in space. However, existing prospective spatiotemporal scanning algorithms do not consider the heterogeneity of influencing factors, making the early warning results inaccurate. Summary of the Invention

[0005] To address the shortcomings of the existing technology, this invention provides an infectious disease early warning method based on spatiotemporal heterogeneity, comprising the following steps:

[0006] S100. Collect the first-time infection data of the predicted location, and process the first-time infection data to obtain the first-season weight matrix;

[0007] S200. Collect geographic data of the first-time infection location based on the first-time infection data, and spatially weight the geographic data to obtain a distance weight matrix;

[0008] S300: Collect the second-time infection data of the first-time infection location at the second time, generate the case number matrix, and organize the first-season weight matrix corresponding to the second-time infection data time series to obtain the second-season weight matrix.

[0009] S400. The second season weight matrix, distance weight matrix and case number matrix are fused to obtain the seasonal weighted case number matrix;

[0010] S500 incorporates a seasonally weighted case count matrix into a prospective spatiotemporal scanning algorithm to provide early warning of infectious diseases in predicted locations.

[0011] Further, S110, the first infection data is organized according to time sequence to obtain the first time series, and the first time series is subjected to Fourier transform to obtain the first frequency domain data;

[0012] S120. Reset the portion of the first frequency domain data that is greater than a preset value to 0, and retain the portion that is less than a preset value to obtain the second frequency domain data;

[0013] S130. Perform an inverse Fourier transform on the second frequency domain data to obtain the seasonal weight values;

[0014] S140. After processing the seasonal weight values, we obtain the first season weight matrix.

[0015] Furthermore, the specific formula for obtaining the first frequency domain data by performing a Fourier transform on the first time series is as follows:

[0016]

[0017] Where n represents the nth data point in the first time series, This represents the nth sampling point of the time-domain signal in the first time series, where j is the imaginary unit and N is the total number of data points in the first time series. This represents the first frequency domain data value of the k-th component after Fourier transform in the frequency domain;

[0018] The formula for calculating the second frequency domain data is:

[0019]

[0020] in, This represents the first frequency domain data value of the k-th component after Fourier transform in the frequency domain. Represents the preset value. This represents the second frequency domain data value of the k-th component;

[0021] The formula for calculating the seasonal weight value is:

[0022]

[0023] To extract the real part of F(n), Re represents the operator for extracting the real part of a complex number. Here, N is the seasonal weight value, j is the total number of data points, and e is the natural constant. This is the second frequency domain data value of the k-th component.

[0024] Furthermore, the specific steps of step S140 are as follows:

[0025] S141. Normalize the seasonal weight values ​​to obtain the seasonal weight coefficients;

[0026] S142. Add 1 to the seasonal weight coefficients to obtain the seasonal weight sequence, and transform the seasonal weight sequence into a diagonal matrix to obtain the first seasonal weight matrix.

[0027] The specific formula for normalizing the seasonal weight values ​​and calculating the seasonal weight coefficient is as follows:

[0028]

[0029] in, Here are the seasonal weight values, and n is... The nth component in the seasonal weighting value, The minimum seasonal weight value, The largest seasonal weight value, This is the seasonal weighting coefficient;

[0030] The formula for calculating the seasonal weighted series is:

[0031]

[0032] T is the seasonal weight sequence. Transforming the seasonal weight sequence into a diagonal matrix yields the first seasonal weight matrix:

[0033]

[0034] Where H is the weight matrix for the first season.

[0035] Furthermore, the specific steps of S200 are as follows:

[0036] The specific steps for S200 are as follows:

[0037] S210. Collect geographical data of the first infection location;

[0038] S220. Calculate the distance between each initial infection location based on geographical data to obtain a distance matrix;

[0039] S230. Process the distance matrix using the Gaussian decay function to obtain the distance weight matrix.

[0040] Furthermore, the geographical data of the first infection sites is collected as the coordinate data of the spatial location points of the first infection sites. The distances between the first infection sites are calculated based on the coordinate data, resulting in the following distance matrix:

[0041]

[0042] Where m represents the total number of initial infection locations. Let m be the distance between the m-th first-time infection location and the 1st first-time infection location;

[0043] Use Gaussian decay function The distance weight matrix is ​​calculated using the formula: e is the natural constant, and α represents the attenuation coefficient as a hyperparameter.

[0044] .

[0045] Furthermore, the seasonal weight matrix, the distance weight matrix, and the case count matrix are multiplied by a dot product to obtain the seasonal weighted case count matrix.

[0046] The present invention also provides an infectious disease early warning device based on spatiotemporal heterogeneity, comprising:

[0047] The seasonal weight matrix construction module is used to collect the first-time infection data of the predicted location and process the first-time infection data to obtain the first-season weight matrix.

[0048] The distance weight matrix construction module is used to collect geographic data of the first infected location based on the first-time infection data, and to spatially weight the geographic data to obtain the distance weight matrix.

[0049] The case number matrix construction module is used to collect the second-time infection data of the first-time infection location at the second time, generate the case number matrix, and organize the first-season weight matrix corresponding to the second-time infection data time series to obtain the second-season weight matrix.

[0050] The seasonally weighted case count matrix construction module is used to fuse the second-season weight matrix, the distance weight matrix, and the case count matrix to obtain the seasonally weighted case count matrix.

[0051] The analysis and early warning module is used to input the seasonally weighted case count matrix into a prospective spatiotemporal scanning algorithm to provide early warning of infectious diseases in the predicted areas.

[0052] The present invention also provides a computer-readable storage medium storing computer instructions, which, when executed by a processor, implement the infectious disease early warning method based on spatiotemporal heterogeneity as described in any of the above embodiments.

[0053] The present invention also provides an electronic device, including at least one processor and a memory communicatively connected to the processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the processor to perform the infectious disease early warning method based on spatiotemporal heterogeneity as described in any of the above embodiments.

[0054] Based on the above, compared with the prior art, the infectious disease early warning method based on spatiotemporal heterogeneity provided by the present invention preprocesses the collected data in terms of time and space, and performs weighted processing on the time and space weight coefficients calculated from historical infection data to obtain a seasonal weight matrix, a distance weight matrix, and a case number matrix. At the same time, the three are fused to obtain a seasonal weighted case number matrix, and then the prospective spatiotemporal rearrangement scanning statistic method is used for calculation, so that the distribution of the collected data is more consistent with the temporal and spatial correlation, and the early warning results are more consistent with the actual laws and more accurate.

[0055] Other features and beneficial effects of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other beneficial effects of the invention can be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings. Attached Figure Description

[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Unless otherwise specified, the positional relationships shown in the drawings in the following description are based on the direction in which the components are drawn in the figure.

[0057] Figure 1 The flowchart shows the infectious disease early warning method based on spatiotemporal heterogeneity provided in Embodiment 1 of the present invention.

[0058] Figure 2 This is a flowchart of constructing a seasonal weight matrix provided in Embodiment 1 of the present invention;

[0059] Figure 3 This is a flowchart of normalizing seasonal weight values ​​provided in Embodiment 1 of the present invention;

[0060] Figure 4 This is a flowchart of constructing the distance weight matrix provided in Embodiment 1 of the present invention;

[0061] Figure 5 The warning result is obtained by using a forward-looking spatiotemporal scanning algorithm on the homogeneous raw data collected in Embodiment 2 of the present invention.

[0062] Figure 6 The warning result obtained by using a prospective spatiotemporal scanning algorithm on the processed spatiotemporal heterogeneous data provided in Embodiment 2 of the present invention;

[0063] Figure 7 This is a schematic diagram of the infectious disease early warning device based on spatiotemporal heterogeneity provided in Embodiment 3 of the present invention;

[0064] Figure 8 This is a schematic diagram of the structure of the electronic device provided in Embodiment 5 of the present invention. Detailed Implementation

[0065] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. The technical features designed in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0066] In the description of this invention, it should be noted that all terms used in this invention (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains, and should not be construed as limiting the invention; it should be further understood that the terms used in this invention should be understood to have the same meaning as those in the context of this specification and in the relevant field, and should not be understood in an idealized or overly formal sense, except as expressly defined in this invention.

[0067] Example 1

[0068] A method for infectious disease early warning based on spatiotemporal heterogeneity, such as Figure 1 As shown, it includes the following steps:

[0069] S100. Collect the first-time infection data of the predicted location, and process the first-time infection data to obtain the first-season weight matrix.

[0070] S200. Collect geographical data of the first infection location based on the first infection data, and spatially weight the geographical data to obtain a distance weight matrix.

[0071] S300: Collect the second-time infection data of the first-time infection location at the second-time location, generate a case number matrix, and organize the first-season weight matrix corresponding to the second-time infection data time series to obtain the second-season weight matrix.

[0072] S400. The second season weight matrix, distance weight matrix and case number matrix are fused to obtain the seasonal weighted case number matrix.

[0073] S500 incorporates a seasonally weighted case count matrix into a prospective spatiotemporal scanning algorithm to provide early warning of infectious diseases in predicted locations.

[0074] In practice, the first-time infection data refers to the historical infection data of the year preceding the year the infectious disease is predicted. This first-time infection data is processed to obtain the first-season weight matrix. Based on this historical data, the location of the infectious disease outbreak (i.e., the first-time infection location) is located, and its geographical data is determined. This geographical data is then spatially weighted to obtain a distance weight matrix. Second-time infection data (i.e., the second time period, the second time period) is collected from the first-time infection location to generate the current case count matrix. The seasonal weight matrix values ​​of the first-season weight matrix corresponding to the second time period at the first time period are determined and processed to obtain the second-season weight matrix. The second-season weight matrix, the distance weight matrix, and the case count matrix are merged to obtain a seasonally weighted case count matrix. This calculated seasonally weighted case count matrix is ​​then fed into a prospective spatiotemporal scanning algorithm to provide infectious disease early warning for the predicted location.

[0075] By determining the first seasonal weight matrix corresponding to the second time in the first time, the second seasonal weight matrix is ​​obtained. This incorporates the influence of seasonal factors at the same time on infectious diseases. By determining the distance weight matrix, the influence of spatial factors on the early warning of infectious diseases in that year is incorporated. Finally, a heterogeneous data seasonal weighted case number matrix incorporating spatiotemporal factors is obtained and introduced into the prospective spatiotemporal scanning algorithm. This allows the early warning of infectious diseases to consider different influencing factors such as time and space, resulting in more accurate early warning results.

[0076] S100. Collect the first-time infection data of the predicted location, and process the first-time infection data to obtain the first-season weight matrix.

[0077] Furthermore, since the first-time infection data is time-related sequence data, where the data exhibits periodic changes over time, Fourier transform can be used to decompose the first-time infection database into periodic components of different frequencies, such as... Figure 2As shown, the specific steps of step S100 are as follows:

[0078] S110. The first time-series infection data is organized according to time sequence to obtain the first time-series data. The first time-series data is then subjected to Fourier transform to obtain the first frequency domain data.

[0079] S120. Reset the portion of the first frequency domain data that is greater than the preset value to 0, and retain the portion that is less than the preset value to obtain the second frequency domain data.

[0080] S130. Perform an inverse Fourier transform on the second frequency domain data to obtain the seasonal weight values.

[0081] S140. After processing the seasonal weight values, we obtain the first season weight matrix.

[0082] In practice, the first infection data is the relevant infectious disease data from the previous year. It is organized into a first time series according to the time order. The first time series is subjected to a Fourier transform, which decomposes it into the sum of sine and cosine functions to obtain the first frequency domain data. The first frequency domain data with values ​​greater than a preset value is reset to 0, and the part with values ​​lower than the preset value is retained to obtain the second frequency domain data. This is because seasonal patterns are usually reflected in the low-frequency part. By retaining the part with values ​​lower than the preset value, the seasonal weights can be extracted more accurately without being affected by noise or other high-frequency interference. This step is similar to a low-pass filter, which only retains the low-frequency part of the spectrum. Then, the second frequency domain data is subjected to an inverse Fourier transform to convert the frequency domain data back to time series data, and the real part of the filtered data is extracted. This real part is the final seasonal weight matrix value.

[0083] Better, such as Figure 3 As shown, the specific steps of step S140 are as follows:

[0084] S141. Normalize the seasonal weight values ​​to obtain the seasonal weight coefficients.

[0085] S142. Add 1 to the seasonal weight coefficients to obtain the seasonal weight sequence, and transform the seasonal weight sequence into a diagonal matrix to obtain the first seasonal weight matrix.

[0086] Preferably, the specific formula for performing a Fourier transform on the first time series to obtain the first frequency domain data is as follows:

[0087]

[0088] Where n represents the nth data point in the first time series, and the value of n ranges from 0 to 364. This represents the nth sampling point of the time-domain signal in the first time series; j is the imaginary unit, and N is the total number of data points in the first time series, with a value of 365. This represents the first frequency domain data value of the k-th component after Fourier transform in the frequency domain.

[0089] The formula for calculating the second frequency domain data is:

[0090]

[0091] in, This represents the first frequency domain data value of the k-th component after Fourier transform in the frequency domain. Represents the preset value. This is the second frequency domain data value of the k-th component.

[0092] The formula for calculating the seasonal weight value is:

[0093]

[0094] To extract the real part of F(n), Re represents the operator for extracting the real part of a complex number. Here, N is the seasonal weight value, j is the total number of data points, and e is the natural constant. This is the second frequency domain data value of the k-th component.

[0095] The specific formula for normalizing the seasonal weight values ​​and calculating the seasonal weight coefficient is as follows:

[0096]

[0097] in, Here are the seasonal weight values, and n is... The nth component in the seasonal weighting value, The minimum seasonal weight value, The largest seasonal weight value, This is the seasonal weighting coefficient.

[0098] The formula for calculating the seasonal weighted series is:

[0099]

[0100] T is the seasonal weight sequence. Transforming the seasonal weight sequence into a diagonal matrix yields the first seasonal weight matrix:

[0101]

[0102] Where H is the weight matrix for the first season.

[0103] S200. Collect geographical data of the first infection location based on the first infection data, and spatially weight the geographical data to obtain a distance weight matrix.

[0104] like Figure 4As shown, the steps for constructing the distance weight matrix are as follows:

[0105] S210. Collect geographical data of the first infection location.

[0106] S220. Calculate the distance between each first-time infection location based on geographical data to obtain a distance matrix.

[0107] S230. Process the distance matrix using the Gaussian decay function to obtain the distance weight matrix.

[0108] The geographical data of the first infection sites were collected as coordinate data of their spatial locations. The distances between these first infection sites were calculated based on the coordinate data, resulting in the following distance matrix:

[0109]

[0110] Where m represents the total number of initial infection locations. Let m be the distance between the m-th first-time infection location and the 1st first-time infection location.

[0111] Use Gaussian decay function The distance weight matrix is ​​calculated using the formula: e is the natural constant, and α represents the attenuation coefficient as a hyperparameter.

[0112] .

[0113] S300: Collect the second-time infection data of the first-time infection location at the second-time location, generate a case number matrix, and organize the first-season weight matrix corresponding to the second-time infection data time series to obtain the second-season weight matrix.

[0114] Furthermore, the steps for constructing the case count matrix are as follows:

[0115] Unit area The case count matrix for day d is as follows:

[0116]

[0117] Where K represents the case count matrix, and the unit region p is the location of infection at the first time corresponding to the second time point. This represents the matrix value of the number of cases in the p-th unit region on day d.

[0118] Furthermore, the weight matrix of the first season corresponding to the second time-series infection data is rearranged to obtain the weight matrix of the second season, and the specific formula is as follows:

[0119]

[0120]

[0121] In practice, by using the date of the second time infection data within d days, the specific date of the first time infection data of the previous year is obtained, and the date is calculated as a certain day of the previous year. At this point, the seasonal weight subset corresponding to the d days starting from that certain day is obtained from the seasonal weight sequence T, represented as follows: Seasonal weighted subset Convert to a diagonal matrix representation as This yields the weight matrix for the second season, where x represents the start date of day d and y represents the end date of day d.

[0122] If d is 5, and the start date is January 1st, year B+1 (which is the first day of year B+1), then the specific date of the first time of the previous year is January 1st, year B (which is the first day of year B). Therefore, the seasonal weight subset consists of the values ​​in the seasonal weight sequence T corresponding to the sequence from 1 to 5. Seasonal weighted subset Convert to a diagonal matrix representation as That is, to obtain the unit area The weight matrix for the second season over 5 days. It should be noted that the specific values ​​mentioned above are merely an illustrative example; the actual values ​​should be determined based on the specific dates in practice.

[0123] S400. The second-season weight matrix, distance weight matrix, and case count matrix are fused to obtain the seasonally weighted case count matrix. The specific formula is as follows:

[0124]

[0125] S400. The second-season weight matrix, distance weight matrix, and case count matrix are fused to obtain the seasonally weighted case count matrix. The specific formula is as follows:

[0126]

[0127] The matrix of seasonally weighted case counts at the first infection location and the second infection location is represented as follows:

[0128]

[0129] S500 incorporates a seasonally weighted case count matrix into a prospective spatiotemporal scanning algorithm to provide early warning of infectious diseases in predicted locations.

[0130] Let the seasonally weighted case count matrix of the location of infection at the first time point in the second time point be: The total number of cases in all regions and all ranges. for:

[0131]

[0132] The expected number of cases per day within the unit area is:

[0133]

[0134] in, The number of cases across the entire study area during day d. This represents the number of cases within the entire study timeframe in region p.

[0135] The expected number of cases U in each bar graph scan window A A for:

[0136] in, This represents the number of cases within the entire study timeframe in the entire unit region p.

[0137] Let the actual number of cases observed within the cylindrical scanning window A be . Without considering changes in time and spatial interactions, then It conforms to the hypergeometric distribution model:

[0138]

[0139] when and When the total number of cases is extremely small, Approximately follows the mean The Poisson distribution can be used to determine whether cases in window A are abnormal. The null hypothesis of the likelihood ratio test is that the temporal and spatial distribution characteristics of the cases are completely random. The GLR for a specific window is:

[0140]

[0141] In the formula This represents the expected number of cases within window A after covariate adjustment, assuming the null hypothesis.

[0142] GLR reflects the probability of window clustering, so the window with the largest GLR value is definitely not randomly generated; its randomness is believable. However, to verify its non-randomness, a confidence analysis is needed. The null hypothesis is that the probability of the event occurring in space and time is completely random. Obtaining the probability distribution of the window's scan statistics is very difficult. Monte Carlo hypothesis testing can be used to calculate the P-value. Randomization detection is performed on potentially anomalous clusters. N randomly distributed datasets are generated based on the total number of clusters. The GLRs of these datasets are compared with the GLRs of the real dataset window. The GLRs of the N randomly generated datasets are sorted in ascending order, with the real GLR ranked in position S. Then, the P-value is... The higher the ranking, the smaller the P-value, and the less random the window.

[0143] The original case count matrix was rearranged using a Monte Carlo randomization method before scanning, and the corresponding probability estimate P under the log-likelihood ratio was calculated. The rearrangement was performed using a correlated rearrangement method. Due to spatiotemporal correlation, a fully random rearrangement could not be used. In correlated rearrangement, cases at each spatial node interact and influence each other. The same random seed, i.e., the entire case count matrix, was used when rearranging each column matrix. The rearrangement moves simultaneously in a column-by-column manner, ensuring that adjacent spatial positions remain close during the rearrangement process.

[0144] Example 2

[0145] In one specific embodiment, a set of first-time infection data of a certain infectious disease in region A is collected. The data spans one year and is granular at the daily level, specifically the infectious disease infection data from January 1, 2018 to December 31, 2018.

[0146] The complete data is: 3, 2, 3, 3, 3, 2, 2, 1, 1, 3, 3, 3, 4, 2, 1, 3, 2, 2, 4, 3, 2, 4, 3, 4, 1, 3, 3, 3, 1, 2, 3, 5, 5, 2, 5, 4, 3, 3, 5, 3, 4, 4, 2, 2, 2, 4, 3, 4, 3, 4, 4, 4, 3, 5, 2, 2, 5, 5, 5, 5, 6, 4, 6, 7, 4, 5, 5, 4, 7, 6, 6, 4, 7, 4, 6, 5, 7, 7, 4, 6, 4, 7, 6, 6, 7, 6, 5, 6, 4, 7, 7, 8, 9, 8, 7, 9, 8, 8, 8, 8, 8, 6, 9, 7, 9, 8, 9, 6, 6, 7, 7, 9, 8, 8, 6, 7, 7, 6, 9, 9, 6, 4, 6, 3, 4, 3, 5, 3, 4, 4, 4, 4, 6, 3, 4, 4, 5, 3, 3, 4, 5, 4, 6, 4, 6, 3, 3, 3, 6, 6, 5, 2, 2, 1, 0, 3, 3, 3, 1, 3, 1, 1, 3, 2, 2, 2, 3, 0, 2, 3, 3, 1, 2, 1, 1, 3, 3, 1, 1, 2, 2, 1, 0, 0, 1, 2, 1, 0, 1, 0, 0, 2, 2, 1, 1, 2, 0, 1, 2, 2, 0, 0, 0, 2, 0, 1, 0, 0, 1, 0, 0, 1, 2, 0, 2, 0, 2, 0, 1, 1, 1, 0, 0, 1, 2, 0, 0, 0, 0, 2, 2, 0, 1, 0, 0, 0, 0, 1, 2, 0, 0, 0, 0, 0, 3, 3, 1, 3, 3, 0, 3,3, 3, 3, 2, 3, 0, 2, 1, 3, 0, 0, 2, 0, 2, 0, 1, 3, 0, 1, 2, 5, 2, 2, 3, 5, 2,5, 2, 3, 2, 3, 2, 2, 4, 5, 5, 4, 4, 4, 4, 5, 5, 2, 4, 2, 5, 4, 5, 4, 3, 5, 6,6, 6, 4, 4, 6, 3, 6, 4, 4, 4, 6, 5, 4, 3, 3, 6, 3, 3, 3, 4, 6, 5, 5, 3, 6, 4,3, 5, 5, 8, 7,5, 7, 5, 6, 6, 5, 8, 6, 5, 6, 7, 5, 8, 7, 7, 6, 8, 5, 7, 6, 8, 8, 5, 8, 6, 8, 7, 8, 7.

[0147] Next, we collected the geographic data of the first infection location corresponding to the first infection data, specifically the latitude and longitude coordinates of the first infection location, as shown in Table 1 below:

[0148]

[0149] Secondary infection data were collected from September 1st to September 5th, B+1, at the aforementioned locations of initial infection, as shown in Table 2:

[0150]

[0151]

[0152] The confirmed case data undergoes spatiotemporal heterogeneity processing, and the calculation method is as follows:

[0153] 1. For a certain infectious disease in region A, the first-time infection data is organized chronologically to obtain the first time series. A Fourier transform is performed on the first time series to obtain the value of the first frequency domain data F(k):

[0154]

[0155] Where N represents the length of the first time-domain sequence, in this embodiment, N=365; n represents the nth data point in the first time-domain sequence, ranging from 0 to 364; k is the kth data point of the first frequency-domain data corresponding to the first time-domain sequence n, ranging from 0 to 364; j is the imaginary unit.

[0156] but:

[0157]

[0158] 2. Set the high-frequency portion of the first frequency domain data that is greater than a preset value to 0, retain the low-frequency portion, and obtain the second frequency domain data. If the preset value... The calculation process for the second frequency domain data U(k) is as follows:

[0159]

[0160] 3. Perform an inverse Fourier transform on the second frequency domain data U(k), and represent the result using f(n):

[0161]

[0162] but:

[0163]

[0164] 4. Take the real part of the result F(n) from the previous step; this real part is the seasonal weight value. This indicates the seasonal weight value:

[0165]

[0166] 5. Regarding the above seasonal weight values Normalization is performed to normalize it to the range [0,1]. Seasonal weighting value The minimum value in; for The maximum value in, The result is expressed as the normalized result:

[0167]

[0168] 6. Because the original data is weighted, all normalized data... Adding 1 gives us the seasonal weight sequence T, which can be transformed into a diagonal matrix H to obtain the seasonal weight matrix.

[0169]

[0170] 7. Spatial weighting is performed using a distance decay weight matrix. The latitude and longitude coordinates of the first infected location are shown in Table 1, and the location point ID is: 1. 9. The distance matrix is ​​obtained by calculating the distances between the first-time infection locations using geographic latitude and longitude coordinate data. The first-time infection locations are represented as follows:

[0171] The distance matrix is ​​represented by G, specifically:

[0172]

[0173] 8. Use the Gaussian decay function Calculate the spatial distance matrix G, where e represents the natural constant. This example indicates that the attenuation coefficient is a hyperparameter. The value is set to 0.5, and g represents the distance between the infected locations at the first moment. The spatial weight matrix S is calculated as follows:

[0174]

[0175] 9. The case count matrix for the second time period infection data from September 1st to September 5th, B+1, in the 9 regions of Region A is as follows, denoted by K:

[0176]

[0177] 10. Second-time infection data: September 1st to September 5th, B+1 year is the [number missing]th [period missing] of B+1 year. For each day, the corresponding seasonal weight subset obtained from the seasonal weight sequence T is represented as follows: Seasonal weight subset Convert to a diagonal matrix representation as The second season weight matrix is ​​obtained. ,but:

[0178]

[0179] 11. Take the dot product of the distance weight matrix S and the case count matrix K for the 9 regions in region A, and simultaneously combine it with the seasonality weight matrix. By performing a product and rounding up, we can obtain the seasonally weighted number of cases. ,but:

[0180]

[0181] By performing spatiotemporal heterogeneity correlation processing on the second-time infection data from September 1st to September 5th, B+1, for 9 areas in region A, the seasonally weighted case count was obtained. The data is shown in Table 3 below:

[0182]

[0183]

[0184] The data in Table 3 is then input into the SaTScan software, which has a built-in prospective spatiotemporal scanning algorithm to calculate prospective spatiotemporal rearrangement statistics.

[0185] Specifically, the data in Table 2 from September 1st to September 5th, B+1 (without undergoing spatiotemporal heterogeneity processing) were input into the SaTScan software, and the results are as follows: Figure 5 As shown, regions numbered 9 and 6 can be detected as having a high risk of clustering.

[0186] The prospective spatiotemporal rearrangement statistics were calculated using the confirmed case data in Table 3 after spatiotemporal heterogeneous weighting according to this invention, and the results are as follows: Figure 6As shown, regions numbered 9, 4, 7, and 1 have a high risk of clustering. After processing the raw data to improve its spatiotemporal heterogeneity, regions with clustering risk can be detected more effectively. This method is particularly useful for certain infectious diseases that exhibit strong temporal and spatial correlations; it allows for more accurate data preprocessing and warning results.

[0187] Example 3

[0188] An infectious disease early warning device based on spatiotemporal heterogeneity, such as Figure 7 As shown, including

[0189] The seasonal weight matrix construction module is used to collect the first-time infection data of the predicted location and process the first-time infection data to obtain the first-season weight matrix.

[0190] The distance weight matrix construction module is used to collect geographic data of the first infected location based on the first-time infection data, and to spatially weight the geographic data to obtain the distance weight matrix.

[0191] The case number matrix construction module is used to collect the second-time infection data of the first-time infection location at the second time, generate the case number matrix, and organize the first-season weight matrix corresponding to the second-time infection data time series to obtain the second-season weight matrix.

[0192] The seasonally weighted case count matrix construction module is used to obtain the seasonally weighted case count matrix by taking the dot product of the second-season weight matrix, the distance weight matrix, and the case count matrix.

[0193] The analysis and early warning module is used to input the seasonally weighted case count matrix into a prospective spatiotemporal scanning algorithm to provide early warning of infectious diseases in the predicted areas.

[0194] Example 4

[0195] A computer-readable storage medium storing computer instructions that, when executed by a processor, implement the infectious disease early warning method based on spatiotemporal heterogeneity as described in any of the above embodiments.

[0196] In specific implementations, computer-readable storage media may include magnetic disks, optical disks, read-only memory (ROM), random access memory (RAM), flash memory, hard disk drives (HDDs), or solid-state drives (SSDs); computer-readable storage media may also include combinations of the above types of memory.

[0197] Example 5

[0198] An electronic device, such as Figure 8 As shown, it includes at least one processor and a memory communicatively connected to the processor, wherein the memory stores instructions executable by at least one processor, which are executed by at least one processor to cause the processor to perform the spatiotemporal heterogeneous infectious disease early warning method as described in any of the above embodiments.

[0199] In practice, the number of processors can be one or more, and the processor can be a Central Processing Unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations of the above types of chips. A general-purpose processor can be a microprocessor or any conventional processor.

[0200] The memory and the processor can be connected to communicate via a bus or other means. The memory stores instructions that can be executed by at least one processor. The instructions are executed by at least one processor to cause the processor to perform the infectious disease early warning method based on spatiotemporal heterogeneity as described in the above method embodiments.

[0201] Furthermore, those skilled in the art should understand that although many problems exist in the prior art, each embodiment or technical solution of the present invention can be improved in only one or a few aspects, without necessarily solving all the technical problems listed in the prior art or the background art simultaneously. Those skilled in the art should understand that any content not mentioned in a claim should not be construed as a limitation on that claim.

[0202] Although this paper frequently uses terms such as first-time infection data, seasonal weight matrix, first-time infection location, geographic data, distance weight matrix, second-time infection data, case number matrix, seasonally weighted case number matrix, prospective spatiotemporal scanning algorithm, first time series, first frequency domain data, second frequency domain data, and seasonal weight value, the possibility of using other terms is not excluded. These terms are used merely for the convenience of describing and explaining the essence of the invention; interpreting them as any additional limitation would contradict the spirit of the invention. The terms "first," "second," etc. (if present) in the specification, claims, and accompanying drawings of the embodiments of the invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0203] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for early warning of infectious diseases based on spatiotemporal heterogeneity, characterized in that, Includes the following steps: S100: Collect the first-time infection data of the predicted location, and process the first-time infection data to obtain the first-season weight matrix; the specific steps of step S100 are as follows: S110. The first time-infection data is organized according to time sequence to obtain a first time series, and the first time series is subjected to Fourier transform to obtain the first frequency domain data. S120. Reset the portion of the first frequency domain data that is greater than a preset value to 0, and retain the portion that is less than a preset value to obtain the second frequency domain data; S130. Perform an inverse Fourier transform on the second frequency domain data to obtain the seasonal weight value; S140. After processing the seasonal weight values, we obtain the first seasonal weight matrix. The specific steps of S140 are as follows: S141. Normalize the seasonal weight values ​​to obtain the seasonal weight coefficients; S142. Add 1 to the seasonal weight coefficients to obtain the seasonal weight sequence, and transform the seasonal weight sequence into a diagonal matrix to obtain the first seasonal weight matrix; S200. Collect geographical data of the first-time infection location based on the first-time infection data, and spatially weight the geographical data to obtain a distance weight matrix; S300. Collect second-time infection data from the first-time infection location at the second time, generate a case count matrix, and organize the first-season weight matrix corresponding to the second-time infection data time series to obtain the second-season weight matrix; wherein, the seasonal weight subset corresponding to the d days starting from a certain day in the seasonal weight sequence is represented as follows. Seasonal weighted subset Convert to a diagonal matrix representation as This yields the weight matrix for the second season, where x represents the start date of day d and y represents the end date of day d. S400. Take the dot product of the second seasonal weight matrix, the distance weight matrix, and the case count matrix to obtain the seasonal weighted case count matrix; S500. The seasonally weighted case count matrix is ​​input into a prospective spatiotemporal scanning algorithm to provide infectious disease early warning for the predicted location.

2. The infectious disease early warning method based on spatiotemporal heterogeneity according to claim 1, characterized in that, The specific formula for performing a Fourier transform on the first time series to obtain the first frequency domain data is as follows: Where n represents the nth data point in the first time series, This represents the nth sampling point of the time-domain signal in the first time series, where j is the imaginary unit and N is the total number of data points in the first time series. This represents the first frequency domain data value of the k-th component after Fourier transform in the frequency domain; The formula for calculating the second frequency domain data is: in, This represents the first frequency domain data value of the k-th component after Fourier transform in the frequency domain. Represents the preset value. This represents the second frequency domain data value of the k-th component; The formula for calculating the seasonal weight value is: To extract the real part of F(n), Re represents the operator for extracting the real part of a complex number. Here, N is the seasonal weight value, j is the total number of data points, and e is the natural constant. This is the second frequency domain data value of the k-th component.

3. The infectious disease early warning method based on spatiotemporal heterogeneity according to claim 2, characterized in that, The specific formula for normalizing the seasonal weight values ​​and calculating the seasonal weight coefficient is as follows: in, Here are the seasonal weight values, and n is... The nth component in the seasonal weighting value, The minimum seasonal weight value, The largest seasonal weight value, This is the seasonal weighting coefficient; The formula for calculating the seasonal weighted series is: T is the seasonal weight sequence. Transforming the seasonal weight sequence into a diagonal matrix yields the first seasonal weight matrix: Where H is the weight matrix for the first season.

4. The infectious disease early warning method based on spatiotemporal heterogeneity according to claim 1, characterized in that, The specific steps for S200 are as follows: S210. Collect geographical data of the first infection location; S220. Calculate the distance between each of the first time-infected locations based on the geographical data to obtain a distance matrix; S230. Process the distance matrix using the Gaussian decay function to obtain the distance weight matrix.

5. The infectious disease early warning method based on spatiotemporal heterogeneity according to claim 4, characterized in that... The geographical data of the first infected location is collected as the coordinate data of the spatial location points of the first infected location. The distance between the first infected locations is calculated based on the coordinate data, and the distance matrix is ​​obtained as follows: Where m represents the total number of initial infection locations. Let m be the distance between the m-th first-time infection location and the 1st first-time infection location; Use Gaussian decay function The distance weight matrix is ​​calculated using the formula: e is the natural constant, and α represents the attenuation coefficient as a hyperparameter. 。 6. An infectious disease early warning device based on spatiotemporal heterogeneity, characterized in that: The seasonal weight matrix construction module is used to collect the first-time infection data of the predicted location and process the first-time infection data to obtain the first seasonal weight matrix. Specifically, the first-time infection data is organized according to time sequence to obtain the first time series, and the first time series is subjected to Fourier transform to obtain the first frequency domain data. The portion of the first frequency domain data that is greater than a preset value is reset to 0, and the portion that is lower than the preset value is retained to obtain the second frequency domain data. Perform an inverse Fourier transform on the second frequency domain data to obtain seasonal weight values; process and organize the seasonal weight values ​​to obtain the first seasonal weight matrix; Specifically, the process of processing and organizing the seasonal weight values ​​to obtain the first seasonal weight matrix involves: normalizing the seasonal weight values ​​to obtain seasonal weight coefficients; adding 1 to the seasonal weight coefficients to obtain a seasonal weight sequence; and converting the seasonal weight sequence into a diagonal matrix to obtain the first seasonal weight matrix. The distance weight matrix construction module is used to collect geographical data of the first-time infection location based on the first-time infection data, and to spatially weight the geographical data to obtain the distance weight matrix; The case count matrix construction module is used to collect second-time infection data from the first-time infection location at the second-time location, generate a case count matrix, and organize the first-season weight matrix corresponding to the time series of the second-time infection data to obtain a second-season weight matrix; wherein, the seasonal weight subset corresponding to the d-day period starting from a certain day in the seasonal weight sequence is represented as follows. Seasonal weighted subset Convert to a diagonal matrix representation as This yields the weight matrix for the second season, where x represents the start date of day d and y represents the end date of day d. The seasonally weighted case count matrix construction module is used to perform a dot product of the second seasonal weight matrix, the distance weight matrix, and the case count matrix to obtain the seasonally weighted case count matrix. The analysis and early warning module is used to input the seasonally weighted case number matrix into a prospective spatiotemporal scanning algorithm to provide early warning of infectious diseases for the predicted location.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, which, when executed by a processor, implement the infectious disease early warning method based on spatiotemporal heterogeneity as described in any one of claims 1-5.

8. An electronic device, characterized in that: The method includes at least one processor and a memory communicatively connected to the processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the processor to perform the spatiotemporal heterogeneous infectious disease early warning method as described in any one of claims 1-5.